Listed July 16, 2026 as a public preview (mistral-ocr-4-1), the updated model refines Mistral's Document AI stack with native paragraph-level bounding boxes, structural block labels, and block-level confidence scores.
2026-08-13 · by Moe Ameen
Mistral surfaced OCR 4.1 on July 16, 2026 as a public preview, served as `mistral-ocr-4-1` on its Premier tier. The French AI lab describes it as the latest OCR service powering its Document AI stack — the layer that turns PDFs, slide decks, scans, and images into clean, structured, machine-readable text instead of a flat character dump.
The update is about the structure of that output. Mistral highlights native paragraph-level bounding box extraction, structural block labels, and block-level confidence scores, with the extracted text formatted as markdown. It exposes OCR, bounding-box extraction, and structured annotations through the `/v1/ocr` endpoint and high-volume batching through `/v1/batch`. This is an iteration on Mistral OCR 4, which launched June 23, 2026; Mistral has not published a granular head-to-head against OCR 4, so 4.1 reads as a refinement of the same document-intelligence approach rather than a leap with new benchmark claims.
Mistral's documentation lists OCR 4.1 at €3.5 per 1,000 pages for standard OCR and €4.38 per 1,000 annotated pages when you request structured annotations. It continues the OCR 4 line, which covers 170 languages across 10 language groups and can be self-hosted on a single container for data residency. Because 4.1 is a preview, pricing and availability may still shift.
The move to make on this news is not a single hot take — it is a backlog. OCR 4's launch was about reacting to a report the day it drops; OCR 4.1's batch endpoint and per-page pricing are about the archive you already own. Point OCR 4.1 at a pile of past reports, decks, case studies, or transcripts, batch-extract them into clean markdown, and drop that library into Kompozy as sources. Then let Kompozy's autopilot fan each document into a carousel, an Infographic Photo, a LinkedIn post and an X thread in your voice via the Persona Brief, a blog article, and an email newsletter — and schedule the whole calendar across your nine connected platforms.
The leverage is turning a one-time extraction into weeks of content. A folder of documents that has been sitting untouched is, after an OCR 4.1 batch run, a queue of source material; Kompozy turns that queue into a scheduled content calendar without you writing each piece by hand. Because 4.1 preserves tables and structure, the generated infographics and carousels track the real figures instead of a garbled paste — so the backlog you mine ships as accurate, on-brand posts, not approximate ones.
Mistral OCR 4.1 surfaced on July 16, 2026 as a public preview, served as mistral-ocr-4-1 on Mistral's Premier tier. It is the latest OCR service powering Mistral's Document AI stack.
OCR 4.1 is a preview iteration on OCR 4 (June 23, 2026), emphasizing native paragraph-level bounding box extraction, structural block labels, and block-level confidence scores, with markdown output. Mistral has not published a granular head-to-head, so treat it as a refinement rather than a benchmarked leap.
Mistral's documentation lists OCR 4.1 at €3.5 per 1,000 pages for standard OCR and €4.38 per 1,000 annotated pages for structured annotations. As a preview model, pricing can change, so confirm current rates in Mistral's docs.
Batch-extract a backlog of documents into clean, structured markdown, then feed that library into a content engine like Kompozy to generate carousels, infographics, blogs, newsletters, and platform-native posts, and schedule them across nine platforms.