// AI NEWS · MODEL RELEASE

Mistral Surfaces OCR 4.1, a Preview Document-AI Model With Paragraph-Level Bounding Boxes and Block Confidence

Listed July 16, 2026 as a public preview (mistral-ocr-4-1), the updated model refines Mistral's Document AI stack with native paragraph-level bounding boxes, structural block labels, and block-level confidence scores.

2026-08-13 · by Moe Ameen

What happened

Mistral surfaced OCR 4.1 on July 16, 2026 as a public preview, served as `mistral-ocr-4-1` on its Premier tier. The French AI lab describes it as the latest OCR service powering its Document AI stack — the layer that turns PDFs, slide decks, scans, and images into clean, structured, machine-readable text instead of a flat character dump.

The update is about the structure of that output. Mistral highlights native paragraph-level bounding box extraction, structural block labels, and block-level confidence scores, with the extracted text formatted as markdown. It exposes OCR, bounding-box extraction, and structured annotations through the `/v1/ocr` endpoint and high-volume batching through `/v1/batch`. This is an iteration on Mistral OCR 4, which launched June 23, 2026; Mistral has not published a granular head-to-head against OCR 4, so 4.1 reads as a refinement of the same document-intelligence approach rather than a leap with new benchmark claims.

Mistral's documentation lists OCR 4.1 at €3.5 per 1,000 pages for standard OCR and €4.38 per 1,000 annotated pages when you request structured annotations. It continues the OCR 4 line, which covers 170 languages across 10 language groups and can be self-hosted on a single container for data residency. Because 4.1 is a preview, pricing and availability may still shift.

Why it matters for creators

  • Cleaner document structure means cleaner source material. Paragraph-level bounding boxes and block labels keep tables and figures intact, so what you extract is closer to usable content and further from a text-cleanup chore.
  • Batch endpoints plus per-page pricing make processing an entire document backlog cheap — an archive of reports or decks you never mined becomes a content library almost for free.
  • It is still extraction only. Nothing here writes a caption, builds a carousel, or publishes a post; the model got better at the upstream half of a content workflow, not the downstream one.
  • Preview status means the numbers and access can change. If you standardize on it, keep an eye on Mistral's docs for the move from preview to general availability.
  • Structured annotations at €4.38 per 1,000 pages are a small line item next to the content you can generate from a well-parsed document — the extraction is the cheap part.

How to act on this with Kompozy

The move to make on this news is not a single hot take — it is a backlog. OCR 4's launch was about reacting to a report the day it drops; OCR 4.1's batch endpoint and per-page pricing are about the archive you already own. Point OCR 4.1 at a pile of past reports, decks, case studies, or transcripts, batch-extract them into clean markdown, and drop that library into Kompozy as sources. Then let Kompozy's autopilot fan each document into a carousel, an Infographic Photo, a LinkedIn post and an X thread in your voice via the Persona Brief, a blog article, and an email newsletter — and schedule the whole calendar across your nine connected platforms.

The leverage is turning a one-time extraction into weeks of content. A folder of documents that has been sitting untouched is, after an OCR 4.1 batch run, a queue of source material; Kompozy turns that queue into a scheduled content calendar without you writing each piece by hand. Because 4.1 preserves tables and structure, the generated infographics and carousels track the real figures instead of a garbled paste — so the backlog you mine ships as accurate, on-brand posts, not approximate ones.

Quick takeaways

  • Mistral OCR 4.1 (mistral-ocr-4-1) surfaced July 16, 2026 as a public preview powering Mistral's Document AI stack.
  • It refines the OCR 4 line with paragraph-level bounding boxes, structural block labels, and block-level confidence scores.
  • Pricing is listed at €3.5 per 1,000 pages standard and €4.38 per 1,000 annotated pages; 170 languages and single-container self-hosting carry over.
  • It extracts text only — pair it with a content engine like Kompozy to batch-turn a document backlog into a scheduled content calendar.

Frequently asked questions

When was Mistral OCR 4.1 released?

Mistral OCR 4.1 surfaced on July 16, 2026 as a public preview, served as mistral-ocr-4-1 on Mistral's Premier tier. It is the latest OCR service powering Mistral's Document AI stack.

What is new in Mistral OCR 4.1 versus OCR 4?

OCR 4.1 is a preview iteration on OCR 4 (June 23, 2026), emphasizing native paragraph-level bounding box extraction, structural block labels, and block-level confidence scores, with markdown output. Mistral has not published a granular head-to-head, so treat it as a refinement rather than a benchmarked leap.

How much does Mistral OCR 4.1 cost?

Mistral's documentation lists OCR 4.1 at €3.5 per 1,000 pages for standard OCR and €4.38 per 1,000 annotated pages for structured annotations. As a preview model, pricing can change, so confirm current rates in Mistral's docs.

How can creators use Mistral OCR 4.1?

Batch-extract a backlog of documents into clean, structured markdown, then feed that library into a content engine like Kompozy to generate carousels, infographics, blogs, newsletters, and platform-native posts, and schedule them across nine platforms.

Related news

← All AI news · Get started →