← All posts / Models

Cohere Parse: The $1.50-Per-1,000-Page Specialist Picking Apart Enterprise Documents

Cohere's Parse (parse-v5.0) is a 2.3B-parameter vision-language model that turns contracts, invoices, and filings into clean Markdown at $1.50 per 1,000 pages — scoring 79.2 on ParseBench, beating Mistral OCR 4, and running at 2,160 pages per minute on an 8x H100 node.

Cohere Parse: The $1.50-Per-1,000-Page Specialist Picking Apart Enterprise Documents

On August 27, 2026, Toronto-based Cohere released Parse, a document-intelligence model with a deceptively simple contract: complex multimodal files go in — PDFs, PowerPoint decks, scanned JPEGs — and clean, structured Markdown comes out. The model, registered as parse-v5.0 and tracked at just 2.3 billion parameters by the independent model database AI/TLDR, reads text, tables, and embedded images in a single pass with no separate OCR step, and returns them in reading order with tables preserved as HTML and visual elements annotated with bounding-box coordinates.

The launch is small in parameter count but large in what it signals: document parsing is splitting off from the “just let the general LLM handle it” approach into an independent layer of small, specialized, aggressively priced models — aimed squarely at the hyperscaler document-AI business that has quietly underpinned enterprise digitization for a decade.

The problem: where enterprise knowledge actually lives

The first mile of enterprise RAG has long been stuck at document parsing. Contracts, insurance policies, invoices, scientific papers — the places where enterprise knowledge actually lives are precisely the hardest to feed into a model. Tables get mangled, reading order breaks, and styles that carry meaning — strike-throughs, italics, footnotes — simply vanish.

The ParseBench evaluation data makes the problem brutally concrete: AWS Textract scores just 2.8 on the Semantic Formatting dimension, and Google Document AI manages only 33.0. Traditional document-intelligence services are essentially blind whenever “formatting is meaning” — when the fact that a clause is struck through, or a number sits in a superscript footnote, changes what the document actually says.

Optical character recognition, the software behind decades of digitization efforts, was never built for this. OCR extracts characters; it cannot reliably tell a heading from a caption or keep a table’s rows and columns aligned through a page break. Vision-language models like Parse are designed to capture that structure — headings, tables, reading order — in one pass. Other vendors in the space, including Mistral and Reducto, have been marrying the two capabilities into single products, and the competition is now genuinely fierce.

The benchmark: 79.2, parked between specialists and frontier LLMs

In the ParseBench evaluation Cohere submitted — averaged across three dimensions — Parse scores 79.2, ahead of Mistral OCR 4 at 74.5, Databricks AI Parse at 72.4, and LlamaParse’s Cost Effective tier at 78.3. The gap against hyperscaler services is far larger: more than 20 points above both AWS Textract and Google Document AI. By dimension, Tables (87.0) and Content Faithfulness (86.6) are the clear strengths, while Semantic Formatting (64.0) still trails the frontier.

Cohere is equally explicit about the boundary. In their evaluation set, only three general-purpose frontier LLMs beat Parse — GPT-5.5 (84.4), Anthropic’s Opus 4.8 (84.3), and Gemini 3.5 Flash (81.8) — and all three are significantly larger general models. A 2.3B specialist pushing to the edge of the frontier at a fraction of the cost is the entire thesis of the release.

Two caveats deserve honest airtime. First, ParseBench was built by LlamaIndex — the company behind LlamaParse, a competing product — so the yardstick itself has a horse in the race. Second, Cohere’s headline figure covers only three of ParseBench’s five dimensions, leaving out chart understanding and layout localization. Cohere frames these as product-scope decisions rather than capability gaps, with chart data extraction planned for the next version, and its footnotes state that scores use the corrected August 2026 evaluation rules — which fixed a bold/heading-detection bug that previously inflated Semantic Formatting — with all competitor models re-scored under the same rules. That level of methodological transparency is rarer than it should be.

Pricing and throughput: parsing as a utility

At $1.50 per 1,000 pages through the Cohere API, Parse is priced like infrastructure, not intelligence. Throughput is production-grade: 36 pages per second — 2,160 pages per minute — on a single 8x H100 node, roughly 1.4x dots.mocr and 2.2x Chandra OCR 2 in the official benchmark, with all models served on vLLM.

For regulated industries, Parse runs in private clouds or fully on-premises, or through Cohere’s single-tenant Model Vault offering: 23% cheaper than the API at 50% GPU utilization, and up to 61% cheaper at full utilization. Cohere does the math on a concrete example: an accounts-payable workflow processing roughly 13 million pages per month saves about $12,000 monthly — $144,000 a year — by moving from the API to Model Vault, and about $1.47 million per year versus a hyperscaler document service priced at $10 per 1,000 pages, for that single workflow alone.

Parse is generally available today via the Cohere API, Model Vault, Microsoft Foundry, and AWS SageMaker, and ships inside the Compass retrieval stack alongside Embed and Rerank as a document-to-answer pipeline. It works across what Cohere calls nine major world languages.

Why a 2.3B model is the story

Cohere was founded in 2019 by Aidan Gomez and Nick Frosst — both former Google Brain researchers, with Gomez a co-author of the Transformer paper “Attention Is All You Need” — together with Ivan Zhang. The company has deliberately positioned itself for regulated industries — finance, healthcare, manufacturing, energy, and the public sector — where documents are dense and mistakes are expensive. Its suite now spans generation (Command), embeddings (Embed), search reranking (Rerank), speech-to-text (Transcribe), and now parsing. In April 2026 Cohere agreed to acquire Germany’s Aleph Alpha, followed by Reliant AI in May, extending its reach into the biopharmaceutical sector.

The real lesson of Parse isn’t one benchmark score. It’s confirmation of a structural trend: as frontier LLMs get more expensive per token, the economic pressure to route work to the smallest capable specialist grows. A 2.3B model that preserves tables at 87.0 and charges per thousand pages changes the arithmetic of every pipeline built above it — because document reading is the plumbing behind most AI at work. Before an assistant can answer a question about a contract, something has to turn that contract into text it can read. A cheaper, faster converter lowers the cost of everything downstream.

The rational enterprise strategy now falls out directly: route the small volume of accuracy-critical documents to frontier LLMs, and hand the million-page pipelines to a $1.50-per-thousand-page specialist. The first mile of RAG finally has a price tag — and it’s a small one.