Cohere releases Parse, a vision-language model for enterprise document processing
Cohere has released Parse, a vision-language model purpose-built to turn complex enterprise documents into structured, machine-readable text. The company calls it "a high-throughput vision parsing model with the strongest price-performance profile on the market," aimed at organizations that need to process large volumes of PDFs, scanned forms, and reports for retrieval and automation pipelines.
What's new
Parse reads tables, forms, diagrams, and embedded images and converts them into clean Markdown while preserving document structure. In Cohere's words, "Parse detects and understands key visual elements - such as tables and embedded images - and returns clean Markdown files for downstream processing and application." The model also returns bounding boxes for visual elements, which lets downstream systems ground retrieved text back to its position on the original page.
Key specs from the announcement:
- Supports nine major world languages
- Throughput of 4.5 pages per second on a single instance, scaling to 36 pages per second on an 8-GPU H100 node
- Scored 79.2 on Cohere's ParseBench evaluation, ahead of AWS Textract (53.3) and Google Document AI (57.3) by the company's own benchmarking
- API pricing of $1.50 per 1,000 pages
- Available through the Cohere API, Cohere's Model Vault, Microsoft Foundry, and AWS SageMaker
Cohere says self-hosting Parse through Model Vault cuts costs further as usage scales — 23% savings at 50% utilization and up to 61% at full utilization compared to calling the hosted API, and roughly $1.47 million a year in savings against a hyperscaler OCR service priced at $10 per 1,000 pages for a customer processing 13 million pages a month.
Context
Document parsing has become a competitive front in enterprise AI as companies race to feed unstructured PDFs, scanned contracts, and reports into retrieval-augmented generation and agent pipelines. Incumbent cloud services like AWS Textract and Google Document AI have long handled this extraction step, but they were built before the current wave of LLM-based document understanding and are priced and tuned for older OCR-style workloads rather than vision-language accuracy. Parse is Cohere's answer aimed squarely at that gap, and its distribution through Microsoft Foundry and AWS SageMaker — alongside Cohere's own API and Model Vault — signals an effort to meet enterprise buyers inside the cloud platforms they already use.
Why it matters
Document ingestion is a bottleneck for nearly every enterprise AI deployment: a customer support agent, a compliance search tool, or a financial-research assistant is only as good as the text it can reliably extract from the PDFs and scans it's fed. By undercutting the two dominant cloud OCR services on Cohere's own accuracy benchmark while also pricing meaningfully below at least one hyperscaler alternative, Parse is a direct challenge to treating document extraction as a commodity utility bundled into AWS or Google Cloud. If the accuracy and cost claims hold up under independent testing, Parse gives enterprises processing large volumes of mixed-format documents a reason to route that traffic through Cohere instead, and further cements document intelligence as a full model category rather than a preprocessing afterthought.
Corroborating sources
- Cohere
https://cohere.com/blog/parse
“Parse detects and understands key visual elements - such as tables and embedded images - and returns clean Markdown files for downstream processing and application.”
- Marktechpost
https://www.marktechpost.com/2026/08/27/cohere-releases-parse-5-parse-v5-0-a-2-3b-vision-language-model-that-turns-enterprise-documents-into-markdown/