Skip to content
Home

OpenParser

Converts PDFs and images into structured data with field-level source, confidence, transforms, and review history

An API that converts PDFs and images into structured data and retains source blocks, confidence scores, transforms, and review history for each extracted field to support verification. It sells to developers and data teams in businesses that need to build document extraction and processing pipelines. It is delivered as a hosted API with nine OCR models accessible through a unified endpoint, including open-weight hosted OCR, with open-source availability.

Key features

  • PDF and image to structured data extraction
  • Field-level source grounding and lineage
  • Confidence scores per field and token
  • Human review for low-confidence values
  • Nine hosted OCR models via unified API
  • Open-weight hosted OCR
  • Typed blocks and markdown output
  • POST /parse API endpoint
  • Switch OCR model via ocr_model parameter
Social posts
  • No social media activity detected
GTM channels
  • Blog
  • API
  • Docs
ICP
  • Software developers
  • Data analytics teams
  • Engineering teams