Skip to content
Home

Hanji

Parses complex documents into structured data with cited fields, including tables and figures, and flags ungrounded values.

Hanji is a document parsing and extraction API that converts complex documents such as medical insurance forms, government paperwork, and scanned faxes into structured, AI-ready data. It targets developers and operations teams in healthcare, insurance, and government sectors, including enterprises with high-volume needs. The service offers a grounded extraction pipeline where every value is cited back to its source, with automatic re-reading of failed pages and flagging of unverifiable fields. It is delivered as a cloud API with a free tier and enterprise options, including HIPAA compliance and uptime SLAs.

Key features

  • Parse any document into grounded structure
  • Extract fields with cited sources
  • Automatic re-read of failed pages
  • Flag unverifiable values for review
  • Bounding boxes on every element
  • Schema-based extraction with null for ungrounded
  • Tune on edge cases for accuracy
  • HIPAA compliance with BAA
  • SOC 2 Type II in progress
  • Rate limiting and secret scanning
Social posts
  • No social media activity within the last 30 days
GTM channels
  • API
  • Docs
ICP
  • Software developers
  • Engineering teams
  • Enterprises