Skip to content
Home

OpenDataLoader

Converts PDFs to LLM-ready Markdown and JSON with reading order, tables, bounding boxes, and auto-tagging for accessibility.

A PDF parsing tool that converts PDFs into structured Markdown and JSON for AI pipelines, solving issues like scrambled reading order, lost table structure, and missing source coordinates. It targets developers and data teams building RAG systems, as well as organizations needing PDF accessibility remediation. The tool is open-source, local-first, and offers hybrid OCR with optional LLM enhancement, achieving top benchmark scores.

Key features

  • Structured JSON output with semantic types
  • Bounding boxes for citations
  • XY-Cut++ reading order algorithm
  • Noise filtering (headers, footers, watermarks)
  • LangChain integration
  • No GPU required, fast rule-based heuristics
  • Local-first processing
  • Multi-language SDK (Python, Node.js, Java)
  • Table detection with merged cells
  • List and heading hierarchy detection
  • Image extraction with captions
  • Tagged PDF support
  • AI safety filters for prompt injection
Social posts
  • No social media activity within the last 30 days
GTM channels
  • Marketplace
  • Community
  • API
  • Docs
ICP
  • Software developers
  • Data analytics teams
  • Engineering teams