Skip to content
Home

PDFCanon

Strips hidden revisions, scripts and metadata from PDFs and rewrites them deterministically so identical logical documents hash to identical bytes

The product is a PDF normalization API and infrastructure service that strips hidden incremental revisions, embedded JavaScript, metadata and non-deterministic structures and rewrites files deterministically so the same logical document produces the same SHA-256 hash bit-for-bit. It sells to developers and operations teams in B2B organizations that need stable hashing for deduplication, integrity checks, auditing and signing. Compared to running qpdf alone, it is positioned as an 11-stage pinned toolchain that delivers deterministic output across regions and hosts with active-content removal, PDF/A validation, audit reports and canonical hashing.

Key features

  • Deterministic PDF structure rewriting
  • Flatten incremental revisions
  • Rebuild XRef table sequentially
  • Metadata scrubbing
  • Active-content removal (JS, embedded files)
  • High throughput REST API
  • HMAC-signed webhooks with retries
  • Detached cryptographic toolchain signature
  • PDF/A compliance validation via veraPDF
  • Canonical SHA-256 hashing and audit report
  • Pinned versioned toolchain
  • Hosted isolated sandbox processing
Social posts
  • No social media activity within the last 30 days
GTM channels
  • API
  • Docs
ICP
  • Software developers
  • DevOps sre teams
  • Engineering teams