Skip to content
Home

Pencil Puzzle Bench

Benchmark for evaluating AI models on multi-step verifiable reasoning using pencil puzzles.

A benchmark dataset of 62,231 pencil puzzles across 20 types, used to evaluate AI models on multi-step verifiable reasoning. It provides a leaderboard of 51 frontier models tested on 300 puzzles, with metrics for direct and agentic performance and cost. The benchmark targets AI developers and researchers, offering a standardized evaluation tool for model reasoning capabilities. It is delivered as an open dataset with associated paper, leaderboard, and puzzle resources.

Key features

  • 62,231 puzzle dataset
  • 20 puzzle types
  • 51 models tested
  • 17k evaluation runs
  • Leaderboard with cost metrics
  • Direct and agentic evaluation modes
Social posts
  • No social media activity detected
GTM channels
  • Blog
ICP
  • Software developers
  • Data analytics teams
  • Educators institutions