Pencil Puzzle Bench
Benchmark for evaluating AI models on multi-step verifiable reasoning using pencil puzzles.
A benchmark dataset of 62,231 pencil puzzles across 20 types, used to evaluate AI models on multi-step verifiable reasoning. It provides a leaderboard of 51 frontier models tested on 300 puzzles, with metrics for direct and agentic performance and cost. The benchmark targets AI developers and researchers, offering a standardized evaluation tool for model reasoning capabilities. It is delivered as an open dataset with associated paper, leaderboard, and puzzle resources.
Key features
- 62,231 puzzle dataset
- 20 puzzle types
- 51 models tested
- 17k evaluation runs
- Leaderboard with cost metrics
- Direct and agentic evaluation modes
Social posts
- No social media activity detected
GTM channels
- Blog
ICP
- Software developers
- Data analytics teams
- Educators institutions