Data Catalog
Auditable task scorecards from Jigsaw’s RL environments, reconstructed from real work and evaluated across frontier models with deterministic grading and full provenance.
ML Engineering
Tasks that evaluate the real day-to-day work of an ML engineer: running GPU training jobs, repairing framework internals, debugging training runs, and tracing data pipelines.
LeaderboardDataTasks
View more →
gpt-5.5
86% ± 7%gpt-5.4
70% ± 8%0%10%20%30%40%50%60%70%80%90%100%
0-1 SWE
Greenfield engineering tasks where agents build complete systems from a blank workspace and a dense specification: binary codecs, constraint engines, and protocol implementations, graded by hidden deterministic test suites.
LeaderboardDataTasks
View more →
claude-opus-4-6
85% ± 5%gpt-5.5
78% ± 3%claude-opus-4-7
75% ± 6%0%10%20%30%40%50%60%70%80%90%100%
Decompilation
Matching-decompilation tasks where agents lift ARM assembly from shipped retail binaries back into C that recompiles byte-identical, underpinning cybersecurity and vulnerability research.
Calibration in progress · scorecards coming soon

