Data/0-1 SWE

0-1 SWE

Greenfield engineering tasks where agents build complete systems from a blank workspace and a dense specification: binary codecs, constraint engines, and protocol implementations, graded by hidden deterministic test suites.

The 0-1 SWE leaderboard

We built the 0-1 SWE collection to evaluate whether agents can build working software from scratch against a detailed specification. The agent starts in an empty workspace with no tests, no skeleton, and no third-party packages; grading is fully deterministic via hidden pytest suites injected only at grade time, each scored all-or-nothing on its exit code.

The tasks measure specification fidelity at scale, cross-module integration discipline, and self-verification without a visible test harness. The collection is positioned as RL training data: frontier-tier discrimination varies by task, and each scorecard reports its own frontier-tier discrimination figure.

The leaderboard aggregates scores across the 1 tasks with published scorecards. Each model was evaluated over 10 independent rollouts per task, graded deterministically with no LLM judge.

Model performance is reported only for tasks with published scorecards. Reach out for the full task listing.

Catalog Tasks
179
full inventory
Published Scorecards
2
full calibration data
Models Evaluated
11
frontier panel
Rollouts Scored
107
across published scorecards
ModelMean Score
claude-opus-4-6
85% ± 5%
gpt-5.5
78% ± 3%
claude-opus-4-7
75% ± 6%
claude-sonnet-4-6
59% ± 9%
glm-5.1
32% ± 10%
claude-haiku-4-5
36% ± 10%
gemini-3.1-pro-preview
31% ± 11%
kimi-k2.5
22% ± 7%
qwen3.6-plus
24% ± 8%
grok-4.3
4% ± 4%
gemini-2.5-pro
3% ± 3%
0%20%40%60%80%100%
179
tasks in catalog

The full task set

The full 0-1 SWE catalog contains 179 greenfield engineering tasks built to the same construction: an original multi-module specification, a reference solution, and hidden deterministic pytest suites scored all-or-nothing at grade time. The 1 tasks above are the published calibration sample; every task that publishes goes through the same panel evaluation and scorecard process before its data appears here.

Request the full task listing

Verifiable data to hillclimb frontier models.

Model progress is moving into the long-horizon work where models are actually deployed, bound by the realism of the data behind it.

We reconstruct real workflows into environments models are trained and measured inside, graded deterministically and traceable end to end. Models rise to the fidelity of their data, and we build the ground truth they climb.