Data/ML Engineering

ML Engineering

Tasks that evaluate the real day-to-day work of an ML engineer: running GPU training jobs, repairing framework internals, debugging training runs, and tracing data pipelines.

The ML Engineering leaderboard

We built the ML Engineering collection to evaluate the real day-to-day work of an ML engineer: running GPU training jobs, repairing deep framework internals, debugging training runs, and tracing data pipelines.

Tasks range from GPU-backed training runs and source-tree repairs in production ML repositories to CPU sandboxes with artifacts extracted from real training workflows. Network access is disabled where not required; grading is deterministic via functional checks or rubrics anchored to ground truth recoverable from the artifacts.

The leaderboard aggregates scores across the 3 tasks with published scorecards. Each model was evaluated over 5-10 independent rollouts per task, graded deterministically with no LLM judge.

Model performance is reported only for tasks with published scorecards. Reach out for the full task listing.

Catalog Tasks
161
full inventory
Published Scorecards
3
full calibration data
Models Evaluated
13
frontier panel
Rollouts Scored
290
across published scorecards
ModelMean Score
gpt-5.5
86% ± 7%
glm-5.2
76% ± 13%
gpt-5.4
70% ± 8%
gemini-3.5-flash
64% ± 12%
claude-sonnet-4-6
63% ± 10%
claude-opus-4-7
59% ± 13%
kimi-k2.5
57% ± 9%
claude-opus-4-5
57% ± 11%
glm-5.1
54% ± 14%
kimi-k2.6
53% ± 15%
qwen3.6-plus
52% ± 14%
claude-haiku-4-5
34% ± 13%
grok-4.3
20% ± 8%
0%20%40%60%80%100%
161
tasks in catalog

The full task set

The full ML Engineering catalog contains 161 tasks spanning GPU training jobs, framework source-tree repair, training-run debugging, and data-pipeline analysis. The 3 tasks above are the published calibration sample; every task that publishes goes through the same panel evaluation and scorecard process before its data appears here.

Request the full task listing

Verifiable data to hillclimb frontier models.

Model progress is moving into the long-horizon work where models are actually deployed, bound by the realism of the data behind it.

We reconstruct real workflows into environments models are trained and measured inside, graded deterministically and traceable end to end. Models rise to the fidelity of their data, and we build the ground truth they climb.