ML Engineering
Tasks that evaluate the real day-to-day work of an ML engineer: running GPU training jobs, repairing framework internals, debugging training runs, and tracing data pipelines.
The ML Engineering leaderboard
We built the ML Engineering collection to evaluate the real day-to-day work of an ML engineer: running GPU training jobs, repairing deep framework internals, debugging training runs, and tracing data pipelines.
Tasks range from GPU-backed training runs and source-tree repairs in production ML repositories to CPU sandboxes with artifacts extracted from real training workflows. Network access is disabled where not required; grading is deterministic via functional checks or rubrics anchored to ground truth recoverable from the artifacts.
The leaderboard aggregates scores across the 3 tasks with published scorecards. Each model was evaluated over 5-10 independent rollouts per task, graded deterministically with no LLM judge.
Model performance is reported only for tasks with published scorecards. Reach out for the full task listing.
The full task set
The full ML Engineering catalog contains 161 tasks spanning GPU training jobs, framework source-tree repair, training-run debugging, and data-pipeline analysis. The 3 tasks above are the published calibration sample; every task that publishes goes through the same panel evaluation and scorecard process before its data appears here.
Request the full task listing →
