Agent-12 — the local agent leaderboard

Which model can actually be your agent on your own hardware? Measured by doing real agent tasks in a sandboxed working directory — judged by the filesystem, never by the model's prose. Temperature 0, fixed caps, fresh sandbox per task, one variable moved per comparison.

Local models

Model Easy (12)time Hard (8)time tok/s Hardware
Qwen3-Coder-30B-A3B (MLX 8-bit)30B MoE, 3B active · native tool calls12/1242.6s7/8391.5s84.3Apple M5, 128 GB unified memory
Qwen3.6-35B-A3B (MLX 8-bit)35B MoE, 3B active · native tool calls12/1264.2s8/8125.2s46.2Apple M5, 128 GB unified memory
Gemma 4 31B (MLX 4-bit)31B dense · prompted-XML dialect11/1292.1s8/8347.9s25.6Apple M5, 128 GB unified memory
DeepSeek V4 Flash (2-bit, 0731 imatrix)284B MoE · ds4.c engine12/12203.2s8/8550.8s8.4Apple M5, 128 GB unified memory
DeepSeek V4 Flash (2-bit, original) · retired 8/10, weights deleted284B MoE · ds4.c engine12/12376.1s8/8842.4s8.4Apple M5, 128 GB unified memory

Cloud reference (not competing — context)

Model Easy (12)time Hard (8)time tok/s Where it runs
Claude Sonnet 5 (cloud reference)same engine, same tasks, remote model12/12122.0s8/8131.0sAnthropic API

How to trust a number here:

Tasks, judges, runner, and the validation gate are open: github.com/nicedreamzapp/agent12 · methodology · contamination policy. Cloud reference times include network/API round-trips — that is the honest end-to-end experience. Updated 2026-08-10.