Agent-12 — the local agent leaderboard

Which model can actually be your agent on your own hardware? Measured by doing real agent tasks in a sandboxed working directory — judged by the filesystem, never by the model's prose. Temperature 0, fixed caps, fresh sandbox per task, one variable moved per comparison.

Local models

Measured in Agent-12's own reference harness (Anvil, a ~550-token native terminal engine with per-model tool-call dialects) — not inside Claude Code. The same model inside Claude Code as installed scores and times differently, because that harness sends far more tool tokens per turn; the harness is held fixed here so models can be compared with each other. See METHODOLOGY and the measured gap in issue #1.

Model Easy (12)time Hard (8)time tok/s Hardware
Qwen3-Coder-30B-A3B (MLX 8-bit)30B MoE, 3B active · native tool calls12/1242.6s7/8391.5s84.3Apple M5, 128 GB unified memory
Qwen3.6-35B-A3B (MLX 8-bit)35B MoE, 3B active · native tool calls12/1264.2s8/8125.2s46.2Apple M5, 128 GB unified memory
Gemma 4 31B (MLX 4-bit)31B dense · prompted-XML dialect11/1292.1s8/8347.9s25.6Apple M5, 128 GB unified memory
Qwen3.8-27B (MLX 8-bit, dense)27B dense · native tool calls · 8/8 hard at an 8000-token budget, 16.9x slower than Qwen3.6 (writeup)12/12350.7s7/81155.9s—Apple M5, 128 GB unified memory
Qwen3.8-27B (MLX 8-bit, dense) + DFlash 2 drafter27B dense · native tool calls · incoai DFlash2 drafter, ~3x faster than without it, still ~3x slower than Qwen3.6; tok/s here includes re-prefill every turn (writeup, 2026-09-19)12/12122.0s7/8401.4s28.7Apple M5, 128 GB unified memory
DeepSeek V4 Flash (2-bit, 0731 imatrix)284B MoE · ds4.c engine12/12203.2s8/8550.8s8.4Apple M5, 128 GB unified memory
DeepSeek V4 Flash (2-bit, original) · retired 8/10, weights deleted284B MoE · ds4.c engine12/12376.1s8/8842.4s8.4Apple M5, 128 GB unified memory

Cloud reference (not competing — context)

Model Easy (12)time Hard (8)time tok/s Where it runs
Claude Sonnet 5 (cloud reference)same engine, same tasks, remote model12/12122.0s8/8131.0s—Anthropic API

Vendor-reported (NOT Agent-12 numbers — credited, not competing)

Models we have not run through Agent-12 yet. These scores are the vendor's own published numbers on the vendor's own benchmarks, linked to the source. They are not comparable to the tables above and move into the local table the day they get a real run.

Model Their published scores Credit Agent-12 status
Meta Muse Glimmer 30B30B dense · multimodal · Apache 2.0MCP-Atlas 75.5 · SWE-bench Verified 76.0 · SWE-bench Pro 51.2 · Terminal-Bench 2.1 51.7 · τ3-Banking 23.5Meta, official model cardAgent-12 run blocked until mlx_lm supports the muse_glimmer architecture
NVIDIA Nemotron 3 Nano Omni 30B-A3B30B MoE, 3B active · text + vision + audioOSWorld (computer use) 47.4NVIDIA, official model cardAgent-12 run queued (our MLX runtime: nemotron-omni-mlx)

How to trust a number here:

Tasks, judges, runner, and the validation gate are open: github.com/nicedreamzapp/agent12 · methodology · contamination policy. Cloud reference times include network/API round-trips — that is the honest end-to-end experience. Updated 2026-09-19.

Want to actually run the winner? Every row here plugs straight into Claude Code Local — Claude Code, 100% on-device on Apple Silicon, no cloud. Ready-made MLX builds on Hugging Face.