Agent-12 — the local agent leaderboard
Which model can actually be your agent on your own hardware? Measured by doing real agent tasks in a sandboxed working directory — judged by the filesystem, never by the model's prose. Temperature 0, fixed caps, fresh sandbox per task, one variable moved per comparison.
Local models
Measured in Agent-12's own reference harness (Anvil, a ~550-token native terminal engine with per-model tool-call dialects) — not inside Claude Code. The same model inside Claude Code as installed scores and times differently, because that harness sends far more tool tokens per turn; the harness is held fixed here so models can be compared with each other. See METHODOLOGY and the measured gap in issue #1.
| Model | Easy (12) | time | Hard (8) | time | tok/s | Hardware |
|---|---|---|---|---|---|---|
| Qwen3-Coder-30B-A3B (MLX 8-bit)30B MoE, 3B active · native tool calls | 12/12 | 42.6s | 7/8 | 391.5s | 84.3 | Apple M5, 128 GB unified memory |
| Qwen3.6-35B-A3B (MLX 8-bit)35B MoE, 3B active · native tool calls | 12/12 | 64.2s | 8/8 | 125.2s | 46.2 | Apple M5, 128 GB unified memory |
| Gemma 4 31B (MLX 4-bit)31B dense · prompted-XML dialect | 11/12 | 92.1s | 8/8 | 347.9s | 25.6 | Apple M5, 128 GB unified memory |
| Qwen3.8-27B (MLX 8-bit, dense)27B dense · native tool calls · 8/8 hard at an 8000-token budget, 16.9x slower than Qwen3.6 (writeup) | 12/12 | 350.7s | 7/8 | 1155.9s | — | Apple M5, 128 GB unified memory |
| Qwen3.8-27B (MLX 8-bit, dense) + DFlash 2 drafter27B dense · native tool calls · incoai DFlash2 drafter, ~3x faster than without it, still ~3x slower than Qwen3.6; tok/s here includes re-prefill every turn (writeup, 2026-09-19) | 12/12 | 122.0s | 7/8 | 401.4s | 28.7 | Apple M5, 128 GB unified memory |
| DeepSeek V4 Flash (2-bit, 0731 imatrix)284B MoE · ds4.c engine | 12/12 | 203.2s | 8/8 | 550.8s | 8.4 | Apple M5, 128 GB unified memory |
| DeepSeek V4 Flash (2-bit, original) · retired 8/10, weights deleted284B MoE · ds4.c engine | 12/12 | 376.1s | 8/8 | 842.4s | 8.4 | Apple M5, 128 GB unified memory |
Cloud reference (not competing — context)
| Model | Easy (12) | time | Hard (8) | time | tok/s | Where it runs |
|---|---|---|---|---|---|---|
| Claude Sonnet 5 (cloud reference)same engine, same tasks, remote model | 12/12 | 122.0s | 8/8 | 131.0s | — | Anthropic API |
Vendor-reported (NOT Agent-12 numbers — credited, not competing)
Models we have not run through Agent-12 yet. These scores are the vendor's own published numbers on the vendor's own benchmarks, linked to the source. They are not comparable to the tables above and move into the local table the day they get a real run.
| Model | Their published scores | Credit | Agent-12 status |
|---|---|---|---|
| Meta Muse Glimmer 30B30B dense · multimodal · Apache 2.0 | MCP-Atlas 75.5 · SWE-bench Verified 76.0 · SWE-bench Pro 51.2 · Terminal-Bench 2.1 51.7 · τ3-Banking 23.5 | Meta, official model card | Agent-12 run blocked until mlx_lm supports the muse_glimmer architecture |
| NVIDIA Nemotron 3 Nano Omni 30B-A3B30B MoE, 3B active · text + vision + audio | OSWorld (computer use) 47.4 | NVIDIA, official model card | Agent-12 run queued (our MLX runtime: nemotron-omni-mlx) |
How to trust a number here:
- Every judge is smoke-tested on a known-good AND a known-bad reference solution before it judges anything real.
- Tasks that ship a test file pin its hash — editing the test scores zero.
- All rows run the same tasks, same caps, same judges, on the stated hardware.
- Cloud rows are labeled reference points, not contestants.
Tasks, judges, runner, and the validation gate are open: github.com/nicedreamzapp/agent12 · methodology · contamination policy. Cloud reference times include network/API round-trips — that is the honest end-to-end experience. Updated 2026-09-19.
Want to actually run the winner? Every row here plugs straight into Claude Code Local — Claude Code, 100% on-device on Apple Silicon, no cloud. Ready-made MLX builds on Hugging Face.