Benchmarks · Overview
Every score. Every run. Numbers you can verify.
b0 is measured on the standard coding-agent benchmarks — SWE-bench Lite, Verified, and Full, plus Terminal-bench. This page collects every run we've done, with links to the full result artefacts and reproduction recipes.
Latest — SWE-bench Lite v8
62.0%
186 / 300 resolved · GLM 4.6 · 94% cache · $48 total
SWE-bench
Lite · Verified · Full
Same task shape (auto-generated GitHub-issue → patch). Lite = 300 hand-selected issues, Verified = the reproducible subset, Full = the entire 2 294. We publish deep-views for each.
SWE-bench LiteTerminal-bench
Real-world CLI tasks
Coding agents solving 10+ real terminal-based tasks — writing scripts, debugging, fixing failing builds. Different flavour of test: no repo context, more agent freedom.
Terminal-benchComparison
b0 across models
Same agent. Same harness. Only the model changes. One column per model: DeepSeek Flash, Tencent HY3, GLM 5.2, MiMo. Compare $/resolved and cache-hit rate side-by-side.
Compare modelsReference
Methodology
What SWE-bench measures, what 'resolved' means (Docker-eval spec), how the b0 harness invokes the CLI. Read this before challenging the numbers.
Read the methodology