Benchmarks · Overview

Every score. Every run. Numbers you can verify.

b0 is measured on the standard coding-agent benchmarks — SWE-bench Lite, Verified, and Full, plus Terminal-bench. This page collects every run we've done, with links to the full result artefacts and reproduction recipes.

Latest — SWE-bench Lite v8

62.0%

186 / 300 resolved · GLM 4.6 · 94% cache · $48 total

SWE-bench Lite deep-view

SWE-bench

Lite · Verified · Full

Same task shape (auto-generated GitHub-issue → patch). Lite = 300 hand-selected issues, Verified = the reproducible subset, Full = the entire 2 294. We publish deep-views for each.

SWE-bench Lite

Terminal-bench

Real-world CLI tasks

Coding agents solving 10+ real terminal-based tasks — writing scripts, debugging, fixing failing builds. Different flavour of test: no repo context, more agent freedom.

Terminal-bench

Comparison

b0 across models

Same agent. Same harness. Only the model changes. One column per model: DeepSeek Flash, Tencent HY3, GLM 5.2, MiMo. Compare $/resolved and cache-hit rate side-by-side.

Compare models

Reference

Methodology

What SWE-bench measures, what 'resolved' means (Docker-eval spec), how the b0 harness invokes the CLI. Read this before challenging the numbers.

Read the methodology