How we benchmark

Every tracked model runs the same versioned prompt suites across five categories — coding, reasoning, creative, factual, and tool use — every 6 hours, at temperature 0 for reproducibility.

Responses are scored 0–100 by an independent judge model that never sees which tool produced the answer. The judge is configurable and rotated to prevent bias. Each row records quality, full-response latency, token counts, and normalized cost per run.

Degradation flags fire when a model's latest run drops more than 10% below its own 7-day rolling average — that threshold is per-user configurable in alert settings.

We take no vendor payments to influence scores, we publish suite versions, and results include sample sizes and timestamps. When a judge fails, we store the raw result without a score rather than guessing.