System One models read a situation and answer typed questions in one fast call: yes or no, pick one option, or score a level, each with probabilities. Every model gets the same cases and the same request body, scored against each dataset's labels.
Mean accuracy across all 16 datasets. Higher is better.
Median time per call, in milliseconds. APIs with 8 calls in flight; Laya locally, one at a time. Lower is better.
USD per 1,000 calls, as billed by the API. Lower is better.
The index is the mean accuracy across datasets, with every dataset weighted equally, on a 0 to 100 scale. A model gets an index only if it ran every dataset in the set, so the Arabic fine-tunes appear in the Arabic index alone.
The shaded corner is the most attractive: more accurate, and faster or cheaper. The line joins the models that nothing else beats on both measures.
System One Index (all datasets) against median latency per call. API models measured through OpenRouter with 8 calls in flight; Laya on a local Apple M5 GPU, so its latency excludes any network.
System One Index (all datasets) against USD per 1,000 calls. Local models have no per-call price.
Calibration error is the average gap between a model's stated confidence and how often it's right, in percentage points. A well-calibrated model's 80% answers are right about 80% of the time.
Top-label ECE over every question, all datasets. Lower is better.
Milliseconds per call. The bar is the median; the tick marks the 95th percentile, shown on hover. Lower is better.
USD for all 3,250 cases, as billed. Lower is better.
Mean billed input tokens. Each tokenizer counts differently, so this explains cost rather than ranks models.
One chart per dataset. Bars are colored by who makes the model.
Every model run with its overall numbers. Overall accuracy here counts every question, so the 5-question typed-decisions dataset weighs more than in the index.
| Model | Index | Accuracy | Calibration error | Median latency | Cost per 1,000 calls | Datasets |
|---|
Models we plan to add. Vendors publish their own scores for several of these; none appear on this page until they've run on System One Bench like every other model.
state plus one to five named questions, sent unchanged to every model.usage.cost the API reports.export OPENROUTER_API_KEY=<your key> python run_decisions.py --model openai/gpt-6-luna-decisions --name luna python run_decisions.py --model typesafe/jev-1.13 --name or_jev python score.py preds_or_jev.jsonl preds_luna.jsonl
To race two models live on the same cases, run python server.py and open localhost:8061. The race needs API keys, so it runs on your machine rather than on this page.