The newest model in the System One benchmarks is Kev-9B.
AIM System One Index
System One models, also called decision models, return a typed decision with a probability for each option in one pass instead of generating text token by token. How AIM System One Index works
System One Models Leaderboard
System One models ranked by the AIM System One Index.
# | Model | Index | AIM-Agentic-RAG accuracy (%) | AIM-AI-Bias score | AIM-Decision (%) | AIM-Decision time per task (s) | Median latency (ms) |
|---|---|---|---|---|---|---|---|
| 1 | Kev-9B System One model | 69 | 97.1 | - | 40 | 38.8 | 952 |
| 2 | Jev 1.13 System One model | 66 | 98.8 | 80.6 | 34 | 5.4 | 344 |
| 3 | Laya System One model | 25 | 50.6 | - | 0 | - | 200 |
AIM System One Index vs time per decision
Higher and further left is better.
Fine-tuned vs stock
Browser tasks completed out of 50, each fine-tuned model against its stock version.
Individual benchmarks
AIM-Decision
Tasks completed out of 50 on the same browser tasks.
AIM-Agentic-RAG
243 database routing questions.
How AIM System One Index works
AIM System One Index is the average of each model's accuracy on AIM-Agentic-RAG and AIM-Decision.
Latency and cost are shown for reference and do not count toward the index.
AIM-AI-Bias results are shown for reference and do not count toward the index.
API costs are what we were billed. Costs for self-hosted models (Kev-9B, Laya) are hardware estimates at full utilization, and Kev-9B's latency includes an SSH tunnel to its GPU.
Benchmarks We Used
Recent Updates
Models most recently added to the System One benchmarks.
Kev-9B
New model added to the System One benchmarks.
Jev 1.13
New model added to the System One benchmarks.
Laya
New model added to the System One benchmarks.