Premium
Services
Premium

AIM System One Index

System One models, also called decision models, return a typed decision with a probability for each option in one pass instead of generating text token by token. How AIM System One Index works

Best model
Kev-9B
AIM System One Index 69
Fastest model
Laya
200 ms median per decision
Cheapest API model
Jev 1.13
$0.034 per 1,000 decisions
AIM System One Index

System One Models Leaderboard

System One models ranked by the AIM System One Index.

#
Model
Index
AIM-Agentic-RAG accuracy (%)
AIM-AI-Bias score
AIM-Decision (%)
AIM-Decision time per task (s)
Median latency (ms)
1
Kev-9B
System One model
69
97.1-4038.8952
2
Jev 1.13
Jev 1.13
System One model
66
98.880.6345.4344
3
Laya
Laya
System One model
25
50.6-0-200

Cost-
Latency952 ms
Context-
AIM-Agentic-RAG accuracy (%)
97.1
AIM-Decision (%)
40

Cost-
Latency344 ms
Context-
AIM-Agentic-RAG accuracy (%)
98.8
AIM-AI-Bias score
80.6
AIM-Decision (%)
34

Cost-
Latency200 ms
Context-
AIM-Agentic-RAG accuracy (%)
50.6
AIM-Decision (%)
0

AIM System One Index vs time per decision

Higher and further left is better.

AIM System One Index vs time per decision
*AIM System One Index, as explained in How AIM System One Index works. Time: median per AIM-Agentic-RAG routing decision; Kev-9B's includes an SSH tunnel to its GPU.

Fine-tuned vs stock

Browser tasks completed out of 50, each fine-tuned model against its stock version.

Fine-tuned vs stock: browser tasks completed out of 50
*Each fine-tuned model against its stock version, run the same day. Kev-9B: p = 0.0034. Laya: p = 0.25, not significant.

Individual benchmarks

AIM-Decision

Tasks completed out of 50 on the same browser tasks.

Browser tasks completed out of 50
*One attempt per task.

AIM-Agentic-RAG

243 database routing questions.

*Laya: local M4 MPS. Kev: later A100 FP32 run, including SSH/network time. Others: hosted APIs.

How AIM System One Index works

AIM System One Index is the average of each model's accuracy on AIM-Agentic-RAG and AIM-Decision.

Latency and cost are shown for reference and do not count toward the index.

AIM-AI-Bias results are shown for reference and do not count toward the index.

API costs are what we were billed. Costs for self-hosted models (Kev-9B, Laya) are hardware estimates at full utilization, and Kev-9B's latency includes an SSH tunnel to its GPU.

Benchmarks We Used

Recent Updates

Models most recently added to the System One benchmarks.

Jared Palmer

Kev-9B

New model added to the System One benchmarks.

TypeSafe

Jev 1.13

New model added to the System One benchmarks.

Convai Innovations

Laya

New model added to the System One benchmarks.