Premium
Services
Premium

Benchmark of 80+ LLMs in Finance: Claude Opus 5.5 & GPT-6 Astra

We evaluated 80+ LLMs in finance on 238 hard questions from the FinanceReasoning benchmark to identify which models excel at complex financial reasoning tasks.

Ekrem Sarı
Ekrem Sarı
updated on Oct 1, 2026
Loading Chart

The test set is the hard subset of the FinanceReasoning benchmark (Tang et al.), with 238 questions.1 Accuracy is the percentage of questions answered correctly. Numerical answers receive a 0.2% relative tolerance.

Output tokens are the tokens generated across the 238 answers. Input tokens are counted separately in the cost calculation.

Cost is the calculated USD spend for the question-answering run, using the token counts and rates recorded in our cost analysis. Some input counts are estimated. These figures exclude answer-extraction costs and do not represent a new price check.

Accuracy, token use and cost findings

Claude Opus 5.5 answers 218 of 238 questions correctly

Claude Opus 5.5 leads the recorded results at 91.60%, followed by GPT-5.6 Sol Pro at 90.76% (216 correct). Two questions separate them. GPT-5.6 Sol, Claude Fable 5.1, Claude Fable 5, Grok 4.6 and GPT-6.1 Sol Pro each answer 215 correctly, scoring 90.34%.

Opus 5.5 generates 144,834 output tokens. Its calculated cost is $3.25, compared with $16.35 for Sol Pro. We estimated Opus input usage from a sibling model, as explained in the methodology. A small score gap in one run provides limited evidence about repeat performance.

GPT-6.1 Sol Pro has the lowest cost among models tied at 90.34%

GPT-6.1 Sol Pro answers 215 questions correctly for $3.28, the lowest calculated cost among the five models tied at 90.34%. GPT-5.6 Sol costs $3.85, Grok 4.6 costs $4.45, Claude Fable 5.1 costs $8.64 and Claude Fable 5 costs $10.05.

GPT-6.1 Sol Pro generates 187,502 output tokens, compared with GPT-5.6 Sol’s 117,735. Its lower recorded token rates offset the longer output and higher input count. Among these tied models, GPT-5.6 Sol still uses the fewest output tokens.

GPT-6 Sol generates the fewest output tokens

GPT-6 Sol uses 85,228 output tokens, the lowest count among the 87 models. It answers 212 questions correctly (89.08%) for $0.98. GPT-6 Astra generates 87,774 tokens, answers 208 correctly (87.39%) and costs $5.02 at its recorded rates.

GPT-6.1 Sol answers 214 questions correctly (89.92%), using 89,153 output tokens for $1.02. Claude Sonnet 5.5 reaches 212 correct (89.08%) with 142,728 tokens at $1.61. Each comparison covers one completed run on the same 238 questions.

GPT-6 Luna exceeds 83% accuracy for $0.0754

GPT-6 Luna answers 207 questions correctly (86.97%) for $0.0754, the lowest recorded cost among models above 83%. DeepSeek V4.1 Flash also answers 207 correctly for $0.2572. GPT-6 Sol increases the count to 212 (89.08%) for $0.9792.

GPT-OSS-120B scores 81.09% for $0.0611, while Llama 4 Maverick scores 75.21% for $0.0995. Mistral Small 3.2 24B has the cheapest run overall at $0.0352, with 41.18% accuracy.

Higher spending does not guarantee a higher score

O1 Pro has the highest calculated cost, $381.28, and scores 80.67%. O3 Pro costs $39.23 for 78.15%, while O1 costs $46.59 for 74.79%. Each scores below Gemini 3 Flash Preview at $0.39.

Hy4 Preview generates 4,119,093 output tokens, the largest count in the comparison, and scores 89.92%. GPT-6.1 Sol reaches the same 214 correct answers with 89,153 tokens. DeepSeek R1 0528 generates 1,251,064 tokens at 62.18%, another example of output length varying independently of accuracy.

Results for all 87 models

Recorded accuracy values use different rounding in some older entries. GPT-5, Grok 4.5 and Kimi K3 each have 210 correct answers, despite the displayed 88.23% and 88.24% values. Model identifiers are preserved, and costs are rounded to four decimal places.

Model
Accuracy (%)
Correct / 238
Output tokens
Cost (USD)
claude-opus-5.5
91.60
218
144,834
3.2532
gpt-5.6-sol-pro
90.76
216
385,886
16.3533
gpt-5.6-sol
90.34
215
117,735
3.8493
claude-fable-5.1
90.34
215
154,919
8.6421
claude-fable-5
90.34
215
183,258
10.0543
grok-4.6
90.34
215
704,035
4.4460
gpt-6.1-sol-pro
90.34
215
187,502
3.2817
claude-opus-5
89.92
214
225,562
6.0847
hy4-preview
89.92
214
4,119,093
10.3574
gpt-6.1-sol
89.92
214
89,153
1.0184

Retrieval experiment: 24 additional correct answers

Our existing experiment report compares standalone generation with retrieval-augmented generation (RAG), which adds retrieved documents to a model’s prompt. It records an increase from 104 to 128 correct answers on the same 238-question set.

Experiment notes label the tested model GPT-4o Mini. The main results table assigns the same baseline accuracy and token count to O4 Mini. RAG model identity needs confirmation against the original run records. Its token column also lacks a verified input/output split, so it cannot be compared directly with the output-token column above.

Recorded time rises from about 3 to 59 minutes. Per-request latency and the separate time spent on retrieval were not established here. Testing covered one reported model, which limits conclusions about RAG on the other models.

Recorded retrieval setup

Historical notes describe a Qdrant index containing financial explanations and Python functions from the benchmark’s knowledge corpus. Text-embedding-3-small produces 1,536-dimensional vectors. Documents are split by section, and each function forms a separate chunk.

For each question and its context, the pipeline searches two collections using cosine similarity, retrieving three document chunks and two functions. Those notes record two embedding calls per query and an increase from roughly 300–500 to 3,000–5,000 input tokens. Retrieved material is included in the generation prompt. These setup details remain part of the historical report. The local retrieval script inspected for this rewrite does not establish that this specific Qdrant run executed.

Get our team to automate one of your business processes with AI agents, free of charge.
Automate a process

Financial reasoning benchmark methodology

Questions and prompting

We used FinanceReasoning’s hard subset of 238 questions. Each question supplies context and a reference answer. Context can include financial tables, time series or the inputs for a financial calculation.

One example asks for the last upper Keltner Channel band from 25 days of stock prices, using a 10-day exponential moving average, a 10-day average true range and a multiplier of 1.5. The answer is requested to two decimal places. Solving it requires calculating the moving average and true range, then applying the upper-band formula.

With chain-of-thought prompting, the model is instructed to work through the calculation and end with a numerical answer. It receives the question and context in a free-form generation task. The 31 models added in September and October 2026 each completed inference and evaluation for all 238 questions through OpenRouter, with no missing outputs. Opus 5.5’s accepted result comes from an execution setup that differed from the main API pipeline. Execution conditions therefore vary across the comparison.

Answer extraction and scoring

Evaluation extracts an answer from the generated response, then compares it with the reference. Our current configuration uses Claude Sonnet 4.5 for answer extraction. Numerical answers pass when their absolute error is at most 0.2% of the reference answer’s absolute value.

Accuracy is the correct-answer count divided by 238. One question changes the score by about 0.42 percentage points. Correct-answer counts make ties explicit where recorded percentages differ in rounding.

Token and cost calculations

Output-token totals come from model usage records. We calculate cost by combining input and output tokens with their recorded per-million-token prices:

Stored prices include historical rates for deprecated models. Estimated input counts are marked in the underlying results table. Opus 5.5’s cost combines its measured 144,834 output tokens with an estimated 89,138 input tokens from a sibling model, using $4 per million input tokens and $20 per million output tokens. That produces $3.2532.

Costs cover the final retained answer-generation responses. Evaluation-model calls, failed or superseded retries, and RAG indexing or embedding charges are outside this calculation. Differences in tokenizers, generation settings and execution setup also affect comparisons.

Benchmark limitations

This public dataset may have appeared in model training data. Scores therefore cannot establish how well a model handles unseen financial problems.

Results cover a fixed question set and the recorded runs. Repeated-run variation and statistical significance have not been established. LLM-based answer extraction can introduce errors, and scoring the final answer does not validate every step in the explanation.

Estimated input counts and stored prices limit cost precision. Our separate RAG report has unresolved model attribution and token accounting, as described above. Its figures remain historical observations pending source reconciliation.

This benchmark measures financial question answering. It does not test investment returns, regulatory compliance, fraud detection or deployment reliability.

Conclusion

Claude Opus 5.5 has the highest recorded score, with 218 correct answers. GPT-6.1 Sol Pro reaches 215 for $3.28, GPT-6 Sol reaches 212 with the fewest output tokens, and GPT-6 Luna reaches 207 for $0.0754. These are observed results on a fixed question set. Repeated-run reliability remains unmeasured.

Don’t miss our benchmarks and data-driven insights. The button opens Google; selecting AIMultiple confirms that you wish to see AIMultiple more often in Google search results.
GoogleAdd as preferred source

Further reading

Cite this benchmark

Pick the format that matches where you're publishing. Pasting the link version into your CMS preserves the backlink.

Ekrem Sarı (2026) - "Benchmark of 80+ LLMs in Finance: Claude Opus 5.5 & GPT-6 Astra". Published online at AIMultiple.com. Retrieved October 1, 2026, from: https://aimultiple.com/finance-llm [Online Resource]

Sarı, E. (2026, October 1). Benchmark of 80+ LLMs in Finance: Claude Opus 5.5 & GPT-6 Astra. AIMultiple. https://aimultiple.com/finance-llm

@misc{sari2026,
  author = {Sarı, Ekrem},
  title  = {{Benchmark of 80+ LLMs in Finance: Claude Opus 5.5 & GPT-6 Astra}},
  year   = {2026},
  month  = oct,
  howpublished    = {\url{https://aimultiple.com/finance-llm}},
  note   = {AIMultiple. Retrieved October 1, 2026}
}
Download all data

Results and timestamps of 90 data points. Download the summary data shown in this article's charts and tables as a ZIP file containing 2 CSV files.

Last updated: September 30, 2026
Download

Want the granular data behind it? Join Premium

Ekrem Sarı
Ekrem Sarı
AI Researcher
Ekrem is an AI Researcher and Data Scientist at AIMultiple. He designs and runs hands-on benchmarks for AI and LLM systems.
View Full Profile

Be the first to comment

Your email address will not be published. All fields are required. Comments are left in their original language.

0/450