Benchmark of 80+ LLMs in Finance: Claude Opus 5.5 & GPT-6 Astra
We evaluated 80+ LLMs in finance on 238 hard questions from the FinanceReasoning benchmark to identify which models excel at complex financial reasoning tasks.
The test set is the hard subset of the FinanceReasoning benchmark (Tang et al.), with 238 questions.1 Accuracy is the percentage of questions answered correctly. Numerical answers receive a 0.2% relative tolerance.
Output tokens are the tokens generated across the 238 answers. Input tokens are counted separately in the cost calculation.
Cost is the calculated USD spend for the question-answering run, using the token counts and rates recorded in our cost analysis. Some input counts are estimated. These figures exclude answer-extraction costs and do not represent a new price check.
Accuracy, token use and cost findings
Claude Opus 5.5 answers 218 of 238 questions correctly
Claude Opus 5.5 leads the recorded results at 91.60%, followed by GPT-5.6 Sol Pro at 90.76% (216 correct). Two questions separate them. GPT-5.6 Sol, Claude Fable 5.1, Claude Fable 5, Grok 4.6 and GPT-6.1 Sol Pro each answer 215 correctly, scoring 90.34%.
Opus 5.5 generates 144,834 output tokens. Its calculated cost is $3.25, compared with $16.35 for Sol Pro. We estimated Opus input usage from a sibling model, as explained in the methodology. A small score gap in one run provides limited evidence about repeat performance.
GPT-6.1 Sol Pro has the lowest cost among models tied at 90.34%
GPT-6.1 Sol Pro answers 215 questions correctly for $3.28, the lowest calculated cost among the five models tied at 90.34%. GPT-5.6 Sol costs $3.85, Grok 4.6 costs $4.45, Claude Fable 5.1 costs $8.64 and Claude Fable 5 costs $10.05.
GPT-6.1 Sol Pro generates 187,502 output tokens, compared with GPT-5.6 Sol’s 117,735. Its lower recorded token rates offset the longer output and higher input count. Among these tied models, GPT-5.6 Sol still uses the fewest output tokens.
GPT-6 Sol generates the fewest output tokens
GPT-6 Sol uses 85,228 output tokens, the lowest count among the 87 models. It answers 212 questions correctly (89.08%) for $0.98. GPT-6 Astra generates 87,774 tokens, answers 208 correctly (87.39%) and costs $5.02 at its recorded rates.
GPT-6.1 Sol answers 214 questions correctly (89.92%), using 89,153 output tokens for $1.02. Claude Sonnet 5.5 reaches 212 correct (89.08%) with 142,728 tokens at $1.61. Each comparison covers one completed run on the same 238 questions.
GPT-6 Luna exceeds 83% accuracy for $0.0754
GPT-6 Luna answers 207 questions correctly (86.97%) for $0.0754, the lowest recorded cost among models above 83%. DeepSeek V4.1 Flash also answers 207 correctly for $0.2572. GPT-6 Sol increases the count to 212 (89.08%) for $0.9792.
GPT-OSS-120B scores 81.09% for $0.0611, while Llama 4 Maverick scores 75.21% for $0.0995. Mistral Small 3.2 24B has the cheapest run overall at $0.0352, with 41.18% accuracy.
Higher spending does not guarantee a higher score
O1 Pro has the highest calculated cost, $381.28, and scores 80.67%. O3 Pro costs $39.23 for 78.15%, while O1 costs $46.59 for 74.79%. Each scores below Gemini 3 Flash Preview at $0.39.
Hy4 Preview generates 4,119,093 output tokens, the largest count in the comparison, and scores 89.92%. GPT-6.1 Sol reaches the same 214 correct answers with 89,153 tokens. DeepSeek R1 0528 generates 1,251,064 tokens at 62.18%, another example of output length varying independently of accuracy.
Results for all 87 models
Recorded accuracy values use different rounding in some older entries. GPT-5, Grok 4.5 and Kimi K3 each have 210 correct answers, despite the displayed 88.23% and 88.24% values. Model identifiers are preserved, and costs are rounded to four decimal places.
Model | Accuracy (%) | Correct / 238 | Output tokens | Cost (USD) |
|---|---|---|---|---|
claude-opus-5.5 | 91.60 | 218 | 144,834 | 3.2532 |
gpt-5.6-sol-pro | 90.76 | 216 | 385,886 | 16.3533 |
gpt-5.6-sol | 90.34 | 215 | 117,735 | 3.8493 |
claude-fable-5.1 | 90.34 | 215 | 154,919 | 8.6421 |
claude-fable-5 | 90.34 | 215 | 183,258 | 10.0543 |
grok-4.6 | 90.34 | 215 | 704,035 | 4.4460 |
gpt-6.1-sol-pro | 90.34 | 215 | 187,502 | 3.2817 |
claude-opus-5 | 89.92 | 214 | 225,562 | 6.0847 |
hy4-preview | 89.92 | 214 | 4,119,093 | 10.3574 |
gpt-6.1-sol | 89.92 | 214 | 89,153 | 1.0184 |
Retrieval experiment: 24 additional correct answers
Our existing experiment report compares standalone generation with retrieval-augmented generation (RAG), which adds retrieved documents to a model’s prompt. It records an increase from 104 to 128 correct answers on the same 238-question set.
Experiment notes label the tested model GPT-4o Mini. The main results table assigns the same baseline accuracy and token count to O4 Mini. RAG model identity needs confirmation against the original run records. Its token column also lacks a verified input/output split, so it cannot be compared directly with the output-token column above.
Recorded time rises from about 3 to 59 minutes. Per-request latency and the separate time spent on retrieval were not established here. Testing covered one reported model, which limits conclusions about RAG on the other models.
Recorded retrieval setup
Historical notes describe a Qdrant index containing financial explanations and Python functions from the benchmark’s knowledge corpus. Text-embedding-3-small produces 1,536-dimensional vectors. Documents are split by section, and each function forms a separate chunk.
For each question and its context, the pipeline searches two collections using cosine similarity, retrieving three document chunks and two functions. Those notes record two embedding calls per query and an increase from roughly 300–500 to 3,000–5,000 input tokens. Retrieved material is included in the generation prompt. These setup details remain part of the historical report. The local retrieval script inspected for this rewrite does not establish that this specific Qdrant run executed.
Financial reasoning benchmark methodology
Questions and prompting
We used FinanceReasoning’s hard subset of 238 questions. Each question supplies context and a reference answer. Context can include financial tables, time series or the inputs for a financial calculation.
One example asks for the last upper Keltner Channel band from 25 days of stock prices, using a 10-day exponential moving average, a 10-day average true range and a multiplier of 1.5. The answer is requested to two decimal places. Solving it requires calculating the moving average and true range, then applying the upper-band formula.
With chain-of-thought prompting, the model is instructed to work through the calculation and end with a numerical answer. It receives the question and context in a free-form generation task. The 31 models added in September and October 2026 each completed inference and evaluation for all 238 questions through OpenRouter, with no missing outputs. Opus 5.5’s accepted result comes from an execution setup that differed from the main API pipeline. Execution conditions therefore vary across the comparison.
Answer extraction and scoring
Evaluation extracts an answer from the generated response, then compares it with the reference. Our current configuration uses Claude Sonnet 4.5 for answer extraction. Numerical answers pass when their absolute error is at most 0.2% of the reference answer’s absolute value.
Accuracy is the correct-answer count divided by 238. One question changes the score by about 0.42 percentage points. Correct-answer counts make ties explicit where recorded percentages differ in rounding.
Token and cost calculations
Output-token totals come from model usage records. We calculate cost by combining input and output tokens with their recorded per-million-token prices:
Stored prices include historical rates for deprecated models. Estimated input counts are marked in the underlying results table. Opus 5.5’s cost combines its measured 144,834 output tokens with an estimated 89,138 input tokens from a sibling model, using $4 per million input tokens and $20 per million output tokens. That produces $3.2532.
Costs cover the final retained answer-generation responses. Evaluation-model calls, failed or superseded retries, and RAG indexing or embedding charges are outside this calculation. Differences in tokenizers, generation settings and execution setup also affect comparisons.
Benchmark limitations
This public dataset may have appeared in model training data. Scores therefore cannot establish how well a model handles unseen financial problems.
Results cover a fixed question set and the recorded runs. Repeated-run variation and statistical significance have not been established. LLM-based answer extraction can introduce errors, and scoring the final answer does not validate every step in the explanation.
Estimated input counts and stored prices limit cost precision. Our separate RAG report has unresolved model attribution and token accounting, as described above. Its figures remain historical observations pending source reconciliation.
This benchmark measures financial question answering. It does not test investment returns, regulatory compliance, fraud detection or deployment reliability.
Conclusion
Claude Opus 5.5 has the highest recorded score, with 218 correct answers. GPT-6.1 Sol Pro reaches 215 for $3.28, GPT-6 Sol reaches 212 with the fewest output tokens, and GPT-6 Luna reaches 207 for $0.0754. These are observed results on a fixed question set. Repeated-run reliability remains unmeasured.
Further reading
- AI-based Stock Trading covers a separate financial application.
- Legal AI Tools covers tools for legal work.
Cite this benchmark
Pick the format that matches where you're publishing. Pasting the link version into your CMS preserves the backlink.
@misc{sari2026,
author = {Sarı, Ekrem},
title = {{Benchmark of 80+ LLMs in Finance: Claude Opus 5.5 & GPT-6 Astra}},
year = {2026},
month = oct,
howpublished = {\url{https://aimultiple.com/finance-llm}},
note = {AIMultiple. Retrieved October 1, 2026}
}Results and timestamps of 90 data points. Download the summary data shown in this article's charts and tables as a ZIP file containing 2 CSV files.
Want the granular data behind it? Join Premium
Be the first to comment
Your email address will not be published. All fields are required. Comments are left in their original language.