Finance AI Benchmark
Finance Benchmark for LLMs
Large language models may prove to the service economy what steam power was to industry: a general-purpose technology that sharply lowers the cost of skilled work. Nowhere are the stakes higher than in financial services, among the world's most expensive and information-intensive businesses. Yet capturing that value depends less on finding one universal winner than on routing each task to the model best equipped to perform it. Which LLM should underwrite a borrower, interpret a Basel rule or price a derivative? Finance Benchmark supplies the evidence.
Open leaderboard for AI models on real finance workflows — Basel III, credit underwriting, derivative pricing, and more.
v2 is the primary leaderboard. Browse the task catalog for v2 and v3 sets.
Leaderboard
93.2%
- Finance Index
- 92.8
- Consistency
- 2.8/3
- Eval cost
- $1.14
- Tasks passed
- 93.2%
93.2%
- Finance Index
- 94.4
- Consistency
- 2.5/3
- Eval cost
- $0.02
- Tasks passed
- 93.2%
91.8%
- Finance Index
- 93.3
- Consistency
- 2.7/3
- Eval cost
- $1.08
- Tasks passed
- 91.8%
91.8%
- Finance Index
- 91.7
- Consistency
- 2.7/3
- Eval cost
- $1.23
- Tasks passed
- 91.8%
90.4%
- Finance Index
- 92.2
- Consistency
- 2.6/3
- Eval cost
- $0.98
- Tasks passed
- 90.4%
90.4%
- Finance Index
- 90.6
- Consistency
- 2.6/3
- Eval cost
- $0.64
- Tasks passed
- 90.4%
90.4%
- Finance Index
- 92.2
- Consistency
- 2.6/3
- Eval cost
- $1.95
- Tasks passed
- 90.4%
89.0%
- Finance Index
- 82.8
- Consistency
- 2.6/3
- Eval cost
- $0.78
- Tasks passed
- 89.0%
89.0%
- Finance Index
- 89.4
- Consistency
- 2.6/3
- Eval cost
- $0.07
- Tasks passed
- 89.0%
89.0%
- Finance Index
- 89.4
- Consistency
- 2.6/3
- Eval cost
- $0.60
- Tasks passed
- 89.0%
89.0%
- Finance Index
- 91.1
- Consistency
- 2.5/3
- Eval cost
- $0.95
- Tasks passed
- 89.0%
89.0%
- Finance Index
- 87.2
- Consistency
- 2.6/3
- Eval cost
- $3.13
- Tasks passed
- 89.0%
87.7%
- Finance Index
- 88.3
- Consistency
- 2.6/3
- Eval cost
- $1.05
- Tasks passed
- 87.7%
87.7%
- Finance Index
- 88.3
- Consistency
- 2.6/3
- Eval cost
- $0.52
- Tasks passed
- 87.7%
86.3%
- Finance Index
- 87.2
- Consistency
- 2.6/3
- Eval cost
- $1.87
- Tasks passed
- 86.3%
84.9%
- Finance Index
- 80.6
- Consistency
- 2.4/3
- Eval cost
- $0.75
- Tasks passed
- 84.9%
84.9%
- Finance Index
- 80.0
- Consistency
- 2.4/3
- Eval cost
- $0.42
- Tasks passed
- 84.9%
84.9%
- Finance Index
- 85.0
- Consistency
- 2.5/3
- Eval cost
- $0.33
- Tasks passed
- 84.9%
84.9%
- Finance Index
- 84.4
- Consistency
- 2.4/3
- Eval cost
- $0.15
- Tasks passed
- 84.9%
84.9%
- Finance Index
- 84.4
- Consistency
- 2.4/3
- Eval cost
- $0.02
- Tasks passed
- 84.9%
83.6%
- Finance Index
- 82.8
- Consistency
- 2.5/3
- Eval cost
- $0.41
- Tasks passed
- 83.6%
80.8%
- Finance Index
- 77.2
- Consistency
- 2.2/3
- Eval cost
- $1.21
- Tasks passed
- 80.8%
65.9%
- Finance Index
- 46.9
- Consistency
- 2.0/3
- Eval cost
- $0.52
- Tasks passed
- 65.9%
63.0%
- Finance Index
- 40.6
- Consistency
- 1.9/3
- Eval cost
- $0.53
- Tasks passed
- 63.0%
61.6%
- Finance Index
- 43.9
- Consistency
- 1.8/3
- Eval cost
- $101.61
- Tasks passed
- 61.6%
60.3%
- Finance Index
- 39.4
- Consistency
- 1.8/3
- Eval cost
- $0.02
- Tasks passed
- 60.3%
| # | Model | Provider | Tasks passed ↓ | Finance Index | Consistency | Eval cost |
|---|---|---|---|---|---|---|
| #1 | claude-opus-4-7 | anthropic | 93.2%86.3%–98.6% | 92.8 | 2.8/3 | $1.14 |
| #2 | deepseek-v4-flash | deepseek | 93.2%87.7%–98.6% | 94.4 | 2.5/3 | $0.02 |
| #3 | gpt-5.6-sol | openai | 91.8%84.9%–97.3% | 93.3 | 2.7/3 | $1.08 |
| #4 | claude-opus-4-8 | anthropic | 91.8%84.9%–97.3% | 91.7 | 2.7/3 | $1.23 |
| #5 | gpt-5.6-luna | openai | 90.4%83.6%–95.9% | 92.2 | 2.6/3 | $0.98 |
| #6 | grok-3 | xai | 90.4%83.6%–95.9% | 90.6 | 2.6/3 | $0.64 |
| #7 | claude-fable-5 | anthropic | 90.4%83.6%–95.9% | 92.2 | 2.6/3 | $1.95 |
| #8 | claude-haiku-4-5-20251001 | anthropic | 89.0%80.8%–95.9% | 82.8 | 2.6/3 | $0.78 |
| #9 | deepseek-v4-pro | deepseek | 89.0%80.8%–95.9% | 89.4 | 2.6/3 | $0.07 |
| #10 | grok-4 | xai | 89.0%82.2%–95.9% | 89.4 | 2.6/3 | $0.60 |
| #11 | gpt-5.5 | openai | 89.0%80.8%–95.9% | 91.1 | 2.5/3 | $0.95 |
| #12 | kimi-k3 | moonshot | 89.0%80.8%–95.9% | 87.2 | 2.6/3 | $3.13 |
| #13 | claude-sonnet-4-6 | anthropic | 87.7%79.5%–94.5% | 88.3 | 2.6/3 | $1.05 |
| #14 | gpt-5.6-terra | openai | 87.7%79.5%–94.5% | 88.3 | 2.6/3 | $0.52 |
| #15 | claude-opus-4-6 | anthropic | 86.3%78.1%–93.2% | 87.2 | 2.6/3 | $1.87 |
| #16 | gemini-2.5-pro | 84.9%76.7%–93.2% | 80.6 | 2.4/3 | $0.75 | |
| #17 | claude-sonnet-5 | anthropic | 84.9%76.7%–93.2% | 80.0 | 2.4/3 | $0.42 |
| #18 | grok-4.5 | xai | 84.9%76.7%–93.2% | 85.0 | 2.5/3 | $0.33 |
| #19 | gemini-2.5-flash | 84.9%76.7%–93.2% | 84.4 | 2.4/3 | $0.15 | |
| #20 | deepseek-reasoner | deepseek | 84.9%76.7%–93.2% | 84.4 | 2.4/3 | $0.02 |
| #21 | gemini-3.5-flash | 83.6%74.0%–91.8% | 82.8 | 2.5/3 | $0.41 | |
| #22 | kimi-k2.7-code | moonshot | 80.8%71.2%–90.4% | 77.2 | 2.2/3 | $1.21 |
| #23 | gpt-5.3-codex | openai | 65.9%54.9%–75.6% | 46.9 | 2.0/3 | $0.52 |
| #24 | gpt-5.4 | openai | 63.0%52.0%–74.0% | 40.6 | 1.9/3 | $0.53 |
| #25 | gpt-5.5-pro | openai | 61.6%50.7%–72.6% | 43.9 | 1.8/3 | $101.61 |
| #26 | deepseek-chat | deepseek | 60.3%49.3%–71.2% | 39.4 | 1.8/3 | $0.02 |
openai/gpt-5.5-pro is omitted from the Time · Cost · Rating plots below: at $101.61 in eval cost it would flatten every other model onto a single edge of the cost axis. It remains in the leaderboard table (Tasks passed 61.6%).
Time · Cost · Rating
Rotating 3D tradeoff: eval wall-clock time, list-price cost ($), and Finance Index. Ideal models sit toward low time, low cost, and high rating. Extreme cost outliers are excluded from this plot (see note above) so the rest of the field stays readable. Hover to pause and inspect.
- deepseek
- openai
- anthropic
- xai
- moonshot
2D tradeoff
Pick a pair of axes. Published models with both values are plotted (cost outliers that would collapse the scale are noted above). Points are colored by provider; hover for the model name and values.
- anthropic
- deepseek
- openai
- xai
- moonshot
Scores by domain
Tasks passed (%) · top 8 models