Tabular Calculations, JSON Schema Validation & Multi-Hop Retrieval

Data Analysis & Structured Reasoning

Evaluating how accurately foundation models interpret unstructured tables, compute financial or statistical metrics, and format strict JSON schemas without hallucinated keys.

Human Baseline75.0% (Data Analyst Baseline)
Contamination StatusModerate on static tables; Low on dynamic synthesis
Top Frontier ModelClaude Fable 5 (81.2%) & Gemini 3 Pro (88.0%)

Forensic Analysis & Methodology

Why Classic Benchmarks Failed

Standard Q&A benchmarks tested text extraction rather than actual numerical aggregation, multi-table joins, and schema enforcement.

The TrustTheBench & LiveBench Approach

LiveBench Data Analysis and TableMWP require models to parse complex CSV/JSON structures and execute programmatic data transformations.

Verified Category Leaderboard

Ranked strictly by verified Data Analysis & Structured Reasoning performance.

RankModelResearch LabAccess TypeCategory ScoreSpeedPricing / 1M
#1GPT-5OpenAIproprietary88.5%85 tok/s$2.50 in
#2Gemini 3 ProGoogle DeepMindproprietary88.0%80 tok/s$2.50 in
#3Grok 4xAIproprietary87.0%75 tok/s$5.00 in
#4Claude Sonnet 5Anthropicproprietary86.5%95 tok/s$3.00 in
#5Claude Fable 5.1 (Hybrid Reasoning)Anthropicproprietary82.1%54 tok/s$5.00 in
#6Llama 4 MaverickMeta AIopen-weight82.0%82 tok/s$0.50 in
#7GPT-4.5OpenAIproprietary81.5%42 tok/s$75.00 in
#8GPT-6 Astra (max)OpenAIproprietary81.4%62 tok/s$4.50 in
#9Claude Fable 5Anthropicproprietary81.2%45 tok/s$10.00 in
#10DeepSeek V4.1 FlashDeepSeekproprietary80.5%96 tok/s$0.18 in
#11GPT-5.6 SolOpenAIproprietary80.1%62 tok/s$5.00 in
#12Gemini 2.0 ProGoogle DeepMindproprietary80.1%60 tok/s$1.50 in
#13Llama 4 ScoutMeta AIopen-weight80.0%140 tok/s$0.20 in
#14o3OpenAIproprietary79.5%68 tok/s$4.00 in
#15Claude 3.7 SonnetAnthropicproprietary79.2%72 tok/s$3.00 in

Deployment Recommendations

🏆 Best Overall Frontier

Claude Fable 5 (81.2%) & Gemini 3 Pro (88.0%)

Maximum reasoning depth and lowest error rate on difficult boundary problems.

🔓 Best Open-Weight

Llama 4 Maverick (82.0%) & Kimi K3 (77.2%)

Host locally or via cost-effective cloud providers with full data privacy.

⚡ Best Value & Speed

Gemini 3.5 Flash-Lite ($0.10/1M)

Optimal price-to-performance ratio for high-volume automated agent pipelines.

Relevant Benchmark References

LiveBench Data Analysis

Metric: Auto-Grade %View Methodology →

TableMWP

Metric: Calculation Accuracy %View Methodology →