Tokens/Second, TTFT Latency & API Cost Efficiency

Inference Speed, Latency & Price Frontier

Benchmarking real-world generation throughput (tokens/sec), time-to-first-token (TTFT ms), and pricing per million tokens across provider APIs and open-weight hostings.

Human BaselineN/A (Hardware & Quantization Metric)
Contamination StatusZero (Empirical runtime telemetry)
Top Frontier ModelGemini 3.7 Flash (182 tok/s at $0.75/1M)

Forensic Analysis & Methodology

Why Classic Benchmarks Failed

Static model scoreboards ignored deployment latency, making models with high latency impractical for real-time agentic tool loops.

The TrustTheBench & LiveBench Approach

Artificial Analysis continuous API probing across 20+ cloud providers under standardized load conditions.

Verified Category Leaderboard

Ranked strictly by verified Inference Speed, Latency & Price Frontier performance.

RankModelResearch LabAccess TypeCategory ScoreSpeedPricing / 1M
#1Gemini 3.5 Flash-LiteGoogle DeepMindproprietary220 tok/s220 tok/s$0.10 in
#2Amazon Nova LiteAmazon AWSproprietary200 tok/s200 tok/s$0.06 in
#3DeepSeek V4-FlashDeepSeekopen-weight190 tok/s190 tok/s$0.14 in
#4Gemini 2.0 FlashGoogle DeepMindproprietary185 tok/s185 tok/s$0.10 in
#5Gemini 3.7 FlashGoogle DeepMindproprietary182 tok/s182 tok/s$0.75 in
#6Gemma 2 2BGoogle DeepMindopen-weight180 tok/s180 tok/s$0.05 in
#7Gemini 3.6 FlashGoogle DeepMindproprietary175 tok/s175 tok/s$0.75 in
#8Gemini 2.5 Flash (Thinking)Google DeepMindproprietary145 tok/s145 tok/s$0.07 in
#9Llama 4 ScoutMeta AIopen-weight140 tok/s140 tok/s$0.20 in
#10Gemini 1.5 FlashGoogle DeepMindproprietary140 tok/s140 tok/s$0.07 in
#11Gemma 2 9BGoogle DeepMindopen-weight130 tok/s130 tok/s$0.10 in
#12GPT-4o miniOpenAIproprietary125 tok/s125 tok/s$0.15 in
#13Amazon Nova ProAmazon AWSproprietary120 tok/s120 tok/s$0.80 in
#14o3-miniOpenAIproprietary115 tok/s115 tok/s$1.10 in
#15Yi-Lightning01.AIproprietary110 tok/s110 tok/s$0.14 in

Deployment Recommendations

🏆 Best Overall Frontier

Gemini 3.7 Flash (182 tok/s at $0.75/1M)

Maximum reasoning depth and lowest error rate on difficult boundary problems.

🔓 Best Open-Weight

DeepSeek V4-Flash (190 tok/s at $0.14/1M)

Host locally or via cost-effective cloud providers with full data privacy.

⚡ Best Value & Speed

Gemini 3.5 Flash-Lite (220 tok/s at $0.10/1M)

Optimal price-to-performance ratio for high-volume automated agent pipelines.

Relevant Benchmark References

Artificial Analysis Speed Index

Metric: Tok/SecView Methodology →

Cost-Per-Task Index

Metric: USDView Methodology →