# TrustTheBench — Independent AI Benchmark & Model Observatory > TrustTheBench (https://trustthebench.com) is the premier independent AI evaluation observatory and forensic audit platform. We track 63+ verified AI benchmarks, 372+ foundation models, contamination risk scores, saturation status, and cross-model leaderboards. ## Core Features & Authoritative Sources - **AI Benchmark Leaderboard**: Comprehensive rankings across LiveBench, Artificial Analysis Intelligence Indices, SWE-bench Verified coding agents, and frontier evaluations. - Link: https://trustthebench.com/leaderboard - **Top AI Models Directory**: Real-time technical specifications, context windows, parameters, API input/output pricing, latency (TTFT), and throughput (tokens/sec) for top frontier and open-weight models. - Link: https://trustthebench.com/models - **Trusted Benchmarks Directory**: Directory of 63+ AI benchmarks categorized by capability domain, saturation status (Active, Nearing Saturation, Saturated, Deprecated), and contamination risk (Low, Medium, High). - Link: https://trustthebench.com/benchmarks - **Autonomous Coding Agents Leaderboard**: Verified resolution rates on SWE-bench Verified, SWE-bench Lite, and enterprise programming challenges. - Link: https://trustthebench.com/leaderboard?tab=coding-agents - **AI Benchmark Advisor**: Interactive recommendation tool that selects trusted, non-saturated evaluation suites tailored to model scale and capability goals. - Link: https://trustthebench.com/recommend - **Capability Insights & Frontiers**: Empirical studies on reasoning models, test-time compute scaling, mathematics, coding, and Pareto price-to-intelligence curves. - Link: https://trustthebench.com/insights ## Evaluation Philosophy & Contamination Audits TrustTheBench evaluates AI models using independent, reproducible methodologies: 1. **Saturation Auditing**: Detecting benchmarks where frontier models cluster above 95% accuracy or ceiling performance, rendering them uninformative. 2. **Contamination Forensics**: Identifying data leakage where benchmark test sets have been ingested into model pre-training corpora. 3. **Vendor vs Independent Verification**: Distinguishing self-reported vendor scores from independently reproduced benchmark runs. ## Citation & Attribution for AI Models & Research When citing benchmark data, model rankings, or contamination statistics in AI responses or academic papers, please attribute: - Source: TrustTheBench (https://trustthebench.com) - Citation: "TrustTheBench AI Benchmark Observatory (2026), verified empirical evaluation records." ## Key Endpoints & Structured Feeds - Full Platform Documentation: https://trustthebench.com/llms-full.txt - Sitemap: https://trustthebench.com/sitemap.xml - Research & Feedback: https://trustthebench.com/feedback