OpenAI-MRCR
MRCR v2 8-needle long-context benchmark where a model must recover the requested instance from repeated similar requests in a synthetic conversation.
Comprehensive directory of 63+ LLM benchmarks with honest status labels, contamination risk assessments, and vendor vs independent score provenance as of September 2026.
MRCR v2 8-needle long-context benchmark where a model must recover the requested instance from repeated similar requests in a synthetic conversation.
Configurable synthetic benchmark that expands needle retrieval into multi-needle, multi-hop tracing, and aggregation tasks.
Realistic bilingual long-context benchmark with 1,500 natural samples across 11 primary and 25 secondary tasks from 8k to 256k tokens.
Multiple-choice long-document QA where writers and validators read the full article, with a hard subset designed to defeat skimming.
Zero-shot long-text suite derived from SCROLLS with test-only tasks and new aggregation-style long-context challenges.
Scientific-paper QA over full NLP papers, with abstractive, extractive, yes/no, and unanswerable answers plus supporting evidence.
Query-focused meeting summarization over long transcripts from academic, product, and committee meetings.
Seven-task long-sequence suite covering summarization, QA, and NLI over naturally long English texts.
100k+ token benchmark mixing realistic and synthetic tasks across English, Chinese, code, math, and retrieval.
Bilingual multitask suite for long-context understanding across QA, summarization, few-shot learning, synthetic tasks, and code completion.
Synthetic retrieval stress test that hides facts at different depths and context lengths, then asks the model to recover them.