QMSum
Query-focused meeting summarization over long transcripts from academic, product, and committee meetings.
QMSum is still useful as a query-focused summarization component, but the public static 2021 standalone task is now mainly a subtask inside SCROLLS, ZeroSCROLLS, and LongBench-style suites.
Performance Timeline
Longitudinal progression of model scores against human baselines.Performance & Historical Trajectory
Empirical score progression across model release dates and evaluation rounds.
| Model | Score | Date | Source Type | Provenance |
|---|---|---|---|---|
| Random (paper Table 3 R-1) | 12.03% | 2021-04-13 | independent | Source ↗ |
| Ext. Oracle (paper Table 3 R-1) | 42.84% | 2021-04-13 | independent | Source ↗ |
| TextRank (paper Table 3 R-1) | 16.27% | 2021-04-13 | independent | Source ↗ |
| PGNet (paper Table 3 R-1) | 31.37% | 2021-04-13 | independent | Source ↗ |
| BART (paper Table 3 R-1) | 31.74% | 2021-04-13 | independent | Source ↗ |
| HMNet* (paper Table 3 R-1) | 32.29% | 2021-04-13 | independent | Source ↗ |
| PGNet (gold spans) (paper Table 3 R-1) | 31.52% | 2021-04-13 | independent | Source ↗ |
| BART (gold spans) (paper Table 3 R-1) | 32.18% | 2021-04-13 | independent | Source ↗ |
| HMNet (gold spans) (paper Table 3 R-1) | 36.06% | 2021-04-13 | independent | Source ↗ |
Human Baseline & Difficulty Horizon
Calibrated human reference points, specialist benchmarks, and ceiling thresholds.Human annotator score for query-based multi-domain meeting transcript summarization.
Metric & Scoring Methodology
Verification protocols, aggregation formulas, and specialized metric variants.ROUGE over query-focused summariesDataset & Compute Cost
Evaluation volume, public availability, API pricing, and local hardware requirements.$5 – $20 USD for full benchmark evaluation run on frontier APIs.
1x NVIDIA RTX 4090 (24GB) or A100 (40GB/80GB) via vLLM / SGLang
How to Run & Reproduce
Standardized evaluation protocols, CLI commands, and reproducible runner templates.lm_eval --model hf --model_args pretrained=<model_path> --tasks qmsum --batch_size autoopencompass --datasets qmsum --models <model_config># Standard API Evaluation Loop
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.0,
)When publishing results for QMSum, always report the exact prompt template, few-shot exemplar ordering, sampling temperature (temperature=0), maximum reasoning budget tokens, and the precise timestamped model snapshot ID.
Contamination & Memorization Analysis
Audit of pretraining exposure risks, memorization vectors, and refresh policies.Static fixed snapshot
Public on web / HuggingFace