Explainers
No explainer here yet. This is the kind of cell that gets one as soon as something in the queue earns it.
Unread — 4
paper · Nov 2025
Measuring what Matters: Construct Validity in LLM Benchmarks
Reviews 445 benchmarks, finds most fail construct validity, and gives eight concrete fixes for eval builders.
arXiv 2511.04703 paper · Jul 2025Best Practices for Building Rigorous Agentic Benchmarks
Documents real scoring bugs in SWE-bench Verified and TAU-bench, then a checklist that cuts overestimation by a third.
arXiv 2507.02825 code · Active 2026Inspect — LLM evaluation framework
The UK AISI production harness: solvers, scorers, model-graded judges, 200+ prebuilt evals. A better base than your own stub.
GitHub — UKGovernmentBEIS/inspect_ai paper · Nov 2024Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Most published eval gaps are reported without sampling error; supplies the standard-error and clustering corrections that say whether a difference is real.
arXiv 2411.00640