Explainers
No explainer here yet. This is the kind of cell that gets one as soon as something in the queue earns it.
Unread — 8
The Art of Scaling Test-Time Compute for LLMs
30B generated tokens across eight models: no single scaling strategy wins, and the pick depends on trace horizon.
arXiv 2512.02008 paper · Apr 2026Scaling Test-Time Compute for Agentic Coding
Recursive Tournament Voting plus Parallel-Distill-Refine lifts SWE-Bench Verified from 70.9 to 77.6.
arXiv 2604.16529 post · Jan 2026Categories of Inference-Time Scaling
Clean taxonomy separating self-consistency, best-of-N, verifier rejection sampling, self-refinement and search.
Ahead of AI — Sebastian Raschka paper · Aug 2024Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Compute-optimal allocation beats best-of-N by over 4×, letting a small model outperform one 14× larger at matched FLOPs.
arXiv 2408.03314 paper · Jul 2024Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Coverage scales log-linearly across four orders of magnitude of samples: 15.9% to 56% on SWE-bench Lite at 250 draws.
arXiv 2407.21787 paper · Mar 2022Self-Consistency Improves Chain of Thought Reasoning in Language Models
Sampling many chains and majority-voting adds 17.9 points on GSM8K over greedy decoding — the origin of parallel test-time scaling.
arXiv 2203.11171 paper · Mar 2025Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
First structured survey of overthinking, sorting the fixes into model-level, output-level and prompt-level ways to cut tokens.
arXiv 2503.16419 paper · Apr 2025Sleep-time Compute: Beyond Inference Scaling at Test-time
Precomputing context offline cuts the test-time compute needed for equal accuracy by about 5×, and lifts Stateful AIME by up to 18%.
arXiv 2504.13171