Skene Agentic
Method

Test-Time Compute


nothing filed yet · 11 unread

‹ Back to Hive

Explainers

No explainer here yet. This is the kind of cell that gets one as soon as something in the queue earns it.

Unread — 8

paper · Dec 2025

The Art of Scaling Test-Time Compute for LLMs

30B generated tokens across eight models: no single scaling strategy wins, and the pick depends on trace horizon.

arXiv 2512.02008
paper · Apr 2026

Scaling Test-Time Compute for Agentic Coding

Recursive Tournament Voting plus Parallel-Distill-Refine lifts SWE-Bench Verified from 70.9 to 77.6.

arXiv 2604.16529
post · Jan 2026

Categories of Inference-Time Scaling

Clean taxonomy separating self-consistency, best-of-N, verifier rejection sampling, self-refinement and search.

Ahead of AI — Sebastian Raschka
paper · Aug 2024

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Compute-optimal allocation beats best-of-N by over 4×, letting a small model outperform one 14× larger at matched FLOPs.

arXiv 2408.03314
paper · Jul 2024

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Coverage scales log-linearly across four orders of magnitude of samples: 15.9% to 56% on SWE-bench Lite at 250 draws.

arXiv 2407.21787
paper · Mar 2022

Self-Consistency Improves Chain of Thought Reasoning in Language Models

Sampling many chains and majority-voting adds 17.9 points on GSM8K over greedy decoding — the origin of parallel test-time scaling.

arXiv 2203.11171
paper · Mar 2025

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

First structured survey of overthinking, sorting the fixes into model-level, output-level and prompt-level ways to cut tokens.

arXiv 2503.16419
paper · Apr 2025

Sleep-time Compute: Beyond Inference Scaling at Test-time

Precomputing context offline cuts the test-time compute needed for equal accuracy by about 5×, and lifts Stateful AIME by up to 18%.

arXiv 2504.13171

Back