Main benchmark
Across 92 sampler–LLM–dataset settings, PreSTO (ours) speeds up PowerMH, EntropyCut, and MultiTryMH by a median of 1.71×, 1.88×, and 2.45×, respectively.
Sample from the LLM’s power-sharpened distribution:
• p0 is the base model.
• x1:n is the continuation, and x0 is the prompt.
Raising sequence probabilities to α favors higher-likelihood continuations.
TL;DR: We present PreSTO, which preserves the same MH chain while running much faster.
Across 92 sampler–LLM–dataset settings, PreSTO (ours) speeds up PowerMH, EntropyCut, and MultiTryMH by a median of 1.71×, 1.88×, and 2.45×, respectively.
At 100 transitions, cumulative time falls from 481.02 to 253.29 seconds. Calls do more work, but the reduction in their number outweighs their higher cost.
Qwen3.5-9B · LiveCodeBench · Prefetch budget 10 · BFS default Traversal
Across seven datasets, approximately 67%–84% of edges are ε-certain in the prefetched trees (PreSTO-PowerMH, Qwen3.5-9B).
Predictive traversal raises mean speedup in all six settings, reaching 2.46×. The largest time reductions occur for EntropyCut at budgets 4–12; the smaller gains fall within one standard deviation of cumulative time.
Higher is better
Higher is better
Seconds · mean ± SD
Qwen3.5-9B · LCB V6 · 100 MH transitions
Paired TOST establishes equivalence in 12 of 16 PowerMH tests and 17 of 18 EntropyCut tests. Both samplers are equivalent when pooled, within ±0.2 baseline standard deviations.
Boxes show per-prompt terminal-trace distributions under p₀ for eight PowerMH pairs and nine EntropyCut pairs, with prefetch budget 20. Shaded rows establish equivalence; paired TOST p-values appear at right.
Speedup peaks at 1.94× with budget 16, using 5.01% of the preallocated KV pool. Larger budgets increase uncached work without a consistent speedup gain.
| Metric | PowerMH | PreSTO · Prefetch budget | ||||
|---|---|---|---|---|---|---|
| 4 | 8 | 12 | 16 | 20 | ||
Qwen3.5-9B · LiveCodeBench · BFS default Traversal · One seed per budget
Charts could not load. Please reload the page to try again.
@misc{jiang2026presto,
title = {{PreSTO}: Predictive Subtree Prefetching for Fast {LLM} Power Sampling},
author = {Jiang, Nan and Theodoropoulos, Panagiotis and Deng, Wei},
year = {2026}
}