[Orchestrator] Initializing autonomous discovery tree...
[Formulator] Hypothesis drafted: Entropy-gated dynamic draft verification.
Live LaTeX Manuscript Pane
Dynamic Entropy-Gated Speculative Decoding for Large Language Models
Abstract
We introduce dynamic entropy gating for speculative decoding. By dynamically truncating candidate draft tokens when target model prefix entropy exceeds a calibrated threshold , we eliminate catastrophic verification rollovers. Our empirical results demonstrate a 2.41x token/s speedup over the Llama-3-70B baseline with a negligible accuracy delta of -0.12% on MMLU.
Methodology: Dynamic Entropy Drafting
Given draft policy and target distribution , the agent modulates sequence length via Shannon entropy:
Empirical Evaluation & Table 1
| Experiment Node | Methodology | Throughput | MMLU Δ | Status |
|---|---|---|---|---|
| ROOT-BASE | Vanilla Autoregressive | 34 tok/s | 0.00% | Baseline |
| EXP-04-ENTROPY-TAU-0.45 | Dynamic Entropy Gating | 82 tok/s (2.41x) | -0.12% | Optimal |
| EXP-02-FIXED-K8 | Fixed Speculation (K=8) | 58 tok/s | -0.84% | Pruned |
| EXP-03-GREEDY-DRAFT | Pure Greedy Draft | 44 tok/s | -1.42% | Pruned |
Ablation Analysis
Sweeping highlights the optimal Pareto efficiency at , minimizing speculative verification latency while maintaining mathematical consistency.