Distillation drives inferencing demand, spurs smaller efficient chips, and allows knowledge diffusion akin to reading textbooks written by others.
Training on frontier outputs without bearing $100M+ pretraining costs free-rides on US capital investments and violates provider Terms of Service.
| Benchmark Metric | Frontier Teacher | Distilled Student | Retention % | Synthetic Residual Watermark |
|---|
The Great AI Distillation Debate: Economics & Policy
What is Model Distillation in Modern LLMs?
Knowledge distillation transfers capabilities from an expensive, high-parameter "Teacher" model (e.g. 70B–671B MoE) to a lightweight "Student" model (e.g. 1.5B–14B). Instead of spending $50M–$200M on raw web crawl pre-training, the student trains on distilled synthetic outputs: reasoning chains, verifier-filtered math/code traces, or soft logits with high temperature.
Why does Jensen Huang view Distillation as "Competition"?
NVIDIA CEO Jensen Huang argues that distillation is the natural engine of technological progress. It democratizes frontier intelligence, enabling lightweight models to run on phones, laptops, and local enterprise data centers. Crucially for hardware vendors, distillation generates billions of inference tokens and expands the total addressable market for inferencing compute.
Why does Scott Bessent (and frontier labs) view Distillation as "Theft"?
US Treasury Secretary Scott Bessent and leading frontier AI labs argue that proprietary AI models represent massive capital investments in training compute, data licensing, and safety engineering. When a competing firm or foreign entity queries a model to train an imitation or competitor model, it breaches Terms of Service, bypasses pretraining expenditures, and free-rides on US frontier innovation.