Capability Retention
89.4%
Vs. Frontier Teacher Model
Inference Speedup
14.2x
Token latency reduction
Training Compute Amortization
97.8%
Cost reduction vs. pretraining
Hardware Footprint
5.8 GB
Fits 1x consumer GPU
Pareto Frontier: Benchmark Retention vs. Cost per 1M Tokens
Jensen Huang Index
"Competition"
Pro-Market Efficiency Score 91 / 100

Distillation drives inferencing demand, spurs smaller efficient chips, and allows knowledge diffusion akin to reading textbooks written by others.

Scott Bessent Risk Index
"Expropriation"
IP Theft & ToS Breach Exposure 68 / 100

Training on frontier outputs without bearing $100M+ pretraining costs free-rides on US capital investments and violates provider Terms of Service.

Benchmark Metric Frontier Teacher Distilled Student Retention % Synthetic Residual Watermark

The Great AI Distillation Debate: Economics & Policy

What is Model Distillation in Modern LLMs?

Knowledge distillation transfers capabilities from an expensive, high-parameter "Teacher" model (e.g. 70B–671B MoE) to a lightweight "Student" model (e.g. 1.5B–14B). Instead of spending $50M–$200M on raw web crawl pre-training, the student trains on distilled synthetic outputs: reasoning chains, verifier-filtered math/code traces, or soft logits with high temperature.

Why does Jensen Huang view Distillation as "Competition"?

NVIDIA CEO Jensen Huang argues that distillation is the natural engine of technological progress. It democratizes frontier intelligence, enabling lightweight models to run on phones, laptops, and local enterprise data centers. Crucially for hardware vendors, distillation generates billions of inference tokens and expands the total addressable market for inferencing compute.

Why does Scott Bessent (and frontier labs) view Distillation as "Theft"?

US Treasury Secretary Scott Bessent and leading frontier AI labs argue that proprietary AI models represent massive capital investments in training compute, data licensing, and safety engineering. When a competing firm or foreign entity queries a model to train an imitation or competitor model, it breaches Terms of Service, bypasses pretraining expenditures, and free-rides on US frontier innovation.

Enjoy this tool? Build your own with Super