Modeled Forensic Ledger
Modeled Pre-training Tokens Retained
Calculated using standard model benchmarks: 370 tokens/page yield, 2.8g paper mass/leaf, and 1,000 pages/batch processing increments.
Physical Print to Training Tokens Forensics
Quantify physical book de-binding, optical scan throughput, character loss, and training dataset mass with zero data leaving your browser.
Ingestion Modeling Engine
Adjust hypothetical acquisition quantities, mechanical scanner speeds, and optical degradation rates to inspect modeled dataset yield and physical waste tonnage.
Modeled Pre-training Tokens Retained
Calculated using standard model benchmarks: 370 tokens/page yield, 2.8g paper mass/leaf, and 1,000 pages/batch processing increments.
High-throughput corporate digitization removes spine bindings with precision industrial guillotines to feed loose leaves through sheet-fed scanners at 200+ pages per minute.
This forensic workbench models physical workflow throughput and unit costs. It does not certify copyright legality, fair use exemptions, or commercial licensing compliance.
Custody Verification
When artificial intelligence developers procure millions of used physical volumes, the physical chain of custody—purchase orders, warehouse freight bills of lading, and recycling manifests—often contradicts claims of purely digital web-crawl sourcing.
The default assumes one hundred thousand books with three hundred twenty pages each: thirty two million gross pages. At three hundred seventy assumed tokens per page that is eleven point eight four billion potential tokens. Rejecting two point five percent leaves thirty one point two million equivalent pages and eleven point five four four billion retained tokens, removing point two nine six billion. At twenty pixels per billion tokens gross potential measures two hundred thirty six point eight, retained two hundred thirty point eight eight and removed five point nine two. This is scalar arithmetic. No book, OCR text, tokenizer or training data is actually inspected, and the loss input is not a measured character error rate. Acquisition assumes three dollars twenty cents per book, totaling three hundred twenty thousand dollars. Scanning assumes four dollars fifty cents per thousand gross pages; thirty two thousand groups of a thousand cost one hundred forty four thousand dollars. Direct cost is four hundred sixty four thousand. At one pixel per two thousand dollars acquisition measures one hundred sixty, scanning seventy two and total two hundred thirty two. The scan input is per thousand pages, not per book or individual page. Assigned paper mass uses two point eight grams per gross page, yielding eighty nine point six metric tons. OCR loss changes token yield but not acquisition, scanning cost or gross paper mass. No physical purchasing, scanning or destruction occurs. Raising rejection from two point five to ten percent reduces retained equivalent pages to twenty eight point eight million. At three hundred seventy assumed tokens per page that is ten point six five six billion, a drop of point eight eight eight billion from the default eleven point five four four. At twenty pixels per billion tokens before measures two hundred thirty point eight eight, after two hundred thirteen point one two and reduction seventeen point seven six. Direct cost remains four hundred sixty four thousand, so cost per million retained tokens rises. Native fields, reset and actual JSON and CSV exports work without the guarded animation libraries. This output records assumptions and calculations; it does not establish copyright permission, provenance, legal compliance or actual training dataset quality.