—

Em Dash Frequency & AI Register Analyzer

Pew Web Baseline: 6.0 (2023) → 11.5 (2026) / 10k words
Editorial Drafting & Tokenizer
Corpus Presets:
Syntactic Token Map (Click highlighted em dash to inspect):
One-Click De-Biasing Transformations Rewrites detected em dashes into human cadences
Cadence Telemetry & Pew Benchmarks
AI Syntactic Marker Density
High LLM Marker Density (exceeds 2026 web baseline of 11.5 / 10k words)
Word Count 62
Em Dash Count (—) 4
Em Dash Rate / 10k words 645.16
En Dash (–) & Hyphen (-) 0 / 1
vs 2023 Pew Baseline (6.0/10k) +10652.7%
vs 2026 Pew Baseline (11.5/10k) +5510.1%
Empirical Benchmark Comparison Rate per 10,000 words
Forensics Artifact Export

Dash frequency normalization arithmetic

Read the explanation

Source counts exact em dash characters and word tokens using its regular expression. One hypothetical em dash among one thousand words normalizes to ten per ten thousand; two normalize to twenty. Bars share twenty-five pixels per dash per ten thousand words. Exact punctuation frequency alone does not identify an author or prove text was generated by a model. Keeping one exact em dash while doubling hypothetical token count from one thousand to two thousand halves normalized frequency from ten to five per ten thousand. Bars share forty pixels per unit. En dashes and intra-word hyphens are separate counters, so visually similar marks do not enter this numerator. For hypothetical rate ten and stored reference six, source subtracts six then divides six, giving about sixty-six point six seven percent above the reference. Bars compare the two rates at fifty pixels per unit. The stored benchmark labels and research attribution are not independently verified, and the heuristic density verdict is not reliable AI detection. Complete raw JSON and native reset/resize/footer checks preserve the existing analysis tool.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.