Neutral AI Benchmark Studio
Benchmark scores alone mislead when cost, latency, hallucination drift, and test set contamination are concealed. Run multi-objective sensitivity audits to find the true Pareto frontier.
Score: 84.6 pts (Pareto Optimal)
$0.28 / M tokens at 79.2 pts
on un-dominated efficiency border
Ranked Multi-Criteria Assessment
| Rank & Model | Raw Accuracy | Hallucination | Latency (TTFT) | Cost / 1M | Neutral Score | Status |
|---|
Why AI Benchmarks Are Broken
Public LLM leaderboards frequently incentivize model builders to train directly on test sets (eval contamination) or cherry-pick evaluation prompts with extreme few-shot scaffolding.
- MMLU Saturation: High scores on standardized benchmarks no longer guarantee real-world reliability.
- Vendor Weighting: Standard leaderboards often omit cost and latency, making $30/M token models seem equivalent to $0.50/M token alternatives.
- Hallucination Penalties: A model that answers with 92% confidence but hallucinates 18% of facts is dangerous for regulated workflows.
How NeutralBench Restores Trust
This studio implements multi-objective Pareto optimization and contamination discounts:
- Pareto Dominance: Identifies models that cannot be beaten in accuracy without paying more or running slower.
- De-Contamination Multiplier: Applies a discount factor based on historical test-leakage flags and synthetic benchmark perplexity shifts.
- Normalized Composite Scoring: Converts disparate scales (tokens/sec, USD, error rate) into normalized percentiles.
Frequently Asked Questions
How is the neutral composite score calculated?
Each metric is normalized into a standard [0, 100] percentile score across the active candidate pool. Inverted metrics (cost, latency, hallucination rate) are transformed so that lower is better. User-selected objective weights then compute the weighted harmonic mean, adjusted by the contamination penalty.
Can I upload my own internal benchmark dataset?
Yes. Use the "Add Model Candidate" form to insert internal fine-tunes or custom endpoints. You can export the full synthesized scorecard as a JSON or CSV report for executive and engineering reviews.
What defines a "Pareto Optimal" model?
A model is Pareto optimal if no other active model achieves higher neutral quality at a lower or equal operational cost and latency. These models represent true efficient frontiers rather than marketing hype.