Agent Harness Skew & Scaffolding Workbench

Isolate whether LLM agent benchmark failures stem from core model reasoning or defective scaffolding (parsing, error reflection, truncation & statelessness).

Diagnostic Finding: Harness B achieves +24.4% higher Pass@1 on identical tasks purely by fixing tool parse retries and adding compiler error reflection.

Harness Scaffolding A (Baseline)Pipeline A

Pass@1 Rate
34.2%
Harness Bottleneck
54.0%
Attribution Breakdown (A)

Harness Scaffolding B (Optimized)Pipeline B

Pass@1 Rate
58.6%
Harness Bottleneck
16.0%
Attribution Breakdown (B)

Failure Attribution & Skew Telemetry

Pass (Resolved) JSON / Tool Call Timeout Unreflected Stderr Trap Truncation Blindness Stateless Lost State True Logic / Reasoning Fail

Sample Agent Execution Traces (Harness A)

Sample Agent Execution Traces (Harness B)

Enjoy this tool? Build your own with Super