Isolate whether LLM agent benchmark failures stem from core model reasoning or defective scaffolding (parsing, error reflection, truncation & statelessness).