Logs and code expose failures hidden by final papers
The authors’ targeted auditor detects methodological pitfalls more accurately with paper, logs and code than with the final paper alone.
Five pitfall categories; balanced audit examples include naturally occurring failures and controlled perturbations. Panels show accuracy, true-positive rate, false-positive rate and F1. Error bars are bootstrapped 95% confidence intervals. The reported overall Fisher test uses n=200, p<0.0001.

View original chart and methods ↗
The more you automate, the less you see: Hidden pitfalls of autonomous AI scientists · Figure 4, page 6; Proposed Remedies and Limitations and Scope
How to interpret this chart
- Tests two systems on controlled symbolic-pattern-reasoning workflows; not a measured failure rate for all AI science.
- Data-leakage and metric-misuse audit positives were introduced through controlled perturbations; dataset fabrication examples include natural and introduced cases.
- The underlying experiments found no direct test-set peeking and no clear deliberate metric gaming.
- Rule-guided forensic prompting tailored to these pitfalls is not a validated general-purpose scientific auditor.
- No exact bar heights extracted from plot geometry.