peer reviewed · 2026-10-05

The more you automate, the less you see: Hidden pitfalls of autonomous AI scientists

Why it matters here

Controlled experiments test Agent Laboratory and AI Scientist v2 on benchmark selection, data handling, metric choice and result selection. They reveal positional or difficulty-based benchmark biases, undisclosed dataset substitutions and reward functions influenced by test results. The study finds no direct test-set peeking and no clear deliberate metric gaming; metric behavior instead shows procedural instability. A separate targeted auditing experiment finds that logs and code expose problems largely missed by manuscript-only review.

Read original source ↗

Evaluation integrityResearch autonomy

What the charts show

Logs and code expose failures hidden by final papers

The authors’ targeted auditor detects methodological pitfalls more accurately with paper, logs and code than with the final paper alone.

Five pitfall categories; balanced audit examples include naturally occurring failures and controlled perturbations. Panels show accuracy, true-positive rate, false-positive rate and F1. Error bars are bootstrapped 95% confidence intervals. The reported overall Fisher test uses n=200, p<0.0001.

Original four-panel Figure 4 comparing paper-only and paper-plus-logs/code auditing.
Original Figure 4 cropped from page 6 of the owner-supplied published PDF at 150 dpi; page download watermark and unrelated text excluded. No marks or values altered. Enlarge chart ↗

View original chart and methods ↗
The more you automate, the less you see: Hidden pitfalls of autonomous AI scientists · Figure 4, page 6; Proposed Remedies and Limitations and Scope

How to interpret this chart
  • Tests two systems on controlled symbolic-pattern-reasoning workflows; not a measured failure rate for all AI science.
  • Data-leakage and metric-misuse audit positives were introduced through controlled perturbations; dataset fabrication examples include natural and introduced cases.
  • The underlying experiments found no direct test-set peeking and no clear deliberate metric gaming.
  • Rule-guided forensic prompting tailored to these pitfalls is not a validated general-purpose scientific auditor.
  • No exact bar heights extracted from plot geometry.

What to keep in mind

  • Tests two systems on controlled symbolic-pattern-reasoning workflows; not a measured failure rate for all AI science.
  • Data-leakage and metric-misuse audit positives were introduced through controlled perturbations; dataset fabrication examples include natural and introduced cases.
  • The underlying experiments found no direct test-set peeking and no clear deliberate metric gaming.
  • Rule-guided forensic prompting tailored to these pitfalls is not a validated general-purpose scientific auditor.
  • No exact bar heights extracted from plot geometry.