preprint · 2026-09-29

AIM: Agentic Idea Management for Automated Research

Why it matters here

AIM organizes research ideas, estimates their promise, allocates experiments and audits whether implementations match the ideas credited for their results. On ten AutoLab tasks using Gemini-3.1-Pro-Preview, it reports average scores of 67.0 versus 65.4 for ScientistOne on five systems tasks, and 55.8 versus 50.9 on five model-development/CUDA tasks. Flash Attention trajectories show faster attainment of strong solutions. This is evidence about a human-designed research-management scaffold operating against fixed verifiers.

Read original source ↗

AI-R&D capabilityAI-system improvementEvaluation integrity

What the charts show

AIM reaches strong Flash Attention solutions earlier

On Flash Attention, AIM reaches 90.5% at 3.3 hours; ScientistOne reaches 86.2% at 3.5 hours. The authors report up to 3.1× faster time to the strongest baseline’s score.

Left: mean best-so-far AutoLab score (%) versus elapsed wall-clock hours; right: labeled endpoint scores/times and AIM times reaching baseline levels. All methods use Gemini-3.1-Pro-Preview; up to 300 executions and six hours for this task. Three independent runs; Table 1 reports mean and standard error.

Two original panels comparing mean score trajectories and endpoint times.
Original source figure faithfully rasterized from its SVG at 150 dpi. Enlarge chart ↗

View original chart and methods ↗
AIM: Agentic Idea Management for Automated Research · Figure 4 (HTML container S5.F5); Section 5.1; Table 1

How to interpret this chart
  • This figure concerns Flash Attention only; the 3.1× claim is not a suite-wide research speedup.
  • Normalized AutoLab score is not percentage runtime reduction. No uncertainty bands are shown in this time plot.
  • Equal nominal execution/time caps do not imply equal tokens or monetary costs.
  • The authors designed the management scaffold; this does not demonstrate autonomous improvement of the optimizer itself.
Published values from the chart

Exact labels in Figure 4. Endpoints differ in quality; do not divide endpoint times to reproduce the matched-score 3.1× claim.

AIM reaches strong Flash Attention solutions earlier
MethodPrinted endpoint scorePrinted time
AIM90.5%3.3 h
ScientistOne86.2%3.5 h
AdaEvolve85.3%5.5 h
AIRA (MCTS)77.9%3.7 h
Arbor (depth=3)77.2%2.2 h
DeepScientist72.3%5.0 h

What to keep in mind

  • Author-reported preprint with three independent runs; source standard errors are not confidence intervals.
  • Task scores and domain averages do not measure the percentage of research automated or percentage progress toward RSI.
  • AIM does not win every task; AdaEvolve reports a higher Data Select IFEval score. Some baseline methods lack Moving MNIST results; Table 2’s displayed averages conflict with its caption about omitted averages, so incomplete-baseline averages are not used here.
  • Fixed tasks, verifiers and base model; no demonstrated recursive improvement of the research manager itself.