How to read this tracker

Methodology

Understand how we choose sources, check reported results, and decide which scores can be compared. This page also explains the labels and reference points used throughout the tracker.

Scope and categories

The tracker follows six evidence categories: AI-R&D capability, AI-system improvement, research autonomy, observed R&D automation, recursive-improvement evidence, and evaluation integrity. These are navigation labels, not a maturity ladder.

Evidence types stay distinct

Benchmark results, operational observations, autonomy measures, qualitative assessments, recursive-improvement studies, and integrity evaluations answer different questions. A general research agent result can inform the picture without directly demonstrating recursive self-improvement.

Source and review status

Records point to original lab reports, papers, official benchmark materials, or data releases when available. “Agent checked” means the extraction was checked against a source; it does not mean the experiment was independently reproduced. Empirical verification and extraction review are recorded separately.

Lab-reported evidence

A lab-reported result is attributed to the reporting lab. Independent evaluation and independent replication are separate labels. Publicly reported results from private evaluations may be included when their evidence is publicly inspectable.

Comparability

Results can be connected only within a version, metric, and defensible comparison group. Unknown settings do not count as matching. Limited-comparability groups retain their caveats and do not support automatic deltas. Different labs’ internal benchmark percentages are not ranked together.

When a valid difference is shown for a percent metric, it is expressed in percentage points. Rounded point estimates alone do not establish statistical significance. A single point remains a point, with no invented trend line.

Thresholds and reference points

Human baselines, policy triggers, normalization anchors, and estimated necessary capabilities are represented separately. An estimated score that may be necessary for a task does not prove that reaching the score is sufficient for researcher substitution or RSI. Reference points appear only when their metric and version applicability are supported.

Dates, missing values, and corrections

Evaluation date, publication date, model release date, and tracker-added date are distinct. Month-only dates retain month precision. “Not reported” and “not evaluated” are missing evidence states, never zero. Superseded or corrected results remain traceable and are excluded from current summaries by default.

Recursive-improvement claims

Studies describe what changed, what stayed fixed, whether an improved system later served as optimizer, human contribution, transfer evaluation, and resource accounting when reported. Repeated gains from a fixed optimizer are not automatically evidence that the optimizer itself improved. Increasing performance is not the same as increasing improvement efficiency.

Limitations

Public evidence is incomplete and uneven. The tracker does not estimate a universal fraction of RSI, predict a timeline, or infer capability from nondisclosure. Source availability can change; a current access failure does not erase a historically supported observation.

Suggest an improvement

Tell us how we could improve the site. Leave an email address if you would like a reply.

JavaScript is required to send a suggestion.