Scope and categories
The tracker follows six evidence categories: AI-R&D capability, AI-system improvement, research autonomy, observed R&D automation, recursive-improvement evidence, and evaluation integrity. These are navigation labels, not a maturity ladder.
Evidence types stay distinct
Benchmark results, operational observations, autonomy measures, qualitative assessments, recursive-improvement studies, and integrity evaluations answer different questions. A general research agent result can inform the picture without directly demonstrating recursive self-improvement.
Source and review status
Records point to original lab reports, papers, official benchmark materials, or data releases when available. “Agent checked” means the extraction was checked against a source; it does not mean the experiment was independently reproduced. Empirical verification and extraction review are recorded separately.
Lab-reported evidence
A lab-reported result is attributed to the reporting lab. Independent evaluation and independent replication are separate labels. Publicly reported results from private evaluations may be included when their evidence is publicly inspectable.
Comparability
Results can be connected only within a version, metric, and defensible comparison group. Unknown settings do not count as matching. Limited-comparability groups retain their caveats and do not support automatic deltas. Different labs’ internal benchmark percentages are not ranked together.
When a valid difference is shown for a percent metric, it is expressed in percentage points. Rounded point estimates alone do not establish statistical significance. A single point remains a point, with no invented trend line.
Thresholds and reference points
Human baselines, policy triggers, normalization anchors, and estimated necessary capabilities are represented separately. An estimated score that may be necessary for a task does not prove that reaching the score is sufficient for researcher substitution or RSI. Reference points appear only when their metric and version applicability are supported.
Dates, missing values, and corrections
Evaluation date, publication date, model release date, and tracker-added date are distinct. Month-only dates retain month precision. “Not reported” and “not evaluated” are missing evidence states, never zero. Superseded or corrected results remain traceable and are excluded from current summaries by default.
Recursive-improvement claims
Studies describe what changed, what stayed fixed, whether an improved system later served as optimizer, human contribution, transfer evaluation, and resource accounting when reported. Repeated gains from a fixed optimizer are not automatically evidence that the optimizer itself improved. Increasing performance is not the same as increasing improvement efficiency.
Limitations
Public evidence is incomplete and uneven. The tracker does not estimate a universal fraction of RSI, predict a timeline, or infer capability from nondisclosure. Source availability can change; a current access failure does not erase a historically supported observation.
Suggest an improvement
Tell us how we could improve the site. Leave an email address if you would like a reply.
JavaScript is required to send a suggestion.