Evidence record
KernelGen 1P
GPT-6.1 Sol card snapshot, September 29, 2026 · GPT-6.1 Sol · 2026-09-29
Correction or update linkedGPT-6.1 Sol adds research-debugging 75.52% and KernelGen 1P 60.44% rubric scores as separate report snapshots.
Reported result · numeric
60.44%
Mean kernel optimization reward · percent
lab reported · extraction review: agent checked
Sources
Published 2026-09-29
Evaluation setup
| aggregate method | Mean rubric score (average reward), percent |
|---|---|
| comparability caveats | Source refers to Astra card methods but exact scored task population, attempts and per-model budgets are not restated.; Older model comparison values can reflect later model versions according to §2. Keep this dated snapshot separate from launch records. |
Not reported: split, task count, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
Not compared with other results.
Limitations
- Mean rubric score, not percentage of tasks solved or RSI achieved.
- Separate dated snapshot; task population and budgets are not fully disclosed.
- OpenAI reports GPT-6.1 Sol remains below its High AI self-improvement threshold; no numeric threshold inferred.
- Not a percentage reduction in runtime or training cost.
- Internal suite; no cross-lab percentage comparison.
- Not a demonstration of recursive improvement.