Evidence record

KernelGen 1P

GPT-6.1 Sol card snapshot, September 29, 2026 · GPT-6.1 Sol · 2026-09-29

Correction or update linkedGPT-6.1 Sol adds research-debugging 75.52% and KernelGen 1P 60.44% rubric scores as separate report snapshots.

Reported result · numeric

60.44%

Mean kernel optimization reward · percent

lab reported · extraction review: agent checked

Sources

Published 2026-09-29

Evaluation setup
aggregate methodMean rubric score (average reward), percent
comparability caveatsSource refers to Astra card methods but exact scored task population, attempts and per-model budgets are not restated.; Older model comparison values can reflect later model versions according to §2. Keep this dated snapshot separate from launch records.

Not reported: split, task count, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

Not compared with other results.

Limitations

  • Mean rubric score, not percentage of tasks solved or RSI achieved.
  • Separate dated snapshot; task population and budgets are not fully disclosed.
  • OpenAI reports GPT-6.1 Sol remains below its High AI self-improvement threshold; no numeric threshold inferred.
  • Not a percentage reduction in runtime or training cost.
  • Internal suite; no cross-lab percentage comparison.
  • Not a demonstration of recursive improvement.