benchmark

NanoGPT

Reduce small-model training time to a target validation objective.

AI-system improvement

What this measure tells us

Targets a specific AI-development task.

  • Internal suite; no cross-lab percentage comparison.
  • Not a demonstration of recursive improvement.

Results by version

Astra card snapshot, September 2026

2026-09-03

Version identity is this disclosure snapshot, not an invented upstream version.

Metric definitions

Fraction of baseline training time saved (normalized score): Reduction in training time to the target objective, divided by baseline training time and limited to 0–1; 0 means no improvement and 1 would mean zero training time.

Published results · Astra card snapshot, September 2026
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—GPT-6 AstraFraction of baseline training time savedResults appear in a chart at several output-token budgets; exact point values are not printed.2026-09-03nanogpt-protocollab reported
—GPT-5.6 SolFraction of baseline training time savedResults appear in a chart at several output-token budgets; exact point values are not printed.2026-09-03nanogpt-protocollab reported

Version lineage

Reference points

human baseline

0.7238 · normalized score

Source prints 72.38%; represented as 0.7238 on its 0–1 reward scale. Best human solution, not average human ability.

What this reference means: measured baseline · Where it applies: established

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksno
public codeunknown
public evaluation serviceunknown

Official sources