benchmark

PostTrainBench Lite

Improve a pretrained model within a five-hour GPU budget.

AI-system improvement

What this measure tells us

Targets a specific AI-development task.

  • Internal suite; no cross-lab percentage comparison.
  • Not a demonstration of recursive improvement.

Results by version

Astra card snapshot, September 2026

2026-09-03 · 12 tasks

Version identity is this disclosure snapshot, not an invented upstream version.

Metric definitions

Gain toward instruction-tuned reference performance (normalized score): Improvement over the base model divided by the gap to an instruction-tuned reference, limited to 0–1. This is not raw downstream accuracy.

Published results · Astra card snapshot, September 2026
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—GPT-6 AstraGain toward instruction-tuned reference performanceResults appear in a chart at several output-token budgets; exact point values are not printed.2026-09-03posttrainbench-lite-protocollab reported
—GPT-5.6 SolGain toward instruction-tuned reference performanceResults appear in a chart at several output-token budgets; exact point values are not printed.2026-09-03posttrainbench-lite-protocollab reported

Version lineage

Reference points

normalization anchor

1 · normalized score

1 represents the instruction-tuned reference model’s performance for each task.

What this reference means: unknown · Where it applies: established

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksno
public codeunknown
public evaluation serviceunknown

Official sources