benchmark creator
PostTrainBench team
Creators of the original full benchmark and its maintained v1.1 leaderboard.
Published evidence
No published results indexed yet.
This means no result is currently represented in this tracker; it does not indicate zero capability or zero disclosure.
Relevant sources
PostTrainBench: Can LLM Agents Automate LLM Post-Training? ↗
original paper · 2026-03-09 · available
PostTrainBench v1.1 leaderboard and methodology ↗
official benchmark page · 2026-07-28 · available
PostTrainBench website scores.json at commit 2270545 ↗
repository · 2026-09-18 · available
Hardening PostTrainBench against reward hacking ↗
organizational research article · 2026-07-28 · available
Our response to Epoch AI’s review of PostTrainBench ↗
organizational research article · 2026-09-17 · available
PostTrainBench website config.js: agent identities and scaffolds ↗
repository · 2026-09-18 · available
PostTrainBench agent run prompt and rules ↗
repository · 2026-08-21 · available
PostTrainBench judge policy and score effects ↗
repository · 2026-08-21 · available
PostTrainBench API use judge prompt ↗
repository · 2026-08-21 · available
PostTrainBench score collection and fallback code ↗
repository · 2026-08-21 · available