Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours, build by @karinanguyen, @ThoughtfulLab_posttrainbench.comJoined March 2026
$PTB
Can AI agents improve performance of base LLMs? We give each agent 4 small target LLMs, an H100 GPU, and 10 hours to post-train them.
build by @thoughtfullab , @karinanguyen
CA : GtQndxV3DxFqMZDkXMdT17BLY4fH8BHM6hqgWThjBAGS
PostTrainBench v1.0!
This benchmark evaluates the ability of frontier AI agents to post-train language models in a simplified setting.
We believe this is a first step toward tracking progress in recursive self-improvement 🧵:
From ImportAI by @jackclarkSF (thank you for the feature):
"Imagine where we’ll be in two years - we’ll certainly have AI models that are smart enough to point themselves at a specific objective, find an open weight model, then autonomously improve it to get better performance
Introducing PostTrainBench
How well can AI agents post-train language models? We built a benchmark to find out.
Post-training is how raw language models become useful, the stage that turns a capable but unsteered base model into a system that follows instruction
Since our initial release, we made our benchmark more robust: - added more tasks (ArenaHard-Writing and HealthBench-Easy are new) - ran more seeds - weighted the tasks by difficulty We also added a lot of agents! 3/3
Findings: 1. There was a lot of progress recently (Sonnet 4.5 was at 9.9%, Opus 4.6 has 23.2%) 2. Agents lack behind instruct tuned models by human engineers (51.1%) 3. Persistent agents are usually more performant 4. Agents like to cheat. E.g. Opus 4.6 and GPT-5.1 Codex Max 2/n
PostTrainBench v1.0 and the accompanying paper are out! We believe this benchmark will be important to measure progress in AI R&D automation. What are our findings? 1/n