Prime Intellect measures how frontier models conduct autonomous AI research
articleRevision 1
No relevant image available
Prime Intellect reports 153 autonomous nanoGPT speedrun runs across 18 frontier models to test how well agents conduct multi-day AI research. The article finds that the best-performing models were not distinguished by wholly new methods, but by stronger experimental judgment: they handled noisy results, revisited earlier findings, and built useful research workflows. The authors also stress that the benchmark is variable and that its lessons may not transfer directly to real model training.