Karpathy Autoresearch is Andrej Karpathy’s three-file open-source loop in which an AI coding agent edits a nanochat training script overnight, runs five-minute experiments on a single GPU, and keeps or discards each one on the measured validation loss.
Autoresearch is the founding artifact of the loop family this category tracks, and its judging rule is the simplest here: the validation curve is the only oracle, so the loop cannot claim a win the measurement did not produce, and its failure mode is overfitting the metric rather than fabricating the result.
What it is #
A GitHub repository whose working surface is deliberately three files: prepare.py (fixed data preparation and evaluation, not modifiable), train.py (the single file the agent edits, holding the GPT model, the Muon plus AdamW optimizers, and the training loop), and program.md (the instructions the human edits, which Karpathy calls the research-org code).
A run works like this: the agent agrees a run tag, branches autoresearch/<tag>, reads the repository, then experiments autonomously, each experiment training for a fixed five minutes on one GPU, with every result appended to a results.tsv ledger and kept or reverted on the numbers.
The training target is a simplified single-GPU implementation of Karpathy’s nanochat (58,537 stars as of 2026-10-10).
Karpathy announced it in March 2026 with a deliberately mythologizing framing: the repo is “the story of how it all began” for a future where agent swarms run research.
Status #
Dormant as a repository, dominant as an influence: 97,621 stars, 13,564 forks, created 2026-03-06, last push 2026-03-26, no releases and no license file, as of 2026-10-10.
The community footprint is the largest in this category: the March 19, 2026 Hacker News thread on scaling it reached 237 points with 94 comments, and the project spawned a dedicated awesome list (created 2026-03-20) that catalogs dozens of descendant skills, ports, and domain adaptations. The strongest independent test came from SkyPilot (March 18, 2026): pointed at the loop with 16 Kubernetes GPUs, Claude Code submitted about 910 experiments in 8 hours, taught itself to screen ideas on H100s and validate winners on H200s, and moved validation loss from 1.003 to 0.974, a 2.87 percent improvement over baseline.
Strengths #
- The judge cannot be argued with: every experiment’s verdict is a measured loss number, the only oracle in this category that no model or vendor mediates.
- The programming surface is Markdown, so the research-org design is legible, diffable, and copyable in a way no harness codebase is.
- It is the smallest complete loop in the category: three files, one GPU, no framework, no API keys.
- The scaling test shows the loop’s strategy changes with compute: with 16 GPUs the agent ran factorial waves of 10 to 13 experiments and caught interaction effects that sequential search misses.
Cautions #
- The practitioner critique cuts at ambition, not accuracy: the top scaling-thread replies read the loop as automated hyperparameter tuning at toy scale, “brute-force search, but guided”, and note that a two-node cluster is the whole test bed.
- No license file ships with the repository, so reuse beyond reading it sits in unclear legal ground.
- Dormant since 2026-03-26: Karpathy moved on, and the descendants (skills, ports, benchmarks) carry the line rather than the original.
- The fixed five-minute budget bounds what an improvement can mean, and the 2.87 percent cluster result is the vendor’s own run, not a replicated one.
Pricing #
Free; the loop costs your own GPU time and model API calls. There is no paid tier and no license under which to sell it, so pricing does not apply.
Compared to #
- Agon: the producer-critic factory that wraps the loop in adversarial review and aims at papers; Autoresearch is the one-agent keep-or-revert ancestor with the harder oracle.
- OpenResearch: the workspace-grade descendant that industrializes the loop on your own agents with git lineage; Autoresearch is the three-file prototype.
- DeepAnalyze: the trained-model descendant where the loop is learned in the weights rather than prompted in a Markdown file.
Bottom line #
Recommended as the first loop to read in this category: everything else here descends from it or competes with its structure, and its judge is the one to steal. Not for anyone who needs a maintained tool, a license that permits reuse, or research beyond a toy training run.
Changes #
- 2026-10-10 - Created.
See also #
- Agon - the producer-critic elaboration of the keep-or-revert loop
- OpenResearch - the workspace descendant running the loop on your own agents
- DeepAnalyze - the trained-model descendant
- Andrej Karpathy - the author’s profile note in this section
- Automated Research Feature Matrix - the category comparison this note joins
References #
https://github.com/karpathy/autoresearch - the repository: three-file design, keep-or-revert loop, nanochat base, and the March 2026 framing (fetched 200, 2026-10-10)
https://api.github.com/repos/karpathy/autoresearch - stars, forks, created and pushed dates, and the no-license status for the as-of line (fetched 200, 2026-10-10)
https://raw.githubusercontent.com/karpathy/autoresearch/master/program.md - the agent-facing loop definition: branch per run, results.tsv ledger, fixed five-minute budget (fetched 200, 2026-10-10)
https://blog.skypilot.co/scaling-autoresearch/ - the independent 16-GPU scaling test: about 910 experiments in 8 hours, the H100-screen and H200-validate strategy, 1.003 to 0.974 validation loss (fetched 200, 2026-10-10)
https://hn.algolia.com/api/v1/items/47442435 - the 237-point, 94-comment scaling thread and its hyperparameter-tuning and brute-force critiques (fetched 200, 2026-10-10)
https://raw.githubusercontent.com/webfuse-com/awesome-autoresearch/main/README.md - the dedicated descendant list the loop spawned (fetched 200, 2026-10-10)
https://api.github.com/repos/karpathy/nanochat - the nanochat base and its star count (fetched 200, 2026-10-10)