nanoGPT: The simplest, fastest repository for training/finetuning medium-sized GPTs
nanoGPT: A Practical Guide to Training and Finetuning Medium-Sized GPTs

nanoGPT represents a compact, fast entry point for building and refining medium-sized language models. Born as a rewrite of minGPT, this project prioritizes practicality and accessibility: a straightforward training loop (train.py, roughly 300 lines) and a lean model definition (model.py, roughly 300 lines) that can even load GPT-2 weights from OpenAI. The goal is to give you something easy to hack, adapt, and experiment with—whether you want to train from scratch or finetune a pretrained checkpoint.
The project emphasizes transparency and readability. With a minimal, well-documented code path, you can quickly understand how data flows through the system, how the transformer layers are configured, and where to hook in your own ideas. On a single high-end GPU, the base setup can reproduce GPT-2 (124M) behavior on OpenWebText in about four days; more modest hardware can still yield useful results by dialing down the size and length. The included images show training dynamics and results, illustrating how a tiny, character-based model can generate surprisingly coherent Shakespeare-like text in a few minutes or hours, depending on resources.
Update Nov 2025: A new direction and a note on deprecation
There is an important update in the ecosystem. nanoGPT has a new and improved cousin called nanochat. It is very likely you meant to use/find nanochat instead. nanoGPT (this repo) is now very old and deprecated but I will leave it up for posterity.
- nanochat expands the same spirit of simplicity but with new features and a refreshed design that aligns with current tooling and best practices.
- If you’re starting a new project, consider using nanochat as your primary entry point.
- This post preserves the essence of nanoGPT’s approach and serves as historical context and a learning resource for those curious about the lineage of open-source GPT tooling.
This update reflects how quickly the field evolves and how lightweight tooling can evolve into more capable, maintainable ecosystems. The core ideas—straightforward data processing, transparent training loops, and flexible finetuning—still resonate, whether you start from nanoGPT or migrate to nanochat.
Quick start: your first steps with a Shakespeare character GPT
If you’re new to the domain, the fastest way to feel the magic is to train a character-level GPT on Shakespeare’s works. The steps below sketch the process, with representative commands you can copy-paste.
Dependencies and setup
Install the essential libraries:
- pip install torch numpy transformers datasets tiktoken wandb tqdm
Core dependencies to keep in mind:
- PyTorch (< 3)
- numpy
- transformers (for loading GPT-2 checkpoints)
- datasets (for OpenWebText and similar corpora)
- tiktoken (OpenAI’s fast BPE tokenization)
- wandb (optional logging)
- tqdm (progress bars)
Prepare Shakespeare data (character-level)
The pipeline first downloads Shakespeare as a single text file and converts it into a stream of integers:
- python data/shakespeare_char/prepare.py
This produces train.bin and val.bin in the data directory.
Train a tiny Shakespeare model on a GPU
With a local GPU, you can train a small GPT with the provided config. For example:
- python train.py config/trainshakespearechar.py
This config sets a context window up to 256 characters, 384 feature channels, 6 Transformer layers, and 6 attention heads per layer.
On an A100, this run typically finishes quickly and yields a validation loss around 1.47.
Sampling from the best Shakespeare model
After training, sample from the best model with:
- python sample.py --out_dir=out-shakespeare-char
Sample outputs demonstrate the model’s ability to generate imaginative and sometimes witty text in the Shakespearean space. For example, you might see lines that resemble classical dialogue, with a modern, playful twist.
If you’re on a Mac (Apple Silicon) or a recent machine
You can leverage on-chip GPUs with Metal Performance Shaders (MPS). Use:
- python train.py config/trainshakespearechar.py --device=mps
PyTorch on Apple Silicon often provides substantial speedups, enabling larger experiments on consumer hardware.
CPU-only or lighter hardware
If you don’t have a GPU, you can still run a smaller setup (though slower). A representative CPU-based run uses:
- python train.py config/trainshakespearechar.py --device=cpu --compile=False --evaliters=20 --loginterval=1 --blocksize=64 --batchsize=12 --nlayer=4 --nhead=4 --nembd=128 --maxiters=2000 --lrdecayiters=2000 --dropout=0.0
The result will be noisier but still entertaining. A smaller network (4 layers, 4 heads, 128 embedding size) can finish in a few minutes, producing a low-cost but instructive preview of how the model behaves.
Notes on sampling quality
Even with modest resources, you’ll observe the model producing very Shakespearean quirks, misplacements, and occasional near-poetic turns. This is not a final model for production, but it shows how simple architectures and fast iteration can reveal core properties of language modeling.
Practical tip
If you want to push bigger results without a mega GPU cluster, finetuning a pretrained GPT-2 checkpoint on Shakespeare (instead of training from scratch) often yields richer, more coherent outputs in a shorter time.
Code blocks are representative; you can adapt paths and config files to your environment. The goal is to get a feel for the workflow: data preparation, training, and sampling against a small, interpretable dataset.
Reproducing GPT-2: a more serious track
For practitioners aiming to reproduce GPT-2 results, the workflow involves tokenization, dataset preparation, and distributed training. The OpenWebText dataset provides a public analogue to the original WebText.
Tokenize and prepare OpenWebText
Prepare the data and convert it into a binary format suitable for fast loading:
- python data/openwebtext/prepare.py
This script downloads and tokenizes the OpenWebText dataset (a public re-creation of OpenAI’s WebText) and stores the resulting token IDs as train.bin and val.bin.
Training at GPT-2 scale
To reproduce GPT-2 124M performance, you’ll want a substantial GPU setup. The recommended approach uses PyTorch Distributed Data Parallel (DDP) across many GPUs:
- torchrun --standalone --nprocpernode=8 train.py config/train_gpt2.py
On an eight-GPU A100 40GB node, the training can take roughly four days, with a learned loss around 2.85 on the validation set after training.
The GPT-2 124M model, evaluated on OpenWebText, tends to show a slightly higher val loss due to the dataset domain gap relative to the original WebText.
Multinode scaling
If you have multiple GPU nodes, you can spread the workload across nodes. Example setup:
- On the master node:
- torchrun --nprocpernode=8 --nnodes=2 --noderank=0 --masteraddr=123.456.123.456 --master_port=1234 train.py
- On the worker node:
- torchrun --nprocpernode=8 --nnodes=2 --noderank=1 --masteraddr=123.456.123.456 --master_port=1234 train.py
Benchmarking the interconnect (e.g., with iperf3) helps you understand potential bottlenecks. If Infiniband isn’t available, you can set:
- NCCLIBDISABLE=1
That setting may slow things down but keeps the multi-node setup functional.
Sampling after training
You can sample from a trained model with:
- python sample.py
Or prompt from a trained directory:
- python sample.py --outdir=pathtotrainedmodel
Baselines and expectations
OpenAI GPT-2 baselines offer reference points for various model sizes. In this setup, you can observe train and val losses across different GPT-2 variants (124M, 350M, 774M, 1558M) and compare to OpenWebText baselines, recognizing the domain gaps between datasets.
This path is more demanding but yields generally stronger, more coherent text at scale. Finetuning can bridge domain gaps and improve performance when you don’t have the full compute budget to train from scratch.
Baselines: GPT-2 checkpoints and what you can expect
Baseline evaluations help anchor expectations when working with nanoGPT. The OpenAI GPT-2 checkpoints provide a spectrum of scale, and the nanoGPT workflow allows you to reproduce and compare these baselines with your own runs.
GPT-2 124M
Parameters: 124M
Train loss: ~3.11
Validation loss: ~3.12
GPT-2 Medium (350M)
Parameters: 350M
Train loss: ~2.85
Validation loss: ~2.84
GPT-2 Large (774M)
Parameters: 774M
Train loss: ~2.66
Validation loss: ~2.67
GPT-2 XL (1.558B)
Parameters: 1558M
Train loss: ~2.56
Validation loss: ~2.54
A key caveat: GPT-2 was originally trained on WebText, a closed dataset not publicly released. OpenWebText is an open reproduction, which can lead to a domain gap. Finetuning GPT-2 on OpenWebText tends to lower losses to around 2.85, aligning the open data results more closely with the published baselines under real-world conditions. This context matters for interpretation of the numbers and for planning your finetuning or replication strategy.
In practice, if you’re starting from a GPT-2 checkpoint and finetuning on a domain-specific corpus (like Shakespeare), you’re often better off starting from a pretrained checkpoint rather than from scratch. This tends to yield faster convergence and stronger text quality with less compute.
Finetuning: adapting a pretrained model to new text
Finetuning is conceptually the same as training but with a pretrained model as the starting point and a smaller learning rate. This is often the fastest route to useful results, especially when you have a target domain or style.
Example: fine-tuning Shakespeare
Prepare the tiny Shakespeare dataset as described earlier.
Run finetuning with a GPT-2 checkpoint initialized as the starting point:
- python train.py config/finetune_shakespeare.py
The finetuning config file overrides in config/finetuneshakespeare.py guide the training. You can load the GPT-2 checkpoint with initfrom and train as usual, but with shorter runs and a smaller learning rate.
Practical tips
If you’re memory-constrained, reduce model size (e.g., {'gpt2', 'gpt2-medium', 'gpt2-large', 'gpt2-xl'}) or shorten the context length (block_size).
The best checkpoint (lowest validation loss) sits in the out_dir corresponding to the chosen dataset, as configured in the finetune script.
After training, sample using:
- python sample.py --out_dir=out-shakespeare
Sample outputs during finetuning
You may see dialogues and monologues that resemble the target style, albeit sometimes with quirky or inconsistent phrasing. This is a natural artifact of small models and limited data, but it demonstrates the model’s capacity to imitate the source style with a relatively modest amount of compute.
Finetuning provides a practical route to quickly tailor language models to new domains or constraints, without the heavy cost of training from scratch.
Sampling and inference: getting text from trained models
Sampling is a focused, interactive way to evaluate and showcase what your model has learned. nanoGPT ships with a sampling utility that works with both pretrained GPT-2 models and models you trained yourself.
Basic sampling from a pretrained model
Example:
- python sample.py --initfrom=gpt2-xl --start="What is the answer to life, the universe, and everything?" --numsamples=5 --maxnewtokens=100
This command seeds the model with a starting prompt and generates multiple samples, up to 100 new tokens per sample.
Sampling from a trained model
If you trained a model and want to generate from it, point the sampler to the directory containing the checkpoints:
- python sample.py --out_dir=out-shakespeare-char
You can also seed with a file prompt:
- python sample.py --start=FILE:prompt.txt
Typical outputs
Generated text can resemble the input style, continuing lines in a way that often feels coherent and stylistically consistent with the training data. For Shakespeare-style training, the samples may echo iambic rhythm, archaic diction, and dramatic dialogue cues, albeit with the model’s own idiosyncrasies.
Practical tips
You can adjust maxnewtokens to control how long each sample runs, or tune sampling temperature and top-k/top-p settings to steer creativity and coherence.
Sampling is inexpensive compared to training, making it a quick way to iterate on prompts and outputs.
Visual reference
The project includes a visual representation of GPT-2 124M loss during training, offering a quick sense of the training progress and stability. (Asset: gpt2124Mloss.png)
Images from the Input are embedded here to illustrate the workflow and results:
- The hero image, nanoGPT banner:

- Training loss visualization for GPT-2 124M:

Efficiency notes: speed, compile, and profiling
nanoGPT includes a few practical notes for efficiency and benchmarking, especially when you’re iterating rapidly.
PyTorch 2.0 and torch.compile
The repository, by default, uses PyTorch 2.0 (which introduces torch.compile for performance). For some platforms (e.g., Windows) this can be experimental or unavailable.
If you encounter issues, you can disable compilation with:
- --compile=False
This trades a bit of speed for stability, but keeps the workflow accessible.
Benchmarking with bench.py
bench.py mirrors the core training loop but omits extraneous logging and data-handling overhead. It’s useful for quick performance profiling.
Practical guidance
The improvement from a single line of PyTorch code can be noticeable, reducing iteration times from roughly 250 ms per iteration to around 135 ms per iteration on some setups.
This kind of benchmarking helps you gauge the impact of hardware, software versions, and minor code changes on training speed.
Deployment considerations
If you’re planning multi-node training, ensure you benchmark network interconnects (e.g., via iperf3) to anticipate potential bottlenecks and to configure NCCL settings appropriately.
To-dos: ideas for future improvements
- Investigate and integrate FSDP (Fully Sharded Data Parallel) as an alternative to DDP for memory efficiency on large GPUs.
- Add robust zero-shot evaluation on standard benchmarks (e.g., LAMBADA, HELM) to quantify zero-shot perplexity and capabilities.
- Finetune the finetuning script with improved hyperparameters and examples to improve usability and results.
- Implement a policy to linearly increase batch size during training, to optimize resource usage and converge faster.
- Explore alternative embeddings (e.g., rotary embeddings, ALiBi) and their impact on performance and training stability.
- Create clearer separation of optimizer state and model parameters in checkpoint files for more flexible resuming and experimentation.
These ideas reflect ongoing exploration and are representative of the nanoGPT ethos: simple, hackable, and open to experimentation.
Troubleshooting: tips for a smoother run
PyTorch 2.0 and environment compatibility
The default setup uses PyTorch 2.0 with torch.compile. If you encounter errors, disable compilation with:
- --compile=False
Some platforms (like Windows) may not support certain features; in those cases, opting out of compile keeps the project usable.
Common questions
If you’re unsure about how to proceed, you can consult the Zero To Hero series for a broader context on GPTs and language modeling. The GPT video is a popular starting point.
For questions and community support, the nanoGPT Discord server is a valuable resource. The community link is here: https://discord.gg/3zy8kqD9Cp
A visual badge for the Discord community is included here as well:
Practical environments
If you’re using a single Mac or CPU-only setup, you can still run small experiments, but you’ll maximize speed by using MPS on Apple Silicon where available, and by keeping the model size modest.
Acknowledgements: thanks to the cloud and the community
All nanoGPT experiments are powered by GPUs on Lambda Labs, the project’s favored cloud GPU provider. The collaboration and infrastructure support from Lambda labs makes these rapid experiments possible, enabling researchers, students, and hobbyists to explore language modeling in a hands-on, approachable way.
This acknowledgment reflects the collaborative spirit of the project and the broader AI community, where cloud providers, open-source maintainers, and researchers share tools and ideas that accelerate learning and discovery.
Community, resources, and getting involved
- If you want to dive deeper, the Zero To Hero series offers helpful context about GPTs and language modeling. The GPT-focused material there is widely referenced and appreciated by newcomers and seasoned practitioners alike.
- The NanoGPT project sits in a lineage that includes minGPT and related efforts aimed at simplifying the training and experimentation pipeline for transformer-based language models.
- For ongoing updates, discussions, and collaborative experiments, joining the project’s Discord community is a great way to stay connected and get help from other builders.
Closing thoughts: a practical, hackable path to GPT-scale exploration
nanoGPT provides a clear, approachable path to understanding and working with language models at a medium scale. It emphasizes readability, transparency, and the ability to tinker—qualities that are invaluable for learning, experimentation, and rapid prototyping. From a Shakespearean character GPT trained in minutes on a single GPU to full GPT-2 reproductions on OpenWebText, the workflow is designed to be intuitive and adaptable.
Though updates in the ecosystem have introduced newer tooling (notably nanochat, the suggested successor), the core ideas endure: start simple, iterate quickly, and learn by doing. The combination of straightforward data preparation, a clean training loop, and flexible sampling provides a fertile ground for exploration, experimentation, and education.
If you’re just getting started, try the Shakespeare character GPT as your first project, observe the sampling output, and then gradually scale up or pivot toward finetuning or larger datasets. As you grow more comfortable, you can experiment with multi-node distributed training, domain adaptation through finetuning, and more sophisticated evaluation. And if you’re curious about the latest developments, consider checking out nanochat as the evolving successor to nanoGPT, while keeping this post as a practical, historical, and instructional reference.
Images included from the input:
- Hero image:

- GPT-2 loss visualization:

- Discord community badge: embedded in the Troubleshooting section
Notes:
- The content above preserves the spirit and specifics of the input while organizing them into a detailed, blog-style narrative with clear sections, practical steps, and accessible explanations.
- All sections avoid tables, as requested, and include relevant images sourced from the provided input.
Enjoying this project?
Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.
Repository:https://github.com/karpathy/nanoGPT
GitHub - karpathy/nanoGPT: nanoGPT: The simplest, fastest repository for training/finetuning medium-sized GPTs
A practical guide to training and finetuning medium-sized GPTs using the nanoGPT framework, emphasizing transparency, readability, and ease of experimentation....
github - karpathy/nanogpt


