nanoGPT: The simplest, fastest repository for training/finetuning medium-sized GPTs
GitHub Repo
MIT
June 30, 2026 at 08:22 AM
0 views

nanoGPT: The simplest, fastest repository for training/finetuning medium-sized GPTs

@karpathyProject Author

nanoGPT: A Practical Guide to Training and Finetuning Medium-Sized GPTs

nanoGPT

nanoGPT represents a compact, fast entry point for building and refining medium-sized language models. Born as a rewrite of minGPT, this project prioritizes practicality and accessibility: a straightforward training loop (train.py, roughly 300 lines) and a lean model definition (model.py, roughly 300 lines) that can even load GPT-2 weights from OpenAI. The goal is to give you something easy to hack, adapt, and experiment with—whether you want to train from scratch or finetune a pretrained checkpoint.

The project emphasizes transparency and readability. With a minimal, well-documented code path, you can quickly understand how data flows through the system, how the transformer layers are configured, and where to hook in your own ideas. On a single high-end GPU, the base setup can reproduce GPT-2 (124M) behavior on OpenWebText in about four days; more modest hardware can still yield useful results by dialing down the size and length. The included images show training dynamics and results, illustrating how a tiny, character-based model can generate surprisingly coherent Shakespeare-like text in a few minutes or hours, depending on resources.


Update Nov 2025: A new direction and a note on deprecation

There is an important update in the ecosystem. nanoGPT has a new and improved cousin called nanochat. It is very likely you meant to use/find nanochat instead. nanoGPT (this repo) is now very old and deprecated but I will leave it up for posterity.

  • nanochat expands the same spirit of simplicity but with new features and a refreshed design that aligns with current tooling and best practices.
  • If you’re starting a new project, consider using nanochat as your primary entry point.
  • This post preserves the essence of nanoGPT’s approach and serves as historical context and a learning resource for those curious about the lineage of open-source GPT tooling.

This update reflects how quickly the field evolves and how lightweight tooling can evolve into more capable, maintainable ecosystems. The core ideas—straightforward data processing, transparent training loops, and flexible finetuning—still resonate, whether you start from nanoGPT or migrate to nanochat.


Quick start: your first steps with a Shakespeare character GPT

If you’re new to the domain, the fastest way to feel the magic is to train a character-level GPT on Shakespeare’s works. The steps below sketch the process, with representative commands you can copy-paste.

  • Dependencies and setup

  • Install the essential libraries:

    • pip install torch numpy transformers datasets tiktoken wandb tqdm
  • Core dependencies to keep in mind:

    • PyTorch (< 3)
    • numpy
    • transformers (for loading GPT-2 checkpoints)
    • datasets (for OpenWebText and similar corpora)
    • tiktoken (OpenAI’s fast BPE tokenization)
    • wandb (optional logging)
    • tqdm (progress bars)
  • Prepare Shakespeare data (character-level)

  • The pipeline first downloads Shakespeare as a single text file and converts it into a stream of integers:

    • python data/shakespeare_char/prepare.py
  • This produces train.bin and val.bin in the data directory.

  • Train a tiny Shakespeare model on a GPU

  • With a local GPU, you can train a small GPT with the provided config. For example:

    • python train.py config/trainshakespearechar.py
  • This config sets a context window up to 256 characters, 384 feature channels, 6 Transformer layers, and 6 attention heads per layer.

  • On an A100, this run typically finishes quickly and yields a validation loss around 1.47.

  • Sampling from the best Shakespeare model

  • After training, sample from the best model with:

    • python sample.py --out_dir=out-shakespeare-char
  • Sample outputs demonstrate the model’s ability to generate imaginative and sometimes witty text in the Shakespearean space. For example, you might see lines that resemble classical dialogue, with a modern, playful twist.

  • If you’re on a Mac (Apple Silicon) or a recent machine

  • You can leverage on-chip GPUs with Metal Performance Shaders (MPS). Use:

    • python train.py config/trainshakespearechar.py --device=mps
  • PyTorch on Apple Silicon often provides substantial speedups, enabling larger experiments on consumer hardware.

  • CPU-only or lighter hardware

  • If you don’t have a GPU, you can still run a smaller setup (though slower). A representative CPU-based run uses:

    • python train.py config/trainshakespearechar.py --device=cpu --compile=False --evaliters=20 --loginterval=1 --blocksize=64 --batchsize=12 --nlayer=4 --nhead=4 --nembd=128 --maxiters=2000 --lrdecayiters=2000 --dropout=0.0
  • The result will be noisier but still entertaining. A smaller network (4 layers, 4 heads, 128 embedding size) can finish in a few minutes, producing a low-cost but instructive preview of how the model behaves.

  • Notes on sampling quality

  • Even with modest resources, you’ll observe the model producing very Shakespearean quirks, misplacements, and occasional near-poetic turns. This is not a final model for production, but it shows how simple architectures and fast iteration can reveal core properties of language modeling.

  • Practical tip

  • If you want to push bigger results without a mega GPU cluster, finetuning a pretrained GPT-2 checkpoint on Shakespeare (instead of training from scratch) often yields richer, more coherent outputs in a shorter time.

Code blocks are representative; you can adapt paths and config files to your environment. The goal is to get a feel for the workflow: data preparation, training, and sampling against a small, interpretable dataset.


Reproducing GPT-2: a more serious track

For practitioners aiming to reproduce GPT-2 results, the workflow involves tokenization, dataset preparation, and distributed training. The OpenWebText dataset provides a public analogue to the original WebText.

  • Tokenize and prepare OpenWebText

  • Prepare the data and convert it into a binary format suitable for fast loading:

    • python data/openwebtext/prepare.py
  • This script downloads and tokenizes the OpenWebText dataset (a public re-creation of OpenAI’s WebText) and stores the resulting token IDs as train.bin and val.bin.

  • Training at GPT-2 scale

  • To reproduce GPT-2 124M performance, you’ll want a substantial GPU setup. The recommended approach uses PyTorch Distributed Data Parallel (DDP) across many GPUs:

    • torchrun --standalone --nprocpernode=8 train.py config/train_gpt2.py
  • On an eight-GPU A100 40GB node, the training can take roughly four days, with a learned loss around 2.85 on the validation set after training.

  • The GPT-2 124M model, evaluated on OpenWebText, tends to show a slightly higher val loss due to the dataset domain gap relative to the original WebText.

  • Multinode scaling

  • If you have multiple GPU nodes, you can spread the workload across nodes. Example setup:

    • On the master node:
    • torchrun --nprocpernode=8 --nnodes=2 --noderank=0 --masteraddr=123.456.123.456 --master_port=1234 train.py
    • On the worker node:
    • torchrun --nprocpernode=8 --nnodes=2 --noderank=1 --masteraddr=123.456.123.456 --master_port=1234 train.py
  • Benchmarking the interconnect (e.g., with iperf3) helps you understand potential bottlenecks. If Infiniband isn’t available, you can set:

    • NCCLIBDISABLE=1
  • That setting may slow things down but keeps the multi-node setup functional.

  • Sampling after training

  • You can sample from a trained model with:

    • python sample.py
  • Or prompt from a trained directory:

    • python sample.py --outdir=pathtotrainedmodel
  • Baselines and expectations

  • OpenAI GPT-2 baselines offer reference points for various model sizes. In this setup, you can observe train and val losses across different GPT-2 variants (124M, 350M, 774M, 1558M) and compare to OpenWebText baselines, recognizing the domain gaps between datasets.

This path is more demanding but yields generally stronger, more coherent text at scale. Finetuning can bridge domain gaps and improve performance when you don’t have the full compute budget to train from scratch.


Baselines: GPT-2 checkpoints and what you can expect

Baseline evaluations help anchor expectations when working with nanoGPT. The OpenAI GPT-2 checkpoints provide a spectrum of scale, and the nanoGPT workflow allows you to reproduce and compare these baselines with your own runs.

  • GPT-2 124M

  • Parameters: 124M

  • Train loss: ~3.11

  • Validation loss: ~3.12

  • GPT-2 Medium (350M)

  • Parameters: 350M

  • Train loss: ~2.85

  • Validation loss: ~2.84

  • GPT-2 Large (774M)

  • Parameters: 774M

  • Train loss: ~2.66

  • Validation loss: ~2.67

  • GPT-2 XL (1.558B)

  • Parameters: 1558M

  • Train loss: ~2.56

  • Validation loss: ~2.54

A key caveat: GPT-2 was originally trained on WebText, a closed dataset not publicly released. OpenWebText is an open reproduction, which can lead to a domain gap. Finetuning GPT-2 on OpenWebText tends to lower losses to around 2.85, aligning the open data results more closely with the published baselines under real-world conditions. This context matters for interpretation of the numbers and for planning your finetuning or replication strategy.

In practice, if you’re starting from a GPT-2 checkpoint and finetuning on a domain-specific corpus (like Shakespeare), you’re often better off starting from a pretrained checkpoint rather than from scratch. This tends to yield faster convergence and stronger text quality with less compute.


Finetuning: adapting a pretrained model to new text

Finetuning is conceptually the same as training but with a pretrained model as the starting point and a smaller learning rate. This is often the fastest route to useful results, especially when you have a target domain or style.

  • Example: fine-tuning Shakespeare

  • Prepare the tiny Shakespeare dataset as described earlier.

  • Run finetuning with a GPT-2 checkpoint initialized as the starting point:

    • python train.py config/finetune_shakespeare.py
  • The finetuning config file overrides in config/finetuneshakespeare.py guide the training. You can load the GPT-2 checkpoint with initfrom and train as usual, but with shorter runs and a smaller learning rate.

  • Practical tips

  • If you’re memory-constrained, reduce model size (e.g., {'gpt2', 'gpt2-medium', 'gpt2-large', 'gpt2-xl'}) or shorten the context length (block_size).

  • The best checkpoint (lowest validation loss) sits in the out_dir corresponding to the chosen dataset, as configured in the finetune script.

  • After training, sample using:

    • python sample.py --out_dir=out-shakespeare
  • Sample outputs during finetuning

  • You may see dialogues and monologues that resemble the target style, albeit sometimes with quirky or inconsistent phrasing. This is a natural artifact of small models and limited data, but it demonstrates the model’s capacity to imitate the source style with a relatively modest amount of compute.

Finetuning provides a practical route to quickly tailor language models to new domains or constraints, without the heavy cost of training from scratch.


Sampling and inference: getting text from trained models

Sampling is a focused, interactive way to evaluate and showcase what your model has learned. nanoGPT ships with a sampling utility that works with both pretrained GPT-2 models and models you trained yourself.

  • Basic sampling from a pretrained model

  • Example:

    • python sample.py --initfrom=gpt2-xl --start="What is the answer to life, the universe, and everything?" --numsamples=5 --maxnewtokens=100
  • This command seeds the model with a starting prompt and generates multiple samples, up to 100 new tokens per sample.

  • Sampling from a trained model

  • If you trained a model and want to generate from it, point the sampler to the directory containing the checkpoints:

    • python sample.py --out_dir=out-shakespeare-char
  • You can also seed with a file prompt:

    • python sample.py --start=FILE:prompt.txt
  • Typical outputs

  • Generated text can resemble the input style, continuing lines in a way that often feels coherent and stylistically consistent with the training data. For Shakespeare-style training, the samples may echo iambic rhythm, archaic diction, and dramatic dialogue cues, albeit with the model’s own idiosyncrasies.

  • Practical tips

  • You can adjust maxnewtokens to control how long each sample runs, or tune sampling temperature and top-k/top-p settings to steer creativity and coherence.

  • Sampling is inexpensive compared to training, making it a quick way to iterate on prompts and outputs.

  • Visual reference

  • The project includes a visual representation of GPT-2 124M loss during training, offering a quick sense of the training progress and stability. (Asset: gpt2124Mloss.png)

Images from the Input are embedded here to illustrate the workflow and results:

  • The hero image, nanoGPT banner: nanoGPT
  • Training loss visualization for GPT-2 124M: repro124m

Efficiency notes: speed, compile, and profiling

nanoGPT includes a few practical notes for efficiency and benchmarking, especially when you’re iterating rapidly.

  • PyTorch 2.0 and torch.compile

  • The repository, by default, uses PyTorch 2.0 (which introduces torch.compile for performance). For some platforms (e.g., Windows) this can be experimental or unavailable.

  • If you encounter issues, you can disable compilation with:

    • --compile=False
  • This trades a bit of speed for stability, but keeps the workflow accessible.

  • Benchmarking with bench.py

  • bench.py mirrors the core training loop but omits extraneous logging and data-handling overhead. It’s useful for quick performance profiling.

  • Practical guidance

  • The improvement from a single line of PyTorch code can be noticeable, reducing iteration times from roughly 250 ms per iteration to around 135 ms per iteration on some setups.

  • This kind of benchmarking helps you gauge the impact of hardware, software versions, and minor code changes on training speed.

  • Deployment considerations

  • If you’re planning multi-node training, ensure you benchmark network interconnects (e.g., via iperf3) to anticipate potential bottlenecks and to configure NCCL settings appropriately.


To-dos: ideas for future improvements

  • Investigate and integrate FSDP (Fully Sharded Data Parallel) as an alternative to DDP for memory efficiency on large GPUs.
  • Add robust zero-shot evaluation on standard benchmarks (e.g., LAMBADA, HELM) to quantify zero-shot perplexity and capabilities.
  • Finetune the finetuning script with improved hyperparameters and examples to improve usability and results.
  • Implement a policy to linearly increase batch size during training, to optimize resource usage and converge faster.
  • Explore alternative embeddings (e.g., rotary embeddings, ALiBi) and their impact on performance and training stability.
  • Create clearer separation of optimizer state and model parameters in checkpoint files for more flexible resuming and experimentation.

These ideas reflect ongoing exploration and are representative of the nanoGPT ethos: simple, hackable, and open to experimentation.


Troubleshooting: tips for a smoother run

  • PyTorch 2.0 and environment compatibility

  • The default setup uses PyTorch 2.0 with torch.compile. If you encounter errors, disable compilation with:

    • --compile=False
  • Some platforms (like Windows) may not support certain features; in those cases, opting out of compile keeps the project usable.

  • Common questions

  • If you’re unsure about how to proceed, you can consult the Zero To Hero series for a broader context on GPTs and language modeling. The GPT video is a popular starting point.

  • For questions and community support, the nanoGPT Discord server is a valuable resource. The community link is here: https://discord.gg/3zy8kqD9Cp

  • A visual badge for the Discord community is included here as well: Discord badge

  • Practical environments

  • If you’re using a single Mac or CPU-only setup, you can still run small experiments, but you’ll maximize speed by using MPS on Apple Silicon where available, and by keeping the model size modest.


Acknowledgements: thanks to the cloud and the community

All nanoGPT experiments are powered by GPUs on Lambda Labs, the project’s favored cloud GPU provider. The collaboration and infrastructure support from Lambda labs makes these rapid experiments possible, enabling researchers, students, and hobbyists to explore language modeling in a hands-on, approachable way.

This acknowledgment reflects the collaborative spirit of the project and the broader AI community, where cloud providers, open-source maintainers, and researchers share tools and ideas that accelerate learning and discovery.


Community, resources, and getting involved

  • If you want to dive deeper, the Zero To Hero series offers helpful context about GPTs and language modeling. The GPT-focused material there is widely referenced and appreciated by newcomers and seasoned practitioners alike.
  • The NanoGPT project sits in a lineage that includes minGPT and related efforts aimed at simplifying the training and experimentation pipeline for transformer-based language models.
  • For ongoing updates, discussions, and collaborative experiments, joining the project’s Discord community is a great way to stay connected and get help from other builders.

Closing thoughts: a practical, hackable path to GPT-scale exploration

nanoGPT provides a clear, approachable path to understanding and working with language models at a medium scale. It emphasizes readability, transparency, and the ability to tinker—qualities that are invaluable for learning, experimentation, and rapid prototyping. From a Shakespearean character GPT trained in minutes on a single GPU to full GPT-2 reproductions on OpenWebText, the workflow is designed to be intuitive and adaptable.

Though updates in the ecosystem have introduced newer tooling (notably nanochat, the suggested successor), the core ideas endure: start simple, iterate quickly, and learn by doing. The combination of straightforward data preparation, a clean training loop, and flexible sampling provides a fertile ground for exploration, experimentation, and education.

If you’re just getting started, try the Shakespeare character GPT as your first project, observe the sampling output, and then gradually scale up or pivot toward finetuning or larger datasets. As you grow more comfortable, you can experiment with multi-node distributed training, domain adaptation through finetuning, and more sophisticated evaluation. And if you’re curious about the latest developments, consider checking out nanochat as the evolving successor to nanoGPT, while keeping this post as a practical, historical, and instructional reference.

Images included from the input:

  • Hero image: nanoGPT
  • GPT-2 loss visualization: repro124m
  • Discord community badge: embedded in the Troubleshooting section

Notes:

  • The content above preserves the spirit and specifics of the input while organizing them into a detailed, blog-style narrative with clear sections, practical steps, and accessible explanations.
  • All sections avoid tables, as requested, and include relevant images sourced from the provided input.

Enjoying this project?

Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.

Project
nanogpt-practical-guide
Created
June 30
Last Updated
June 30, 2026 at 08:22 AM

Find more projects like this

One email a week: new and trending developer tools, fresh comparisons, and what shipped. Unsubscribe in one click.