LLaMA Factory

LLaMA Factory: A Detailed Guide to Fine-Tuning 100+ Language Models with Zero-Code Accessibility
Introduction
LLaMA Factory is a comprehensive toolkit designed to democratize the fine-tuning and deployment of large language models (LLMs). It brings together a rich ecosystem of models, training methodologies, and deployment options under one roof, offering both a command-line interface (CLI) and a web-based GUI to satisfy developers at every level. From pre-training and supervised fine-tuning to reward modeling and advanced optimization tricks, LLaMA Factory provides a scalable, flexible framework for researchers and practitioners who want to push the boundaries of language modeling without getting lost in the plumbing.
The project is widely adopted across major tech players and research labs, with a growing community behind it. It emphasizes zero-code experiences for common workflows, but remains highly configurable for power users who want to tailor training templates, data pipelines, and inference backends to their specific needs. The collaborative nature of the project is reflected in its extensive list of supported models, training approaches, and integrations with popular tools for experiment tracking, visualization, and deployment.
Images that accompany this journey:
- LLaMA Factory logo serves as a visual anchor for the post.
- Warped sponsorships and community badges show the ecosystem’s support network and collaborations.
- Quick-access badges point to Colab and other platforms that help accelerate hands-on experimentation.
Highlights: What LLaMA Factory Brings to Your Workspace
- Broad model coverage: LLaMA Factory supports a wide spectrum of large and multimodal models. Expect compatibility with LLaMA, LLaVA, Mistral, Mixtral-MoE, Qwen series (Qwen3, Qwen3-VL, Qwen2, and more), DeepSeek, Gemma, GLM, Phi, Llama series, and beyond. This breadth means you can experiment with diverse architectures and capabilities without migrating between disparate tools.
- Integrated methods for end-to-end training: The framework unifies continuous pre-training, multimodal supervised fine-tuning, and reward-based methods such as PPO, DPO, KTO, ORPO, and more. This enables you to explore advanced training paradigms in a single, cohesive environment.
- Scalable resource strategies: Full-tuning in 16-bit precision, freeze-tuning, and a hierarchy of parameter-efficient techniques—including LoRA and quantization variants (QLoRA and beyond) across 2/3/4/5/6/8-bit spectra—are supported. This makes it practical to train and fine-tune large models on a range of hardware budgets.
- Advanced optimization and accelerators: GaLore, BAdam, APOLLO, Adam-mini, Muon, OFT, DoRA, LongLoRA, LLaMA Pro, Mixture-of-Depths, LoRA+, LoftQ, and PiSSA are among the cutting-edge methods integrated into the workflow. Practical tricks like FlashAttention-2, Liger Kernel, and dynamic techniques for context lengths help maximize efficiency and performance.
- Broad task support: From multi-turn dialogue to tool usage, image understanding, visual grounding, video and audio understanding, the platform is designed to accommodate a wide range of real-world tasks. This makes it easier to build end-to-end systems that combine language with perception and control.
- Rich experiment-tracking and monitoring: LlamaBoard, TensorBoard, Wandb, MLflow, SwanLab, and similar tools are compatible for vigilant experiment monitoring, enabling you to compare runs, visualize metrics, and share results with teammates.
- Inference acceleration and accessible interfaces: The project emphasizes faster deployment with an OpenAI-style API, Gradio-based UI, and CLI workflows that leverage vLLM or SGLang workers for high-concurrency inference. This makes it practical to deploy LLM-powered apps with a familiar developer experience.
- Day-N readiness for the latest models: The project maintains fast iteration cycles, with explicit support timelines for new models (e.g., day-0 to day-1 coverage for Qwen3, Llama 3/4, GLM, Mistral Small, and more). This ensures you can stay current with evolving architectures.
- Open workflow for data and templates: A broad set of datasets is provided for pre-training, supervised fine-tuning, and preference modeling, plus templates for chatting, coding, and reasoning. You can also add custom chat templates, enabling rapid experimentation with new data formats and alignment targets.
Needed images in this section:
- Warp sponsorship image to emphasize collaboration with developer tooling partners.
- SerpAPI sponsorship badge to highlight external tooling support.
Supported Models: A World of Choices
LLaMA Factory maintains a vast catalog of supported models, spanning classic large-language models to modern multimodal ensembles. Examples include:
- BLOOM and BLOOMZ across multiple scales
- DeepSeek family (LLM/Code/MoE) with multiple size variants
- Falcon/H1 and its variants
- Gemma family, including Gemma 2 and Gemma 3n
- GLM-4 and GLM-Z1
- GPT-OSS family and GPT-2 lineage
- InternLM, InternVL, InternVL2.5–3.5
- Kimi-VL and Ling 2.0 variants
- Llama family (Llama, Llama 2, Llama 3, Llama 4 and related vision-capable variants)
- LLaVA and LLaVA-NeXT (multimodal LLMs)
- Qwen series (Qwen2, Qwen3, Qwen3-VL, Qwen3-Omni, Qwen2.5 Omni, Qwen 2-Audio)
- Qwen Mix and MoE variants
- StarCoder 2 and other multi-task code-oriented models
- MiMo, Pixtral, Pixtral family
- Mistral and Mixtral variants
- PaliGemma and PaliGemma2
- Phi family, including Phi-3 and Phi-4
- QLoRA-enabled templates for efficient fine-tuning
- LFM, InternLM, LLaVA-multimodal, and many more
If the model has base, instruct, or chat variants, remember to choose the corresponding template to ensure training and inference compatibility. The developer-friendly guidelines emphasize consistency of the same template across training and inference.
Supported Training Approaches: From Foundation to Fine-Tuning
- Pre-Training and Full-Tuning: End-to-end training to build robust capabilities or adapt from scratch.
- Supervised Fine-Tuning (SFT): Tailor models to follow specific instructions and align with domain-specific tasks.
- Reward Modeling and RL-based Fine-Tuning: PPO, DPO, KTO, ORPO, SimPO, and other reinforcement-learning-based strategies for alignment and preference shaping.
- Parameter-Efficient Tuning: LoRA, QLoRA, and related quantization-aware techniques to maximize efficiency on limited hardware.
- Hybrid and Advanced Methods: GaLore, BAdam, APOLLO, Adam-mini, Muon, OFT, QOFT, DoRA, and LongLoRA to explore new optimization frontiers.
- Inference-Aware Optimizations: RoPE scaling, FlashAttention-2 acceleration, vLLM and SGLang backends for high-throughput inference, and dynamic context management.
- Data-Centric Techniques: NEFTune-inspired data strategies, data streaming, and synthetic data generation to accelerate model adaptation.
- Multi-Modal and Tool-Usage Scenarios: Training tasks that integrate image understanding, tool usage, and multi-turn dialogues to build richer assistants.
Day-N Support for Cutting-Edge Models
- Day 0 through Day 1 coverage notes: The platform tracks model releases like Qwen3, Llama 3/4, GLM-4, Mistral Small, and Gemma 3 across rapid release cycles, guiding you on templates and training settings for the newest models.
Blogs and Learning: A Dedicated Space for Insights
- LLaMA Factory maintains a growing blog ecosystem with a dedicated site: blog.llamafactory.net
- Blog posts cover practical tuning recipes, domain adaptation strategies, data preparation pipelines, and architecture-specific notes. Topics include:
- Fine-tuning 1000-billion-scale models with two high-end GPUs
- Efficient domain knowledge acquisition through curated datasets
- DataFlow-based data preparation for quality improvements
- Data-centric approaches using DataFlex and similar data-centric pipelines
- Data-centric training pipelines for enhanced data control
- Some posts are available in multiple languages, including English and Chinese, facilitating global accessibility.
- There are also practical guides for Colab-based fine-tuning and Colab notebooks that demonstrate end-to-end workflows.
Changelog: A Glimpse into Ongoing Improvements
- The project publishes frequent updates, including:
- Megatron-core training backends with mcore_adapter
- OFT and OFTv2 updates
- Intern-S1-mini model fine-tuning
- GPT-OSS and Qwen model-family support
- MoE model support such as Mixtral 8x7B
- Unsloth’s long-sequence improvements and patching
- vLLM-based inference acceleration
- Labeled notes on dataset streaming, oracle-based training, and patch-level enhancements
- The changelog reveals the team’s focus on expanding model compatibility, improving training speed, and introducing novel optimization strategies.
Provided and Supported Datasets: Building Blocks for Training
- Pre-training datasets include Wiki Demo, RefinedWeb, RedPajama V2, Wikipedia, The Stack, Pile, and large zh-language corpora such as SkyPile.
- Fine-tuning datasets include a broad suite of identity data, Alpaca-like datasets, and domain-specific corpora across multiple languages.
- Preference datasets span DPO mixtures, RLHF-driven data profiles, UltraFeedback, and specialized RLHF collections for nuanced alignment.
- The data ecosystem emphasizes flexible dataset access via Hugging Face, ModelScope, and Modelers Hub, enabling you to mix local datasets with cloud-hosted resources.
- A note on usage: Some datasets require user confirmation or authentication. Logging in to Hugging Face may be necessary for certain datasets, with commands provided to facilitate access.
Requirement and Hardware: What You Need to Know
- A detailed matrix outlines minimum and recommended hardware settings across model sizes and precision modes. There are clear guidelines for 7B, 14B, 30B, and up to 70B–109B-scale models, with estimates for memory footprints under bf16/pure bf16 and various training configurations (Full, Freeze/LoRA, QLoRA, etc.).
- Hardware considerations span GPUs (memory, bandwidth), storage, and throughput requirements, with notes on x86_64 support, CUDA versions, and accelerator environments.
- Optional dependencies include DeepSpeed and bitsandbytes for quantized training, plus vLLM or SGLang for efficient inference.
Getting Started: Step-by-Step Installation
- Install from Source:
- Clone the repository and install the package in editable mode.
- Install required metrics dependencies to enable evaluation features.
- Optional: install deepspeed and metrics for extended capabilities.
- Install from Docker:
- Use the provided Docker image, which is built on Ubuntu 22.04, CUDA 12.4, Python 3.11, PyTorch 2.6.0, Flash-attn 2.7.4.
- Pre-built images are available on Docker Hub for quick setup.
- Windows and GPU considerations:
- For Windows, you need to install PyTorch with CUDA support manually, following official guidelines.
- BitsAndBytes for Windows enables quantized LoRA on supported CUDA versions.
- FlashAttention-2 can be installed via a Windows wheel from a dedicated repository.
- Ascend NPU users:
- Upgrade Python to 3.10+ and follow the Ascend NPU-specific installation steps, including kernel and toolkit requirements.
- Docker images tailored for Ascend NPUs are provided.
- Data and environment setup:
- A data preparation guide is provided to ensure datasets follow the required formats.
- Easy Dataset, DataFlow, GraphGen, and synthetic data generation templates help you bootstrap data for finetuning.
Data Preparation: How to Ready Your Data
- The data preparation guide describes dataset formats and how to structure training data for LLaMA Factory.
- You can combine local datasets with remote datasets from Hugging Face, ModelScope, and Modelers Hub, or point the system to cloud storage locations.
- You can customize dataset_info.json to reflect your own data, and you can leverage synthetic data pipelines to augment scarce-domain data.
Quickstart: Three Core Commands to Begin LoRA Fine-Tuning
- Train: llamafactory-cli train [your LoRA SFT YAML]
- Chat / Inference: llamafactory-cli chat [your SFT YAML]
- Export / Merge LoRA: llamafactory-cli export [merge YAML]
- These three commands let you begin, test, and merge LoRA-based adapters for quick iteration.
- A tip: Use llamafactory-cli help to see available options; consult the FAQs if you hit roadblocks.
Fine-Tuning with LLaMA Board GUI (Powered by Gradio)
- Web-based fine-tuning and evaluation provide a user-friendly alternative to the CLI.
- The web UI enables interactive training and evaluation in your browser, lowering the barrier to entry for practitioners who prefer visual controls.
- Command to launch the GUI: llamafactory-cli webui
Build Docker: Environment Setup for Speed
- CUDA workflows:
- docker-compose up in the CUDA-specific directory starts a GPU-enabled service; you can attach to the container to monitor training.
- Ascend NPU workflows:
- Similar docker-compose approach with NPU-specific configurations.
- AMD ROCm workflows:
- A ROCm-based setup is provided to leverage AMD hardware for training and inference.
- Build without Docker Compose:
- You can build standalone Docker images for CUDA, Ascend, and ROCm, then run containers with appropriate GPU/device mappings.
- Data volumes:
- The Dockerfile supports data volumes for Hugging Face caches, shared data directories, and export outputs, enabling convenient data persistence across runs.
Deploy with OpenAI-Style API and vLLM
- A common path to production is exposing a REST-like OpenAI-compatible API while leveraging vLLM for high-throughput inference.
- Example:
- APIPORT=8000 llamafactory-cli api examples/inference/qwen3.yaml inferbackend=vllm vllmenforceeager=true
- This setup provides a familiar API surface for integrating LLaMA Factory-backed models into existing apps that rely on OpenAI-compatible endpoints.
Downloading from ModelScope Hub and Modelers Hub
- If Hugging Face downloads become difficult, you can switch to ModelScope or Modelers Hub for model weights and datasets.
- ModelScope Hub example: export USEMODELSCOPEHUB=1
- Modelers Hub example: export USEOPENMINDHUB=1
- You can then reference model IDs such as LLM-Research/Meta-Llama-3-8B-Instruct or TeleAI/TeleChat-7B-pt.
W&B Logger and SwanLab Logger
- To log experiments with Weights & Biases (W&B), add:
- report_to: wandb
- runname: yourrun_name
- For SwanLab, enable:
- use_swanlab: true
- swanlabrunname: optionalrunname
- Authentication keys must be supplied via environment variables or YAML configuration to access these logging services.
Projects Using LLaMA Factory: A Glimpse into Real-World Applications
- The ecosystem hosts a roster of projects that leverage LLaMA Factory for fine-tuning, deployment, or research exploration. Examples include efficient sampling-based reinforcement learning for sequence generation, large-scale dataset curation and alignment, and multi-modal model integration.
- The community-driven nature of these projects fosters cross-pollination of ideas and accelerates the development of robust LLM pipelines.
License and Citation
- The repository is licensed under the Apache-2.0 License.
- Responsible use of model weights requires attention to the licenses of individual models (e.g., BLOOM, DeepSeek, Falcon, Gemma, GLM, Llama, Llama 2/3/4, etc.).
- If this work proves helpful, it’s encouraged to cite the LlamaFactory paper:
- Zheng, Y., et al. LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models. ACL 2024.
- Acknowledgments highlight inspirations from PEFT, TRL, QLoRA, and FastChat, among others, and gratitude to contributors who advanced this project.
Images and Visuals: A Quick Gallery of the Ecosystem
- Warp sponsorship banner shows industry support for developer tooling and productivity enhancements.
- Warp, the agentic terminal for developers
- SerpAPI sponsorship banner highlights search tooling integration for data collection and automation.
- WeChat and NPU groups:
- WeChat group image for English and Chinese-speaking communities
- NPU-specific group image for clear collaboration among AI developers working on NPUs
- Colab and Spaces:
- Colab badge, enabling quick Colab-based experiments
- Star History:
- A dynamic star-history chart to visualize community enthusiasm and adoption over time
A Community-Driven Path Forward
- The LLaMA Factory ecosystem continues to grow through active community engagement, contributions, and open collaborations with industry partners and research teams.
- The project encourages new contributors to propose features, add templates, or extend support to additional models and backends.
- The documentation emphasizes careful use of external resources and clear licensing for model weights to ensure compliance and responsible usage.
Conclusion: Making Fine-Tuning Accessible, Scalable, and Open
LLaMA Factory stands out as a holistic platform that lowers the barrier to entry for fine-tuning and deploying large language models. Its emphasis on zero-code experiences, combined with robust, scalable tooling for both research and production, makes it an attractive option for teams of all sizes. Whether you’re exploring domain-specific GPT-like assistants, building multimodal systems, or iterating on sophisticated alignment strategies, LLaMA Factory offers a coherent, feature-rich workflow that unifies modeling, data, and deployment.
If you’re ready to embark on a journey toward more capable, adaptable LLMs, LLaMA Factory provides a structured, community-supported path. Install from source or pull a Docker image, prepare your data, and begin your first LoRA fine-tuning run. Try the GUI for a hands-on approach, or jump straight into API deployment for real-world applications. The ecosystem’s breadth means you’re unlikely to outgrow it quickly; instead, you’ll continuously discover new models, training paradigms, and deployment strategies to keep your AI initiatives at the leading edge.
Images to explore as you read and experiment:
- LLaMA Factory Logo
- Warp Sponsorship Banner
- SerpAPI Sponsorship Banner
- WeChat and NPU Group Images
- Colab Badge
- Star History Chart
References and Quick Links
- Official Blog: blog.llamafactory.net
- Documentation: llamafactory.readthedocs.io
- Model and template references: constants.py and template.py in the LLaMA Factory repository
- Model Scope Hub and Modelers Hub entries for alternative model/download options
- Open Colab notebook and Gradio-based GUI showcase
Star History
- For a sense of the community momentum, explore the star history chart at: https://api.star-history.com/svg?repos=hiyouga/LLaMA-Factory&type=Date
Acknowledgments
- The project acknowledges contributions from PEFT, TRL, QLoRA, and FastChat communities, whose ideas and implementations helped shape LLaMA Factory’s capabilities.
- Thanks to the broader AI community for ongoing feedback, bug reports, and feature requests that drive the project forward.
Footnotes
- Remember to review licenses for each model you intend to use and to respect the terms of use for model weights and datasets.
- When in doubt about latest features, pull the latest code and reinstall to ensure compatibility with new backends and templates.
This guide aims to give you a detailed, practical sense of what LLaMA Factory offers and how to get started quickly. May your experiments be productive, your models be robust, and your deployments reliable. Happy fine-tuning!
Enjoying this project?
Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.
Repository:https://github.com/hiyouga/LLaMA-Factory
GitHub - hiyouga/LLaMA-Factory: LLaMA Factory
LLaMA Factory is a comprehensive toolkit designed to democratize the fine-tuning and deployment of large language models (LLMs). It brings together a rich ecosy...
github - hiyouga/llama-factory


