FlagEmbedding — Retrieval and Retrieval-augmented LLMs

BGE: One-Stop Retrieval Toolkit For Search and RAG
A comprehensive look at a one-stop retrieval toolkit designed to empower search and retrieval-augmented generation (RAG) workflows. Born from FlagOpen’s FlagEmbedding ecosystem, BGE (BAAI General Embedding) brings together embedding models, cross-encoder rerankers, and a curated set of tools to streamline building powerful, multilingual, and multi-functional retrieval systems. This post offers a detailed tour of what BGE is, what has recently changed, how to get started, and where to look for deeper knowledge. To accompany the narrative, you’ll find visual references from the Input images that illustrate the project’s branding, community, and roadmap.

Introduction: What BGE Seeks to Solve
BGE is a one-stop retrieval toolkit crafted to support search and RAG applications across languages, modalities, and task requirements. It sits at the intersection of dense embeddings, cross-encoder reranking, and retrieval strategies that span multiple granularities and vector spaces. The core idea is simple: encode queries and documents into meaningful representations, rank results efficiently, and integrate these capabilities into LLM-powered pipelines for enhanced accuracy and relevance.
Key pillars of BGE include:
- Dense embedding models that map textual content into vector spaces where semantic proximity equates to relevance.
- Reranking mechanisms built on cross-encoder architectures, providing higher-precision reordering of candidate documents.
- Multilingual focus with support for languages beyond English, enabling cross-lingual retrieval and multilingual benchmarks.
- A unified workflow that accommodates different retrieval methods (dense, sparse, multi-vector) to adapt to varied data landscapes.
- An active, growing community with tutorials, documentation, benchmarks, and example pipelines.
Inspiration and branding come through a collection of visuals, including the FlagOpen branding and the BGE logo, which you’ll see reflected in the imagery that accompanies this narrative. The images below capture some of the visual context around BGE’s ecosystem and community engagement.
Community, News, and Milestones: A Timeline of Growth
News updates illuminate a vibrant and evolving project with frequent releases, new capabilities, and open collaboration. Here is a distillation of notable milestones, organized as a narrative timeline rather than a strict changelog, so you can sense the through-line of BGE’s development.
- The latest wave of multimodal embedding models arrives with BGE-VL, a state-of-the-art family designed to support a range of visual search tasks. This includes text-to-image, image-to-text, image-and-prompt-to-image, and more, all under a permissive MIT license. The release opens doors for both academic and commercial use, inviting developers to experiment with end-to-end visual search pipelines powered by BGE-VL, and the MegaPairs synthetic dataset that fuels its capabilities.
- Documentation centralization: A dedicated BGE documentation hub consolidates information, tutorials, and materials, acting as a single source of truth for users who want to learn, deploy, and extend BGE in real projects.
- Community growth: A WeChat group has been created to facilitate quick, real-time collaboration, updates, and questions. The QR-based join flow invites practitioners to connect, share ideas, and stay in the loop about updates and releases.
- OmniGen: A unified image-generation model that removes the friction of needing additional plugins or specialized detectors to accomplish complex generation tasks. This aligns with BGE’s broader mission of unifying capabilities to deliver end-to-end AI tooling.
- MemoRAG: A memory-inspired approach that advances retrieval-augmented generation by leveraging memory discovery, moving toward RAG 2.0 paradigms that emphasize persistent knowledge integration.
- Tutorials advancement: A dedicated tutorials section began to take shape, with ongoing content creation to help users master retrieval techniques, evaluation, and practical deployment.
- In-context learning and multilingual expansion: New embedding models such as bge-en-icl and bge-multilingual-gemma2 broaden the semantic and cross-lingual reach, enabling richer query representations and multilingual retrieval performance.
- Lightweight, efficient optimizers: Reranking models with efficiency-focused designs (e.g., v2.5 gemma2-lightweight) bring speed and resource savings without sacrificing accuracy for many use cases.
- Benchmarks and long-context advances: New benchmarks for long video understanding and neural IR evaluation frameworks highlight BGE’s intent to address complex, real-world tasks where context length and domain variety matter.
- Cross-encoder and embedding integration: The platform’s rerankers and embedder work in harmony, including integration with Langchain, making BGE accessible through familiar toolchains.
- Long-context and large models: Inference improvements, larger context support, and high-performing base-to-large scale embeddings demonstrate BGE’s ambition to scale with the needs of modern LLM assistants and retrieval pipelines.
Throughout 2023 and 2024, BGE’s release cadence has emphasized both performance and practicality. The project has shipped a nested ecosystem that includes:
- Inference mechanisms: Embedding and reranking inference pipelines that plug into typical retrieval and augmentation workflows.
- Finetuning options: Fine-tuning tools and scripts to tailor embeddings for domain-specific retrieval objectives.
- Evaluation and datasets: Dedicated datasets and evaluation scripts to quantify performance on standard benchmarks, plus a focus on specialized Chinese benchmarks (C-MTEB) for broader benchmarking coverage.
- Community and tutorials: An active tutorial program, training materials, and collaborative channels that invite new contributors and practitioners to participate.
The visual map of Projects within the BGE ecosystem (as shown in the provided image) hints at the breadth of tools from embeddings to evaluation to datasets and tutorials. This multi-faceted approach reflects a deliberate design choice: to minimize the friction between building a retrieval system and achieving production-grade performance.
What’s New in 2024–2025: A Focus on Versatility and Access
A few notable threads stand out when surveying the recent updates:
- Multimodal and multi-task embeddings: BGE-VL and related models showcase an expansive view of retrieval that isn’t constrained to text alone. The emphasis on multimodal search aligns with real-world needs where images, prompts, and text interact to produce relevant results.
- Accessibility and licensing: Free for both academic and commercial use, with MIT licensing that makes experimentation and deployment straightforward.
- Documentation-first approach: The BGE documentation site anchors the knowledge base, helping new users ascend from quick-start usage to deeper customization.
- Open community channels: The WeChat group and tutorials promise ongoing, collaborative learning, with a roadmap that includes evaluation tooling, model updates, and tutorials that cover end-to-end workflows.
- Efficiency improvements: Lightweight rerankers and multilingual capabilities reduce the barrier to production, enabling faster inference while maintaining or improving ranking quality across languages.
In the spirit of openness and collaboration, the BGE project also highlights contributor engagement, with banners and badges signaling welcoming attitudes toward new developers who want to join and help sustain the ecosystem.
Getting Started: Installation and Quick Setup
A practical entry point is the installation path that allows you to start using BGE with minimal friction. There are two primary routes:
Quick installation (no finetuning required)
Install via pip:
- Command: pip install -U FlagEmbedding
This path provides access to the embedding and inference capabilities without pulling in heavy finetuning dependencies.
Full installation with finetuning support
Install via pip with the finetune extras:
- Command: pip install -U FlagEmbedding[finetune]
Or install from source if you want to contribute or run in editable mode:
- Clone the repository: git clone https://github.com/FlagOpen/FlagEmbedding.git
- Change into the directory: cd FlagEmbedding
- If you don’t need finetune: pip install .
- If you want finetuning: pip install .[finetune]
For editable development:
- If you don’t need finetune: pip install -e .
- If you want finetune: pip install -e .[finetune]
A quick start example demonstrates how to load a BGE embedding model and compute embeddings:
Load a prefinetuned BGE model:
from FlagEmbedding import FlagAutoModel
model = FlagAutoModel.fromfinetuned('BAAI/bge-base-en-v1.5', queryinstructionforretrieval="Represent this sentence for searching relevant passages:", use_fp16=True)
Encode sample sentences:
sentences_1 = ["I love NLP", "I love machine learning"]
sentences_2 = ["I love BGE", "I love text retrieval"]
embeddings1 = model.encode(sentences1)
embeddings2 = model.encode(sentences2)
Compute similarity (inner product) and inspect the results:
similarity = embeddings1 @ embeddings2.T
print(similarity)
Beyond the quick-start code, the project hosts a suite of resources to support learning and experimentation, including:
- embedder inference
- reranker inference
- embedder finetune
- reranker finetune
- evaluation
- datasets
- tutorials
- broader research topics
A note on practical usage: if you are new to retrieval and RAG, begin with the tutorials to build a concrete mental model of how embeddings populate vector spaces, how to select an appropriate reranker, and how to integrate these components with an LLM in a retrieval-augmented prompt.
Visual and Community Context: Images that Guide the Experience
The included images serve as a visual narrative of BGE’s ecosystem. The project’s branding, the community, and the roadmap are illustrated by:
- The FlagOpen banner and the BGE logo, anchoring the project identity.
- A “Projects” graphic that enumerates the major components in the BGE/FlagEmbedding ecosystem, from Inference to Datasets and Tutorials.
- The WeChat group QR code, inviting readers to join a live, ongoing discussion with the BGE community.
- The Tutorials roadmap image, offering a visual glance at planned educational materials and the evolution of learning tracks for new users.
For readers who want to see these visuals, they are embedded in the post at their respective places to provide a cohesive, picture-filled narrative of the journey.
Model List: A Catalogue of Models and Capabilities
BGE’s model ecosystem is diverse, encompassing English, Chinese, multilingual variants, and models designed for mixed retrieval tasks. The goal is to provide a catalog that helps practitioners choose the right tool for their domain and language needs. Here is a narrative mapping of the major models, rendered in a concise, readable format (without tabular constraints):
BAAI/bge-en-icl
Language: English
Description: An embedding model with in-context learning capabilities. It leverages task-relevant examples to enrich query representations, enabling richer semantic encoding for retrieval.
Query instruction for retrieval: Provide instructions and few-shot examples freely based on the given task.
BAAI/bge-multilingual-gemma2
Language: Multilingual
Description: A multilingual embedding model built on Gemma-2 with strong cross-language transfer properties, designed to support diverse tasks.
Query instruction for retrieval: Provide instructions based on the given task.
BAAI/bge-m3
Language: Multilingual
Description: A multi-function model feature set that includes dense retrieval, sparse retrieval, and multi-vector (ColBERT-like) retrieval. It supports high degrees of multilinguality and granularity (up to 8192 tokens).
Notes: A versatile backbone for complex retrieval workloads.
LM-Cocktail (Shitao)
Language: English
Description: Fine-tuned models to reproduce LM-Cocktail results, illustrating practical re-use in retrieval augmentation.
BAAI/llm-embedder
Language: English
Description: A unified embedding model designed to support diverse retrieval augmentation needs for LLMs.
BAAI/bge-reranker-v2-m3
Language: Multilingual
Description: A lightweight cross-encoder model with strong multilingual capabilities that is easy to deploy and supports fast inference.
BAAI/bge-reranker-v2-gemma
Language: Multilingual
Description: A cross-encoder optimized for multilingual contexts; excels in both English and multilingual retrieval settings.
BAAI/bge-reranker-v2-minicpm-layerwise
Language: Multilingual
Description: A cross-encoder with layerwise configurability for output, enabling accelerated inference and flexible resource use.
BAAI/bge-reranker-v2.5-gemma2-lightweight
Language: Multilingual
Description: A lightweight cross-encoder designed for multilingual contexts with adjustable compression and layer selection to speed up inference.
BAAI/bge-reranker-large
Language: Chinese and English
Description: A higher-accuracy cross-encoder that trades some efficiency for stronger ranking performance.
BAAI/bge-reranker-base
Language: Chinese and English
Description: A robust cross-encoder that balances accuracy and efficiency for broader use cases.
BAAI/bge-large-en-v1.5
Language: English
Description: Version 1.5 with improved similarity distribution for English embeddings; includes a retrieval instruction header such as “Represent this sentence for searching relevant passages:”.
BAAI/bge-base-en-v1.5
Language: English
Description: Version 1.5 with improved similarity distribution; common retrieval instruction included for English.
BAAI/bge-small-en-v1.5
Language: English
Description: A smaller model with improved similarity distribution for English retrieval tasks.
BAAI/bge-large-zh-v1.5
Language: Chinese
Description: Chinese variant with stable similarity distribution improvements and a retrieval-oriented instruction prefix in Chinese (for example, “为这个句子生成表示以用于检索相关文章:”).
BAAI/bge-base-zh-v1.5
Language: Chinese
Description: Base-scale Chinese model with enhanced retrieval quality, comparable to larger configurations.
BAAI/bge-small-zh-v1.5
Language: Chinese
Description: A compact Chinese model tuned for cost-effective retrieval tasks.
BAAI/bge-large-en
Language: English
Description: Large English embedding model mapping text to vector space for robust retrieval.
BAAI/bge-base-en
Language: English
Description: A solid base-scale English embedding model with strong retrieval capability similar to larger peers.
BAAI/bge-small-en
Language: English
Description: A compact English embedding model balancing performance and efficiency.
BAAI/bge-large-zh
Language: Chinese
Description: Large Chinese embedding model with strong retrieval performance.
BAAI/bge-base-zh
Language: Chinese
Description: Base Chinese embedding model with competitive retrieval performance.
BAAI/bge-small-zh
Language: Chinese
Description: Small but efficient Chinese embedding model.
Contributors: A Community of Builders
FlagEmbedding and BGE thrive on a community of researchers, engineers, and enthusiasts who contribute code, documentation, benchmarks, and tutorials. The project celebrates its contributors, with a reminder that new members are welcome to join in and help advance the ecosystem. A visual badge from the contributor graph is included to highlight this collaborative spirit and to invite participation from more developers.
- Contributors: A note of thanks to everyone who has contributed so far, and a call for new members to join the effort. If you’d like to explore the contributor landscape, you can view the contribution graph on GitHub.
Citations and Acknowledgments: Foundational References
BGE rests on a lineage of research and practical engineering in the field of retrieval-augmented language models. The project maintains a small set of citations that reflect its technical roots and inspiration:
- BGE M3 Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. Chen, Jianlv, Xiao, Shitao, Zhang, Peitian, Luo, Kun, Lian, Defu, Liu, Zheng. 2023. arXiv:2309.07597.
- LM-Cocktail: Resilient Tuning of Language Models via Model Merging. Xiao, Shitao; Liu, Zheng; Zhang, Peitian; Xing, etc. 2023. arXiv:2311.13534.
- Retrieve Anything To Augment Large Language Models (LLM-Embedder). Zhang, Peitian; Xiao, Shitao; Liu, Zheng; Dou, Zhicheng; Nie, Jian-Yun. 2023. arXiv:2310.07554.
- C-Pack: Packaged Resources To Advance General Chinese Embedding. Xiao, Shitao; Liu, Zheng; Zhang, Peitian; Muennighoff, Niklas. 2023. arXiv:2309.07597.
These citations frame the lineage of techniques and ideas that inform BGE’s approach to embedding, retrieval, and model merging for robust information access in LLM-assisted workflows.
Licensing and Accessibility
FlagEmbedding, including BGE, is licensed under the MIT License, reflecting a commitment to openness and broad usability. The MIT license is widely used in the research and developer communities, enabling re-use, modification, and distribution with minimal restrictions. This permissive stance aligns with BGE’s goal of fostering an ecosystem where researchers, developers, and organizations—both academic and commercial—can build, deploy, and experiment with retrieval-enhanced AI systems.
Practical Takeaways: How to Employ BGE in Real Projects
- Start with the quick-start embedding path to validate your data and use case. This is ideal for prototyping ideas quickly and evaluating fundamental retrieval quality.
- When moving toward production or domain-specific needs, consider finetuning or choosing a model that aligns with your language requirements (English, Chinese, multilingual) and your resource constraints.
- Combine embedding models with rerankers to achieve higher precision in top-k retrieval results. Cross-encoder rerankers often yield more accurate top results than embeddings alone, at the cost of some additional compute.
- Leverage multilingual models to support cross-lingual search capabilities, which can be particularly valuable for global datasets or multilingual user bases.
- Use the tutorials and documentation as you scale from a small prototype to a full-fledged retrieval-augmented pipeline integrated with LLMs.
Visual Roadmap and Tutorial Roadmap
The Tutorials section, alongside the public roadmap, outlines a path from introductory concepts to advanced techniques. The visuals illustrate how the tutorials map to practical tasks (embedding, fine-tuning, evaluation) and how they fit into an end-to-end retrieval workflow. The roadmap is designed to be updated iteratively as new content is added and as the community’s needs evolve.
What to Expect Next
The ongoing trajectory points to richer evaluation tools, expanded model coverage (including English, Chinese, and multilingual variants), and deeper tutorials that cover end-to-end retrieval augmentation with LLMs. Expect more tutorials on in-context learning, retrieval evaluation under different domain shifts, and practical deployment patterns that help you move from research into production with confidence.
Images embedded in this narrative
- FlagOpen branding and project identity to anchor the post visually
- FlagOpen banner
- BGE brand and core models
- BGE Logo
- Projects and ecosystem overview
- Projects Diagram
- WeChat community QR code for BGE updates
- BGE WeChat Group QR
- Tutorials roadmap illustrating the training and learning path
- Tutorial Roadmap
Conclusion: A Mature Yet Growing Retrieval Toolkit
BGE represents a mature, practical, and extensible approach to retrieval in the era of large language models. By combining dense embeddings, cross-encoder rerankers, multilingual capabilities, and a strong emphasis on documentation and tutorials, the project lowers the barriers to building robust RAG systems. The ecosystem is designed to welcome new contributors, encourage experimentation, and support real-world deployments—whether your focus is English-centric, Chinese-focused, or multilingual.
If you’re embarking on a project that requires strong semantic search, cross-lingual retrieval, and integration with LLMs, BGE offers a ready-to-use set of tools and a path toward scalable, production-grade retrieval systems. The ongoing updates, documentation, and community channels promise continuous improvement and practical guidance as you navigate the complexities of modern retrieval and augmentation pipelines.
To get involved, explore the installations, sample code, and tutorials, and consider joining the community channels to stay close to the latest releases and discussions. The journey from embedding to retrieval to RAG is a collaborative one, and BGE is positioned to be a central part of that journey for many teams and researchers.
Enjoying this project?
Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.
Repository:https://github.com/FlagOpen/FlagEmbedding
GitHub - FlagOpen/FlagEmbedding: FlagEmbedding — Retrieval and Retrieval-augmented LLMs
Retrieval and Retrieval-augmented LLMs...
github - flagopen/flagembedding

