Docling: Document Parsing for Generative AI
GitHub Repo
MIT
October 2, 2026 at 09:19 AM
0 views

Docling: Document Parsing for Generative AI

@docling-projectProject Author

What Docling is

Docling is a document conversion toolkit that turns messy real-world files into structured data that language models can work with. The README's one-line summary: it simplifies document processing by parsing diverse formats, including advanced PDF understanding, and providing integrations with the generative AI ecosystem. In practice that means taking a PDF, Word file, slide deck, spreadsheet, web page, e-book, email or even an audio recording, and producing clean Markdown, HTML or lossless JSON that preserves headings, reading order, tables, code and formulas.

The project was started by the AI for Knowledge team at IBM Research Zurich and is now hosted by the LF AI & Data Foundation. The code is MIT-licensed and written in Python. Its audience is anyone building RAG systems, agents or document analytics who needs a reliable, locally runnable front end for ingestion rather than a hosted parsing API.

How it works

Docling's architecture docs describe a small set of concepts. For each input format, the document converter knows which format-specific backend to use for parsing and which pipeline orchestrates execution, along with any options. That mapping has defaults but is configurable, so a PDF can be processed with different backends or pipeline options depending on your quality and speed needs.

Every conversion produces a DoclingDocument, the project's unified representation format. From there you can call export methods (Markdown, HTML, dictionary and others), pass it to a serializer, or hand it to a chunker. Chunkers implement a BaseChunker interface with chunk() and contextualize() methods. The built-in HybridChunker starts from hierarchical, document-structure-based chunks and then applies tokenizer-aware refinements: it splits chunks that exceed the token limit and merges undersized neighbors that share headings and captions. Table chunks can repeat their headers when a table spans several chunks. Integrations with frameworks like LlamaIndex are built on this same interface, and base classes can be subclassed for specialized implementations.

For PDFs specifically, the pipeline handles page layout, reading order, table structure, code, formulas and image classification, with OCR for scanned pages. An alternative VLM pipeline runs vision-language models such as IBM's GraniteDocling to convert pages end to end. The full design is described in the Docling Technical Report on arXiv.

Key features

  • Wide input coverage: PDF, DOCX, XLSX, PPTX, legacy DOC/XLS/PPT and RTF (via LibreOffice), ODT/ODS/ODP, EPUB, Apple Pages and Keynote, Markdown, AsciiDoc, LaTeX, HTML, MHTML, CSV, images, WebVTT, Box Notes, EML and MSG email, and IBM AFP.
  • Domain XML schemas: DocLang, USPTO patents, JATS articles and XBRL financial reports, plus EBCDIC mainframe files given a COBOL record layout.
  • Advanced PDF understanding: layout, reading order, table structure, code, formulas and image classification.
  • Multiple export formats: Markdown, HTML, plain text, lossless JSON, DocLang XML, DocTags, WebVTT, LaTeX and chunked JSONL for RAG.
  • OCR: extensive support for scanned PDFs and images.
  • VLM support: run visual language models such as GraniteDocling as the conversion pipeline.
  • Audio and video: automatic speech recognition for WAV, MP3 and other audio, and video files whose audio is transcribed with representative keyframes.
  • Chart understanding: bar, pie and line charts converted into tables or code with descriptions, listed under recent additions.
  • Local execution: designed to run on your own hardware, including air-gapped environments for sensitive data.
  • Integrations: LangChain, LlamaIndex, CrewAI and Haystack, an MCP server for agents, and docling-serve for running Docling as an API service.

Getting started

Install from PyPI. Python 3.10 or newer is required; Python 3.9 support was dropped in version 2.70.0.

pip install docling

Convert a document from the command line; this writes a .md file to the current directory:

docling https://arxiv.org/pdf/2206.01062

Use the VLM pipeline with GraniteDocling instead of the default pipeline:

docling --pipeline vlm --vlm-model granite_docling https://arxiv.org/pdf/2206.01062

The recommended route is the Python API:

from docling.document_converter import DocumentConverter

source = "https://arxiv.org/pdf/2408.09869"  # a document via a local path or URL
converter = DocumentConverter()
result = converter.convert(source)
print(result.document.export_to_markdown())  # output: "## Docling Technical Report[...]"

Some formats need extras: the supported-formats page notes that audio and video require the asr extra (and ffmpeg for video), Apple iWork files require format-iwork, and legacy Office formats require LibreOffice.

Use cases

  • RAG ingestion: convert a document corpus into structure-aware chunks with headings and captions attached, ready for embedding.
  • Table extraction: pull tables out of PDFs and reports with their structure intact rather than as flattened text.
  • Scientific and patent processing: parse arXiv PDFs, JATS articles and USPTO patents into a consistent format, including formulas and code.
  • Financial reporting: ingest XBRL filings and annual-report PDFs into JSON for downstream analysis.
  • Agent tools: expose conversion to an agent through the MCP server so it can read arbitrary documents on demand.
  • Meeting and media transcripts: transcribe recordings and videos into the same document model as text sources.
  • Air-gapped pipelines: process sensitive legal or medical documents without sending them to a third-party API.

How it compares

The README does not benchmark itself against named tools. Docling sits in the same space as other open-source document parsers used ahead of RAG pipelines, and as hosted parsing APIs offered by cloud and LLM vendors. Its distinguishing points, from the docs, are the breadth of formats behind a single document model, structure-aware native chunkers, the choice between a classic layout pipeline and a VLM pipeline, and the ability to run fully locally. Projects that only ever handle clean digital PDFs may find a lighter text extractor sufficient; Docling earns its weight when layout, tables and mixed formats matter, and when you want one library to cover PDFs, office files, web pages and recordings instead of stitching together a parser per format.

Things to know before adopting

  • Model licenses differ: the codebase is MIT, but the README says that for individual model usage you should check the model licenses in the original packages.
  • Python version: 3.10 or higher is required for current releases.
  • Platform support: macOS, Linux and Windows on both x86_64 and arm64.
  • Optional dependencies: audio, video, iWork and legacy Office formats depend on extras or external tools like ffmpeg and LibreOffice.
  • Compute: layout analysis, OCR and especially VLM pipelines run models locally, so throughput depends on your hardware.
  • Governance: hosted by the LF AI & Data Foundation, with an OpenSSF Best Practices badge linked from the README.
  • Roadmap: metadata extraction (title, authors, references, language) and chemistry understanding are listed as coming soon, not shipped.

Project activity

As of October 2026 the repository has about 68,300 stars on GitHub. It was created on 2024-07-09, is written in Python, and is released under the MIT license. It is published on PyPI as docling. Source code is at github.com/docling-project/docling and documentation is at docling-project.github.io/docling.

Enjoying this project?

Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.

Project
docling
Created
October 2
Last Updated
October 2, 2026 at 09:19 AM

Find more projects like this

One email a week: new and trending developer tools, fresh comparisons, and what shipped. Unsubscribe in one click.