Chunkr
GitHub Repo
AGPL-3.0
July 1, 2026 at 02:22 AM
0 views

Chunkr

@lumina-ai-incProject Author

Chunkr Logo

Chunkr: Open Source Document Intelligence for Modern AI Workflows

Chunkr is a production-ready service designed to transform how you process documents. It combines layout analysis, optical character recognition (OCR) with precise bounding boxes, and semantic chunking to produce machine-readable fragments that are ready for retrieval-augmented generation (RAG) and large language model (LLM) workflows. Whether you’re converting PDFs, PowerPoint decks, Word documents, or images, Chunkr breaks down complex documents into structured, searchable chunks that your AI models can understand and act upon.

This blog post walks you through what Chunkr offers, how the open-source version differs from the cloud API and enterprise options, plus practical guidance to get started, configure LLMs, and connect with the Chunkr ecosystem. We’ll also share key considerations around licensing, support, and how to choose the right deployment for your needs.

Include: Visuals and quick references to input assets that accompany Chunkr’s ecosystem, including the Chunkr logo and the Cloud API overview image.

A visual overview of Chunkr’s capabilities

  • Layout Analysis: Chunkr understands the structural organization of documents, identifying headings, tables, figures, columns, and other layout cues that matter for downstream processing.
  • OCR with Bounding Boxes: It detects text and associates precise bounding boxes, enabling pixel-perfect extraction and alignment with the original document.
  • Structured HTML and Markdown: Output can be bundled into clean HTML or Markdown, preserving semantic structure for web display or downstream ingestion.
  • Vision-Language Model Processing: Beyond pure OCR, Chunkr can incorporate vision-language insights to enhance interpretation, extraction accuracy, and semantic chunking.

In practice, this means you can take a multi-page PDF or a slide deck, generate a robust set of chunks, and feed them into an LLM-based pipeline for question answering, summarization, extraction, or semantic search. The pipeline is designed to be ready for RAG workflows, enabling easier indexing, retrieval, and synthesis from large document collections.

Open Source vs Cloud API vs Enterprise: What’s the right choice for you?

Chunkr offers three deployment paradigms, each with its own strengths. Although the feature set is shared in spirit across editions, the underlying models, performance characteristics, and deployment options differ. Below is a guided, narrative comparison to help you decide which path matches your needs.

  • Open Source Release (AGPL-3.0)

  • Perfect for development and testing, experimentation, or local hosting where transparency and control are paramount.

  • Layout analysis relies on community or open-source models, offering a transparent view into the pipeline.

  • OCR and VLM processing leverage compatible open-source engines and models.

  • Excel is not natively supported in the open-source release; users who require Excel-like parsing can implement custom parsers or use added tools in their stack.

  • Infrastructure is self-hosted, giving you full control over deployment, data, and security postures.

  • Community-driven support via Discord, community forums, and public issues.

  • Migration path: If you later need enterprise-grade reliability or in-house model tuning, you can consider moving to the Cloud API or Enterprise editions.

  • Cloud API (Chunkr Cloud)

  • Best for production workloads requiring performance, reliability, and rapid time-to-value.

  • In-house, proprietary models deliver higher accuracy and faster response times, with enterprise-grade reliability and scale.

  • Enhanced vision-language processing capabilities optimized for real-world, large-scale usage.

  • Native Excel-style parsers and rich tooling to accelerate business workflows.

  • Cloud-hosted infrastructure with managed security, monitoring, and updates.

  • Dedicated support and ongoing improvements from the vendor.

  • Migration path: You can start in the cloud and migrate toward on-prem or dedicated enterprise deployments if needed.

  • Enterprise

  • Designed for high-security or regulated industries with strict data sovereignty requirements.

  • Offers on-premises or VPC deployments, providing maximum control over data flow, access, and compliance.

  • Access to in-house tuned models and specialized configurations to meet industry-specific needs.

  • Comprehensive support, SLAs, and structured migration programs.

  • Ideal for organizations with rigorous governance, privacy, and audit requirements.

While the open-source release provides transparency and local hosting, Chunkr Cloud delivers best-in-class performance and enterprise-grade reliability for most production environments. For organizations with stringent security or regulatory constraints, the Enterprise edition provides on-prem or tightly controlled VPC deployments.

Quick Start with Docker Compose: Getting started quickly

Prerequisites

  • Docker and Docker Compose installed on your machine.
  • NVIDIA Container Toolkit installed if you plan to use GPU acceleration (optional but recommended for better performance).

Steps to get up and running

  • Clone the repository
  • Run:
    • git clone https://github.com/lumina-ai-inc/chunkr
    • cd chunkr
  • Set up environment variables
  • Copy the example environment file:
    • cp .env.example .env
  • Configure your LLMs by copying the example models file:
    • cp models.example.yaml models.yaml
  • Start the services
  • For GPU deployment:
    • docker compose up -d
  • For CPU-only deployment:
    • docker compose -f compose.yaml -f compose.cpu.yaml up -d
  • For Mac ARM (M1, M2, M3, etc.):
    • docker compose -f compose.yaml -f compose.cpu.yaml -f compose.mac.yaml up -d
  • Access and verify
  • Web UI: http://localhost:5173
  • API: http://localhost:8000
  • Stop the services when finished
  • For GPU deployment:
    • docker compose down
  • For CPU-only deployment:
    • docker compose -f compose.yaml -f compose.cpu.yaml down
  • For Mac ARM:
    • docker compose -f compose.yaml -f compose.cpu.yaml -f compose.mac.yaml down

This streamlined flow makes it straightforward to experience Chunkr locally, experiment with different LLM configurations, and observe how document chunks are produced and surfaced for downstream AI tasks.

LLM Configuration: Two paths to configure models

Chunkr provides two principal approaches to configure LLMs, each with its own trade-offs. You can use either or both depending on your needs for flexibility, multi-provider setups, and governance.

  • Benefits:
  • Flexible configuration for multiple LLM providers.
  • Ability to set default and fallback models.
  • Per-model rate limits (useful for budgeting and throttling).
  • Reference models by ID in API requests to simplify client configuration.
  • How to use:
  • Copy the example file to create your configuration:
    • cp models.example.yaml models.yaml
  • Edit models.yaml with your provider details, for example:
    • id: gpt-4o
    • model: gpt-4o
    • provider_url: https://api.openai.com/v1/chat/completions
    • apikey: "youropenaiapikey_here"
    • default: true
    • rate-limit: 200
  • Why this approach shines:
  • You can configure multiple providers side-by-side, select default and fallback models, and enforce distribution of requests across models with rate limits.
  • It enables more resilient deployments, where if one provider is slow or unavailable, another can seamlessly take over.

2) Environment Variables (Basic)

  • Use this approach if you have a single LLM endpoint and want a straightforward setup without the complexity of multiple providers.
  • You can configure a single provider by setting the following variables in your .env file:
  • LLM__KEY
  • LLM__MODEL
  • LLM__URL
  • This path is ideal for simple workflows or for teams that want a quick-start, minimal configuration, and predictable behavior.

Common LLM API Providers: Getting started quickly

Chunkr’s flexibility shines through its support for a range of LLM providers. Below are some of the common options you may consider, along with their typical endpoints and documentation references:

  • OpenAI
  • API URL: https://api.openai.com/v1/chat/completions
  • Documentation: OpenAI Docs (https://platform.openai.com/docs)
  • Google AI Studio (Gemini/OpenAI-compatible endpoints)
  • API URL: https://generativelanguage.googleapis.com/v1beta/openai/chat/completions
  • Documentation: Google AI Docs (https://ai.google.dev/gemini-api/docs/openai)
  • OpenRouter
  • API URL: https://openrouter.ai/api/v1/chat/completions
  • Documentation: OpenRouter Models (https://openrouter.ai/models)
  • Self-Hosted
  • API URL: http://localhost:8000/v1
  • Documentation: VLLM/Open Source options (https://docs.vllm.ai/en/latest/serving/openaicompatibleserver.html) or Ollama (https://ollama.com/blog/openai-compatibility)

These providers illustrate the spectrum of deployment choices: managed cloud services, hosted alternatives, and self-hosted solutions. Depending on your policy constraints, data governance requirements, and performance targets, you can mix and match providers, all while keeping your API surface consistent.

Licensing: Dual-licensing and how to use Chunkr legally

Chunkr’s core is dual-licensed to balance open collaboration with commercial flexibility:

  • GNU Affero General Public License v3.0 (AGPL-3.0)
  • Commercial License

If you want to use Chunkr without complying with AGPL-3.0 terms, you can contact the Chunkr team or explore the commercial license options. This dual-licensing approach is designed to maximize openness for experimentation and transparency, while offering a path to production-grade usage under commercial terms for organizations that require it.

Connections and support

Chunkr’s ecosystem includes multiple channels to stay in touch, seek help, and explore opportunities:

  • Email: [email protected]
  • Schedule a call: Book a 30-minute meeting
  • Website: chunkr.ai
  • Community and support: Discord channel for quick help and discussions

Engaging with the community, official channels, and the vendor’s support structure helps teams accelerate adoption, resolve issues faster, and align Chunkr usage with deployment goals.

Visuals to accompany your deployment narrative

To give readers a sense of the Chunkr branding and its product imagery, the post features two key visuals that reflect the product’s identity and its cloud narrative:

  • Chunkr Logo: A visual anchor for the project, illustrating the brand identity and its open-source roots.
  • Chunkr Cloud API Overview: A visual representation of how Chunkr’s cloud API integrates with modern AI pipelines, highlighting end-to-end processing from layout analysis to RAG-ready chunks.

You can also explore the Chunkr Cloud API landing page and its OG image for additional context about deployment models and usage expectations.

In-depth notes on structure, output, and integration

  • Output formats: When processing documents, Chunkr can produce structured HTML and Markdown outputs that preserve document semantics. This makes it easier to ingest the results into knowledge bases, content management systems, or downstream AI pipelines without reformatting.
  • Bounding boxes and searchability: The OCR layer emits bounding boxes for extracted text, enabling precise alignment with source pages. This is critical when linking extracted data to specific document regions, like tables or figures, and when performing pixel-perfect reconstructions in downstream tools.
  • Vision-language capabilities: The VLM processing component helps with nuanced interpretation, such as understanding captions for figures or cross-referencing textual content with visual cues. This improves the quality of semantic chunks and enhances downstream tasks like question answering and summarization.
  • Document types: Chunkr supports common document types such as PDFs, PPTs, Word documents, and images. The open-source edition focuses on core functionality, while the cloud and enterprise variants extend support to additional formats and specialized parsers, including native Excel-related tooling in the higher tiers.
  • Infrastructure considerations: The open-source release is self-hosted, giving you full control over security, data residency, and network boundaries. The cloud edition abstracts infrastructure management away from you, while the enterprise edition provides deployment options that align with regulated environments.
  • Security and governance: For regulated industries or environments with strict data controls, the enterprise option’s on-prem or VPC deployments offer the strongest guarantees. It’s important to align deployment choices with your data privacy and compliance programs.

Putting it all together: A practical path forward

  • If you’re in the exploration phase and want maximum transparency, run Chunkr locally or in a controlled environment using the AGPL-3.0 open-source release. This is ideal for developers who want to validate the pipeline, experiment with models, and build custom integrations.
  • If you’re seeking scale, reliability, and lower time-to-value, start with Chunkr Cloud. You’ll benefit from optimized models, robust infrastructure, and enterprise-grade features without managing the underlying stack.
  • If your organization operates in a highly regulated domain with data sovereignty requirements, pursue the Enterprise edition for on-prem or private cloud deployments, coupled with tailored support and governance controls.

Sample configuration snippet: a glimpse into models.yaml (illustrative)

  • id: gpt-4o
  • model: gpt-4o
  • provider_url: https://api.openai.com/v1/chat/completions
  • apikey: "youropenaiapikey_here"
  • default: true
  • rate-limit: 200

This simplified snippet demonstrates how multiple providers can be orchestrated within models.yaml. The real configuration offers additional options, thresholds, and fallback rules to fit your operational needs.

A final note on the journey with Chunkr

Chunkr is more than a tool; it’s a framework for turning unstructured documents into structured, actionable intelligence. By combining layout-aware analysis, precise text extraction, and semantic chunking that resonates with modern LLM workflows, Chunkr unlocks new efficiencies in knowledge extraction, content modernization, and AI-powered decision support.

Whether you’re building a document-centric RAG system, organizing large knowledge bases, or enabling robust content pipelines for enterprise AI, Chunkr provides the foundation to transform messy source material into easily searchable, semantically meaningful chunks. The open-source base keeps you honest and collaborative, while the cloud and enterprise paths give you the performance, security, and governance your organization requires.

Connect with us and explore how Chunkr can fit your architecture

  • Email: [email protected]
  • Schedule a 30-minute call: Book a meeting
  • Website: chunkr.ai
  • Community: Discord

Images used in this narrative

  • Chunkr Logo: Chunkr Logo
  • Chunkr Cloud API overview: Chunkr Cloud API

Appendix: Quick reference at a glance (for readers who want a summary)

  • What Chunkr does: Layout analysis, OCR with bounding boxes, structured HTML/Markdown, and vision-language processing to create RAG-ready document chunks.
  • Deployment options: Open Source AGPL (self-hosted), Cloud API (proprietary in-house models), and Enterprise (on-prem or VPC with tailored deployments).
  • Getting started: Docker Compose-based quick start with GPU or CPU paths, Mac ARM support, web UI and API endpoints.
  • LLM configuration: models.yaml (recommended for multi-model setups) or environment variables (basic single-model setups).
  • Providers: OpenAI, Google AI Studio, OpenRouter, Self-Hosted (VLLM / Ollama) with corresponding URLs and docs.
  • Licensing: AGPL-3.0 and Commercial License; license terms in the project.
  • Connect with Chunkr: Mehul and the team are reachable via email, scheduling, and the website, with a community channel for ongoing support.

This narrative presents a detailed, practical overview of Chunkr’s capabilities, deployment models, and how teams can get started. It’s designed to be a living guide that you can reference as you evaluate Chunkr for your document intelligence workflows.

Enjoying this project?

Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.

Project
chunkr
Created
July 1
Last Updated
July 1, 2026 at 02:22 AM

Find more projects like this

One email a week: new and trending developer tools, fresh comparisons, and what shipped. Unsubscribe in one click.