PageIndex: Vectorless, Reasoning-Based RAG
GitHub Repo
MIT
October 2, 2026 at 09:19 AM
0 views

PageIndex: Vectorless, Reasoning-Based RAG

@VectifyAIProject Author

What PageIndex is

PageIndex is a retrieval engine for long documents that drops the two staples of conventional RAG: the vector database and chunking. Its README makes the argument in one line: vector-based RAG retrieves by semantic similarity, but similarity is not relevance, and relevance requires reasoning. On professional documents that demand domain expertise and multi-step reasoning, similarity search can miss passages that are relevant but not similar, and return ones that are similar but not relevant.

PageIndex's answer, which the README says was inspired by AlphaGo, is to build a hierarchical tree index of each document and let an LLM reason its way through that tree, the way a human expert flips to the right section of a long report. It is built by VectifyAI (PageIndex AI), released under the MIT license, and distributed as a Python SDK. The README lists financial reports, legal documents, regulatory filings, technical manuals, medical literature and academic textbooks as the documents it is designed for.

How it works

Retrieval happens in two steps:

  1. Index: generate a tree-structure index for each document.
  2. Retrieve: agentically search that tree with LLM reasoning.

According to the README, the tree structure itself is extracted from the document layout without an LLM; an index model then summarizes and refines the nodes. At query time a chat model navigates the tree, reading only the nodes its reasoning reaches, and answers with references back to the source. Because the search is done by an LLM rather than an embedding lookup, it can take in the full context of a question (conversation history, domain knowledge) instead of only a query embedding.

The README summarizes the difference from vector RAG in a small table:

Vector RAGPageIndex
Indexvector indextree index
Retrievalsemantic similarity searchLLM reasoning over the tree
Resultopaque, "vibe retrieval"traceable to explicit references
Contextquery embedding onlyfull context: conversation history, domain knowledge, etc.

A recent addition, PageIndex Flash, provides fast tree index generation for text-based PDFs and is now the default indexing method in local mode.

Key features

  • No vector database: indexes are trees stored in a local directory, so there is no embedding pipeline or vector store to operate.
  • No chunking: documents keep their natural section structure instead of being split into fixed-size pieces.
  • Traceable retrieval: answers point to explicit references; local mode gives page-level citations.
  • Local mode: the SDK can index, retrieve and chat entirely on your machine with your own LLM key.
  • Separate index and chat models: a cheap model builds the index, while the best model you can afford does the searching.
  • Agent integration: PageIndex tools can be dropped into the OpenAI Agents SDK, the Claude Agent SDK or other frameworks, per the linked docs.
  • Same client for cloud: switching index="cloud" moves indexing and storage to PageIndex Cloud while chat still uses your own model provider.

Getting started

Install the SDK:

pip install -U pageindex

Index a PDF and ask a question in local mode:

import os
from pageindex import PageIndexClient

os.environ["OPENAI_API_KEY"] = "your-openai-key"

client = PageIndexClient(
    index="gpt-5.6-luna",               # model to build the tree index
    chat="gpt-5.6-sol",                 # model to search the tree
)
doc_id = client.submit_document("report.pdf")["doc_id"]

answer = client.chat("What was the 2023 operating margin?", doc_id=doc_id)
print(answer)

The README's model advice: for index= a basic model is sufficient, since it only summarizes and refines a structure already extracted from layout; for chat=, use the best model you can afford, because that model does the actual retrieval. The SDK documentation covers other models, streaming, multi-document search and citations.

Cost and accuracy claims

The README publishes several numbers, all from the project's own experiments:

  • Indexing cost: it claims local tree building runs about $0.001 per page with gpt-5.6-luna, so a 1,000-page textbook costs a little over a dollar, once. Benchmark documents of 9 to 1,098 pages reportedly indexed in roughly 13 seconds to 4.5 minutes.
  • Query accuracy: a separate PageIndex-OSS-Benchmark repo evaluates the quickstart setup (local mode, flash indexing, no OCR) on 62 lookup questions over 34 PDFs (1,945 pages) drawn from MMLongBench-Doc-V2.
  • Cost vs. native PDF input: on documents where both approaches return the same answer, the README says sending the whole PDF to the model costs 2.1x more at 52 pages and 16.6x more at 420 pages, and at 805 pages the document no longer fits in the context window.
  • FinanceBench: the README claims 98.7% accuracy on FinanceBench, linking evaluation results from VectifyAI's Mafin 2.5 system.

Use cases

  • Financial analysis: answering specific questions over 10-Ks, annual reports and earnings filings where the answer lives in one section of hundreds of pages.
  • Legal and regulatory review: locating clauses or requirements in long contracts and filings with traceable page references.
  • Technical manuals: support or field-engineering assistants that need the exact procedure from a structured manual.
  • Academic and medical literature: question answering over textbooks and long papers where section hierarchy carries meaning.
  • Agent tooling: giving an agent a document-reading tool that navigates structure instead of returning loosely related chunks.

How it compares

The README positions PageIndex against conventional vector RAG and against passing the full PDF into a long-context model. Compared with vector RAG, the trade-off is more LLM work at query time in exchange for structure-aware, explainable retrieval and no embedding infrastructure. Compared with native PDF input, the README argues retrieval cost stays flat as documents grow, since only visited nodes are read. The approach suits single long, well-structured documents best; for large corpora the README points to PageIndex File System, a Cloud-only file-level tree layer for reasoning across many documents.

Things to know before adopting

  • Open-source scope: the README says the open-source version is ideal for text-heavy PDFs and local workflows. Local mode handles text-based PDFs only, with no OCR or image understanding.
  • Cloud-only features: OCR, scanned and image-rich documents, block-level citations, metadata, folders, the MCP server and PageIndex File System require PageIndex Cloud and an API key.
  • Query-time LLM cost: every query runs a reasoning search, so answer quality and cost depend heavily on the chat model you choose.
  • Vendor-reported benchmarks: all published numbers come from PageIndex's own experiments and benchmark repos.
  • Deployment: dedicated VPC or on-premises deployments of the cloud product are available by contacting the company.

Project activity

As of October 2026 the repository has about 38,500 stars on GitHub. It was created on 2025-04-01, is written in Python, and is released under the MIT license. The SDK is published on PyPI as pageindex, with local mode added in August 2026. Source code is at github.com/VectifyAI/PageIndex and the project site is pageindex.ai.

Enjoying this project?

Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.

Project
pageindex
Created
October 2
Last Updated
October 2, 2026 at 09:19 AM

Find more projects like this

One email a week: new and trending developer tools, fresh comparisons, and what shipped. Unsubscribe in one click.