Newsletter image

Subscribe to the Newsletter

Join 10k+ people to get notified about new posts, news and tips.

Do not worry we don't spam!

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

Search

GDPR Compliance

We use cookies to ensure you get the best experience on our website. By continuing to use our site, you accept our use of cookies, Privacy Policy, and Terms of Service.

VectifyAI - Document Understanding, Inference, Reasoning

PageIndex

PageIndex replaces vector search with a tree index an LLM navigates. How it really works, why the headline benchmark is 90.7% strict, and how to run it against a local model.

License MIT
License MIT
TL;DR
  • Tree index built from PDF layout, no chunking and no embeddings
  • Headline 98.7% on FinanceBench is adjudicated; strict agreement is 90.7%
  • Routes through LiteLLM, so Ollama works and Flash can index with no LLM

PageIndex has 36,000 stars and one very good argument: vector search retrieves what is similar, and similar is not the same as relevant. Instead of chopping a PDF into 500-token windows and embedding them, it reads the document's own table of contents into a tree and lets an LLM navigate to the right pages. The headline claim is 98.7% accuracy on a financial QA benchmark where vector RAG supposedly scores 50%. Both of those numbers need an asterisk, and the honest version of this tool is more interesting than the marketing version.

Similarity is not relevance

The pitch, in the project's own words: "Vector-based RAG retrieves by semantic similarity. But similarity is not relevance, what retrieval actually needs is relevance, and relevance requires reasoning." The elaboration is the part worth keeping: "similarity search misses what is relevant but not similar, and returns what is similar but not relevant."

Anyone who has shipped a RAG system over long structured documents knows the failure. You ask about the 2023 operating margin and get back three chunks that all contain the phrase "operating margin" from the wrong fiscal year, because cosine distance has no idea which section of a 200-page filing it is standing in. Chunking throws away the structure that would have answered the question, then embeddings try to reconstruct meaning from the debris.

We have covered two responses to this complaint already. PixelRAG keeps retrieval but drops the text parser, rendering pages to images so tables survive. PageIndex keeps the text and drops the embeddings. The Supabase pgvector stack we walked through is the thing both are arguing against, and it is worth saying up front that it remains the right answer for a lot of workloads. This article is about when it is not.

What it actually builds

Indexing produces a JSON tree shaped like a table of contents. Each node carries a title, a node_id, a start_index and end_index giving its physical page range, and an LLM-written summary. Note what is missing: there is no node text and no embedding anywhere in the structure. The shipped default is literally if_add_node_text: "no".

Here is the detail that undercuts the branding, and it comes from the project's own documentation. The default indexer since August 2026 is called Flash, and Flash does not use an LLM to build the tree: "The tree structure itself is extracted from the document layout without an LLM; the index model only summarizes and refines it."

Flash is roughly forty modules of classical PDF engineering. A character-level PDFium parser with font and glyph tables, column and gutter detection, header and footer classification, style-based heading detection, clique-based outline assembly. It reports where the structure came from: bookmarks if the PDF had an embedded outline, detected if it inferred one from layout, hybrid, or pages if it found no hierarchy at all. So the flagship reasoning-based retrieval system builds its index with heuristics, and the reasoning happens later. That is not a criticism. Deterministic layout analysis is cheap, fast and reproducible, which are three things LLM calls are not. It is just not what the word "reasoning" in the tagline implies.

What happens when you ask a question

Most write-ups about PageIndex quote a prompt that hands the LLM the whole tree and asks it to return a list of node IDs. That prompt is real, it is from the original 2025 implementation, and it is no longer what ships.

The current version is an agent tool loop. The model gets two functions and drives itself. The instruction is explicit: for documents over 20 pages, call get_document_structure() first to locate relevant sections, then call get_page_content() with targeted page ranges. Under 20 pages, skip straight to reading. Page arguments accept "5", "3,7,10", "5-10" or "1-3,7,9-12".

This matters for how you describe the thing. The model is not descending the tree branch by branch under some traversal algorithm. It is shown the entire outline, paginated if it exceeds a 100,000 character budget, and then asked to request page ranges. There is no PageIndex-side search. The tree is a cheap, precomputed, page-addressed map that makes an agent's read loop converge quickly, and grounding is enforced by prompt rather than by code: "Answer only from the user's PageIndex documents. Never fill a gap from general knowledge."

About that 98.7%

The number is real, it is published with its working, and it does not mean what the headline suggests.

First, provenance. The 98.7% belongs to Mafin 2.5, a closed commercial financial RAG product built on PageIndex, and it was published on 19 February 2025, roughly six weeks before the PageIndex repository existed. It is a result for a whole product, not for the retriever you can pip install.

Second, and more important, what it counts. The evaluation labels every one of the 150 questions in the FinanceBench public set into one of five buckets, and the published result files break down like this.

LabelMeaningCount
ALAligned with the benchmark's gold answer136
BEBenchmark Error, judged to be wrong in the benchmark itself6
MVAMultiple Valid Approaches5
SEDCSame Evidence, Different Conclusion1
NALNot aligned, an outright miss2

98.7% is 148 out of 150, which is everything except the two outright failures. Strict agreement with FinanceBench's own gold answers is 136 out of 150, or 90.7%. Twelve questions, eight percent of the benchmark, count as correct because the vendor's own human annotators decided the benchmark was wrong, ambiguous, or that a different conclusion was equally defensible.

That is a legitimate methodology and they published the per-question labels so you can check it, which is more than their listed competitors did. But 90.7% strict and 98.7% adjudicated are two different claims, and only one of them is on the front page.

The "vector RAG scores 50%" comparison deserves less patience. It appears only as alt text on a chart image. No embedding model, no vector database, no chunk size, no top-k, no citation anywhere in the repository or the blog post. For scale, the FinanceBench paper itself reported that GPT-4-Turbo with a retrieval system got roughly 19% of questions right, which is a different number arrived at by a different route. Treat the 50% as an uncited vendor figure.

Then there is the awkward one. In issue 518, opened on 18 September 2026 and still open, a user ran Flash against FinanceBench and reported "coverage only 50%", with pages missing from filings like 3M's 10-K. The maintainer's reply: "Yeah, for large and complex PDF, we recommend cloud instead." The open-source path, on the benchmark the project is famous for, is not what produced the famous number.

The cost model is inverted, and that is the real trade

Vector RAG is expensive to set up and nearly free to query. You pay once to embed a corpus and run a database, then each lookup is a millisecond of arithmetic. PageIndex flips both halves. Indexing is cheap, around $0.001 per page by their measurement, so a 1,000 page textbook costs a little over a dollar. Querying costs LLM tokens, every single time.

Their own open-source benchmark, 62 questions over 34 PDFs, publishes the per-question cost alongside accuracy.

Chat modelReasoning effortAccuracyCost per question
gpt-5.6-lunanone85.5%$0.0031
gpt-5.6-lunahigh96.8%$0.0036
gpt-5.6-terramedium98.4%$0.0303
gpt-5.6-solmedium100.0%$0.0810

Between one third of a cent and eight cents per question, against effectively zero for a vector lookup. At a thousand queries a day the cheap configuration costs about $3 and the expensive one about $81. That is the decision, stated plainly: you are buying accuracy with a per-query bill that scales with traffic instead of with corpus size.

Be careful with these numbers in the other direction too. The benchmark is 62 questions written by the authors, judged by the authors' own evaluator, with no baseline system of any kind, and documents the indexer refuses are excluded rather than scored as failures. To their credit they say all of that themselves.

You can point it at a local model

This is the finding that matters most if you self-host, and the README's OpenAI-flavoured examples bury it. PageIndex routes every LLM call through LiteLLM, so the model string is a LiteLLM model string. Bare names are OpenAI, and anything else takes a provider prefix. The test suite asserts that ollama/llama3 and bedrock/anthropic.claude-sonnet both work.

from pageindex import PageIndexClient

client = PageIndexClient(
    index="ollama/llama3",   # builds the tree summaries
    chat="ollama/llama3",    # reasons over it at query time
)
doc_id = client.submit_document("report.pdf")["doc_id"]
print(client.chat("What was the 2023 operating margin?", doc_id=doc_id))

Better still, you can index with no LLM at all. Flash builds the tree from layout, and the summary and optimize passes are the only parts that call a model, so turning both off gives you a pure local structural index.

pip install -U pageindex
python3 run_pageindex.py --mode flash --pdf_path document.pdf --no-summary --optimize off

That writes a tree to results/<name>_structure.json without a single API call. If you have been avoiding this project because it looked like an OpenAI wrapper, it is not one.

What is free and what is the product

The open-source repository is a real implementation, not a thin client for a paid API. It contains the full Flash indexer, the older LLM-based one, markdown support, tree optimization, a local document store, the chat loop, and integrations for the OpenAI Agents SDK and the Claude Agent SDK. It is MIT. But the line between free and paid is drawn in a specific place and you should know where.

CapabilityLocal, MITCloud
Text-based PDFsYesYes
Scanned documents, OCR, image understandingNoYes
CitationsPage levelBlock level with bounding boxes
Multi-document corpus reasoningNoYes, PageIndex File System
MCP serverNoYes

So you get single-document, text-PDF, page-cited retrieval for free. Reasoning across a whole corpus, reading anything scanned, and the MCP server are the commercial product. Cloud pricing is $0.01 per page to index, one time, plus $0.001 per page per month to keep a page queryable, with the first 1,000 active pages free and $10 of credit on signup. Retrieval carries no platform fee because you are paying your own model provider for it.

One more caveat on the word vectorless. It is accurate about passage retrieval. The Cloud document browser exposes a sort mode that "orders documents by semantic relevance to query", and whether that is embedding-backed is not stated in the open-source code. The claim is about how it finds the passage, not necessarily how it picks the document.

PageIndex, RAPTOR, and the stack you already run

Tree-structured retrieval is not new. RAPTOR proposed it in 2024, and the comparison is the most useful way to understand what is actually novel here.

RAPTORPageIndexpgvector stack
Where the hierarchy comes fromInduced by clusteringRead from document layoutNone, flat chunks
EmbeddingsYes, build and queryNo, at passage levelYes
Query-time selectionSimilarity over nodesLLM tool callsSimilarity over chunks
Cost per queryLowHighLowest
Citable page numbersNoYesDepends on metadata

RAPTOR recursively embeds, clusters and summarizes chunks to construct a tree from the bottom up. The one-line difference: RAPTOR builds a hierarchy the document does not have, while PageIndex reads the one it does. That is PageIndex's advantage and its single biggest liability, because it is entirely parasitic on good document structure. Give it a well-bookmarked filing and it is excellent. Give it a scanned fax, and local mode refuses the document outright once it exceeds ten pages with no hierarchy detected.

It also makes the marketing dichotomy a false binary. Tree versus vectors was never the choice, because RAPTOR is tree and vectors, and it came first. The genuinely interesting claim is narrower: that a page-addressed structural map plus an agent loop beats similarity search on documents whose structure carries meaning.

Who should use it

Reach for PageIndex when your documents are long, structured and professional, when you query one at a time, and when a citation you can point a human at matters more than the per-query bill. Financial filings, legal contracts, regulatory submissions, technical manuals. The page-level citations and the traceable path from question to pages are worth real money in those settings.

Stay on your pgvector stack when you have a large corpus rather than a few documents, when query volume is high enough that cents per question matter, or when your content is unstructured prose where there is no table of contents to read. And if your problem is that scanned pages and tables are destroying your parser, PixelRAG or an OCR step is closer to the actual fix than swapping your retriever.

Two small things that say a lot. VectifyAI's first public repository, from 2023, was a tool for fine-tuning OpenAI embedding models, so the company now selling vectorless started on vectors. And the README's jab at vector search as "vibe retrieval" is almost word for word the criticism its own top Hacker News commenter aimed at PageIndex when it launched.

Sources and further reading

Ten minutes: install it, run the Flash indexer on a PDF you know well with summaries and optimization switched off, and open the resulting JSON. No API key, no cost, no model. If the tree matches the document's real structure you have learned that this approach will work for your corpus, and if it comes back as a flat list of pages you have learned the opposite, which is the more valuable answer.

Tested on: not independently tested. Architecture details, node structure, the Flash indexer behaviour and all install commands are quoted from the project's own repository and documentation. Every benchmark figure is author-reported: the 98.7% and its label breakdown are from VectifyAI's published FinanceBench result files, which we tallied, and the per-question costs are from VectifyAI's own open-source benchmark. No independent reproduction of either exists at the time of writing. Repository statistics were read from the GitHub API on the date below, when the project stood at 36,116 stars.
Date checked: 2026-09-28

Prev Article
TensorFold vs MTPLX vs mlx-dspark
Next Article
Pinokio Computer

Related to this topic: