Newsletter image

Subscribe to the Newsletter

Join 10k+ people to get notified about new posts, news and tips.

Do not worry we don't spam!

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

Search

GDPR Compliance

We use cookies to ensure you get the best experience on our website. By continuing to use our site, you accept our use of cookies, Privacy Policy, and Terms of Service.

Firecrawl - Document Understanding

anydoc

Firecrawl's Rust converter handles 14 formats including legacy Office binaries, at a 4.4 ms median. No local OCR, which is the catch worth knowing.

License MIT
License MIT
TL;DR
  • 14 formats across Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV and PDF
  • 4.4 ms median conversion, about 100x faster than driving LibreOffice
  • No local OCR: scanned pages error out and route to a hosted service

Every retrieval system in this series starts with the same unglamorous step: turning whatever the user gave you into text you can index. PageIndex reads text PDFs and refuses the rest. PixelRAG sidesteps parsing by rendering to pixels. Anydoc is the piece upstream of both: a Rust library that converts fourteen document formats to clean Markdown at a median of 4.4 milliseconds, which is roughly a hundred times faster than driving LibreOffice. It is MIT, it is from Firecrawl, and it is the least exciting and most useful thing we have covered this week.

Fourteen formats, one output

The format list is the product. Word in .doc, .docx and .docm. PowerPoint across .ppt, .pps, .pot, .pptx, .pptm, .ppsx and .ppsm. Excel in .xls, .xlsx, .xlsm and .xlsb. OpenDocument text, spreadsheet and presentation. Plus RTF, EPUB, CSV and PDF.

Note how much of that is legacy. The pre-2007 binary Office formats are still everywhere in corporate document stores, and they are exactly what most converters quietly skip. Output is GitHub-Flavored Markdown, the same shape regardless of what went in.

The benchmark, and it is a real one

Firecrawl published a comparison across 100 real-world documents in 14 formats, scored on completeness, structure, formatting and cleanliness.

ToolFormats coveredOverall score
anydoc14 of 1481
unstructured840 to 70 band
markitdown640 to 70 band
pandoc540 to 70 band

The sub-scores are completeness 87, structure 79, formatting 78, cleanliness 81. On speed, a 4.4 ms median against LibreOffice's 1,129.5 ms on a Ryzen 9 9950X3D. This is a vendor benchmark like every other one we audit, but it is a more checkable species than most: the corpus is described, the competitors are named and freely available, and you can rerun it on your own documents in an afternoon.

The coverage column is the honest headline rather than the score. Beating pandoc on quality is arguable. Handling nine formats pandoc does not handle at all is not.

The catch is OCR

Anydoc does not do OCR locally. A scanned or image-only PDF fails with a NeedsOcr error rather than returning mush, which is the right behaviour, but it means the hard half of document ingestion is out of scope.

The escape hatch routes to Firecrawl's hosted service: ocr: 'hosted' in Node, ocr="hosted" in Python, --ocr hosted on the CLI. So the MIT library covers text-bearing documents and the scanned ones go to a network service run by the company that wrote it. That is a reasonable split and worth knowing before you build a pipeline on the assumption that anydoc handles everything you throw at it.

If scanned documents are your actual problem, an OCR model is the piece you need, and anydoc sits beside it rather than replacing it.

Using it

Three bindings plus WebAssembly for the browser, which is unusual reach for a converter.

cargo add anydoc
npm install @firecrawl/anydoc
pip install firecrawl-anydoc

The API is one function in every language.

import anydoc

markdown = anydoc.to_markdown("report.docx")
markdown = anydoc.to_markdown_bytes(data)

The bytes variant matters more than it looks. If you are ingesting uploads in a web service you never want to touch the filesystem, and a converter that only takes paths forces you to write temp files you then have to clean up.

Where it fits in the retrieval stack

This series has been about what happens after ingestion: how you index, how you retrieve, whether you need embeddings. Anydoc is the step before all of it, and it is the step where most real pipelines lose their tables.

StageToolHandles
Ingestanydoc14 text-bearing formats to Markdown
Ingest, scannedOCR model or hosted serviceImages and scans
IndexPageIndexTree over document structure
Retrieve without parsingPixelRAGPixels, when parsing destroys the content

Note the interaction with PageIndex specifically. PageIndex local mode takes text PDFs and Markdown, and refuses documents with no detectable structure. Anydoc produces Markdown from a Word file with its heading hierarchy intact. Chaining them gives you a tree index over a .docx, which neither tool does alone.

Who should use it

Use it if you ingest documents users send you, especially in an enterprise setting where half the corpus is .doc and .xls from 2009. Use it if conversion latency shows up in your traces, because 4.4 ms effectively removes that stage from the budget. Use it if you want one converter instead of a pipeline that branches on file extension.

Do not expect it to solve scanned documents, and do not assume the benchmark transfers to your corpus without checking. One caveat on project health: the repository has not been pushed since 28 August, which is a month of quiet on a 22,000-star project. Not alarming for a library with a narrow job, but worth noticing before you depend on it.

Sources and further reading

Ten minutes: pip install it, point it at the ugliest document in your corpus, the one with merged cells and a header nobody can explain, and read the Markdown that comes out. Conversion quality is not something a benchmark can tell you about your own files, and this is a ten-second test per document.

Tested on: not independently tested. Format coverage, benchmark scores, speed figures and the competitor comparison are Firecrawl's own published numbers, measured on their hardware and corpus. The OCR limitation and install commands are quoted from the project documentation. Repository statistics were read from the GitHub API on the date below.
Date checked: 2026-09-30

Prev Article
Jeff
Next Article
CodeGraph

Related to this topic: