Indexify
Open-source data extraction and indexing engine for RAG applications.
About
Indexify is an open source data extraction and indexing engine from Tensorlake, built for teams assembling retrieval-augmented generation systems. It ingests unstructured content such as documents, images, and audio and runs extraction pipelines over it, producing structured data and embeddings that downstream applications can query, with a focus on keeping indexes updated continuously as new content arrives rather than through periodic batch jobs. The engine is designed to scale out on Kubernetes for production workloads, placing it in the infrastructure layer between raw data sources and the vector databases or LLM applications that consume the results. Released under the Apache 2.0 license, it targets engineers who need ongoing ingestion and extraction rather than one-off preprocessing scripts. Tensorlake, the company behind the project, has since shifted its public focus toward sandbox infrastructure for AI agents, and the original Indexify repository is no longer available at its previous address.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- RAG & Document Retrieval
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Added
- Apr 3, 2026
Related Tools
All-in-one desktop and Docker app for private LLM chat with your documents.
Modular open-source RAG framework for building production document retrieval applications.
Open-source embedding database for AI applications
Web scraping API that turns websites into clean LLM-ready markdown.
Local RAG capabilities in GPT4All for chatting with documents privately.
Cloud-native vector database for scalable similarity search