quivr is a self-hosted application that durably ingests continuous content streams and makes them searchable within seconds, with background enrichment and topic subscriptions. It combines lexical, semantic and hybrid search over ingested corpora with alerting on newly searchable content.
Project overview
It addresses durable ingestion of continuous content streams with writes acknowledged only once committed, search results available within seconds, and subscription alerts that turn newly searchable Versions into unique Matches — an end-to-end pipeline from ingestion to monitoring.
Project type
RAG · AI Search
Use cases
Knowledge Q&A · Search & Research
Deployment
Refer to project documentation
License
MIT
Best for
Developers who need to durably ingest continuous content streams and make them searchable and monitorable over time, and who can run and operate the stack themselves.
Self-hosters comfortable with a medium-difficulty setup involving Go, Docker Compose v2, Python 3, Node.js 22+, and jq.
Key capabilities
Lexical, semantic, and hybrid search with canonical rehydration and access rechecks on every hit.
Inline text, bounded batches with per-entry outcomes, verified uploads, and structured Manifests; writes are acknowledged only once committed.
Saved Queries and Subscriptions turn newly searchable Versions into unique Matches.
Limitations and risks
The project is at an evaluation stage; its API is v0 and may change. It should not be run in production yet.
Setup and operation require coding ability; the documented path uses make targets and a toolchain of Go, Docker Compose v2, Python 3, Node.js 22+, and jq.
If the optional TypeSafe Jev classifier is enabled for the described alerts, article text is sent to TypeSafe — an external data flow to consider before adoption.
Getting started
Install Go 1.27.1, Docker with Compose v2, Python 3, Node.js 22+, and jq. Then run `make dev`, followed by `make demo`, and open http://127.0.0.1:5183. Note that the first run downloads roughly 1 GB of images and model. Setup difficulty is medium.
Embeddings are served locally by TEI with the pinned multilingual-e5-small model (384 dimensions).
The TypeSafe Jev classifier is optional and used for the described alerts; when used, article text is sent to TypeSafe. The system is not local-only end-to-end when this optional classifier is enabled.
Evidence and sources
README: An open-source engine that turns continuous content streams into search and monitoring.
README: Status: evaluation stage. The API is `v0` and may change without notice. Do not > run it in production yet.
README: **Durable before fast.** A write is acknowledged only once it is committed; outages delay processing, they never lose or silently fail content.
README: **Durable, idempotent ingestion**: inline text, bounded batches with per-entry outcomes, verified uploads (presigned PUT + checksum confirm), structured Manifests,
README: **Search**: lexical, semantic and hybrid, with canonical rehydration and access rechecks on every hit