A high-performance RAG system built in Rust.
Chunk, embed, store, and query your documents — then stream grounded LLM answers in real time.
Rustriever is a complete Retrieval-Augmented Generation system written in Rust. It ships as two binaries that share a common library:
rustriever-server— A production-ready REST API built on Actix-web. Handles document ingestion, semantic search, JWT auth, and streaming LLM responses over SSE.rustriever— An interactive CLI REPL for chatting and querying your Qdrant collections directly from the terminal, no server required.
The entire pipeline — upload, chunk, embed, store, search, generate — runs in a single async Rust process with no Python dependencies.
- Full RAG pipeline — upload, chunk, embed, store, search, and stream answers from one binary
- Two API keys to get started — Groq for LLM completions, Cohere for embeddings
- SSE streaming — tokens stream to the client as they are generated; sources are delivered as a leading SSE event before generation begins
- Cohere Embed v3 — asymmetric encoding with separate
search_documentandsearch_queryinput types for higher retrieval accuracy - Concurrent ingestion — ZIP archives are unpacked and processed across 8 parallel workers with batched Cohere calls (100 chunks/batch) and batched Qdrant inserts (50 chunks/batch)
- Qdrant HNSW search — cosine similarity with approximate nearest neighbor indexing
- JWT auth + Argon2id — stateless authentication with secure password hashing
- Interactive Swagger UI — auto-generated OpenAPI 3.0 docs at
/swagger-ui/ - Built-in chat frontend — a minimal SSE-powered UI at
/static/chat.htmlfor testing RAG and plain chat - Interactive CLI REPL — switch between plain chat and RAG mode, change collections, and adjust retrieval limits at runtime
┌──────────────┐ ┌──────────────────────────────────────────────────┐
│ │ SSE │ Rustriever │
│ Client / │◄─────►│ │
│ Frontend │ │ ┌──────────┐ ┌────────────┐ ┌──────────────┐ │
│ │ │ │ Actix-web│ │ Cohere │ │ Groq │ │
└──────────────┘ │ │ Router │─►│ Embeddings │ │ LLM Client │ │
│ └──────────┘ └──────┬─────┘ └──────┬───────┘ │
│ │ │ │
│ ┌───────────────▼───────────────┘ │
│ │ │
│ ┌────▼────────┐ ┌──────────────┐ │
│ │ Qdrant │ │ PostgreSQL │ │
│ │ HNSW Index │ │ Users/Auth │ │
│ └─────────────┘ └──────────────┘ │
└──────────────────────────────────────────────────┘
- Rust 1.70+ — install via rustup
- Docker & Docker Compose — for PostgreSQL and Qdrant
- Groq API key — console.groq.com
- Cohere API key — dashboard.cohere.com
git clone https://github.com/your-org/rustriever
cd rustriever
cp .env.example .envEdit .env with your credentials:
# ── Required ──────────────────────────────────────────────
GROQ_API_KEY=gsk_your-groq-key-here
COHERE_API_KEY=your-cohere-key-here
# ── Infrastructure (defaults work with docker-compose) ────
DATABASE_URL=postgres://rustriever:rustriever@localhost:5432/rustriever
JWT_SECRET=change-me-to-a-long-random-string
QDRANT_URL=http://localhost:6334
# ── Embedding model ───────────────────────────────────────
EMBEDDING_MODEL=embed-english-light-v3.0
EMBEDDING_DIMENSION=384The CLI only requires
GROQ_API_KEY(plusCOHERE_API_KEYfor RAG mode).DATABASE_URLandJWT_SECRETare only needed by the server.
docker compose up -dStarts PostgreSQL 16 and Qdrant 1.13.
API server:
cargo run --bin rustriever-serverThe server starts at http://127.0.0.1:8080. Database migrations run automatically on first boot.
CLI REPL:
cargo run --bin rustriever # Plain chat mode
cargo run --bin rustriever -- --collection docs # Start in RAG mode- Chat UI → http://localhost:8080/static/chat.html
- Swagger UI → http://localhost:8080/swagger-ui/
The CLI provides an interactive shell for chatting and querying Qdrant collections without running the full server.
Commands:
/chat Switch to plain chat mode
/rag <collection> Switch to RAG mode using the given collection
/limit <n> Set number of retrieved chunks (default: 5)
/help Show available commands
/quit Exit
File upload (up to 2 GB, streamed to disk)
→ Text extraction (PDF via pdf-extract / TXT via UTF-8)
→ ZIP? Unpack and process entries concurrently (8 workers)
→ Word-level chunking (configurable size + overlap)
→ Batch embed via Cohere (100 chunks/call, input_type=search_document)
→ Batch insert into Qdrant (50 chunks/insert)
User question
→ Embed query via Cohere (input_type=search_query)
→ Qdrant HNSW search (top-K, cosine similarity)
→ Inject retrieved chunks as system context
→ Stream LLM answer via SSE (Groq)
→ Sources emitted as leading "event: sources" SSE event
All endpoints are under /api/v1.
| Method | Endpoint | Description |
|---|---|---|
POST |
/documents/upload |
Upload .txt, .pdf, or .zip — chunks, embeds, stores in Qdrant |
POST |
/documents/search |
Semantic search across embedded documents |
| Method | Endpoint | Description |
|---|---|---|
POST |
/chat |
Single-turn LLM completion |
POST |
/chat/stream |
SSE-streamed LLM completion |
POST |
/chat-rag |
RAG: retrieve context then generate answer |
POST |
/chat-rag/stream |
SSE-streamed RAG (sources event + LLM tokens) |
| Method | Endpoint | Description |
|---|---|---|
POST |
/users/register |
Create account (Argon2id-hashed password) |
POST |
/users/login |
Get a JWT bearer token |
GET |
/users/me |
Current user profile (requires Bearer token) |
| Method | Endpoint | Description |
|---|---|---|
GET |
/health |
Liveness check |
Full request/response schemas with examples are in the interactive Swagger UI.
| Variable | Description | Default |
|---|---|---|
GROQ_API_KEY |
Groq API key for LLM completions | required |
COHERE_API_KEY |
Cohere API key for embeddings | required |
DATABASE_URL |
PostgreSQL connection string | required (server) |
JWT_SECRET |
Secret for signing JWT tokens | required (server) |
QDRANT_URL |
Qdrant gRPC endpoint | http://localhost:6334 |
LLM_MODEL |
Groq model for chat completions | openai/gpt-oss-20b |
EMBEDDING_MODEL |
Cohere embedding model name | embed-english-light-v3.0 |
EMBEDDING_DIMENSION |
Vector dimensionality (must match model) | 384 |
CHUNK_SIZE |
Words per chunk | 500 |
CHUNK_OVERLAP |
Overlap words between consecutive chunks | 50 |
HOST |
Server bind address | 127.0.0.1 |
PORT |
Server port | 8080 |
RUST_LOG |
Log level | info |
src/
├── bin/
│ ├── server.rs # API server entry point
│ └── cli.rs # Interactive REPL entry point
├── lib.rs # Shared library re-exports
├── config.rs # Env-based config via serde + envy
├── routes.rs # Route registration (public + JWT-protected)
├── errors.rs # Unified AppError → HTTP response mapping
├── handlers/
│ ├── chat.rs # /chat, /chat/stream, /chat-rag, /chat-rag/stream
│ ├── documents.rs # /documents/upload, /documents/search
│ ├── users.rs # /users/register, /users/login, /users/me
│ └── health.rs # /health
├── schemas/
│ ├── requests.rs # Validated request DTOs (serde + validator)
│ └── responses.rs # Response DTOs with utoipa OpenAPI schemas
├── services/
│ ├── llm.rs # Groq LLM client (sync + SSE streaming)
│ ├── embeddings.rs # Cohere Embed v3 client
│ ├── qdrant.rs # Qdrant vector DB client
│ ├── document.rs # Text extraction, ZIP unpacking, chunking
│ └── password.rs # Argon2id hashing & verification
├── middleware/
│ └── auth.rs # JWT decode + Claims injection
└── db/
├── models.rs # SQLx row types
└── repositories/
└── users.rs # User CRUD queries
migrations/ # Auto-applied SQL migrations
static/
└── chat.html # Built-in SSE chat + RAG frontend
docker-compose.yml # PostgreSQL 16 + Qdrant 1.13
| Layer | Technology | Role |
|---|---|---|
| Runtime | Rust + Tokio + Actix-web 4 | Async web server |
| LLM | Groq (openai/gpt-oss-20b) |
Chat completions + SSE streaming |
| Embeddings | Cohere Embed v3 | Asymmetric document + query vectorization |
| Vector DB | Qdrant 1.13 (HNSW, cosine similarity) | Approximate nearest neighbor search |
| Database | PostgreSQL 16 + SQLx | Users, auth, compile-time checked queries |
| Auth | JWT + Argon2id | Stateless auth with secure password storage |
| Docs | utoipa → OpenAPI 3.0 → Swagger UI | Auto-generated interactive API docs |
| Ingestion | pdf-extract + zip crate | PDF text extraction, ZIP archive processing |
| Infra | Docker Compose | One-command PostgreSQL + Qdrant setup |
# Run the API server
cargo run --bin rustriever-server
# Run the CLI REPL
cargo run --bin rustriever
# Debug logging
RUST_LOG=debug cargo run --bin rustriever-server
# Run tests
cargo test
# Production build
cargo build --releaseMIT