Engineering Hybrid RAG with MongoDB Atlas Vector Search and Voyage AI
This engineering pattern uses MongoDB Atlas Vector Search with Voyage AI embeddings and optional reranking to build a reusable retrieval API for enterprise knowledge workloads.
Published Oct 2, 2026
A production Retrieval-Augmented Generation (RAG) system needs more than a vector database and an LLM. The retrieval layer must ingest changing enterprise content, combine semantic and keyword signals, enforce metadata-based access controls, evaluate retrieval quality, and return context with predictable latency. This engineering pattern uses MongoDB Atlas Vector Search with Voyage AI embeddings and optional reranking to build a reusable retrieval API for enterprise knowledge workloads.
In a reference pilot using more than 100,000 chunks, the retrieval API maintained P95 latency below 1.5 seconds, excluding LLM generation time, while Recall@5 reached at least 70% on the pilot evaluation set. The ingestion pipeline also supported incremental checksum-based synchronization and processed the pilot corpus with a processing error rate below 2%. These measurements are workload-specific engineering results, not universal product guarantees.
1. The engineering problem: enterprise knowledge is fragmented
Enterprise knowledge rarely lives in one system. Documents may sit in object storage or collaboration platforms while operational context remains in application databases. A RAG implementation therefore has to solve two connected problems: create a reliable retrieval index across heterogeneous sources, and expose that index through a controlled API that applications and language models can use consistently.
A weak retrieval layer can return incomplete or irrelevant context, ignore authorization boundaries, or require expensive full re-indexing whenever source data changes. The architecture should treat retrieval as an independent production service rather than as a small helper inside a chatbot.
2. Retrieval as an independent AI data layer
The core pattern separates retrieval from the downstream language model. MongoDB Atlas stores chunks, metadata and embeddings and provides the vector-search index. Voyage AI provides the embedding layer and can optionally rerank top candidates. A retrieval API then coordinates vector retrieval, keyword signals, metadata filters and access-control rules before context is returned to the model.
| Layer | Technology / pattern | Engineering role |
|---|---|---|
| Data and vector store | MongoDB Atlas | Store chunks, metadata and embeddings; serve vector-search indexes. |
| Embedding | Voyage AI embeddings | Generate semantic vectors for multilingual or domain content. |
| Reranking | Voyage AI Reranker (optional) | Reorder top candidates when the workload benefits from a second relevance stage. |
| Ingestion | Incremental sync pipeline | Extract, chunk, checksum, deduplicate and update changed content. |
| Retrieval API | FastAPI / NestJS pattern | Coordinate vector + keyword retrieval, metadata filters, ACL checks and response shaping. |
3. Why hybrid retrieval matters
Semantic retrieval is useful when users express concepts differently from the source text, while keyword signals remain valuable for identifiers, product names, policy terms and other exact vocabulary. A hybrid retrieval layer allows the application to combine these signals instead of forcing every query through one search mode.
The retrieval API is also the correct place to normalize filters and return a consistent result contract to downstream applications. This keeps model integration independent from the details of indexing and retrieval tuning.
4. Enforce metadata ACLs before context reaches the LLM
Authorization should be enforced during retrieval, not after sensitive context has already been assembled. In this pattern, metadata-based ACL filters are applied directly in the vector-query path so that only authorized chunks can be returned for prompt or context construction.
This does not replace the enterprise identity and authorization model. It extends those controls into the retrieval layer so application permissions and RAG context selection remain aligned.
5. Incremental ingestion instead of full re-indexing
Knowledge sources change continuously. The ingestion pipeline therefore tracks source versions or checksums and updates only content that has changed. Extraction, chunking and deduplication are treated as repeatable pipeline steps, while failed items can be retried or routed to a dead-letter path for investigation.
The pilot processed more than 100,000 chunks with a processing error rate below 2% and used incremental synchronization to avoid rebuilding the complete corpus for unchanged content.
6. Measure retrieval separately from generation
RAG evaluation becomes difficult when retrieval latency and model-generation latency are mixed into one number. The reference pilot measured the retrieval API separately: P95 remained below 1.5 seconds for the pilot workload, excluding the time required by the language model to generate a final answer.
Retrieval quality was evaluated with Recall@5 on a workload-specific question set and reached at least 70%. That figure describes whether relevant source material was retrieved into the top five results on the pilot evaluation set. It is not a measure of final answer correctness and should not be generalized to another corpus without re-evaluation.
| Pilot criterion | Observed / accepted result | Interpretation |
|---|---|---|
| Corpus | 100,000+ chunks | Reference pilot scale. |
| Retrieval latency | P95 < 1.5 s | Retrieval API only; LLM generation excluded. |
| Retrieval quality | Recall@5 >= 70% | Pilot evaluation set; workload-specific. |
| Ingestion reliability | Processing error rate < 2% | Pipeline includes retry / dead-letter handling. |
| Data synchronization | Incremental sync | Checksum/version-based update path. |
7. Where reranking fits
Reranking is useful when first-stage retrieval returns a candidate set that still needs a stronger relevance model. Voyage AI Reranker can be inserted after initial retrieval to reorder the top candidates before context assembly. Because reranking adds another network and compute step, it should be evaluated against latency and relevance requirements rather than enabled automatically.
8. Production considerations beyond the pilot
- Define a representative evaluation set before tuning retrieval.
- Measure retrieval latency separately from model-generation latency.
- Track chunking, embedding and index versions so experiments remain reproducible.
- Apply authorization filters before context assembly.
- Use incremental ingestion and operational visibility for failed or stale content.
- Evaluate whether reranking materially improves relevance for the workload.
- Treat benchmark numbers as workload-specific; re-test when corpus size, filters, embedding strategy or network path changes.
9. Key takeaways
- RAG quality depends on the retrieval system, not only the language model.
- MongoDB Atlas can combine operational data, metadata, embeddings and vector retrieval in one data-platform pattern.
- A separate retrieval API creates a reusable interface for different LLMs and applications.
- Metadata ACLs should constrain retrieval before context is exposed to the model.
- Incremental synchronization is essential when enterprise knowledge changes continuously.
- Recall and latency need to be measured explicitly and interpreted within the tested workload.
Designing a production RAG retrieval layer?
TitanBases can help assess data sources, retrieval architecture, MongoDB Atlas Vector Search, evaluation strategy, metadata controls and the path from pilot validation to production.
Discuss a RAG Retrieval Architecture