Vector Databases
Managed retrieval

A vector database with the retrieval pipeline built in.

Extraction, embedding, hybrid search, reranking, and cited answers: one managed layer. You choose the models and tune every stage. You don’t host or scale any of it.

No sign-up. Ask the corpus.
What it is, technically

Six stages, each a named component.

Not one model doing all the work. Every stage below runs on named infrastructure. Some of it yours to choose, all of it managed for you.

Extraction

Each page is read to text, figures and tables included, so everything on it can be searched.

MistralGoogleLlamaIndex

Structural chunking

Documents are split along their own structure, boilerplate dropped and tables kept whole.

StructuralOverlap-awareTables kept whole

Embedding

Each chunk becomes a vector that captures its meaning, using the model you choose.

Voyage AIGoogleQwenOpenAICohereBAAI

Hybrid search

Semantic and keyword search run together for a wider candidate pool, or either can run alone.

Qdrant

Reranking

A cross-encoder rereads every candidate against your question and reorders them, the single biggest lever on quality.

CohereQwenVoyage AINVIDIA

Cited answer

The answer is generated only from the retrieved passages, each citation pointing to its exact source region.

GroundedInline citationsRefuses when unsure
Live demo

Ask the corpus. Watch the whole machine work.

Fourteen enterprise documents, indexed and searchable. Every score, rank, and millisecond below is what the system returned for your query.

Ask the corpus…
Answer
Suggested
Retrieval
TOP-K08
Run a query to retrieve and rank passages.
Vector space · PCA
Your question will appear as a point in this space.
Rate card

Priced like infrastructure.

Two charges, and they stack. The pipeline is metered per unit, so you pay for what you index and query. Dedicated capacity is optional: a leased instance sized by chunk count, billed by the hour, and free while hibernated.

Metered · the pipeline
Extraction
MistralMistral OCR$1.00 / 1K pages
GoogleDocument AI$1.50 / 1K pages
LlamaIndexLlamaParse$12.50 / 1K pages
Embedding
CohereEmbed 4$0.12 / 1M tokens
OpenAIEmbedding 3 Large$0.13 / 1M tokens
GoogleGemini Embedding$0.15 / 1M tokens
Voyage AIVoyage 4$0.06 / 1M tokens
QwenQwen3 Embedding 8B$0.01 / 1M tokens
BAAIBGE-M3$0.01 / 1M tokens
Reranking
CohereRerank 4 Pro$2.00 / 1K queries
CohereRerank 4 Fast$1.00 / 1K queries
CohereRerank 3.5$2.00 / 1K queries
Voyage AIVoyage Rerank 2.5$0.05 / 1M tokens
QwenQwen3 Reranker 8B$0.05 / 1M tokens
NVIDIALlama Nemotron Rerank$0.01 / 1M tokens
Answer generation

Write the cited answer with any language model, from fast and cheap to frontier. Browse every model on Remy →

Capacity · shared or dedicated
SizeCapacityPer hour~ Monthly
Sharedup to ~1M chunks$0included
Small2.4M chunks$0.11~$78
Medium6.4M chunks$0.21~$156
Large24M chunks$0.47~$343
XL48M chunks$0.94~$686

Shared capacity is the default, included in the metered pricing above at no extra charge, and covers most corpora. Dedicated tiers are optional: a leased instance isolated from the shared pool, billed hourly, never unloaded, and $0 while hibernating; more capacity than XL is available on request.