Reranking
Models that reorder search results by relevance to a query.
The cross-encoder models that rerank retrieved passages against a query, the single biggest lever on retrieval quality. Each entry shows how many documents it can score at once and how much text it reads, with a per-million-token price. The vector database applies one by default. Name a specific model when accuracy or language coverage matters.
9 models
CCohere Rerank 3.5Reranking
Cohere
$0.0020/ search
1,000
Max docs
4.1K tok
Context
CCohere Rerank 4 FastReranking
Cohere
$0.0010/ search
1,000
Max docs
4.1K tok
Context
CCohere Rerank 4 ProReranking
Cohere
$0.0020/ search
1,000
Max docs
4.1K tok
Context
NLlama Nemotron Rerank VL 1B v2Reranking
Nvidia· viaDeepInfra
$0.010/ 1M tokens
1,000
Max docs
10.2K tok
Context
QQwen3 Reranker 0.6BReranking
Qwen· viaDeepInfra
$0.010/ 1M tokens
1,000
Max docs
32.8K tok
Context
QQwen3 Reranker 4BReranking
Qwen· viaDeepInfra
$0.025/ 1M tokens
1,000
Max docs
32.8K tok
Context
QQwen3 Reranker 8BReranking
Qwen· viaDeepInfra
$0.050/ 1M tokens
1,000
Max docs
32.8K tok
Context
VVoyage Rerank 2.5Reranking
Voyage AI
$0.050/ 1M tokens
1,000
Max docs
32K tok
Context
VVoyage Rerank 2.5 LiteReranking
Voyage AI
$0.020/ 1M tokens
1,000
Max docs
32K tok
Context