Reranking
Cross-encoders improving precision more than a bigger embedding model.
3 to work through
-
intermediate
A retrieval system's first-stage results are mediocre. Is reranking the right investment, and what does it cost?
2 min answer -
intermediate
Retrieval returns 100 candidates and a cross-encoder reranks all of them. The service must hold 300 queries per second at a 400 ms p95 budget for the whole retrieval stage. Roughly what does that rerank cost in hardware and time and does it change the design?
3 min answer -
advanced
A support assistant adds a cross-encoder reranker. Offline NDCG at 10 rises from 0.61 to 0.74 and it ships. Two weeks later, answers drawn from long runbook pages are worse than they were before reranking, while short FAQ answers improved. Nothing errors and no alert fires. What happened?
3 min answer
3 terms in this topic
Candidate Recall Ceiling
The share of queries whose correct passage appears anywhere in the first-stage candidate set - the hard upper bound on every later stage, since reran…
patternReranking
Retrieving a wide candidate set cheaply, then reordering it with a more expensive model that scores each candidate against the query directly.
patternTwo-Stage Retrieval
Retrieving a broad candidate set cheaply and then reordering it with an expensive, more accurate model.
Neighbouring topics
AI-Era Architecture
General material on architecting systems that include models.
LLM Application Architecture
The shape of a production system with a model in the request path.
RAG Architecture
Retrieval, grounding, citation and the permissions RAG can enforce.
Vector Databases
Approximate nearest-neighbour search, filtering and re-indexing.
Embeddings
Dense representations, model coupling and the migration they imply.
Chunking & Retrieval
Structure-aware splitting, hybrid search and why chunking dominates quality.
Model Selection
Capability, latency, cost and the evaluation that decides between them.
AI Gateways
Centralised routing, keys, quotas, caching, logging and safety policy.
Prompt & Version Management
Prompts as reviewed, versioned, evaluated production configuration.
Agent Architectures
Loops, planning, memory and the boundaries an agent must not cross.
Tool Calling
Typed tool interfaces, narrow parameters and per-tool authorisation.
Multi-Agent Systems
Coordination, hand-off and whether more agents actually help.
LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
AI Observability
Logging prompts, versions, retrieved context and cost per request.
Guardrails
Deterministic checks on input and output that fail closed.
Prompt Injection Defence
Breaking the private-data, untrusted-input, outbound-channel combination.
AI Cost Management
Token accounting, routing, caching and the context-window budget.
Human in the Loop
Gating by reversibility and blast radius, and avoiding approval fatigue.
ML Platform
Feature stores, training pipelines, registries and deployment.