Article

Reranker Production Evaluation: Guarding Retrieval Quality with Hard Negatives, NDCG@10, and Shadow Re-ranking

Rerankers often boost retrieval quality, but model upgrades can introduce domain regressions, tail latency, and candidate truncation issues. This article presents a production evaluation and release methodology—from hard negative construction and NDCG@10 offline gates to shadow re-ranking validation—to safely ship new ranking models.

Reranker Production Evaluation: Guarding Retrieval Quality with Hard Negatives, NDCG@10, and Shadow Re-ranking

Executive Summary

This article focuses on production evaluation and safe release of Rerankers. The core topic isn’t embeddings, vector indexes, or RAG context organization—it’s a more specific question: after the first-stage Retriever has recalled candidates, how do you determine whether a new Cross-Encoder / Reranker is truly worth shipping, and how do you avoid the classic “offline metrics improve, online search gets worse” trap?

One-sentence summary: Treat the Reranker as a production system—fix regression data, layer your gates, validate with shadow traffic—rather than cutting over traffic based on a single impressive offline score.

Why Rerankers Are Most Prone to “Great Offline, Broken Online”

Modern search and RAG retrieval typically don’t rank in a single pass. They use a two-stage or even multi-stage architecture:

  1. First stage (Retriever): Use BM25, dense vectors, or hybrid retrieval to quickly recall tens to hundreds of candidates;
  2. Second stage (Reranker): Use a more expensive Cross-Encoder to re-rank those candidates.

The benefit is obvious: confine the expensive model to a small candidate set and trade controlled cost for better relevance. Vespa’s phased ranking follows the same idea—a cheap first phase handles a larger set, while the complex model runs a second/global phase only on top candidates, with parameters like rerank-count capping compute.

The problem is that Reranker quality is highly dependent on the candidate distribution it sees. A model that performs well on “positive + random irrelevant documents” doesn’t necessarily distinguish the hardest real-world candidates. For example, if a user searches “how long does Apple refund take,” the truly hard negatives might be “Apple Pay refund,” “App Store refund policy,” or “bank card refund arrival time”—not an unrelated sports article.

So the production evaluation focus for Rerankers should shift from “can the model judge relevance” to:

Can it consistently rank the genuinely useful documents to the top among the similar candidates already selected by the real Retriever?

Core Principle 1: Hard Negatives Must Come from the Real First-Stage Retriever

Random Negatives Overestimate Model Capability

Randomly sampled negative documents are usually far from the query in terms of surface form, topic, and semantics, making them easy for the model to distinguish. The truly valuable negatives are Hard Negatives: documents that the first-stage retriever scored highly but human annotators labeled as irrelevant or only partially relevant.

Sentence Transformers’ CrossEncoder training and evaluation docs provide a dedicated mine_hard_negatives() function and recommend feeding the mined candidates directly to CrossEncoderRerankingEvaluator. This data has a key advantage: you can not only evaluate the new Reranker’s ranking quality but also compare the Base Retriever → Reranked gain.

A production dataset should retain at least the following fields:

{
  "query_id": "q-10293",
  "query": "how to change company invoice header",
  "retriever_version": "hybrid-v7",
  "candidates": [
    {"doc_id": "d1", "retriever_rank": 1, "label": 3},
    {"doc_id": "d2", "retriever_rank": 2, "label": 1},
    {"doc_id": "d3", "retriever_rank": 3, "label": 0}
  ]
}

The label field should ideally not be a simple 0/1 but a graded relevance scale aligned with your business, for example:

LabelMeaning
3Directly solves the query
2Partially relevant
1Weakly relevant
0Irrelevant

Only graded labels support metrics like NDCG that handle graded relevance.

Don’t Sneak “Unrecalled Positives” into the Candidate Set

This is an easy evaluation trap to overlook. CrossEncoderRerankingEvaluator can force positive documents into the reranking set by default, which stabilizes the evaluation signal but may also create a more optimistic upper bound than reality—because in production, the Reranker can only re-rank what the Retriever already recalled.

Therefore, track both metrics:

  • Oracle rerank: Ensure positives are in the candidate set to measure the Reranker’s intrinsic ranking ability;
  • End-to-end rerank: Strictly use the Retriever’s original top-K to measure real system performance.

If Oracle is high but End-to-end is low, the problem is usually not the Reranker—it’s first-stage recall.

Core Principle 2: Use NDCG@10 as the Primary Gate, but Don’t Rely on a Single Average

Sentence Transformers’ CrossEncoder Reranking Evaluator computes MRR@K, NDCG@K, and MAP, and its NanoBEIR evaluator defaults to NDCG@10 as the primary metric. Elasticsearch’s _rank_eval API also supports classic IR metrics like reciprocal rank, precision, and discounted cumulative gain.

For search systems with graded relevance labels, NDCG@10 is well-suited as the primary metric because it accounts for:

  1. Whether highly relevant documents appear;
  2. Whether highly relevant documents are ranked high enough;
  3. The discount applied to gains at lower positions.

But a production gate can’t rely on a single “overall average NDCG@10.” At minimum, slice the results by:

  • Query language: Chinese, English, mixed;
  • Query length: short keywords, natural language questions, long queries;
  • Business domain: after-sales, product, contract, technical, policy, etc.;
  • Freshness: recently added vs. long-stable documents;
  • Difficulty: single-answer, multi-document answers, ambiguous queries;
  • Retriever type: BM25, dense, hybrid;
  • Candidate size: top-20, top-50, top-100, etc.

Averages can mask severe regressions. A model might improve overall by 2% but drop significantly on the most important “Chinese after-sales” segment—such a version should not go straight to full production.

Core Principle 3: Candidate Count Is the First Control Knob for Quality, Latency, and Cost

More candidates isn’t always better for a Reranker. Cohere’s current Rerank API accepts query + documents and returns ranked results with relevance_score; the official docs explicitly recommend against sending too many documents in a single request. Long documents may also be truncated by max_tokens_per_doc. Additionally, relevance scores are query-dependent—you can’t treat absolute scores across different queries as a unified probability.

This leads to three production takeaways:

1. candidate_k Must Be Versioned

Don’t just log reranker_model=v4; log the full pipeline:

retrieval_pipeline:
  retriever_version: hybrid-v7
  candidate_k: 100
reranker_version: rerank-v4
output_k: 10
max_tokens_per_doc: 4096

The same Reranker can have completely different latency and quality on top-20 vs. top-200 candidate sets.

2. Monitor Long-Document Truncation Separately

If documents exceed the model’s input limit, the system may truncate, chunk, or run multiple inference passes. Don’t just monitor total request latency—also track:

  • truncated_document_rate;
  • avg/max tokens per candidate;
  • chunks per document;
  • actual number of candidates re-ranked per query.

3. Don’t Use a Fixed Relevance Score as a Global Threshold

A 0.8 on one query and a 0.3 on another can’t be compared as “the former is more relevant.” If your business needs relevance filtering, calibrate the threshold on your own representative query/document samples.

Core Principle 4: Add Shadow Re-ranking Before Shipping, Instead of Directly A/B Testing 5% of Users

Traditional canary testing exposes a small fraction of real users to the new version. For ranking systems, this is still too early—ranking changes can immediately impact search clicks, knowledge citations, and final answers.

A safer approach is to start with Shadow Re-ranking:

User Query
   |
   +--> Retriever --> Production Reranker --> Returned to user
   |
   +--> Shadow Reranker --> Logged only, not returned

The shadow path should reuse the same query and candidate list so that any differences can be attributed to the Reranker, not the Retriever. Log the following diff metrics:

MetricMeaning
shadow_top10_overlapTop-10 overlap
shadow_rank_correlationRanking correlation
shadow_top1_changed_rateTop-1 change rate
shadow_relevant_doc_promote_rateRate of relevant docs promoted
shadow_relevant_doc_demote_rateRate of relevant docs demoted
shadow_p95_latencyP95 latency
shadow_timeout_rateTimeout rate
shadow_truncation_rateTruncation rate

A critical caveat: Shadow alone cannot prove quality improvement. Without human relevance labels, you can only see “what changed in the ranking”—you can’t automatically know whether the change is better.

So the most effective practice is to automatically funnel high-diff queries from Shadow into the next round of human annotation:

  1. Online Shadow identifies queries with the largest Production vs. Candidate differences;
  2. Sample those query + candidate pairs;
  3. Human annotators assign graded relevance;
  4. Add to the regression dataset;
  5. The next version goes through offline gates again.

This creates a truly self-improving Reranker regression testing loop.

Engineering Implementation: A Four-Layer Release Gate

Gate 1: Data Gate

The release package must be bound to a fixed dataset version:

evaluation_dataset:
  version: search-regression-2026-08-30
  query_count: <recorded-value>
  label_schema: graded-0-3
  retriever_snapshot: hybrid-v7
  hard_negative_source: production-top-k

Never keep a single overwritten latest.csv.

Gate 2: Relevance Gate

At minimum, compare:

  • Base Retriever NDCG@10;
  • Production Reranker NDCG@10;
  • Candidate Reranker NDCG@10;
  • MRR@10;
  • Delta per core business segment.

The release condition should be written as “Candidate must not regress beyond an agreed threshold on any core segment,” not just “overall average improvement.”

Gate 3: Performance Gate

Use production-consistent candidate_k, document length distribution, and concurrency model to measure:

  • P50 / P95 / P99 latency;
  • timeout / 429 / 5xx;
  • tokens/doc;
  • candidates/query;
  • batch size;
  • GPU/CPU utilization;
  • per-query rerank cost.

Gate 4: Shadow Gate

The Shadow phase doesn’t change user results, but you must confirm:

  • No significant increase in error rate;
  • No systematic anomalies on long-document and multilingual requests;
  • Top-K ranking changes are within expectations;
  • Extreme-diff queries pass sampled human review.

Only after passing these gates should you proceed to a real canary.

Applicable Scenarios

This methodology is especially suited to:

  • Enterprise knowledge search;
  • Second-stage retrieval ranking in RAG;
  • E-commerce product search;
  • Customer service knowledge bases;
  • Code and API documentation retrieval;
  • Multilingual content search;
  • Cross-Encoder reranking after hybrid search.

If your system has only a few dozen documents, low query volume, and you can manually review all results, a complex Shadow platform isn’t necessary—but Hard Negatives + a fixed regression set + NDCG slices are still worth keeping.

Common Pitfalls

PitfallCorrect Understanding
Only look at the Reranker’s own benchmarkPublic benchmarks reveal model capability but don’t replace your own production query distribution
Higher Reranker score means more trustworthy resultsThe score’s primary use is ranking within a single query; don’t repackage it as a “confidence percentage” without calibration
Test the model without freezing the RetrieverCandidate set changes mean you’re not measuring pure Reranker differences; offline A/B must freeze the Retriever snapshot
Only measure quality, not tail latencyCross-Encoders are far more compute-intensive than the first stage; expensive computation must be kept within predictable bounds
More Shadow traffic is always betterShadow has extra cost; prioritize sampling high-value, high-risk, and distribution-representative queries

Pre-Release Checklist

  • Retriever version and candidate_k are frozen and logged;
  • Dataset includes Hard Negatives from the real Retriever;
  • Oracle rerank and End-to-end rerank metrics are separated;
  • NDCG@10 and MRR@10 are sliced by business domain and language;
  • Long-document truncation strategy for the new model is validated;
  • Relevance scores are not mistakenly treated as a cross-query unified probability;
  • P95/P99 latency, timeout, and cost are load-tested;
  • Shadow requests don’t block the user’s main path;
  • High-diff Shadow samples go through human review;
  • Production / Candidate can be switched quickly;
  • Rollback doesn’t require rebuilding the index;
  • Evaluation data, model version, parameters, and results are all traceable.

References

  1. Sentence Transformers — CrossEncoder Evaluation: https://sbert.net/docs/package_reference/cross_encoder/evaluation.html
  2. Sentence Transformers — CrossEncoder Training Overview / Hard Negative Mining: https://www.sbert.net/docs/cross_encoder/training_overview.html
  3. Cohere — Rerank API v2: https://docs.cohere.com/reference/rerank
  4. Cohere — Best Practices for using Rerank: https://docs.cohere.com/docs/reranking-best-practices
  5. Elasticsearch — Ranking evaluation API: https://www.elastic.co/docs/reference/elasticsearch/rest-apis/search-rank-eval
  6. Vespa — Phased Ranking: https://docs.vespa.ai/en/ranking/phased-ranking.html
  7. Thakur et al. — BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models: https://arxiv.org/abs/2104.08663

FAQ

Why can't reranker evaluation rely only on accuracy over random negatives?
Random negatives are usually too easy and don't represent the high-similarity but incorrect candidates returned by the real first-stage retriever. Production evaluation should prioritize hard negatives mined from the actual retriever and preserve the real candidate ranking distribution.
Can a new reranker go live directly if NDCG@10 improves?
No. NDCG@10 only covers relevance. You also need to check candidate recall limits, P95/P99 latency, error rates, long-document truncation, language and business segment slices, and ranking drift in shadow traffic.
Does shadow re-ranking affect live user results?
No, if implemented correctly. The shadow path reuses live queries and candidate sets, but the new reranker's outputs are only logged and never returned to users, so you can observe latency, anomalies, and result differences without changing production ranking.