The problem with incomplete answer keys

Cohere introduced RCP-nDCG@10 on 30 September as a new way to measure whether search systems return useful documents in the right order. Traditional nDCG compares ranked results with a set of relevance judgements, often created by people reviewing documents retrieved by earlier systems. Because exhaustively labelling every query-document pair is expensive, many genuinely relevant results never enter that answer key.

As retrieval models improve, they can discover useful documents that older systems did not surface. Conventional evaluation may then mark the new result as irrelevant simply because nobody judged it. Cohere argues that this creates a growing coverage gap: benchmark scores can reflect the limits of an old pool more than the usefulness of a modern search result.

Rubrics, comparisons and calibration

Rubric-Calibrated Preferences nDCG@10 uses an AI judge to assess every retrieved document against explicit relevance criteria. The judge answers five yes-or-no questions of increasing difficulty for each document and also compares documents in groups. Calibration combines those signals so pairwise preferences establish order while the rubric places results on a scale intended to be comparable across queries.

This is more elaborate than asking a language model for a relevance score. Cohere says the shared rubric is designed to make the meaning of a score more stable from one query to another. The resulting document grades are then used within the familiar nDCG framework, retaining its emphasis on both relevance and position among the top ten results.

Human validation offers encouraging evidence

Cohere commissioned 46 annotators with at least a relevant bachelor’s degree to grade top-five results without seeing system names, rankings or answer keys. Three people reviewed each head-to-head contest, covering 289 decided contests across 273 queries. When RCP-nDCG and conventional nDCG selected different winners, human reviewers sided with the new method 70 per cent of the time.

Across a deliberately disagreement-heavy pool, RCP-nDCG selected the human-preferred system 77 per cent of the time, compared with 52 per cent for conventional nDCG. Cohere reports that wide margins under the new metric aligned with reviewers more often, while wider conventional-nDCG margins did not help. These results support the method, although replication outside Cohere’s study will be important.

An AI judge introduces its own risks

Replacing sparse labels with model judgements can fill gaps, but it also introduces dependence on the judge, prompt, calibration data and possible model bias. A judge may favour writing styles or concepts similar to its training data, and future model versions could change scores. The method therefore needs versioned evaluation code, frozen judge settings and periodic human checks if it is used to guide consequential model development.

Cohere acknowledges that very small score differences are weak evidence: in its study, agreement below a margin of about 0.02 was no better than chance. That is a useful operational warning. Teams should not turn a narrow leaderboard difference into a confident purchasing decision, especially when the test corpus differs from their own documents, languages and queries.

Open materials make independent testing possible

Cohere has released a paper plus code and data for RCP-nDCG@10, and says the approach informed development of its fifth-generation Embed and Rerank models. The MTEB project is considering the method as a way to add resolution to existing datasets and broaden the languages that can maintain high-quality retrieval benchmarks.

For enterprise search teams, the method is most useful as an additional lens rather than a universal replacement. Conventional recall, human task success, citation quality and production feedback still answer different questions. Applying RCP-nDCG to a private corpus could reveal relevant documents an old answer key missed, but teams should inspect disagreements and keep humans involved in validating the judge. Better evaluation matters because retrieval errors propagate directly into the context supplied to downstream assistants and agents.

A practical trial could begin with queries where search specialists already know the existing labels are incomplete. Evaluators can compare conventional and RCP rankings, review the documents receiving new credit and record whether the judge rewards genuinely useful evidence or superficial topical overlap. Results should be segmented by language, document type and query difficulty, since an average may hide systematic weakness. If the metric influences training or model selection, a separate human-reviewed holdout set is essential to reduce optimisation against the judge itself. Cohere’s public materials make this testing possible; broad adoption should follow independent replication, not the novelty of a higher-resolution score alone.

The cost of evaluation also matters. Judging every candidate document with multiple questions and pairwise comparisons can be more expensive than looking up fixed labels. Cohere argues that the method can improve signal and potentially extend benchmarks to more languages, but teams should measure runtime and judge spend at their own scale. A sensible workflow may use RCP-nDCG for periodic model comparisons while retaining cheaper production metrics for continuous monitoring. That division preserves deeper analysis without turning every search-quality check into a large model-evaluation job.