Evaluation results are easiest to misunderstand when a number is treated as a property of a model rather than an observation from a particular test.

Ask what the examples represent

Look for the source of the test set, the time period it covers, the languages included, and the cases that were excluded. A score can be useful and still fail to describe your users' work.

Separate capability from fit

A model may perform well on a public benchmark and still be a poor fit because it is too slow, too expensive, difficult to monitor, or inconsistent on the documents that matter in practice.

Keep the comparison fair

Record the model version, prompt, retrieval context, tools, sampling settings, and human review process. Without that record, a result is hard to reproduce or challenge.

Report uncertainty

Small differences may be noise. Use a holdout set, repeat the test where possible, and include concrete examples of both successes and failures.

A benchmark is an instrument, not a verdict

Every evaluation makes choices before a model produces an answer. Someone chooses the task, writes or collects examples, determines what counts as correct, selects a prompt, and decides how to score ambiguous results. Those choices do not make an evaluation useless. They define what the result means. A reading-comprehension benchmark can be appropriate for one research question and nearly irrelevant to a team that needs reliable extraction from invoices in two languages.

Begin by asking whether the examples resemble the decision at hand. Look for the source of the dataset, collection dates, languages, licensing, sample size, and exclusion rules. Public benchmarks can become easier to optimise for once they are widely used; that is another reason to combine them with a smaller test set that has not been used for development. A result should be described as performance on a stated test, not as a permanent ranking of a model.

Define the unit of success

Accuracy is not always the useful measure. A routing system may need to minimise missed urgent requests. A drafting assistant may be judged by whether a reviewer can verify citations quickly. A summarisation tool may need to preserve a critical qualification even when most of the summary reads well. Before testing, state the error that is most costly, the reviewer who can recognise it, and the threshold at which the workflow remains worthwhile.

Evaluation design should preserve the conditions of real use. If a product will use retrieval, tools, structured output, or a system instruction, include those elements in the test record. If humans edit the output, record the review process and the time it takes. Comparing a bare model prompt with a fully integrated product can be useful for research, but it does not answer a deployment question without that context.

Use an error review, not just a leaderboard

After calculating a score, examine a sample of correct and incorrect results. Group failures by type: missing source, wrong field, unsupported assertion, refusal when an answer was possible, or an answer that looks fluent but does not follow the required format. This changes the next decision. A failure caused by incomplete reference material may call for a better retrieval collection; a failure caused by ambiguous labels may call for a revised task definition; a failure caused by a particular model limitation may call for a different model or a mandatory human check.

Repeatability matters as much as the first result. Save the examples, annotation guidance, model revision, prompt, sampling settings, tools, retrieval corpus version, scoring code, and date. For generative tasks, a second run can show whether variation is large enough to affect a decision. Where a test is small, report a range or explain that the result is directional rather than precise.

A practical evaluation sequence

  1. Write the actual user task and the decision a result will support.
  2. Build a representative, permissioned test set and reserve a portion for final checking.
  3. Set a review rubric before viewing outputs, including critical errors and acceptable uncertainty.
  4. Run the candidate and baseline under the same documented conditions.
  5. Review failure categories with domain readers, then decide whether to improve, limit, or stop the use case.

Sources and further reading

The NIST AI Risk Management Framework and the Hugging Face Evaluate documentation provide useful starting points for evaluation planning. For a research framing, see The False Promise of Imitating Proprietary LLMs and its discussion of evaluation context.