01 · Outcome
What decision will the answer support?
Define the user, task and consequence rather than beginning with a general chatbot brief.
Technical evaluation
A useful evaluation starts with a real decision, an approved source set and a failure condition. We demonstrate what the system answers, what it refuses and what evidence remains afterwards.
Evaluation path
01 · Outcome
Define the user, task and consequence rather than beginning with a general chatbot brief.
02 · Corpus
Identify ownership, versions, access boundaries and the conditions under which a source becomes current.
03 · Refusal
Agree the unsupported, contradictory and out-of-scope cases before measuring answer quality.
04 · Evidence
Specify the claim, citation, version, verdict and record needed for review.
05 · Success
Use a test set and explicit thresholds for support, refusal, citation fidelity and reproducibility.
Next step