Responsible AI
AI Evaluation
Definition
AI evaluation uses criteria and representative cases to assess behaviour such as factual support, task completion, appropriate refusals and respect for permissions. It can include human review and automated checks. A single successful demonstration is not a complete evaluation.
Why it matters
It helps teams find failures, compare changes and decide whether a system is suitable for a specific use. Criteria should reflect the actual workflow, including difficult and out-of-scope cases, rather than only fluent answers.
Example
Hypothetical example: reviewers test a bilingual knowledge assistant with answerable, ambiguous and restricted questions, checking sources, meaning and whether it correctly asks for help.
Related Ai MindUp
Reviewing a knowledge assistant