Skip to content

Responsible AI

AI Evaluation

Definition

AI evaluation uses criteria and representative cases to assess behaviour such as factual support, task completion, appropriate refusals and respect for permissions. It can include human review and automated checks. A single successful demonstration is not a complete evaluation.

Why it matters

It helps teams find failures, compare changes and decide whether a system is suitable for a specific use. Criteria should reflect the actual workflow, including difficult and out-of-scope cases, rather than only fluent answers.

Example

Hypothetical example: reviewers test a bilingual knowledge assistant with answerable, ambiguous and restricted questions, checking sources, meaning and whether it correctly asks for help.

Related Ai MindUp

Reviewing a knowledge assistant
Back to Glossary