15. Evaluation and Rubrics

Judge AI output with measurable criteria instead of personal feeling.

By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.

The lesson

Professional AI users define what good means before asking for work. A rubric can score correctness, completeness, safety, clarity, cost, maintainability, and test coverage.

For coding, require compile results, tests, edge cases, security impact, and a concise change summary.

For business, require assumptions, risks, alternatives, numbers, and decision criteria.

A junior QA analyst at a software firm scores AI-generated test plans against a rubric covering coverage, edge cases and clarity.

A procurement manager builds a rubric that grades supplier comparison drafts on sourcing evidence, completeness and risk notes.

A nurse educator scores AI-drafted lesson content against a rubric for accuracy, completeness and safe wording.

An underwriter at an insurance firm grades AI claim notes on policy reference, completeness and tone before approval.

A maintenance planner at a mining operation uses a rubric to judge AI-drafted inspection checklists on safety items and completeness.

A university lecturer scores AI feedback drafts against criteria for specificity, fairness and alignment with the marking guide.

A logistics analyst grades AI route summaries on accuracy, missing stops and clarity for drivers.

A quality technician at a food plant scores AI shift reports against a rubric for accuracy, food-safety items and completeness.

A surveyor grades AI-drafted field notes on measurement clarity, completeness and correct referencing of the plan.

A news sub-editor scores AI headline drafts on accuracy, tone and length against a short rubric.

Check yourself

Question 1: What does a rubric provide?
  1. A window size
  2. Measurable criteria for judging output — correct
  3. A secret API key
  4. A random style

Answer: Measurable criteria for judging output

Rubrics define quality before work begins.

Question 2: For code, which gate belongs in a rubric?
  1. Compile and test results — correct
  2. Favorite color
  3. No acceptance criteria
  4. Ignore edge cases

Answer: Compile and test results

Code should prove itself with builds and tests.

Question 3: For business output, what should be checked?
  1. Only grammar
  2. Only emojis
  3. Nothing
  4. Assumptions, risks, alternatives, numbers, and decision criteria — correct

Answer: Assumptions, risks, alternatives, numbers, and decision criteria

Business decisions need evidence and trade-offs.

← Previous lesson · All 91 lessons · Next lesson →

The full course — 91 lessons and 273 quiz questions — ships inside the app. Get TPEE to study it offline.