15. Evaluation and Rubrics
Judge AI output with measurable criteria instead of personal feeling.
By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.
The lesson
Professional AI users define what good means before asking for work. A rubric can score correctness, completeness, safety, clarity, cost, maintainability, and test coverage.
For coding, require compile results, tests, edge cases, security impact, and a concise change summary.
For business, require assumptions, risks, alternatives, numbers, and decision criteria.
A junior QA analyst at a software firm scores AI-generated test plans against a rubric covering coverage, edge cases and clarity.
A procurement manager builds a rubric that grades supplier comparison drafts on sourcing evidence, completeness and risk notes.
A nurse educator scores AI-drafted lesson content against a rubric for accuracy, completeness and safe wording.
An underwriter at an insurance firm grades AI claim notes on policy reference, completeness and tone before approval.
A maintenance planner at a mining operation uses a rubric to judge AI-drafted inspection checklists on safety items and completeness.
A university lecturer scores AI feedback drafts against criteria for specificity, fairness and alignment with the marking guide.
A logistics analyst grades AI route summaries on accuracy, missing stops and clarity for drivers.
A quality technician at a food plant scores AI shift reports against a rubric for accuracy, food-safety items and completeness.
A surveyor grades AI-drafted field notes on measurement clarity, completeness and correct referencing of the plan.
A news sub-editor scores AI headline drafts on accuracy, tone and length against a short rubric.
Check yourself
Question 1: What does a rubric provide?
- A window size
- Measurable criteria for judging output — correct
- A secret API key
- A random style
Answer: Measurable criteria for judging output
Rubrics define quality before work begins.
Question 2: For code, which gate belongs in a rubric?
- Compile and test results — correct
- Favorite color
- No acceptance criteria
- Ignore edge cases
Answer: Compile and test results
Code should prove itself with builds and tests.
Question 3: For business output, what should be checked?
- Only grammar
- Only emojis
- Nothing
- Assumptions, risks, alternatives, numbers, and decision criteria — correct
Answer: Assumptions, risks, alternatives, numbers, and decision criteria
Business decisions need evidence and trade-offs.
← Previous lesson · All 91 lessons · Next lesson →
The full course — 91 lessons and 273 quiz questions — ships inside the app. Get TPEE to study it offline.