Evaluation checklists
Quality gates for judging AI output before it reaches production code, documentation, examples, or customer-facing workflows.
Takeaway
Evaluate AI output with the same standards used for human work: correctness, fit, security, accessibility, tests, and maintainability.
01
Check correctness first
Confirm the output matches the current codebase, APIs, package versions, and product behavior. Plausible but outdated advice should be corrected before style polish begins.
Correctness includes omissions. A generated answer can be wrong because it leaves out a limit, privacy boundary, route behavior, or required verification step.
- Compare claims against source files and command output.
- Check current framework and package versions before accepting API advice.
- Reject suggestions that depend on routes, services, or settings that do not exist.
02
Review user impact
Look for confusing labels, missing error states, accessibility regressions, privacy leaks, and misleading claims. AI output can be syntactically valid while still making the product worse.
For content, user impact often means promise mismatch. The page should not claim depth, coverage, or automation examples that the product does not actually provide.
- Check headings, buttons, and metadata for overclaims.
- Verify keyboard and screen-reader basics when UI changes.
- Look for data exposure risks in examples, logs, and copied snippets.
03
Require verification evidence
Before accepting output, run the relevant checks and capture what passed. For content, verify links and claims; for code, run tests, lint, build, and focused manual flows.
Verification evidence keeps AI output accountable. It also gives future maintainers a compact reason to trust the change.
- Record lint, test, build, and smoke results for code changes.
- Check all linked pages after content or navigation changes.
- List any checks skipped and the reason they were skipped.