Use this blank worksheet to compare AI systems on the same business tasks. It records expected behaviour, evidence, corrections and cost rather than assigning a universal model score.
Download the AI evaluation template (CSV). No email or account is required. Open it in a spreadsheet and save your own copy.
How to use it
- Give each test case a stable ID and write the expected result before running any model.
- Keep the task and evidence the same across the candidates you compare.
- Record the exact model or application version and test date.
- Mark correctness separately from tone, response time and reviewer effort.
- Classify each failure and decide whether it blocks deployment.
What the columns mean
| Field | Purpose |
|---|---|
| Case, task, language, expected behaviour | Define what the test requires. |
| Evidence reference | Point to the allowed source, without putting sensitive data in the sheet. |
| Candidate, version, date | Make the run identifiable. |
| Correctness and reviewer notes | Record why an answer passed or failed. |
| Response time, cost, review minutes | Keep operational measures separate. |
| Failure class and decision | Make follow-up work explicit. |
Interpret results carefully
For pass rate, divide accepted cases by all completed cases in the defined test set. List timeouts and uncompleted cases separately rather than quietly removing them. For cost per accepted task, divide all measured run costs, including failed attempts, by accepted tasks; if none were accepted, the ratio is undefined.
This sheet is a starting structure, not a certification or a claim that a particular sample size is sufficient. Set acceptance criteria for your application and keep a held-out set of cases for the final check. Do not enter customer identifiers, credentials or confidential text into a shared copy.
For language-specific examples, read evaluating Arabic AI for a UAE business. For project decisions, start with measuring an AI pilot.
Try six illustrative cases
Download six synthetic Arabic and bilingual test cases. They cover mixed-language requests, missing evidence, amounts, instructions embedded in documents and ambiguity. These are fictional starter cases, not a validated benchmark; no model has been scored in this file.
