Organisations who licensed an AI assistant, pointed it at their reporting and found the answers plausible but unverifiable. The question is not whether the assistant is good. It is whether the model underneath it says the same thing twice.
- A scored answer sheet
- All 20 questions, each marked correct, wrong or refused, with the answer that was given.
- A cause per failure
- Why each wrong answer was wrong, traced to the model rather than to the assistant.
- A written verdict
- Ready, ready with named fixes, or not ready, dated and tied to the version of the question set behind it.
- An access-rules finding
- Whether the assistant respects who is allowed to see what, tested rather than assumed.
- The use case itself
- Where the model carried it, the working use case stays with you.