The AI Go/No-Go Gate: Informed rather than surprised.
An AI product assistant was ready to ship and I was asked to sign it off. One approval, yes or no. I refused. The same prompt does not return the same answer twice, so a pass mark on the day I tested it would have said nothing about the week after.
I built an evaluation suite instead: adversarial prompt injection, hallucination bounds, sensitivity checks. It does not return a verdict. It shows how the assistant fails and where, which is what a launch decision actually needs.
The decision went to a panel who accepted the risk in writing, with the failure modes in front of them. I instrumented follow-up question rate as the product metric, on the reasoning that a user asking again is a user who was not answered. That suite is still the gate every AI feature there goes through.