A release decision backed by evidence.
When I took responsibility for quality and release assurance on Reward Gateway’s first customer-facing AI product, the application had already been built. The evaluation infrastructure had not. There was no defensible basis for deciding whether it was ready to launch.
Because the model’s wording varied between runs, conventional exact-output checks were insufficient. I built an evaluation and security approach around bounded properties: whether responses completed the task, followed the system’s rules and remained safe across repeated attempts.
The results exposed material gaps before launch and made the remaining uncertainty explicit. The launch panel could then accept the residual risk with a clear view of both the evidence and its limits.
The suite became the release gate for later AI work. I also introduced follow-up question rate as a product signal: a second question often indicated that the first response had not resolved the user’s need.