Robin Miklinski · independent AI and software engineer · London

I make AI systems dependable enough to use in production.

I build AI products and the controls around them: evaluation, access boundaries, observability and human oversight. Together, these determine what a system may do, when it should stop and whether its output can be relied on.

01 / What I do

AI products rarely fail at the model boundary alone. They fail because context is incomplete, permissions are too broad, outputs cannot be explained, or a release decision rests on a convincing demonstration rather than repeatable evidence.

  1. 01 Build dependable AI systems

    • evaluation
    • permissions
    • failure behaviour
    • human control
    • agents
    • retrieval
    I design and build AI-enabled software with the controls in place from the start. The model is one component in a complete product.
  2. 02 Evaluate and harden existing products

    • acceptance criteria
    • repeatable evaluations
    • adversarial tests
    • release evidence
    I turn broad expectations into measurable criteria. The result shows what the system can do, where it breaks and what remains uncertain.
  3. 03 Find where confidence breaks down

    • observability
    • security
    • cost
    • latency
    • accuracy against autonomy
    I investigate failures across models, prompts, tools, data and the software around them.

Output is abundant. Confidence is not.

Models can produce code, analysis and plausible answers. The difficult work is deciding whether an output is correct, whether the system has enough context, what it is permitted to do and who remains accountable when it fails. That is the layer I engineer.

  1. Understand how the work gets done today
  2. Find the parts that could be simpler
  3. Remove what isn’t earning its place
  4. Build the smallest complete system
02 / The story

A release decision backed by evidence.

When I took responsibility for quality and release assurance on Reward Gateway’s first customer-facing AI product, the application had already been built. The evaluation infrastructure had not. There was no defensible basis for deciding whether it was ready to launch.

Because the model’s wording varied between runs, conventional exact-output checks were insufficient. I built an evaluation and security approach around bounded properties: whether responses completed the task, followed the system’s rules and remained safe across repeated attempts.

The results exposed material gaps before launch and made the remaining uncertainty explicit. The launch panel could then accept the residual risk with a clear view of both the evidence and its limits.

The suite became the release gate for later AI work. I also introduced follow-up question rate as a product signal: a second question often indicated that the first response had not resolved the user’s need.

03 / Evidence

Selected evidence.

Three records that can be opened and checked. The complete register lists the rest, and labels every claim by what backs it.

Open the complete evidence register →

04 / About

Robin Miklinski

Background

I began in defence data systems, where traceability, controlled change and clear system boundaries were part of the engineering standard. Since then I have worked across retail, collaboration software and global SaaS, from product decisions and architecture through to implementation, evaluation and production reliability.

I write Python, JavaScript, TypeScript and C#. I can build both the product and the evidence needed to judge it, which matters most when failure can be fluent, plausible and wrong.

Outside work

I produce electronic music and DJ as Eidetic. It is a different medium, but it draws on similar habits: close listening, deliberate iteration and knowing which details matter.

05 / Contact

Tell me what the system must do, and what must never happen.