You shipped an AI feature and don't trust it
It works in the demo. Nobody has load-tested it, nobody knows what happens when it half-fails, and an important customer is about to ask security questions.
Fixed fee, two weeks, one report. Two senior engineers, UK-based.
Production readiness review is an established practice from site reliability engineering, the discipline Google popularised for infrastructure. We apply the same review to systems that call language models.
It works in the demo. Nobody has load-tested it, nobody knows what happens when it half-fails, and an important customer is about to ask security questions.
A security questionnaire lands, or a customer security questionnaire or assessor asks how AI-generated code is reviewed and tested before production. That is the usual trigger, and it is happening now.
It was built fast, possibly by AI, and it will not survive contact with production. Your team is busy shipping the next thing, not hardening this one.
The proof
It found two critical vulnerabilities. We published it in full, including a correction where our first version got something wrong.
Read the reportWhat you get
Every finding is ranked by severity and comes with a fix estimate, so you know what to do next and roughly what it will cost.
See the offerWhy now
Customer security questionnaires and assessors increasingly ask how AI-generated code is reviewed and tested before production. That is happening now, not on a future compliance calendar.
Under the enacted EU AI Act, Annex III high-risk obligations apply from 2 August 2026. A deferral to 2 December 2027 was politically agreed in May 2026 and is expected to be adopted, but is not yet in force. Check the current position with your own advisers. The transparency duties in Article 50 and the AI literacy duty in Article 4 were not deferred. Separately, EU Cyber Resilience Act vulnerability reporting applies from 11 September 2026 to manufacturers of products with digital elements — which may include you if you ship software, and has nothing to do with whether your system decides anything about people.
We track these because they shape what evidence a system needs to produce. We do not give legal advice on whether any of them applies to you.
Works With Agents is two senior engineers, UK-based. Read about us.