Case study · my own app
Before selling this check to anyone else, I pointed it at an AI receptionist and booking SaaS I built. Here is honestly what came back — including the six findings I had to delete, because they turned out to be wrong.
I audited an AI receptionist and booking SaaS I built myself, against the same checklist I sell — 36 points, as it stood then. The method was static source review plus active verification against an isolated instance seeded with two synthetic tenants and fake data — no production database, no real records, at any point. Then I ran an adversarial pass over my own findings and deleted six of them, because they did not survive a second look.
Nothing critical. Total time to find all of it: one focused day.
This goes first on purpose. A report that only lists problems tells you nothing about what is already solid, and the solid half is the part people quietly want to know. Each of these was actively tested against a running instance, not merely read in the source.
Two high, seven medium, eight low. I have described both highs the way you would want them described if you were checking your own app, which means no file paths and no reproduction steps.
The booking assistant would create a confirmed booking against an email address or phone number that had never been verified. The one-time-code requirement did exist. It existed as English prose, inside a prompt. No code path anywhere checked that verification had actually succeeded before the booking was written.
That is worse than calendar clutter, because the confirmation then goes out from the business's own sender to somebody who never asked for it.
Check in yours: take every rule you wrote into a prompt and ask where the code enforces it. If the answer is "the model has been told to", it is not enforced. It is a suggestion with good manners.
A single environment string, left at the value it ships with, made the webhook verifiers fail open. Unsigned messages accepted. Every signature check above it, all four of them correct, rendered decorative by one line of configuration.
I proved it by inversion: the same request is rejected under production-like settings and accepted under the shipped default. Five minutes to fix.
Check in yours: search for any verification that skips itself in development or test mode, then go and look at what your deployed environment is genuinely set to. Those two facts live in different places, which is the whole reason this survives.
The product's headline promise is that the AI never invents a slot — every time it offers is checked against the live calendar first. So I tested the promise. I sent the booking assistant 22 adversarial scheduling messages.
It handled 12 of the 22 categories correctly: already-booked slots, weekends, before opening and after closing, an appointment that would run past closing time, dates in the past, impossible dates like 30 February, and three contradictory reschedules inside one message. In three languages.
Same-day requests are where it fell over. Across those it offered roughly 20 individual times that were not bookable — already in the past, or inside the configured one-hour lead window. Asked at 16:21 for the earliest appointment, it answered "09:00 today" and put a confirm button underneath. Zero future-dated requests failed. Every single failure was same-day.
Two root causes, and they are unrelated to each other:
The past-time filter ran at day granularity instead of time-of-day. The engine knew what day it was — it correctly refused a request for last Monday — and then offered this morning.
Separately, a cap on the first page of generated slots meant that an almost-empty Friday came back as "completely full".
The second one is a revenue bug, and it is the one I would raise with whoever owns the P&L. A customer told you are full does not complain. They book somewhere else, and nothing in your analytics will ever show it happened.
Two of the medium findings — data retention and erasing a single customer record — are cheap to build before real customer data exists and expensive afterwards. That gap only widens.
The tenant boundary — the thing everybody worries about, the thing every "is my AI-built app secure" thread is about — was solid under direct attack. What broke was the layer above it: business logic the AI tools skip precisely because the app still works without it.
Nothing errors. Nothing returns a 500. The thing just quietly does the wrong thing, politely, with a confirm button under it. That is the gap this kind of check exists for, and it is why "it works" and "it is ready" are two different sentences.
Someone read this and asked the question I should have asked myself: once the verified state is stored properly, where does it actually live? Then, when I answered, they asked a better one.
The first question meant writing out every path in the app that can create a booking, rather than the one the assistant drives. That list is longer than it looks — an ordinary form, an API endpoint, a messenger webhook, the staff dashboard, an import. Each needs the same gate, and a gate you write for the path you built first is remarkably easy to not write for the one added later.
The second question was sharper: what happens when the thing you verified changes afterwards? A gate that compares stored state is only as good as the moment it compared. If the value is allowed to move after the check passes, the check can be satisfied against one thing and applied to another — and the failure mode is not a refusal, it is permission pointed somewhere it was never granted.
Enumerate every path that writes the same record, and test the gate on each. Then ask what happens when the value that gate compares changes after it passed.
Neither of those is about AI, and neither is about booking. The first is what happens when a feature adds a second front door to writes that already existed. The second is time-of-check/time-of-use, which is older than most of the people shipping it. They sit in their own section of the list for exactly that reason — the AI-specific block is a different set of points.
Worth saying where they came from: not from me. A stranger in a comment thread asked two sharper questions than my own checklist did, which is the whole argument for writing these up in public rather than filing them.
The checklist I used on my own app is free, and it is the same one I use on other people's. It has 38 points now — see just above for why. If you would rather I ran it, that is the $199 check.
— Slavik. I ran it on my own app before I ran it on anyone else's.
Questions: hello@itworksbut.com