The app runs. People who are not you have signed up, used it, and come back. Nothing has broken, nobody has complained, and the bill from the build tool arrives every month like a utility bill. Underneath all of that sits a question that does not go away: my vibe coded app works, why shouldn’t I trust it?
You are not asking because something went wrong. You are asking because nothing has, and you have started to notice that nobody ever told you what fine would look like. There was no moment when a person who can read code used the word finished. The app simply kept working, and the absence of bad news slowly became the only evidence you own.
You have already decided to keep the app. So the job on this page is narrower than a verdict: say precisely what months of normal use have proved, what they could never have proved, and what would close the gap between the two.
A working vibe coded app is evidence for exactly one path: yours. It proves your account, on your device, doing the thing you built it to do. Trust it for that. Five conditions it has never been under decide the rest, and none of them requires reading code to check.
Should I trust an app I built with AI?
The question has the wrong shape as a yes or no. An app you built with AI is trustworthy for a particular job and a particular group of people, and the evidence differs for each. Right now you hold evidence for one group of one: you.
So the useful form of the question is trust it for what, and trust it for whom. Those two have different answers, and answering them separately is most of the work.
There is a piece of outside evidence worth putting next to your own. The 2025 Stack Overflow Developer Survey asked how much people trust the accuracy of the output from AI tools as part of their development workflow. Read at the survey’s AI section on 25 August 2026, 33,244 people had answered that question: 3.1% highly trust the output, 29.6% somewhat trust it, 26.1% somewhat distrust it, and 19.6% highly distrust it. Those four figures do not add up to 100, so at least one further answer was on the list and roughly a fifth of respondents picked none of the four.
Read it the right way round. The people best placed to judge AI output are the least willing to vouch for it, and every one of them can open the file and look. You cannot. So the standard you have been using, that it works, was never the standard anyone with the ability to check was applying in the first place.
Owners ask the same thing in plainer words. A thread posted to the r/vibecoding subreddit on 13 August 2026 carries the whole worry in its title: “Are you guys actually deploying vibe coded stuff?” Someone building this way described how the worry changes shape once the tool starts working:
I spend less time asking “how do I code this?” and much more time asking: Is the AI actually doing this correctly? Do I understand enough to know when it’s wrong?
The second of those questions is the one this page is about. Would you know?
What your own use has already proved
Start with the part that usually gets skipped, because it is real. You have working software. People who owe you nothing sign up and use it. That is further than most ideas ever get, and it is genuine evidence: every time you or a customer walks the main path and it holds, that path collects one more observation in its favor.
Now name the four things that were true almost every time that evidence was collected. Mostly one account, yours, or one you can see into; your customers add their own accounts, on the paths they happen to walk. Mostly one device, the one on your desk, with your browser and your saved logins. Mostly one network, fast and stable, in one country. And money moving in one direction, forward, from a customer to you, with no refund, dispute or reversal yet tested.
None of that is a criticism of how you tested. Normal use is narrow by definition. You use the app the way you meant it to be used, because you are the person who meant it, and the app was assembled to survive exactly that walk. It does. That is the finding, stated as generously as it deserves and no further.
The evidence has a shape, and the shape is a line rather than an area. A month of that line is a longer line. It is not a wider one.
It is worth being concrete about what stayed outside the line, because the list is shorter than it feels. You have never signed in as somebody else and tried to reach your own records. You have never opened the app while signed out of everything, in a browser that has never seen it before. You have never sent the same request forty times in a minute. You have never had the connection die halfway through a save. You have never taken money back out of your own app. Every one of those is an ordinary Tuesday for a customer, and none of them is a thing an owner does by accident.
That is why the phrase a working app is not proof lands so badly when somebody says it to an owner. It arrives sounding like a judgment on the app when what it actually describes is the size of the sample.
Someone who built an app this way put it the other way round in public: the fact that it runs is the reason they are wary of it, not the reason they are relaxed.
One more piece of evidence usually arrives here, so it deserves a straight answer rather than a shrug. A green test run is the thing owners offer most often, and why a passing suite and a broken app sit together so comfortably has a page of its own; the short version is that the suite came out of the same session as the code, so it agrees with it.
Five conditions your app has never been under
The numbers below come from one fixed set of 26 AI-built applications reviewed at the code level in June and July 2026, and from the 420-finding ledger kept for the 21 of those apps that other people built. For this page, on 25 August 2026, I went back through that ledger and re-sorted it by one thing only: what has to happen before a given finding can show itself. Nothing here was re-run, re-scanned or re-tested for this page.
Sorted that way, the ledger stops reading as a list of defects and starts reading as a list of doors nobody has opened. Five conditions account for the findings in it that would have changed an owner’s mind about whether their app was finished, and an owner never creates any of the five on purpose. Every count below states the group it was measured across, and the full set of denominators behind those 26 apps sits on one page rather than being restated in pieces.
The inventory itself, fourteen recurring problems counted across the same 26 apps, is a separate page, and this one does not re-list them: it asks what each of them is waiting for.
A second person signs in
The condition is one other human being with their own account, using the app at the same time as somebody else, on purpose or by accident reaching for something that is not theirs.
In 7 of the 21 third-party apps in that ledger, a signed-in user could read or write another customer’s data, and each of those was confirmed against the code rather than inferred from a scanner alert. In 9 of the same 21, the database rules meant to keep one customer’s rows away from another had gaps in them.
The distinction the ledger keeps making is small and it decides everything. Signing in proves you are somebody, and a great many of these apps stop there, without a second question asking whether the account making this request is the one that owns the record it is about to hand back. Both checks look identical from the outside, because from the outside you only ever see the one account you are signed into, and your own rows are the rows you are allowed to see.
What you would have noticed while this was true of your app: nothing. There is no screen for it. The first version of this that reaches you usually arrives as a support message about a display problem, where a customer mentions seeing a name they did not recognize in a list.
Someone who never signs in at all
The condition is a stranger. Not an attacker in a hoodie, just anyone on the internet sending a request to an address in your app without ever making an account.
In 11 of the 21 third-party apps, an endpoint that did privileged work answered requests from callers who had never logged in. Fourteen of those 21 apps had an AI feature, and in 12 of the 14 there was a confirmed path for a stranger or a free account to spend the owner’s paid model budget. In 8 of the same 14, untrusted text could reach the model’s instructions, which means the person supplying the text gets a say in what the model does next.
The money version of this is the one that reaches owners first, and it reaches them late. A stranger burning your model budget produces no error, no failed page, and no complaint. It produces an invoice, and that invoice turns up whenever the model provider’s month happens to end, which is why the usual first symptom here is a bill that looks wrong and a fortnight of wondering why.
What you would have noticed: nothing, until the invoice. You have an account, so you have never once used your own app as a person without one.
Two people at once, or one person repeating
The condition is repetition. Two customers doing the same thing in the same second, or one determined person sending the same request two hundred times.
Of the 21 third-party apps in that ledger, 13 put no cap at all on how often one caller could trigger their costliest operation. Expensive here means whatever costs real money per use: a model call, an email send, a file conversion, a login code.
This is not a story about scale, and reading it as one is how owners talk themselves out of it. You do not need traffic for this. One person with a loop and an afternoon is the whole condition, and your app has never met that person because your customers are ordinary people using it once, patiently, at human speed.
Repetition also breaks a second thing that has nothing to do with cost. Two people doing the same thing in the same second is the moment an app finds out whether it can tell the two of them apart: two orders taking the same last item, two saves landing on the same record, two signups claiming the same name. Those failures are quiet as well, because the loser of the race usually gets a page that looks like it worked.
What you would have noticed: nothing, and this is the condition where the aftermath is most likely to be a number in an invoice or a wrong row in a table rather than a broken screen.
A request that does not finish
The condition is an interruption. A phone dropping from wifi to mobile data mid-request, a tab closed while a page was saving, a payment that got halfway. Your desk has fast, stable internet. Your customers have trains, lifts, and basements.
Of the twelve areas the review scored across those same 21 third-party apps, the one that came last was Reliability and Correctness, and the one that came first was Secrets and Credentials. That ordering inverts what most people expect of an AI-built app. Keeping keys out of the browser is the thing these apps are best at. Behaving correctly when something goes sideways is the thing they are worst at, and only one of those two has a public reputation.
Underneath that ranking sits the finding that makes this condition invisible. Across the 21 third-party apps, 17 kept no record anywhere that a person would ever look at when a user hit an error, so a failure that happened to a customer simply ended with that customer. The only reporting channel the app had was somebody caring enough to write in.
What you would have noticed: nothing, and here that word carries more weight than it does anywhere else on this list. Nobody has complained is not evidence that nothing failed. In an app with no error tracking, nobody has complained is the expected reading whether things are fine or not, which makes it the one sentence in this whole subject that carries no information at all.
Money going backwards
The condition is a refund, a cancellation, a downgrade, or a chargeback. Every transaction your app has ever processed moved money forwards, from a person to you, because that is the direction you built and the direction you tested.
Here is a denominator that says more than any percentage on this page. Of the 21 third-party apps in that ledger, the review’s billing and revenue area applied to only 3, because the other 18 had no money path substantial enough to score. Most apps built this way reach real users before they reach a real financial surface, and then the financial surface gets added in an afternoon, on top of an app nobody has re-examined since.
Two other counts land squarely on that afternoon. Of the same 21, in 10 the server believed a fact handed to it by the browser instead of establishing that fact on its own. In one of them, the amount charged for an order came from the browser rather than from the seller’s own price list. And 6 of the 21 shipped a real secret; in three of those, the secret sits permanently in the repository’s history, and two of those three were webhook signing secrets, one of them belonging to the payment provider. A webhook is the message a payment company sends your app to say a payment happened. If somebody else can sign that message, they can tell your app a payment happened when it did not.
An owner posted the feeling that goes with this condition better than a statistic can:
… I had my first paying customer which made me happy and nervous at the same time. Will they refund? Will they be satisfied?
What you would have noticed: nothing, because you have never refunded yourself. The first real refund in a lot of these apps is also the first test of the refund path, run by the person least able to tolerate it failing.
Meeting the first three of those conditions on purpose takes about twenty minutes and a browser, and the exact things to open, with the exact thing you should not see, are written out step by step.
Trust is not one decision
Trust reads like a single switch, and it behaves like four. It is settled once per group of people who depend on the app, and each group needs different evidence before the answer is yes.
Measuring what this code will cost you to change is a different axis from measuring what it has proved, and an app can be cheap to change and still completely unexamined. Trust is the second axis. It moves when evidence arrives, not when the code gets tidier.
| Who depends on it | What breaks for them if you are wrong | The evidence that rung actually needs |
|---|---|---|
| Just you | Your afternoon | That it works, which you have |
| A handful of people you can phone | An hour of theirs, and one apology | The second-person condition, checked once |
| Strangers who pay | Their money, and your ability to refund it | The money-backwards condition, plus somebody being told when a path fails |
| People whose records you cannot afford to leak | Their data, and your business | All five, and a person who has read the code |
Most owners are standing on the third rung with evidence that only supports the first. That is not carelessness. The rungs arrive quietly, one signup at a time, and nothing in the app announces that you crossed one. There is no email from your build tool the morning after the first stranger pays you.
Read the table by its middle column rather than its left one and placing yourself gets easier. What decides your rung is what a wrong answer costs the person on the other end, and whether you could put it right afterwards, rather than how many people have signed up. Ten friends testing a booking tool sit on the second rung. Ten customers whose payment details, medical notes or client lists are in your database sit on the fourth, at ten users, on day one. Rung four is about what you are holding, and it can arrive before your first invoice does.
The honest number belongs here, because it applies to me as well as to you. Across all 26 apps in that fixed set, the review scored between 29 and 81 out of 100, mean 52.1 and median 51, with 22 in the red band, 4 amber and none green. Five of the 26 are production apps of mine, put through the same review, and none of them reached the top rung on their own evidence either.
The reassuring half of the same ledger is worth stating just as plainly. Across the 21 third-party apps there were 958 confirmed findings, about 46 per app, and 58 of them were critical. Most of what a review of an app like yours turns up is not urgent. Knowing which handful is urgent is the entire job, and it is a much smaller job than the total makes it sound.
None of this is a claim about vibe coded apps as a category. Whether the category is safe, and the six security risks that keep recurring inside it, is a question about vibe coded apps in general and it is answered in full on the page about whether vibe coded apps are safe; this page is only about yours, and only from the point where you have already decided to keep it.
One more thing the rungs explain. Owners rarely lose confidence gradually; they lose it at a threshold. The tool carried one founder’s product as far as beta and then became the reason the launch nearly did not happen, which is what a rung change feels like from the inside when nobody warned you it was coming.
What would actually change your mind
Three things move the answer, and none of them requires you to read code.
The first is meeting three of the five conditions on purpose, in a browser, in about the time it takes to have a coffee. Make a second account and try to reach the first one’s records. Open a page that should need a login while signed out of everything, in a private window. Do the most expensive thing your app does, twice in a row, quickly. What you are looking for in each case is one specific wrong thing appearing on screen. If it does not appear, that condition has evidence behind it for the first time in the app’s life.
The second is arranging to be told. Of everything on this page, the single standing change with the best return is that a failure your customer hits produces a message somewhere you will actually see, rather than ending inside their browser. It converts every future problem from something you find out about socially into something you find out about immediately, and it is the reason the fourth condition sits so low on most of these apps. It is also the only item here that keeps paying after today, because it applies to code that has not been written yet.
The third is somebody reading the parts a browser cannot reach. Refund handling, what the server decides versus what it accepts, whether a leaked secret is still live, whether the database rules say what the screens imply. There are six specific things that can be wrong underneath an app that runs, and which of them a person has to look at rather than a browser is the question the rest of these pages start from.
There is a fourth question that turns up alongside these and belongs to a different page. How many defects is normal for a product this young, as opposed to how many mean the foundation is wrong, is a benchmark question with a number behind it.
What none of this produces is a certificate. A working app is not proof, and neither is a clean afternoon of checks; what both give you is evidence for specific paths, named out loud, so that the next time someone asks whether the app is trustworthy you can say which parts of it have been looked at and which have not.
Common questions about trusting a vibe coded app
Is a working app a finished app?
No. A working app is a product that does its job on the paths somebody has walked. A finished app also has the parts that prove it keeps doing that job when you are not watching: an account boundary that has been tested, a limit on the expensive things, a record when something fails, and a money path that works in both directions. Those parts have no screens, so nothing in the app tells you whether they exist.
My app has paying customers already. Does that settle it?
It settles the question of whether people want it, which is the harder question and you have answered it. Paying customers change the stakes on everything else rather than resolving it: money moving means refunds, chargebacks and cancellations become real paths, and each of those has to work backwards through code that was only ever tested forwards. Revenue is evidence about demand, and it is silent about the app.
There is a second effect worth knowing about. Paying customers are usually the point at which an owner stops being willing to experiment on the live app, which is exactly when checking it starts feeling risky. The three browser checks described above run on an account you made yourself, and none of them touches a real customer’s records, which is why they are the ones worth doing first.
Nothing has broken in six months. Is that evidence?
It is weak evidence, and its strength depends entirely on one thing you can check today: whether the app would have told you. Across the 21 third-party apps in that fixed 2026 set, 17 had nothing anywhere that would tell a human being a customer had just hit an error. In an app like those, six quiet months and six broken months look exactly the same from where you sit.
My tests are green. Does that count?
It counts for the questions somebody wrote down, and for nothing else. A passing run means the checks that exist still get the answers they got when they were written, and on an app assembled by prompting those checks were usually written in the same session as the code they check. The green is real and its subject is narrow.
The practical version: ask which check covers the exact screen a customer was on when something went wrong. If nobody can name it, the suite has no opinion about that screen, and it never did.
If I ask the AI whether the code is fine, does that count?
Not as independent evidence, for the same reason the tests do not. Asking the same assistant that wrote the code whether the code is fine is its own question, and the answer depends on whether a model is able to notice the mistake it made. An assistant can genuinely find things when you point it at a specific behavior and a specific file; what it cannot do is tell you what it never thought to consider the first time.
Does it matter which tool built it?
Barely, for this question. The five conditions do not change by vendor, and neither do the findings underneath them: whether the app came out of Lovable, Base44, Bolt, Replit, Cursor or Claude Code, the ledger shows the same shapes, because the mechanism is that nobody asked for the invisible half rather than that a particular tool refused to build it. What does change by vendor is how much of the hosting, database and access rules the platform handles for you, and that shifts which of the five is likeliest to bite.
Should I stop taking payments until I know?
Almost never, and switching payments off has costs of its own that a worry does not automatically outweigh. Whether to switch payments on at all, before any of this is settled, has its own order of operations, and it is not the same question as whether the checkout works. If you already take payments, the useful move is checking the reverse direction first: process one real refund yourself and watch what the app does with it.
Who can actually tell me whether to trust it?
Somebody who reads the code and then shows you the thing itself: the request that came back with another customer’s row in it, the screen that let a free account through, the endpoint that answered a stranger. The point where the answer stops being something anyone can see from outside and starts needing somebody to read the code is exactly what a code review is for when you cannot read code yourself. What that should get you is the broken path put in front of you on a screen and then fixed. A general verdict on your app’s overall health would be worth about as much as the feeling you started with, which is the whole trouble with can I trust my AI built app as a question: it only becomes answerable once it is broken into paths, and every one of those has a name.
Built it with AI. Can’t get the last part right?
That’s the normal state of an AI-built app, and it’s fixable. I trace what the app actually does, explain what needs changing, and build it if you want me to.
Talk about your app →
Free 20-minute video call with Bilal.