Most AI projects don't fail during the build
They fail three months later, when nobody knows how to maintain them.
Everyone is worried about the wrong moment.
The moment people fear is the build: will the agent work, will the automation run, will the demo land. And to be fair, the demo almost always lands. AI has made the impressive demo the easiest part of the whole project.
The moment that actually decides whether the project was worth the money comes about three months later. Something changes. The API returns a slightly different shape. The model gets updated. A new hire feeds it input nobody thought to test. And the question that decides everything is not "does it work" anymore. It is "does anyone here understand this well enough to fix it."
In most projects I have seen, the honest answer is no.
Agents that demo well and agents that work are different things
The gap between them is almost never in the clever part. It is in the boring parts.
- What happens when the input is malformed
- What happens when the model is confidently wrong
- What happens when a human needs to step in
- What happens at 2am when nobody is watching
A demo skips all four. Production is mostly those four. That is where the real engineering time goes, and it is exactly the part that gets cut when the goal is to ship something impressive by Friday.
Why this got worse, not better
I started out as a no-code developer in 2021, so I say this with love: the tools have made building so easy that the building is no longer the filter.
It used to be that if you could ship something, you probably understood it. That link is broken now. AI will happily generate a system that its own builder cannot explain. It runs, it looks right, and it is a black box to the one person responsible for it.
Building something fast and building something you can still reason about later are two different skills. The second one is the one that pays for itself, and it is the one I have been chasing since I stopped trusting my own drag-and-drop blocks and went to learn what was actually underneath them.
What maintainable actually looks like
None of this is glamorous. That is rather the point.
- Decisions written down. Not documentation theatre, just a plain record of why the system works the way it does, so the reasoning survives the person.
- Guardrails and an escalation path. A clear list of what the agent must never decide alone, enforced in the system, not in a prompt comment.
- Logging you can actually read. When it misbehaves, the answer to "what happened" should take minutes, not a forensic weekend.
- Tested against real cases, especially the messy ones. Happy-path testing is how you ship a demo twice.
- A parallel run before the switchover. The agent works alongside the human process until the real usage, not the test plan, says it is ready.
- Documentation your team can act on without me. If the builder is a single point of failure, the project is not done.
The question to ask before you build
Not "can AI do this." It usually can. Ask instead: who maintains this in six months, and would they understand it today?
If the answer is a specific person who was in the room when it was built, you have a plan. If the answer is a shrug, you are about to buy a very expensive demo.
When I built SubTrada, a procurement platform for UK construction with Claude integrated throughout, the AI features were honestly the fun part. The part that made it a real product was everything around them: the approval workflows, the role-based access, the hundred boring decisions that had to still make sense to whoever touched the system next. That ratio, mostly boring parts around a small clever core, is what a working AI system actually looks like.
If you aren't shipping, you aren't learning. But shipping is not the finish line. Shipping is the point where the system starts meeting reality, and reality is undefeated.
Build for month three.
Building something with AI?
I'll send you a structured plan with scope, phases and a timeline, built to survive month three.