← Back to all articles
Guides

You Built the MVP With Claude Code. Production Is the Part the Demo Skipped

By Marc Molas·October 2, 2026·9 min read

A founder wrote to us this week with a sentence I expect to read many more times: «We built the MVP for our startup with Claude Code, and we don't think it can go to production.» The MVP works. Users click through it. Investors have seen it. And the person who built it already knows it will not hold.

They are right to doubt it, and I think they are wrong about why. I say that from the seat that reviews these codebases: Conectia runs Claude Code, Cursor and review agents inside our own engineering harness, and we have shipped regulated identity and payments products with agents in the loop. This is not a post against building with AI. It is a post about what the demo leaves out.

The code is rarely the problem. The system around the code is missing. A vibe-coded MVP is a house that already has walls and a roof and no wiring, no plumbing and no permit. Nobody built those because nobody asked for them, and the agent that wrote the walls does exactly what it is asked.

«Vibe coding» produced the fastest prototypes in history, and the thinnest

Andrej Karpathy named the practice in February 2025: describe, accept, run, paste the error back, repeat, and «forget that the code even exists». By November, Collins had made vibe coding its word of the year. In between, the tools got good enough that a non-technical founder can now reach a working product in a weekend. That is a real change, and I would not trade it back.

What the speed hides is where the time went. Two measurements from 2025 bracket it. Veracode tested AI-generated code across more than a hundred models and found that 45% of the samples failed security tests, with Java failing 72% of the time and cross-site scripting defences missing in 86% of the relevant cases. The models wrote code that compiled and ran; they did not write code that defended itself. In the same summer, METR ran a randomised trial with experienced open-source developers and found they were 19% slower with AI tools while believing they were 20% faster. The gap between felt speed and measured speed is exactly the gap a founder feels when the MVP works and still cannot ship.

Where an AI-built MVP actually breaks

I have read enough of these repositories to know the list in advance. None of it is exotic; all of it is work the prompt never requested.

  • Authorization that is a UI, not a rule. The app hides the admin button; the API serves the admin data to anyone who asks. In May 2025 a researcher scanned 1,645 apps built with Lovable and found 170 of them exposed user data through missing row-level security on Supabase. The apps were not broken. They were open.
  • Secrets in the repository. API keys committed because the agent needed them to run, then pushed to a public GitHub.
  • No separation between environments. One database, one set of credentials, and an agent with write access to it. That is how Replit's agent deleted a production database in July 2025 during a live test, after being told to freeze.
  • A data model shaped by prompts, not by the business. Each feature added its own table; nothing migrates; the schema is whatever the last conversation left behind.
  • Zero tests on the paths that move money or data. Payment, sign-up, permissions. If a test exists, it tests the happy path the agent used to check itself.
  • No observability and no cost ceiling. If the product calls an LLM, nobody can see what one request costs, and nothing stops a loop from running all night.
  • Dependencies pinned to nothing and a deploy that only works from the founder's laptop.

Each item is a day or two of senior work. The founder who wrote to us can see the pile without being able to name it, and that is why «it won't go to production» feels like a verdict on the whole codebase. It is a verdict on seven missing pieces.

The strongest counter-argument, and it is right: the MVP did its job

Someone will say the proper move is to throw it away and rebuild «properly». I have watched that advice destroy twelve months of runway more than once, long before AI was involved. Joel Spolsky called the full rewrite the single worst strategic mistake a software company can make, in 2000, and the reasoning has not aged: the old code has been tested by real users, and the new code has not.

An AI-built MVP has already done the expensive thing. It proved that someone wants the product, it fixed the vocabulary of the domain, and it produced a working reference for every screen and flow. That is a specification written in code, and it is worth more than the slide deck it replaced. The UI usually survives. The flows survive. The data model survives as documentation of what the business actually needs, even when the tables themselves do not.

The MVP is not the mistake. Treating it as the production system is.

We have seen this movie, with a different cast

In 2012 the founder came to us with a prototype built by an agency for a fixed fee, or by a cousin who knew PHP. Same shape: it demoed well, the first paying customer found the first security hole, and the question was always the same: rescue or rewrite? The teams that won did neither. They put a perimeter around the prototype, replaced the parts that carried risk one at a time, and kept shipping features throughout. Martin Fowler named the pattern the strangler fig in 2004.

What AI changed is the speed on both sides. The prototype arrives in days instead of months, and the fix arrives faster too, if the agent is given the one thing the original build never had: a specification it can execute against. An agent with a written data model, acceptance criteria, and a rule that says «every table has row-level security» produces a different codebase than an agent told to «add a dashboard». We see it daily in our own work: the same model, the same tool, and the difference is entirely in what the engineer wrote down first.

The first 30 days, in the order that holds

If the founder who wrote to us were my client, this is the order I would work in. The order matters more than the list.

  1. Read-only audit, two days. Inventory the repository: dependencies, routes, tables, where secrets live, what the deploy script actually does. No changes. The output is a one-page map of what exists and a list of what carries risk.
  2. Lock the perimeter before touching features. Rotate every key that has ever been committed. Enforce authorization at the data layer (row-level security, not hidden buttons). Split production from everything else and take write access away from any agent.
  3. Make the build reproducible. Pinned dependencies, one command to run locally, a CI pipeline that fails on a failing test. Until this exists, every fix is a bet.
  4. Tests on the money paths only. Sign-up, payment, permissions, the one query that would be a breach if it leaked. Not coverage; insurance. If the product calls a model, the equivalent is an evaluation suite that runs on every prompt change.
  5. Observability and cost caps. Logs with request IDs, error tracking, and a hard ceiling on what the model layer can spend per day.
  6. Write the specification the MVP never had. Data model, acceptance criteria per feature, the phased plan, and the context files (CLAUDE.md, AGENTS.md) that turn the agent from a weekend builder into a disciplined junior. This is the step founders skip because it feels like paperwork. It is the step that makes the next six weeks predictable, and it is why we sell it on its own as a Blueprint a founder can execute with us, with their own hire or with the same agent that wrote the first version.
  7. Then, and only then, features. On a codebase that can now take them.

A senior engineer who has done this before covers the first five steps in two to three weeks. Done in a different order (features first, perimeter later) the same work takes a quarter, because every feature lands on sand and gets rebuilt.

Three questions that decide keep versus rewrite

  • Does the data model describe the business, or the conversation? If the entities match how your customers talk, keep it and migrate. If the tables are named after prompts, rebuild the schema and port the data.
  • Can you draw the request path for your riskiest action on a whiteboard? If nobody can, the audit comes before any decision.
  • Is the framework one a senior engineer would choose today? Next.js, Django, a managed Postgres: keep. A generated stack nobody on the market hires for: plan the exit now, execute it later.

Most AI-built MVPs I have seen pass the first and third question. They fail the second, and that failure is a two-day audit away from being fixed.

The demo skipped production. You do not have to

The founder's sentence was half right. The MVP as it stands should not go to production, and that was never its job. Its job was to prove the product, and it did that faster than any method that existed three years ago. What it needs now is the wiring, the plumbing and the permit, in a known order, from someone who has installed them before.

If that is the question in front of you, this is the team we put on startup products: senior engineers, vetted by working CTOs at a 3% acceptance rate, who take over a codebase without a spec handed to them and start on the perimeter in the first week. Or start with the Blueprint, and keep building with whoever you like.

Work with the authors

The CTOs who write this also build

Fractional CTO engagements, AI engineering squads and an AI development platform we install on your codebase in 24 hours. If this piece matched how you think, the conversation is short.

Teams we have embedded engineers in
MintIDCNN InternationalSony Music