ivinco
The Vibe Coding Cleanup Playbook — From Vulnerable to Production-Ready

The Vibe Coding Cleanup Playbook — From Vulnerable to Production-Ready

Ivinco Team·

Every cleanup we run starts the same way. A Lovable app with 10,000 signups. A Stripe secret in the client bundle. A founder who can't tell what's broken from what's working.

We run the same seven-phase playbook on every engagement like this. The sequence is deterministic. The cost is bounded. The outcome is an app that passes the same technical due diligence the Series A process will run on it — not an app that looks like a greenfield rewrite. That matters, because startups pursuing full rewrites often see no measurable ROI relative to incremental cleanup. The product keeps working. The debt gets paid off in order.

This post walks through the seven phases, what each one costs in engineering hours, and where the sharp edges are that turn a $30K cleanup into an $80K one.

Phase 1: Audit (Week 1)

Before we touch code, we generate a report. The audit is a fixed-scope deliverable: 20–40 engineering hours, $3K–$8K at standard rates, and it produces a 10–15 page document covering five areas.

Security scan. Run Semgrep with the default ruleset, TruffleHog against Git history, npm audit for dependency CVEs, and Snyk for a second opinion. Veracode's testing of over 100 LLMs across 80 coding tasks found 45% of AI-generated code fails OWASP Top 10 — expect findings, triage by severity.

Architecture review. Draw the system as-is on one page. Which services talk to which. Where state lives. What's multi-tenant versus single-tenant. Where the external dependencies are. This is the artifact the founder has never seen and always needed.

Database audit. Schema shape, index coverage, query performance under pg_stat_statements, connection pool configuration, backup and restore posture. This is where the "scaling cliff at 1,000 users" is hiding.

Infrastructure posture. Cloud account ownership, secrets management, IAM policies, deployment pipeline (if it exists), monitoring stack (usually zero). Most vibe-coded apps live on the agency or platform's account rather than the founder's — this phase surfaces it.

Priority ranked findings. Every finding gets a severity (critical/high/medium/low) and an effort estimate. Beesoul charges from $2,500 for a 12-page structured audit; the market for just-audit engagements exists because the audit is the deliverable that decides whether cleanup is possible or a rewrite is required.

Phase 2: Critical Security Fixes (Weeks 2–3)

This phase is triage. Nothing else happens until the bleeding stops.

Rotate every exposed secret. If TruffleHog found a Stripe key in Git history, that key is compromised forever — rotate it. Same for database credentials, OpenAI tokens, webhook signing secrets. The rotation is a 2-hour task per service; the mistake is skipping it and hoping nobody noticed. Someone always noticed.

Move secrets to a real store. Pick one of AWS Secrets Manager, Doppler, Infisical, or HashiCorp Vault. Migrate every .env value. Add Secretlint to CI so the pattern can't come back.

Close the obvious injection paths. Semgrep's findings sorted by CWE-79 (XSS), CWE-89 (SQL injection), CWE-94 (code injection). Each one gets a Zod schema at the API boundary and an ORM migration for any raw SQL calls. Time cost: 2–8 hours per finding depending on how many routes are affected.

Add auth middleware. Every API route gets a middleware check before the handler runs. For a Next.js app, this is a single middleware.ts file matching /api/*. For Express, a requireAuth middleware on protected route groups. We covered this pattern in detail in 7 production problems vibe-coded MVPs ship with; it's usually the single highest-impact fix.

Cost and duration. 40–80 hours for a typical Lovable/Bolt-generated app. Two engineers can close the critical findings in one sprint. Delaying this phase to chase architecture work is the most expensive mistake we see — the incidents happen during the cleanup if the bleeding isn't stopped first.

Phase 3: Data Integrity (Week 3–4)

This is where cleanups go sideways when the sequencing is wrong. Touching the schema before security is closed leaves windows open during active data work, and we've watched teams get burned by that ordering mistake. Vibe-coded apps often ship with a data model that works for the demo and fails for everything else. Three specific patterns to check:

1. Missing multi-tenancy. The users, orders, files tables have no org_id column, but the product has teams or workspaces. Every query is implicitly scoped to the logged-in user rather than to their organization. Fix: add org_id columns, backfill from existing ownership relationships, rewrite queries. This is week 3 of a cleanup if it's needed.

2. Migration history that doesn't work. Schema changes were made directly in production with no migration files. The codebase's migration folder doesn't match the actual database. Fix: dump the live schema, introduce Atlas or Prisma Migrate, baseline from current state.

3. Backup posture. Most vibe-coded apps rely on the platform's default backups and have never run a restore. Fix: document RTO/RPO, run a full restore drill on a staging environment, document the procedure in the runbook. Takes a day if the infrastructure supports it, a week if it doesn't.

Phase 4: Observability (Week 4)

You can't fix what you can't see. Before writing any new features or scaling work, install three things.

Sentry for error tracking. SDK install takes 30 minutes for Next.js or Express. Add source map upload. Route alerts to Slack or the on-call rotation. Free tier handles 5,000 events/month — enough for most MVPs under 10,000 users.

Axiom or BetterStack for structured logging. Migrate every console.log to a structured logger (Pino, Winston, or the platform's native option). Index by request ID so a single user's flow through the system can be reconstructed after the fact.

A /api/health endpoint that verifies database connectivity and returns HTTP 200 only if the app is actually healthy. Configure the hosting platform (Vercel, Railway, Fly.io, Kubernetes) to use it for routing decisions. The one hour this takes pays back the first time the app falls over and needs to auto-recover without paging a human.

Phase 5: CI/CD (Week 5)

By this point the app is safe and visible. CI/CD is what keeps it that way. The mistake we see repeated: teams skip this phase because "it works now, we'll add tests later." Later is the next incident.

The minimum viable pipeline is a GitHub Actions workflow running on every pull request: TypeScript compile, unit tests (Vitest), Semgrep security scan, npm audit, Playwright E2E test on the critical user journey. Production deploy requires a green build and a PR approval — no direct pushes to main.

Branch protection rules on main. CODEOWNERS file for sensitive paths (auth, payments, database). Automated deploy previews per pull request so every change can be reviewed in a running environment before merge.

Total setup time: 16–30 hours for an app that started with no CI. The payback is that every subsequent change carries the full test and security check automatically, which is what prevents this same cleanup from being needed again in six months.

Phase 6: Scale Hardening (Week 6)

The app is safe, observed, and shipping through CI. Now we ask the load question — which is where we've watched cleanups fail most predictably. Teams skip the load test and ship. The first marketing campaign hits. The database pool saturates at 400 concurrent users and the app returns 500s for an hour. The team spends the rest of the sprint firefighting what the load test would have surfaced in a morning.

Run k6 or Artillery against staging at 10x expected peak concurrent users. Whatever breaks first gets fixed:

  • Database connection pool exhaustion. Configure PgBouncer or the managed pooler (Supabase, Neon include it).
  • Memory leaks or unbounded caches. Profile with Node's --heap-prof or equivalent; bound every cache.
  • External API rate limits. Queue calls to Stripe, OpenAI, SendGrid through Inngest or BullMQ.
  • Static asset throughput. Put Cloudflare or CloudFront in front.

The load test is the forcing function. You can skip it and hope, or you can run it and fix the three things it surfaces. AWS's own scaling guide calls out each of these as standard failure modes under 10K users — we covered the full architecture progression in From 10 Users to 10,000.

Phase 7: Handoff (Week 7–8)

Cleanup isn't done until the founder's next engineer can run the product without us.

Runbook. Written document covering deploy, rollback, the top five production failure modes, and how to contact on-call. A 20-page Google Doc is fine; a 2-page checklist is better.

Infrastructure as code. Everything in Terraform or Pulumi, committed to the same repository as the application code. The founder's next engineer should be able to stand up a staging environment from scratch in under two hours.

Architecture diagram. One page, current state. The same diagram we drew in Phase 1, updated to match the cleaned-up system.

Credentials transfer. Every external service account — Stripe, Sentry, AWS, SendGrid, OpenAI — in the founder's name, with individual logins for each team member. Shared credentials removed.

30-day warranty. We stay on retainer for 30 days after handoff for bug fixes and clarification questions. This is a standard industry practice and makes the handoff credible.

The Total Cost

A typical vibe-coding cleanup takes 6–10 weeks of work from 2 engineers. At senior rates ($120–$200/hour), that's $100K–$250K. At mid-market EU rates ($65–$90/hour), $55K–$140K. seanconnolly.dev cites rescue engagement ranges of $50K–$500K per startup depending on what was built — the variance is mostly about how much structural rewrite the data model requires.

The alternative is a full greenfield rewrite. Rewrite cost runs $80K–$150K for the same app, plus 3–6 months of delayed roadmap. For apps with real users and real data, the rewrite is almost always the wrong call — the migration cost alone (data, URLs, integrations, SSO, billing state) exceeds the cleanup cost.

What the Playbook Assumes

Two assumptions make this playbook fit:

1. The product is validated. If the app has 10 users and isn't sure they want to keep using it, the right move is not a cleanup — it's to validate the product first. Cleanup is infrastructure spend for a product that needs infrastructure. Pre-PMF products don't.

2. The original architecture is salvageable. We've seen a small number of cleanups cross what we call the 60% Cliff — when Phase 1 finds that more than 60% of code paths have security or correctness issues severe enough to be in the "Kill" category of MVP technical debt. Above the cliff, rewriting from the product spec costs less than cleaning up what's there. Below it, cleanup is cheaper. The audit tells you which side of the line the app is on.

The cleanup costs less than the rewrite in the vast majority of cases. Phase 1 answers the question.

What We Keep

The cleanup playbook keeps everything that works and replaces only what doesn't. The product's user-facing behavior doesn't change during a cleanup. The URL structure doesn't change. The data doesn't move. The customers don't notice the work is happening. All they notice is that the app gets faster, errors get rarer, and the founder stops losing sleep over the Stripe dashboard.

Rewrite the product, or pay off the debt in order. Phase 1 decides which.

Need help? Talk to an engineer.

Frequently Asked Questions

How long does a vibe coding cleanup take?

Six to ten weeks for a typical MVP from Lovable, Bolt.new, Replit, or Cursor. The audit (week 1) is fixed scope. Phases 2 through 7 run sequentially because each depends on the one before — security before observability, observability before load work, load work before handoff. Rushing this order is where cleanups overrun budget.

How much does it cost to clean up a vibe-coded MVP?

Between $50K and $250K for most apps, per rescue engagement data across the industry. Variance comes from three factors: how much of the data model needs restructuring, how many security findings the audit surfaces, and whether the app needs to be migrated to the founder's infrastructure from an agency or platform account. Audit-only engagements run $2,500–$8,000 and produce the scope document the full cleanup is priced against.

Is it cheaper to rewrite my MVP or clean it up?

Cleanup is cheaper for the large majority of apps with real users. Rewrite is cheaper only if the audit finds that more than about 60% of code paths have severe issues — at that point the existing code isn't providing enough baseline to build on. Rewrites also carry migration costs (data, URLs, integrations, customer re-onboarding) that are typically underestimated and sometimes exceed the direct engineering cost.

What's the first thing to fix in a vibe-coded app?

Secrets. If TruffleHog or Semgrep finds a live API key or database credential in Git history, that credential is compromised permanently — rotate it before any other work. The rotation is a two-hour task per service. Skipping it to work on architecture or features first is the highest-risk shortcut in a cleanup.

Does vibe coding cleanup include new features?

No. The cleanup scope stops at production-readiness — audit, security, data integrity, observability, CI/CD, scale hardening, handoff. New features start after cleanup on a separate engagement or with the founder's own team. Combining the two scopes is what turns a 6–8 week cleanup into a 16-week project with unclear exit criteria.