ivinco
From 10 Users to 10,000 — What Changes in Your Stack and Why

From 10 Users to 10,000 — What Changes in Your Stack and Why

Ivinco Team·

MVPs hit the same wall at 1,000 users. Different stacks, different teams, same failure mode — random 500 errors that look like a flaky database and turn out to be connection pool exhaustion. We've rebuilt enough of these to spot the pattern on the first audit call.

The architecture that works at 10 users fails at 1,000 for reasons that have nothing to do with feature velocity. The architecture that works at 1,000 fails at 10,000 for a different set of reasons. Every order of magnitude adds one architectural concern the step before was blind to.

We call this the Order-of-Magnitude Rule. You don't need to build for 100,000 on day one — that's premature optimization, and it's the mistake the other side of the industry makes. You do need to know which concern is coming next, because the wrong architecture at 1,000 users is a rewrite by the time you hit 10,000.

This post walks through four orders of magnitude — 10 → 100, 100 → 1,000, 1,000 → 10,000, 10,000 → 100,000 — and names what breaks at each step, what changes, and how much lead time you have before the cliff becomes visible.

10 to 100 Users: The Starting Architecture

Single process. Single database. Monolith deployed to one cloud region. At this scale, almost anything works. Vercel plus a Postgres database. Railway plus a Docker container. Fly.io plus Neon. The decisions don't matter because the load doesn't stress any of them.

What matters at 10–100 users is not the architecture. It's the shape of the architecture that lets you change it later without a rewrite:

  • Code in source control. Obvious. Still skipped. GitHub or GitLab from day one, not a shared folder.
  • Separate config from code. Secrets in environment variables or a secrets store, not hardcoded.
  • Use a managed database. Neon, Supabase, PlanetScale, Aurora Serverless. Don't run your own Postgres on a single VM unless you're willing to run backups, replicas, and upgrades yourself.
  • One region, one replica. Multi-region at this scale is engineering cost with zero benefit. AWS's own scaling guide is explicit that multi-AZ starts paying off around 10,000 users, not before.

What breaks: essentially nothing, if the features work. If the features don't work, no architecture saves you. This is the stage where the product-market-fit question lives, not the architecture question.

100 to 1,000 Users: Where the First Cracks Appear

Three things show up around 500–1,000 concurrent users, all of them database-related.

1. Connection exhaustion. This is the one we see most often on audit. Postgres defaults to 100 connections on most managed platforms. Serverless functions open a new connection per invocation. A Next.js app with 50 concurrent users and server-rendered pages can blow through the connection limit in one busy minute. The errors look random. They aren't. PgBouncer or Supabase's built-in pooler handles this at the free tier. Neon includes it by default.

2. Slow queries that were fast at 10 rows. A SELECT * FROM orders WHERE user_id = ? without an index runs in 2ms with 50 rows and 200ms with 50,000. At 500 users generating 10 orders each, that's suddenly a database bottleneck. Turn on pg_stat_statements, watch the top 10 slowest queries weekly, and add indexes as the working set grows. The diagnostic tool is more important than any specific fix — you can't optimize what you're not measuring.

3. External service rate limits. Stripe, SendGrid, OpenAI — all have per-minute limits that you never hit at 10 users and always hit at 1,000. Queue the calls (BullMQ, Inngest, Trigger.dev) and process them at a controlled rate. Handle 429 responses with exponential backoff or the retry storm cascades through your application and takes down unrelated endpoints.

The lead time from 100 to the 1,000-user wall is typically 3 to 6 months for a product with early traction. Connection pooling is a day of work. Indexing is a weekly habit. Queues are a 2–3 day engagement. All three are cheap to add as reactions to signals — and painful to add while a customer is on the phone about a broken checkout.

1,000 to 10,000 Users: The Order-of-Magnitude Jump

This is the jump where MVPs fail structurally. The single database that served 1,000 users serves 10,000 only if four things are true that probably weren't true before:

1. Read replicas offloading the read traffic. A fintech scaling writeup on Medium documents two read replicas reducing read latency by 50% and buying six months before sharding became necessary. At 10,000 users the same math holds: most queries are reads, and most reads don't need to hit the primary. Route them to replicas.

2. Caching layer. Redis or Memcached sitting in front of the database, catching the top 20% of queries that produce 80% of load. ElastiCache, Upstash Redis, or Vercel KV for serverless-compatible deployments. Cache the user session. Cache the product catalog. Cache anything that changes once a day and is read a thousand times.

3. CDN for static assets. Images, JS bundles, CSS — none of this belongs on your application servers at 10,000 users. Move it to S3 (or Cloudflare R2) and put CloudFront, Cloudflare, or Bunny in front of it. AWS documentation reports 60–80% origin load reduction from a CDN configured correctly — that's headroom your application servers didn't have before.

4. Background jobs for anything that isn't user-facing. Email sending, report generation, LLM calls, data imports — all off the request path. If a user's HTTP request waits for an LLM, your p99 latency is whatever the LLM's worst day looks like. Queue it.

The load test at this scale is non-negotiable. Run k6, Artillery, or Locust against your staging environment at 10x expected peak concurrent users. Whatever breaks first is the thing to fix before the users find it. Connection pool exhaustion, memory pressure, cache misses on a cold Redis — all of them have signatures in a load test that they don't have in manual testing.

Lead time from 1,000 to 10,000: 6 to 12 months for a product with steady 20%/month growth. The four items above can be added incrementally — replicas first, caching second, CDN third, queues fourth — or all at once as a "production-ready milestone" pass before a major launch. Budget 2–4 weeks of platform engineering, plus the ongoing running cost of the new infrastructure layers.

10,000 to 100,000 Users: Horizontal Scaling Territory

This is where single-node architectures give way to distributed ones. Three transitions define it:

1. Multi-region or multi-AZ for availability. A single region, single replica can give you 99.5% uptime. For 99.9% and above, you need redundancy across failure domains. AWS RDS Multi-AZ (synchronous replication to a standby in a different availability zone) is the entry point. Multi-region (for latency, not just availability) is the next step — and the cost and complexity jump is real.

2. Horizontal application scaling. One server doesn't serve 100,000 users. You need a load balancer (ALB, Cloudflare Load Balancer, nginx in front of multiple app instances) distributing requests across horizontally scaled application servers. For a container-orchestrated deployment, Kubernetes with a Horizontal Pod Autoscaler is the default pattern; for simpler stacks, a Vercel scale-to-zero deployment or a Railway horizontal scale does the job.

3. Database partitioning or sharding. At 100,000+ users, a single Postgres instance starts hitting hard limits: storage ceilings, write throughput ceilings, backup windows that exceed maintenance windows. Options: partition by time (orders_2026_q1, orders_2026_q2) for time-series-heavy workloads; shard by tenant (orders_tenant_1, orders_tenant_2) for multi-tenant SaaS; or move specific hot tables to a different database engine entirely (ClickHouse for analytics, DynamoDB for session data).

What breaks at this stage is politics, not architecture. By 100,000 users you likely have a team of more than three engineers. Decisions about sharding, region topology, and data migration take weeks to align on, because they're no longer reversible in a sprint. The architectural work is the easy part; the organizational coordination is what the lead time is mostly spent on.

Honest boundary on this section: we've run replicas, caching, CDN, and queues across many production systems — the 1,000-to-10,000 work is our core competency. Live shard rebalancing at 100,000+ users is a different operational discipline, one we've seen from the outside more than from the inside. The principles above are correct; the execution at that scale involves edge cases (write amplification during reshard, cross-shard transaction anomalies) that need specialists who've done it before. For a founder at that scale, bring in someone whose last job was running a 100K+ user production system, not someone whose last job was building the MVP.

What the Order-of-Magnitude Rule Doesn't Cover

Two cases this framework ignores. First, products that spike — a news app that goes from 100 users to 100,000 during a breaking event, then returns to 100 the next day. These don't follow the order-of-magnitude progression; they need serverless or auto-scaling infrastructure from day one. Second, products with unusual traffic shapes — a conferencing product where 100 users generate 100,000 messages per minute, or a gaming backend where traffic is bursty and stateful. These violate the "users as the load unit" assumption the rule relies on.

For everything in between — the standard SaaS curve where users grow steadily and traffic scales with them — the rule holds: one new architectural concern per 10x, and the concern is predictable in advance.

The 4-Item Pre-Launch Checklist

Before any MVP goes live at any scale, four things protect against surprises:

  1. Connection pooling configured. Even at 10 users. Takes an hour. Makes the 1,000-user cliff disappear.
  2. Slow query logging enabled. pg_stat_statements on, baseline established, weekly review set up. Finds the problems before users do.
  3. One synthetic load test recorded. 10x peak traffic on staging. Tell-truth whether the architecture holds.
  4. A named next-step plan. "At 2,000 DAU we add read replicas. At 5,000 we add Redis. At 10,000 we add a CDN and background jobs." Written, not remembered.

Teams that write those four items down before launch ship infrastructure that grows with the product. Teams that don't end up rewriting at each scale boundary, which costs 2 to 6 weeks per boundary and throws the roadmap off for a quarter.

What This Actually Reduces To

Every order of magnitude names one thing the previous architecture was wrong about. Connection limits. Read traffic concentration. Static asset throughput. Failure domain dependency. The architecture follows the load, not the other way around.

The product works at 10 users or it doesn't. The business works at 10,000 users or it doesn't. The architecture only has to not be in the way at each step in between.

Frequently Asked Questions

At what user count do I need read replicas?

Between 1,000 and 5,000 active users for most SaaS products, triggered by symptoms rather than a fixed threshold: p95 query latency trending up, read-heavy queries competing with writes, or replication lag appearing in monitoring. For products with heavy read-to-write ratios (content sites, analytics dashboards), replicas help earlier; for write-heavy systems they help less. The honest signal is the database dashboard showing sustained CPU above 70%, not a user count.

When does an MVP need a CDN?

When static asset traffic starts competing with application traffic on the same server, or when users in regions far from your origin complain about page load times. AWS documents 60–80% origin load reduction from a correctly configured CDN. Practically, add Cloudflare (free tier is enough) between 1,000 and 10,000 users; before that, the Vercel or Netlify edge network handles it for you if you're deployed there.

Should I use Kubernetes for an MVP with 1,000 users?

No, unless you already run Kubernetes for other workloads and adding one more is cheap. For a new product at 1,000 users, Vercel, Railway, Fly.io, or Render deliver the same user outcome with a fraction of the operational burden. Kubernetes becomes the right answer around 5–10 engineers working on the product, or when you need workload primitives (StatefulSets, CronJobs, fine-grained networking) that PaaS platforms don't expose.

What's the biggest architectural mistake MVPs make when they try to scale?

Assuming connection pooling is a scaling feature you add later. It isn't — it's a correctness feature that prevents random 500 errors at the 500–1,000 concurrent user range. Without it, your app hits a cliff that looks like a flaky database rather than a scaling problem, and the team spends weeks debugging the wrong thing. Add connection pooling on day one, even at 10 users.

Is serverless still relevant at 10,000 users?

Yes, for certain workloads — especially bursty or infrequent ones (webhook handlers, background batch jobs, scheduled tasks). For the hot path of a 10,000-user SaaS app, serverless cold-starts and per-invocation connection overhead often make dedicated servers or long-running containers more efficient. The transition from "everything serverless" to "hot path on containers, bursty work on serverless" usually happens between 5,000 and 20,000 users depending on traffic pattern.

Need help? Talk to an engineer.