
The Engineering Checklist VCs Use for Series A Technical Due Diligence
When a VC opens technical due diligence on a Series A target, they're not reading your code. They're looking for a specific set of artifacts — runbooks, DORA numbers, infrastructure diagrams, backup test logs, access audit reports. The artifacts either exist or they don't. The build either passes diligence or it doesn't.
We run this checklist on every MVP we ship. The version a founder needs 12 months later costs roughly 10x to retrofit. That's the Diligence Gap — the distance between a product that works and a product a technical diligence reviewer will pass without flagging risk.
This post breaks down the actual 20-item checklist VCs use in 2026, organized into five categories. Each item is either present in your build or it's a deal risk. The checklist is consolidated from HyperNest Labs's 50-point diligence guide, Nemausat's Series A risk analysis, and the 2026 DORA benchmarks that define what "good" looks like on delivery performance.
Category 1: Operational Readiness
This is the category that kills rounds most often. The product works. The team can't prove it will keep working.
1. DORA metrics measured and trending. The 2026 benchmark for elite teams: deployment frequency multiple times per day, lead time from commit to production under one day, change failure rate under 0.5%, and mean time to recovery under one hour (DORA 2026 benchmarks). A Series A target doesn't need to hit elite on all four, but it needs to measure all four. "We don't track that" is a diligence flag.
2. Infrastructure reproducible from code. Terraform, Pulumi, or CloudFormation in version control that provisions the full stack. Nemausat calls the opposite pattern "Snowflake Infrastructure" — hand-crafted cloud environments that prevent reliable scaling and DR. If your engineering lead has to click through the AWS console to explain production, the diligence reviewer notes a scaling risk.
3. Runbooks for incident response. Written procedures for deploy, rollback, the top 5 production failure modes, and the path to contact on-call. The test is behavioral: can an engineer who joined yesterday execute a rollback from the runbook alone? If not, your product has a single-person dependency that Nemausat calls the "Bus Factor" — a direct investor-flagged risk.
4. On-call rotation and alert routing. Named engineers responsible for production at defined hours. Alerts routed to humans, not an email inbox nobody reads. If the answer to "who gets paged at 3 AM?" is "we'll figure it out when it happens," the answer for the diligence reviewer is "this team has not operated production."
Category 2: Observability and Monitoring
Diligence reviewers want to see the dashboards. If you can't show them, the underlying systems probably aren't there either.
5. Structured logging in production. Centralized log aggregation (Datadog, Axiom, BetterStack, or self-hosted Loki) with queryable logs, not stdout streams. Structured JSON logs, not console.log. Retention policy documented.
6. Error tracking with actionable alerts. Sentry or equivalent catching production errors with source maps, release tracking, and user context. The diligence test: can you show the reviewer the error rate for last Tuesday without anyone logging into a server?
7. Metrics and dashboards for critical paths. Prometheus + Grafana, Datadog, or equivalent. At minimum: request latency p50/p95/p99, error rate, and a business-metric dashboard showing DAU or equivalent primary KPI. Per HyperNest Labs's diligence checklist, "monitoring and alerting covers critical paths" is item 48 — dashboards and alert rules for components that could fail silently.
8. Uptime SLO with measured compliance. A stated target (e.g., 99.9% monthly) and real measurement against it. The HyperNest checklist explicitly calls for "uptime documentation showing 99.9%+ performance over 12 months with explanations of major outages." No SLO document means the team hasn't thought about reliability as a commitment.
Category 3: Security and Access Control
This is where most MVPs fail cold. The diligence firm typically runs Snyk, GitHub Dependabot, and a basic OWASP scan against the codebase. What they find defines the first risk paragraph of their report.
9. No secrets in Git, ever. The HyperNest checklist specifies "no hardcoded credentials or secrets in codebase — verified through automated scanning tools like gitleaks or trufflehog." A single leaked key is a finding; three is a conclusion the reviewer writes into the report.
10. Secrets in a real secrets store. AWS Secrets Manager, HashiCorp Vault, Doppler, or Infisical — not .env files. HyperNest item 30: "Secrets management uses proper tooling (not .env files in git)." This is non-negotiable for SOC 2 readiness.
11. IAM with least privilege. Named roles for humans and services. No long-lived AWS root credentials. No shared developer passwords to production. Every access should be traceable to an individual identity. Per the HyperNest checklist item 27, "access controls must follow the principle of least privilege across production systems."
12. Dependency vulnerability scanning. Dependabot, Snyk, or npm audit running on every build with a policy for response time on critical CVEs. Patch timeline documented. The CSA Research Note on AI-generated code found organizations using AI-assisted development saw security findings surge 10x — the diligence reviewer is looking for your defense against that noise.
13. Third-party pentest within 12 months. HyperNest item 24: "third-party penetration testing completed within 12 months with findings documented." A $15K–$40K engagement with a firm like Trail of Bits, Cure53, or Bishop Fox. Deferred findings tracked with remediation dates.
Category 4: Data, Backup, and Disaster Recovery
This is the category where founders discover their backup plan was a pg_dump cron job nobody had tested.
14. Documented backup strategy with RTO/RPO targets. HyperNest item 45: "disaster recovery plan with defined RTO/RPO targets and documented testing history." Recovery Time Objective (how long until we're back up) and Recovery Point Objective (how much data can we lose) stated as numbers, not adjectives.
15. Backup restores actually tested. HyperNest item 46: "backup restoration procedures verified through actual testing exercises." A backup you've never restored isn't a backup — it's a hope. Quarterly restore tests with documented results.
16. Database migration strategy documented. How schema changes reach production without downtime. Tools like Atlas, Flyway, or Django/Rails migrations with rollback policies. "We run migrations manually" fails diligence because the reviewer now has to assume every migration is a potential outage.
17. Data retention and deletion policy. Especially material under GDPR and CCPA. When does user data get purged? How does the right-to-delete request flow through the system? For SaaS targeting enterprise or EU customers, this is a blocker at diligence if undefined.
Category 5: CI/CD and Code Governance
The pipeline tells the reviewer how the team actually works. A weak pipeline is a leading indicator for every other category.
18. Every merge triggers automated checks. Tests, type checks, security scans (Semgrep or equivalent), lint, build. No manual testing as the last line of defense before production. HyperNest item 36 documents this as deployment frequency and lead time data — you can't measure lead time if your pipeline isn't the path code actually travels.
19. Production deploy requires a PR. No direct pushes to main. Branch protection rules on the main branch. CODEOWNERS file defining review requirements for sensitive paths (auth, payments, infrastructure).
20. License compliance and third-party code audit. The Nemausat "License Trap" — copyleft licenses (GPL, AGPL) can force unintended open-sourcing of your IP if bundled improperly. FOSSA or Snyk License Compliance scanning in CI. This becomes expensive to remediate late in legal diligence; catching it in engineering diligence is cheaper.
The Mystery Cloud Bill
Between categories, one additional diligence flag deserves its own mention: can you map your monthly cloud spend to specific features or customers? Nemausat calls the opposite the "Mystery Cloud Bill" — an inability to explain margin at scale. A $40K/month AWS bill that nobody can attribute by product surface is a $3M/year financial question mark with no owner.
The fix is mechanical: tag every resource with a product/customer/environment label, roll the bill up by tag monthly, and review with finance. But it has to exist before the diligence pack gets assembled, because there's no retroactive path.
What the Checklist Doesn't Catch
The 20 items above pass a technical diligence review. They don't prove the product is good. They say nothing about whether the team ships in two weeks or two months. And they're silent on whether the architecture holds at 100x load.
We scope to pass this checklist because it's the measurable, defensible minimum — the floor below which a Series A round stalls. What sits above the floor is judgment: architecture trajectory, team scalability, whether product-market fit is real. VCs pay diligence firms to answer the checklist question because they themselves are busy answering the judgment one.
The Founder's Move
The expensive version of this checklist is retrofit. An MVP shipped by an agency that wasn't measuring DORA numbers, didn't hook up structured logging, never wrote a runbook, kept secrets in .env, and deployed by git push requires approximately 8–16 weeks of dedicated platform engineering to pass diligence. That's $60K–$150K of delayed Series A funding cost.
The cheap version is building to pass it from day one. Almost every item on the list costs the same to include upfront as to skip — the difference is whether the team writing your MVP has operated production before, or has only delivered demos.
If the team writing your build can't name the tool they'll use for each of these 20 items without looking it up, the Diligence Gap is already widening. The artifacts either exist or they don't. Diligence only tells you which.
Need help? Talk to an engineer.
Frequently Asked Questions
What do VCs actually check in Series A technical due diligence?
VCs typically hire a diligence firm to audit across five categories: operational readiness (DORA metrics, runbooks, IaC), observability (logging, error tracking, dashboards), security (secrets management, IAM, pentests), data/DR (backups, RTO/RPO, migrations), and CI/CD (automated checks, branch protection, license compliance). The audit produces a risk report that materially affects valuation and term sheet conditions.
What's the difference between technical due diligence and a code review?
A code review evaluates quality, style, and correctness of code. Technical due diligence evaluates the engineering organization's ability to build, operate, and secure production systems over time. Diligence covers infrastructure, processes, security posture, backup testing, and compliance — things that can exist independent of any line of code. A clean codebase with no runbooks, backups, or monitoring still fails diligence.
How long does Series A technical due diligence take?
Two to six weeks for a standard engagement. The diligence firm typically spends the first week requesting artifacts (infrastructure diagrams, access audits, backup test logs, pentest reports), then runs automated scans against the codebase, then interviews the engineering lead. If artifacts are missing, the clock stops while the team creates them — which is where rounds slip.
What's the cost of failing technical due diligence?
Rarely a full deal kill. More commonly: reduced valuation, deferred close, or specific remediation requirements as closing conditions. Severe findings (ongoing security incidents, GPL contamination of proprietary IP) can trigger re-diligence after fixes. The real cost is time — a diligence redo adds 4–12 weeks to a round in a market where speed matters.
Can I retrofit these items after raising Series A?
You will either way. The question is whether you do it with term-sheet pressure or before it. Retrofitting observability, CI/CD, and backup testing after the round is the common path and costs 2–4 engineers for 8–16 weeks. Doing it before diligence costs the same engineering hours but keeps them out of the critical path of the raise itself.
Related posts

Fractional CTO vs. MVP Agency vs. Build In-House — A 2026 Comparison
Fractional CTO solves "who decides?" MVP agency solves "who builds?" In-house solves "who stays?" Three options, six dimensions, and the combinations that actually work. The single-option pick is the wrong default.

What Does $80K Actually Buy From an MVP Development Company in 2026?
At $150/hour, $80K buys 533 engineering hours. At $70/hour, 1,143. Same visible scope, 2x labor spread. Here's the math most founders never see — and the Invisible Line Item that tells the difference.

The Technical Debt You Should Keep in Your MVP (And the Kind That Kills)
Good debt trades speed now for bounded cost later. Bad debt hides a step function that fires at the worst moment. Five shortcuts to keep forever, five that compound silently, and a better rule than the 20% sprint tax.