
CDC at Scale: Debezium, Flink CDC, and the Real-Time Replication Problem
Change data capture looks simple on a whiteboard. The database has a write-ahead log. Read it, stream the events, replay them into the warehouse. Real-time replication without nightly ETL.
Then production happens. Someone runs an ALTER TABLE on a 4TB table. The CDC connector hits a schema change it can't resolve, falls into resync mode, and spends the next 18 hours replaying the full table while lag climbs into the millions of events. At hour 20, the PostgreSQL replication slot hits its WAL retention limit, and the DBA gets paged because disk usage on the primary is climbing fast enough to threaten writes.
This is the Schema Change Tax: the real cost of change data capture at scale, paid in emergency engineering hours every time a source schema evolves. It is the most consistent production pain in the CDC category, and the one buyers underestimate most.
CDC is not a simple streaming pattern. It is a distributed system with latency-sensitive state and an operational surface that rivals a stream processor. "Read the log, stream the rows" is the mental model that produces that failure.
What CDC Actually Is
CDC systems read the database's native replication log — PostgreSQL WAL, MySQL binlog, Oracle redo log, SQL Server transaction log — and convert each committed transaction into a stream of row-level change events. An insert produces a create event with the new row. An update produces an event with before and after row images. A delete produces a delete event with the old row.
The downstream consumer is typically Kafka, a warehouse like Snowflake or BigQuery, or a lakehouse table format like Iceberg or Delta. The consumer replays the stream to maintain a near-real-time replica of the source.
The main approaches:
- Debezium on Kafka Connect. The de facto open-source standard. Kafka Connect worker pool runs connector instances for MySQL, PostgreSQL, Oracle, SQL Server, MongoDB, and others. Events land in Kafka topics per table. Consumers then write those topics to a warehouse or lakehouse. This is the pattern Debezium documents and the most common production shape.
- Flink CDC. Flink CDC embeds the change capture directly into a Flink job, skipping Kafka for the transport layer. Sources include MySQL, PostgreSQL, Oracle, SQL Server, TiDB, MongoDB, Vitess, Db2, OceanBase. Latency is typically sub-second for binlog synchronization. The trade-off is that Flink CDC inherits Flink's full operational model.
- Managed CDC services. Estuary Flow, AWS Zero-ETL integrations, and Airbyte (for simpler cases) remove the Debezium + Kafka Connect operational burden at the cost of per-event pricing and vendor-specific feature gaps. Estuary claims millisecond latency; AWS zero-ETL is tiered per source (Aurora under 15 seconds, DynamoDB 15 minutes, Salesforce 1 hour).
Who Actually Needs CDC
The short list of use cases where CDC pays for its operational cost:
- Real-time data warehouse replication at high change volume. When nightly batch loads can't keep up with source change rate, or when analytical latency SLAs require sub-hour freshness on specific tables.
- Event-driven architectures built on database state. Downstream services that react to every order, every user update, every payment. These consumers fundamentally need the stream.
- Audit logs and compliance. Regulated industries that need to capture every change for forensics, with immutable records.
- Multi-region replication with transformation. When the destination is a different database engine, a different cloud, or a warehouse schema that doesn't match the source.
- Feeding streaming analytics. Flink or Spark Streaming jobs computing real-time aggregations on transactional data.
The use cases where CDC is usually wrong:
- Moving 10 tables into a warehouse for daily reporting. Use AWS zero-ETL or scheduled extracts — running Kafka Connect for this is infrastructure theater.
- Low-change-rate operational tables. A 500-row customers table that updates twice a day does not warrant a replication slot; an incremental query on a 5-minute cadence is simpler and just as fresh for the business consumer.
- Small teams with competing priorities. CDC has enough operational complexity that three engineers running it alongside four other priorities will spend more hours on replication incidents than on data value.
The Schema Change Tax
Every CDC system handles schema changes badly. The specific failure modes vary, but the pattern is universal.
When a source schema changes — ADD COLUMN, ALTER COLUMN TYPE, RENAME COLUMN, DROP COLUMN — the CDC connector has to reconcile the change. If the change is backward-compatible and the connector supports incremental snapshot, some systems handle it gracefully. If it's not, the connector typically needs a full resync.
Full resync on a large table is where the operational crisis lives:
- A 4TB table needs to be read from the source, serialized through the connector, published to Kafka (or a stream equivalent), and replayed by every downstream consumer.
- During the resync, the replication slot keeps accumulating WAL. If WAL retention fills disk faster than the resync can drain it, the source database runs out of disk and stops accepting writes.
- Consumers see lag grow. Downstream analytical queries against the warehouse return partial data. Some systems route that partial data to the dashboard anyway.
- Recovery windows range from hours (small tables) to days (multi-TB tables) depending on network throughput and downstream write parallelism.
Zero-ETL has the same underlying problem but hides it better. AWS zero-ETL documents 20-90 minute DDL outages on large tables. The outage is documented. It is still 20-90 minutes of stale data in the warehouse.
Mitigation patterns that work:
- Schema change coordination. Source-side DDL goes through a review process that includes the CDC pipeline owner. No schema changes without a migration plan for the CDC layer.
- Snapshot.mode = never for routine restarts. Avoid accidentally triggering full snapshots by misconfiguring
snapshot.mode. Production CDC connectors almost always wantschema_onlyorno_dataafter the initial snapshot. - Per-table topic fanout with selective resync. When one table schema changes, you only want to resync that table's topic, not every table the connector manages. This requires clean per-table topic naming and ownership.
- Debezium's incremental snapshot feature lets you snapshot a subset of tables without stopping the connector. The Debezium 3.0 release added improvements to PostgreSQL snapshot isolation modes (serializable, repeatable_read, read_committed, read_uncommitted) that reduce lock contention during incremental snapshots.
The Replication Slot Problem
PostgreSQL CDC specifically has a class of problems around replication slots that is worth calling out.
A PostgreSQL logical replication slot is a server-side object that tracks how far a subscriber has consumed the WAL. The slot holds a reference to the oldest WAL that the subscriber still needs. If the subscriber stops consuming (Kafka Connect worker dies, network partition, connector bug), the slot stays at its current position. PostgreSQL cannot clean up WAL newer than the slot's position.
Failure mode: a Debezium connector dies on a Friday night. Monday morning, the PostgreSQL primary has accumulated 72 hours of WAL that couldn't be cleaned. Disk is 95% full. WAL archive jobs are failing. Writes to the primary are about to block.
Operational discipline that prevents this:
- Alerting on slot lag. Monitor
pg_replication_slots.confirmed_flush_lsnagainstpg_current_wal_lsn()and alert when lag exceeds a threshold (typically 1GB or 1 hour). - Alerting on slot age.
age(slot_xmin)exceeding a threshold indicates a stuck slot. - Dedicated WAL volume. Never let the CDC-consuming WAL share disk with the primary's data directory.
- Slot cleanup automation. If the slot is clearly abandoned (owner dead, no consumer for 24+ hours), drop it. Losing CDC continuity is better than losing the primary.
MySQL binlog-based CDC has its own failure modes but the slot problem is PostgreSQL-specific. Oracle redo logs have a different model again. If you run CDC across multiple source engines, each needs its own operational playbook.
The Primary Key Requirement
Most CDC systems fail silently on tables without primary keys or with non-unique primary keys.
Debezium, Flink CDC, and most managed offerings assume that primary key uniqueness lets you map CDC events to destination records. When a table has no primary key:
- Updates look like delete-then-insert pairs, which is wrong.
- Destination upserts produce duplicates.
- Row-level ordering becomes unreliable.
Legacy databases often have tables with surrogate keys that aren't primary keys at the database level, or with composite keys that aren't declared as unique. CDC adoption typically requires auditing every source table and declaring missing primary keys before the pipeline is reliable.
The "Real-Time" Expectation Gap
CDC markets on sub-second latency. Production latency is almost always higher, and the sources of delay are structural.
End-to-end CDC latency is the sum of:
- Source commit to WAL visibility. PostgreSQL commits data and writes WAL; the WAL becomes visible to logical decoding with small delay.
- WAL read to connector. Network round-trip, connector parsing.
- Event serialization. Avro or JSON encoding, Kafka producer batching, Schema Registry compatibility check.
- Kafka commit acknowledgment. Replication factor and acks setting determine how long this takes.
- Consumer read latency. Downstream consumer polling interval, consumer group rebalances.
- Destination write latency. Warehouse staging, transformation, commit.
A Debezium connector doing MySQL binlog replication into Kafka into a Spark job into Iceberg is not a sub-second pipeline. It is typically a 5-30 second pipeline when healthy, and a 5-30 minute pipeline when one component falls behind.
For sub-second latency, the architectural choice is Flink CDC writing directly to the destination (bypassing Kafka) or an in-database replication feature (PostgreSQL logical replication to PostgreSQL, for example). Adding Kafka adds hops; hops add latency.
When CDC Is Not the Answer
Three specific scenarios where teams pick CDC and later regret it:
- The source is an API, not a database. You cannot CDC from Salesforce's REST API; you get API pagination and webhook streams, which have different guarantees. Reverse-engineering "CDC" from API polling is a different problem.
- The requirement is "I need fresh data in the warehouse." This is an ETL latency question, not a CDC question. If daily batch hits your freshness SLA, keep daily batch. If hourly works, use hourly. CDC is the right answer only when the latency budget is below what batch can deliver.
- The business doesn't actually act on sub-second changes. Most analytical dashboards are checked on a cadence of minutes or hours. CDC into a warehouse that powers a Tableau dashboard refreshing hourly is infrastructure that serves no business consumer.
See our streaming vs. batch post for the longer treatment of the Freshness Gap — the distance between the architecture's capability and what the business actually consumes.
We don't have systematic failure-rate data across Debezium, Flink CDC, and managed services. The Debezium GitHub issue tracker is the best public source and shows a consistent pattern of WAL slot management issues, OOM from long-running transactions blocking WAL cleanup, and performance limits on logical decoding. These are reports from the people who hit problems, not systematic sampling — the true failure-rate picture lives inside vendor support channels.
What is clear: CDC is operationally harder than batch ingestion at the same scale, and the failure modes are less familiar to most teams. Debezium 3.0 and later releases reduced some pain points around snapshot isolation and TOAST column handling. The category's operational surface has not shrunk.
CDC works when schemas don't change. Schemas change.
Need help designing a CDC architecture that survives schema changes, or rescuing a Debezium cluster stuck on slot lag? Talk to an engineer — we'll tell you honestly if we can help.
Frequently Asked Questions
What is change data capture (CDC)?
CDC reads a database's native replication log — PostgreSQL WAL, MySQL binlog, Oracle redo log — and converts each committed transaction into a stream of row-level change events. Downstream consumers (Kafka, warehouses, lakehouses) replay those events to maintain a near-real-time replica. CDC enables sub-minute replication latency that scheduled batch ingestion cannot achieve.
What is the difference between Debezium and Flink CDC?
Debezium runs on Kafka Connect and produces events into Kafka topics for downstream consumers to read. Flink CDC embeds change capture directly into a Flink job, skipping Kafka as a transport layer. Debezium is the more common production pattern because Kafka acts as a durable buffer. Flink CDC gives lower latency at the cost of inheriting Flink's full operational model.
Why do CDC pipelines fail on schema changes?
When a source schema changes (ALTER TABLE, ADD COLUMN, etc.), the CDC connector must reconcile downstream schema. Non-backward-compatible changes typically trigger a full resync. On a large table this can take hours to days, during which replication slots accumulate WAL and downstream data is stale. This Schema Change Tax is the most common production failure mode across Debezium, Flink CDC, and managed services.
How do you monitor Debezium replication slots in PostgreSQL?
Alert on pg_replication_slots.confirmed_flush_lsn lag against pg_current_wal_lsn() — typically 1GB or 1 hour is the warning threshold. Alert on age(slot_xmin) to catch stuck slots. Use a dedicated WAL volume so slot accumulation cannot fill the primary's data disk. If a slot is clearly abandoned, drop it — losing CDC continuity is better than losing the primary database.
Is AWS zero-ETL a replacement for Debezium?
For moving transactional database data into Redshift, yes, for many use cases. AWS zero-ETL delivers Aurora changes with p50 under 15 seconds and no operational overhead. It does not replace Debezium for custom downstream processing, multi-destination fan-out, or event-driven architectures that need Kafka-style durability. It does replace Debezium for the "move tables to warehouse" use case that drives most CDC adoption.
Related posts

Airflow vs. Dagster vs. Prefect: A 2026 Orchestration Comparison
Airflow 3 shipped breaking changes, Dagster bet on assets over DAGs, Prefect cut 90% of runtime overhead. A decision framework based on team size, not feature lists.

Data Pipeline Observability: Monitoring Airflow + dbt Without Drowning in Alerts
200 tests, 80 weekly alerts, and a Slack channel nobody reads. That's alert debt. The monitoring stack that catches real failures without burying your team.

Data Quality at Scale: Great Expectations vs. Soda vs. dbt Tests
dbt tests catch transformation bugs. Great Expectations catches source corruption. Soda monitors everything between. A decision framework for choosing data quality tools in 2026.