ivinco
Streaming vs. Batch: The Real Trade-Offs Nobody Talks About

Streaming vs. Batch: The Real Trade-Offs Nobody Talks About

Ivinco Team·

The most expensive data stack we see is a streaming stack built for workloads that aren't streaming. A six-person team, a three-node Kafka cluster, a two-node Flink cluster, a Schema Registry, a Kafka Connect cluster — feeding a Tableau dashboard opened at 9 AM. The engineering cost of running those four moving pieces is real. The business benefit over a micro-batch alternative is zero.

The original motivation is usually sound. Modern reference architectures stream. Netflix, Uber, OpenAI — the companies everyone reads about — all run Kafka and Flink at scale. "Modernizing" gets interpreted as "we should stream."

Two years later, the team is spending most of its engineering cycles managing rebalances, checkpoint failures, and schema drift instead of shipping data products. Nothing downstream of the pipeline would notice if it ran on a 5-minute Spark micro-batch job instead.

There is a name for this mismatch: the Freshness Gap. It is the distance between the latency your architecture delivers and the latency your business can actually consume. When that gap is wide — when you pipe data with sub-second freshness into a report that gets opened at 9 AM tomorrow — the streaming infrastructure is pure overhead.

What "Streaming" Actually Means in 2026

Streaming means four different things in 2026, and teams evaluating "should we go streaming?" rarely separate them. Each has a different operational and cost profile.

Flink treats data as an unbounded stream with stateful operators. State is kept local to each task, snapshots are taken asynchronously via checkpointing, and exactly-once semantics are achieved by rewinding to recorded offsets on failure. Event-time latency is commonly sub-second and can be single-digit milliseconds when the pipeline is tuned.

This is what "real streaming" looks like in production. It also has the highest operational surface area: checkpoint tuning, state backend selection (RocksDB vs. heap), savepoint management during upgrades, backpressure diagnosis, and watermark logic that has to match your actual event-time semantics.

Micro-batch streaming (Spark Structured Streaming)

Spark Structured Streaming processes data in small batches — typically 1 second to 1 minute — and calls it a stream. For most analytical workloads this is indistinguishable from true streaming, and the lower operational complexity makes it the right choice more often than teams assume. The trade-off is latency: sub-second end-to-end freshness is achievable but not the default, and stateful operations have different correctness guarantees than Flink's.

The key honest read: if your downstream needs "near-real-time" in the marketing sense (a dashboard that feels live), Spark micro-batch at a 30-second trigger is likely enough. If you need sub-100ms for operational decisions, it isn't.

Embedded stream processing (Kafka Streams)

Kafka Streams is a client library, not a cluster. A Kafka Streams application is just a JVM process that reads from Kafka topics, processes state locally (in RocksDB or an in-memory store), and writes back to Kafka. There is no cluster to operate beyond Kafka itself.

This is the lowest-complexity option when Kafka is already in your stack and the processing logic can fit in service-owned code. The cost is coupling: your stream processing is now a property of the consuming services, which is great for service boundaries and terrible when you need centralized governance over 50+ transformations.

Zero-ETL integrations (managed latency bands)

The managed cloud offerings have created a fourth category that doesn't fit the classic streaming definition but now dominates the "we want data moved faster than nightly batch" conversation.

AWS's zero-ETL integrations have explicit latency tiers documented by source:

  • Aurora MySQL/PostgreSQL → Redshift: p50 under 15 seconds
  • RDS for MySQL → Redshift: minutes-range
  • DynamoDB → Redshift: 15-minute minimum
  • Salesforce → Redshift: 1-hour minimum

These numbers come from AWS's official zero-ETL documentation and the per-source integration guides. They are latency floors, not averages, and they hide real trade-offs — DDL changes trigger 20-90 minute resyncs on large tables, and each Redshift cluster is capped at 50 active integrations. But for most teams asking "do we need Flink?", the honest answer is "no, you need zero-ETL from your transactional database and a daily dbt run."

The Freshness Gap in Practice

Before picking tools, answer one question: what is the latency the business actually consumes?

"Real-time" is a word buyers use for anything from milliseconds (fraud detection, bid optimization) to hours (an operations dashboard that currently updates nightly). The architectures that serve those two ends differ by an order of magnitude in both capability and cost.

A common pattern: an operations team wants "real-time" visibility into system status. Interrogating the requirement shows that the dispatcher checks the dashboard every 15 minutes during their shift. The functional latency requirement is 5 minutes, not 5 seconds. That choice — 5 minutes vs. 5 seconds — is the difference between a micro-batch Spark job writing to Snowflake and a full Kafka + Flink deployment with a low-latency serving layer.

The Freshness Gap formula is not subtle: if the consumer-side polling interval is longer than the pipeline's freshness, every millisecond below that interval is operational cost with zero business value.

The gap can go the other way. Nightly batch for pricing data when ad auctions require sub-second freshness — that is a real streaming use case. But it is not the common one in the mid-market analytics workloads that drive most ETL decisions.

Who Actually Needs True Streaming

The industry's canonical streaming case studies are worth reading, because they illustrate the scale at which the operational complexity pays off — and the narrow shape of the problems for which it does.

OpenAI disclosed at Current 2025 that it uses Kafka and Flink for real-time AI data pipelines: ingestion from services and users across multiple cloud regions, stateful transformations, and continuous feedback loops supporting ChatGPT. The disclosed technical pieces — PyFlink in production with multi-region failover, a custom Kafka Forwarder that converts pull-based consumption into gRPC push — are the kind of engineering that only makes sense when the business genuinely cannot tolerate batch latency.

Kai Waehner's 2026 streaming trends post cites Siemens, Uber, CrowdStrike, Santander Bank, and Robinhood as companies that built reliable real-time error handling at extreme scale. Fraud detection. Rate limiting. Operational alerting. These are genuinely streaming-shaped problems: the business decision being driven by the data has to happen faster than batch can turn around.

Netflix's Keystone platform processes event streams at internet scale through Kafka, with Flink jobs consuming topics for downstream analytics and ML features. That is not a small-team architecture. That is a system built by dozens of platform engineers to support an organization where data pipeline latency directly affects product decisions.

The pattern across these cases: streaming is worth its cost when the downstream action is automated, time-bounded, and cannot wait for batch. Human dashboards are almost never in that set.

When Streaming Is the Wrong Answer

Here is the mirror list — workloads where streaming is a net cost, not a net benefit:

  • BI dashboards with daily or hourly refresh. If the consumer is a human who checks it on a cadence longer than 5 minutes, micro-batch at a 5-minute trigger is strictly better than Flink. The dashboard doesn't care which one wrote the data.
  • ML training features updated weekly or monthly. Feature engineering for batch model retraining does not need streaming ingestion. Streaming a feature that is recomputed every Monday is wasted infrastructure.
  • Financial reporting. Daily or month-end close is by definition a batch problem. Streaming the general ledger does not make the close faster; it just adds fragility to a regulated workflow.
  • Data warehouse ELT for analytics. If the warehouse is the system of truth and analysts query it, the right latency is usually driven by when analysts need fresh data — which is hours, not milliseconds.
  • Small-scale CDC into a warehouse. For low-volume operational tables, zero-ETL managed integrations or a scheduled extraction every 15 minutes are both simpler than running Debezium + Kafka + Flink yourself.

The failure mode is consistent: a team picks Flink because the hyperscaler reference architecture used Flink. Two years later, they are operating Kafka + Flink + Schema Registry + Kafka Connect to move ten tables into a warehouse that serves a daily dashboard. The alternative — dbt + Airflow + managed ingestion — would have shipped in a quarter of the time with a quarter of the on-call surface, and no one would know the difference downstream.

The Decision Framework

Two questions get most teams to the right answer.

Question 1: What is the longest latency the business can tolerate for this data, without changing outcomes?

Under 1 second puts you in operational automation territory — fraud, bidding, anomaly response. True streaming is likely the right answer; Flink or Kafka Streams.

Between 5 seconds and 5 minutes, the field is open. Zero-ETL, Spark Structured Streaming with micro-batching, or scheduled extractions at a fast cadence will all work. Complexity preferences should drive the choice.

Hours to days: batch. dbt on a schedule. Airflow with a daily or hourly cadence. Do not build a streaming system.

Question 2: Is the consumer a system or a human?

  • Systems that take automated action on data can consume sub-second streams productively. Dashboards cannot — humans check them on their schedule, not yours.
  • If the answer is "a system that takes action in less than 5 minutes," streaming is on the table. Otherwise it rarely is.

These two questions filter out most "we should stream it" proposals that come from the architectural aesthetics side rather than the business requirement side.

What Breaks When You Over-Engineer for Streaming

The operational cost of streaming-by-default is real and compounding:

  • Schema evolution is harder. Kafka topic schemas must be backwards-compatible across producers and consumers. Schema Registry enforces this, but it adds a deployment coordination step that batch pipelines skip entirely.
  • Exactly-once is a design constraint, not a checkbox. Flink provides exactly-once semantics through state snapshots and stream replay, but only when every sink supports transactional writes. If your consumer is a warehouse that doesn't handle transactional upserts natively, you are back to at-least-once and deduplication logic.
  • Backfill is an incident. Rerunning a batch job for last Tuesday is a --start-date=2026-04-08 flag. Backfilling a Flink job against historical data means replaying Kafka topics, re-computing state, and praying your watermark logic doesn't misbehave. The Scribe paper from Meta documents the scale at which these operational patterns are routine. At smaller scale, they are a crisis each time.
  • Debugging is different. Batch failures produce a stack trace, a start time, and an end time. Streaming failures produce metrics graphs, backpressure alerts, and lag that may have started 48 hours earlier.
  • Upgrades require savepoints. Flink state versioning across job upgrades requires savepoint management. Spark's batch jobs don't have this problem; they don't have persistent state between runs.

None of these are showstoppers for teams that need streaming. They are all pure tax for teams that don't.

We don't have a clean dollar threshold for when Flink operational cost justifies itself. Industry guidance — "when batch can't meet your latency requirement" — is correct and useless. The one test that has held up: whether there is a dedicated platform engineer whose job description includes the streaming layer. Without one, a Flink cluster operated on the side of other work becomes the on-call problem that consumes the most hours, and teams quietly revert to simpler stacks within 12-18 months.

Vendor consolidation matters too. Decodable was acquired by Redis, Google retired BigQuery Engine for Flink, and the number of independent managed-Flink vendors has shrunk. If you adopt a specific managed streaming service in 2026, budget for a migration within three years.

The Shift-Left Argument, Briefly

There is a legitimate streaming trend worth naming. Waehner calls it the Shift Left Architecture: moving analytical transformations earlier in the data lifecycle, from warehouse queries into the stream layer. The goal is to compute once and serve many consumers, rather than repeatedly running heavy SQL against a warehouse.

This argument is strongest when:

  • The same transformation feeds multiple downstream systems
  • The source data volume is high enough that warehouse query cost exceeds streaming compute
  • The transformation is genuinely stateless or uses state that fits in memory

When those conditions don't hold, a well-partitioned warehouse with incremental dbt models does the same work with lower operational cost.

If your dashboard refreshes once a day, your pipeline doesn't need to stream.


Need help deciding whether your stack needs streaming or a better batch architecture? Talk to an engineer — we'll tell you honestly if we can help.

Frequently Asked Questions

When should you use streaming vs. batch data processing?

Use streaming when a downstream system takes automated action on data within 5 minutes or less — fraud detection, rate limiting, operational alerting, feedback loops. Use batch when the consumer is a human checking a dashboard, a weekly ML retrain, or scheduled reporting. The latency the business can actually consume determines the architecture, not the word "real-time."

Flink is a true stream processor with stateful operators, asynchronous checkpointing, and sub-second event-time latency. Spark Structured Streaming runs micro-batches at 1-second to 1-minute intervals. For most analytical workloads, Spark micro-batching is operationally simpler and indistinguishable from streaming. For sub-100ms automated decisions, Flink is required.

For moving transactional database data into a warehouse, yes, for many use cases. AWS zero-ETL delivers Aurora changes to Redshift in under 15 seconds with no operational overhead. It doesn't replace Flink for stateful stream processing, multi-source joins, or sub-second operational decisions. It does replace Debezium + Kafka + Flink stacks built only to move tables into a warehouse.

What is Kappa architecture and when does it apply?

Kappa architecture unifies batch and streaming into a single stream-processing layer, typically Kafka plus Flink. It applies when the same data transformations serve both historical and real-time consumers and the team has the platform engineering capacity to run a production streaming stack. For most mid-market teams, Lambda architecture or batch-only is lower operational cost with equivalent business outcomes.

How do you tell if your team actually needs streaming?

Two filters: ask what the longest tolerable latency is before the business outcome changes — under 1 second means streaming, hours means batch, minutes is ambiguous. Then ask whether the consumer is a human dashboard or an automated system. Humans consume data on their own cadence, not yours, so dashboards almost never justify streaming infrastructure.