
Lakehouse Architecture Compared: Iceberg vs. Delta Lake vs. Hudi
Most lakehouse comparisons compare the wrong thing. They rank formats on features; the actual decision is which compute engine you're marrying.
Delta Lake pulls analytical queries toward Databricks; its automation features run best there. Iceberg buys multi-engine flexibility at the cost of a catalog selection Delta doesn't require. Hudi commits you to Spark-centric ops in exchange for the category's best record-level upsert performance. The format sits underneath your compute stack; it doesn't replace it.
This is the Engine Gravity: every lakehouse table format pulls the rest of your stack toward its preferred compute engine. Recognizing which gravity you're choosing matters more than any feature-by-feature comparison — and most public comparisons read as feature matrices that obscure exactly this.
The Lakehouse Premise
A lakehouse, per Databricks' 2021 CIDR paper, is open storage on object stores (S3, GCS, Azure Blob) with ACID transactions, schema enforcement, and SQL query support. The goal is to combine the cost profile of a data lake with the consistency guarantees of a warehouse — one layer of data, many engines that query it.
Three table formats have captured this category:
- Apache Iceberg. Originated at Netflix, now Apache-governed with the most diverse contributor base across the three. The current stable release is 1.10.1 with the Iceberg spec reaching Version 3 for extended types and capabilities. Design priority: vendor-neutral multi-engine support.
- Delta Lake. Built at Databricks, open-sourced, managed under the Linux Foundation. Dremio's comparison notes Delta is "Databricks-steered" governance-wise, which matters when you evaluate long-term direction.
- Apache Hudi. Originated at Uber to handle high-volume upsert workloads (ride events, trip mutations). Apache-governed with commercial backing from Onehouse. Record-level indexing with 8+ index types is Hudi's differentiator — nothing else in the category ships this.
All three support ACID transactions on object storage, schema evolution, time travel, and partition pruning. The differences are not whether they have those features; it's how they implement them and what that implementation implies about your compute choices.
The Engine Gravity in Detail
Each format pulls in a different direction. Here's where.
Delta Lake pulls toward Databricks
Delta has the best Databricks integration by design — it's the native table format. Photon (Databricks' vectorized engine) treats Delta as a first-class citizen. Unity Catalog governs Delta tables natively. Liquid Clustering, Z-ordering, and deletion vectors have open-source implementations but auto-optimization features remain proprietary to Databricks' commercial product.
If your stack is Databricks-centered, Delta is the lowest-friction choice. You get faster writes, automatic optimization, and a vendor-supported path for every feature. The trade-off: building a non-Databricks compute layer (Flink, Trino, Spark outside Databricks) on Delta works, but you lose the performance features that make Databricks + Delta feel fast.
Delta UniForm — which exposes Delta tables as Iceberg-compatible — partially addresses the multi-engine concern, but only one direction. Reading Iceberg tables from Databricks is another path.
Iceberg spreads across engines
Iceberg's multi-engine support is the broadest in the category. Spark, Flink, Trino, Dremio, Snowflake, AWS Athena, Hive, Impala, ClickHouse — all have read support; most have write. This is by design: Iceberg's spec is engine-neutral, and engines implement their own readers.
The gravity Iceberg imposes is catalog choice, not engine choice. An Iceberg table needs a catalog — Nessie, AWS Glue, Databricks Unity Catalog (now supports Iceberg), Polaris (Snowflake's open catalog), or Tabular (the commercial entity formerly driving Iceberg, acquired by Databricks in 2024). Your catalog decision determines how tables are governed, which features (branching, tagging) you get, and what vendor relationship you enter.
For teams that want to keep compute options open across Snowflake, Databricks, and open-source engines simultaneously, Iceberg is the only realistic choice. For teams fully committed to one compute vendor, Iceberg's catalog overhead is added complexity.
Hudi pulls toward Spark and upsert workloads
Hudi's record-level indexing with 8+ index types and native managed compaction make it the best choice for high-upsert workloads. CDC into a Hudi table with merge-on-read storage and async compaction is the pattern Uber originally built Hudi for.
The compute reality: Hudi is Spark-first. Flink and Trino have Hudi support, but the full feature set (async clustering, multi-modal indexing, incremental query) assumes Spark. Organizations running a Spark-based processing layer benefit; organizations trying to query Hudi from Trino as a primary workload often hit feature gaps.
Hudi's community is smaller than Iceberg's or Delta's — Dremio counts roughly 200 contributors vs 800+ for Iceberg and 300+ for Delta. This is not a quality signal directly, but it does affect feature velocity and the depth of third-party integrations.
What Each Format Does Better
Iceberg: partition evolution and multi-engine
Iceberg is the only format with true partition evolution. You can change your partitioning strategy (year → year+month, for example) without rewriting any data. The partition spec is metadata; queries adapt automatically.
The multi-engine story is the other main advantage. Organizations with Spark for ETL, Trino for BI, and Snowflake for some analytical workloads can write from Spark and read from all three without data duplication.
Where Iceberg is weaker: manual compaction. Small-file accumulation from high-frequency writes requires explicit compaction jobs. The Iceberg community has discussed compaction automation for years; the current state requires operational attention.
Delta: Databricks integration and auto-optimization
Delta's strongest features — Liquid Clustering (evolving partition strategy without rewrites, similar to Iceberg's partition evolution), auto-compaction, and deletion vectors — shine specifically on Databricks. The open-source Delta Lake project ships the table format and basic operations; the operational automation lives in Databricks commercial product.
Change Data Feed is Delta's CDC output — downstream consumers can read row-level change events directly from the Delta transaction log. For CDC-shaped workloads where the source is a Delta table, this is cleaner than an external CDC pipeline.
Where Delta is weaker: multi-engine without UniForm, the auto-optimization being tied to commercial Databricks, and the Linux Foundation governance being "Databricks-steered" in practice.
Hudi: record-level upserts and streaming ingestion
Hudi's record-level index types (Bloom, global, hash, expression) and DeltaStreamer (its CDC ingestion tool) are purpose-built for high-volume mutation workloads. A Hudi table can absorb millions of upserts per hour with sub-minute freshness while compaction runs asynchronously.
Merge-on-read tables let you tune the read vs write cost trade-off per workload. Copy-on-write tables prioritize read speed. The choice is per-table.
Where Hudi is weaker: pure analytical read performance lags Iceberg and Delta for most workload shapes. The indexing overhead that makes upserts fast adds cost to range scans and aggregations.
The Feature Table That Actually Matters
Most feature comparisons blur together because all three formats have most features. The features that actually differentiate them:
| Capability | Iceberg | Delta | Hudi | |---|---|---|---| | Partition evolution (metadata-only) | Yes | Via Liquid Clustering on Databricks | No | | Record-level indexing | No | Proprietary Bloom filter | 8+ types native | | Auto-compaction | Manual | Proprietary on Databricks | Native managed | | Multi-engine read/write | Broadest (Spark, Flink, Trino, Snowflake, Athena) | Spark-first, others read-only | Spark-first, Flink/Trino gaps | | Time travel | Snapshot ID or timestamp | Version-based | Timeline-based | | Governance | Apache, diverse contributors | Linux Foundation, Databricks-steered | Apache, Onehouse-backed | | CDC output | Via Flink CDC or external | Change Data Feed native | DeltaStreamer native |
Equal features exist (all do ACID, schema evolution, time travel). What matters is the implementation depth and the surrounding engine ecosystem.
Migration Between Formats
Moving from one format to another is rarely painless. All three support bootstrap/migration tooling:
- Iceberg Table Migration — register existing Parquet/Hive tables as Iceberg without rewriting data
- Delta Convert to Delta — same pattern, for Delta
- Hudi Bootstrap — register existing tables, skip the data-copy step
These help when upgrading from plain Parquet-on-S3 to a table format. They help less when migrating between formats, because the metadata layouts differ enough that a straight conversion is not available. Cross-format migration typically means dual-write during transition, read-from-new-write-to-both for a period, and eventual cutover.
Delta UniForm and Iceberg-Delta interoperability features are narrowing this gap but don't eliminate it.
When Each Format Is the Right Answer
Choose Iceberg when:
- Your compute stack spans multiple engines (Spark + Trino + Snowflake, for example)
- You need true partition evolution as the data model matures
- Vendor neutrality is a strategic priority
- Your catalog direction is Nessie, Glue, Polaris, or Unity (any will work)
Choose Delta when:
- Databricks is your primary analytical engine
- You want the smoothest path to Databricks' auto-optimization features
- Change Data Feed simplifies your CDC story
- You can tolerate Databricks-steered governance
Choose Hudi when:
- Your primary workload is high-frequency upserts (CDC ingestion, mutation-heavy streams)
- Spark is already your core processing engine
- Record-level indexing materially changes your query patterns
- You accept slower raw analytical reads in exchange for fast mutation handling
Choose none of the three when:
- Your entire workload fits in DuckDB or a small warehouse. Lakehouse overhead is only worth it above a scale threshold where the compute/storage separation pays for itself. Sub-TB warehouses often do not clear that bar.
The Benchmark Honesty Gap
Almost every published lakehouse benchmark is vendor-produced. Databricks benchmarks show Delta winning; Dremio benchmarks show Iceberg with Trino winning; Onehouse benchmarks show Hudi winning on upsert workloads. The TPC-DS benchmark is the closest to a neutral standard, but TPC-DS doesn't capture upsert-heavy or high-mutation workloads where the three formats actually diverge.
For production workloads, the benchmark that matters is your own — on your actual data volumes, query patterns, and update rates. Vendor-published numbers are useful for order-of-magnitude estimates and useless for tie-breaking. Teams that pick a format on a benchmark blog post re-evaluate once the real workload shape emerges and the initial benchmark no longer looks like the workload they run.
We don't have published neutral benchmarks that compare all three on realistic mixed workloads. The closest independent work comes from Onehouse's feature comparison and Dremio's analysis, both of which have commercial angles but document their methodology. Cross-reading both gives a more honest picture than trusting either alone.
What We're Watching in 2026
Three things to watch that could change the calculus:
- Catalog consolidation. Unity Catalog now supports Iceberg. Polaris (Snowflake's open catalog) has shipped. Nessie continues. Catalog choice, not format choice, may be where vendor lock-in actually sits in 2026. Cross-catalog compatibility is still fragmented.
- UniForm and interop features. Delta UniForm lets Iceberg readers query Delta tables. The reverse (Delta reading Iceberg) is improving. If interop reaches full read-write parity, the format decision becomes less load-bearing.
- Databricks' Tabular acquisition. Databricks acquired Tabular (the commercial entity behind Iceberg) in 2024. The long-term impact on Iceberg's vendor neutrality is open. If Databricks prioritizes Delta over Iceberg, the governance picture shifts.
Your table format is a compute decision disguised as a storage decision.
Need help evaluating lakehouse formats or designing a migration between Iceberg, Delta, and Hudi? Talk to an engineer — we'll tell you honestly if we can help.
Frequently Asked Questions
What is the difference between Iceberg, Delta Lake, and Hudi?
All three are open table formats providing ACID transactions on object storage. Iceberg has the broadest multi-engine support and the only true partition evolution. Delta Lake has the deepest Databricks integration and native Change Data Feed for CDC output. Hudi has record-level indexing with 8+ index types and native managed compaction, making it the strongest for high-volume upsert workloads.
Which lakehouse format is best for Databricks?
Delta Lake, by a significant margin. Delta is the native format on Databricks with Photon-optimized reads, Liquid Clustering, auto-compaction, and Unity Catalog integration. Iceberg works on Databricks via Unity Catalog's Iceberg support, but the performance automation features remain Delta-first. If Databricks is your primary compute, Delta reduces friction.
When should you use Apache Hudi over Iceberg or Delta?
Use Hudi when your workload is upsert-heavy — CDC ingestion, mutation streams, or high-frequency transactional changes. Hudi's record-level indexing and managed compaction handle millions of upserts per hour with sub-minute freshness. For primarily read-heavy analytical workloads, Iceberg or Delta typically outperform Hudi because their metadata models are optimized for scan-and-aggregate patterns.
Do you need a catalog for an Iceberg table?
Yes. Iceberg requires a catalog to track table metadata across engines. Options include Nessie (open-source, git-like branching), AWS Glue Data Catalog, Databricks Unity Catalog (now supports Iceberg), Snowflake's Polaris catalog, or Hive Metastore for legacy stacks. The catalog choice affects governance features (branching, tagging, access control) and which engines can participate.
What is Delta UniForm?
Delta UniForm is a Delta Lake feature that exposes Delta tables as Iceberg-compatible for read access from Iceberg-aware engines. It enables interoperability without data duplication — a single Delta table can be read by both Databricks-native engines and Iceberg-compatible engines like Snowflake or Trino. UniForm is one-directional; Delta engines reading Iceberg tables is a separate path under development.
Related posts

Airflow vs. Dagster vs. Prefect: A 2026 Orchestration Comparison
Airflow 3 shipped breaking changes, Dagster bet on assets over DAGs, Prefect cut 90% of runtime overhead. A decision framework based on team size, not feature lists.

CDC at Scale: Debezium, Flink CDC, and the Real-Time Replication Problem
CDC looks simple — stream row-level changes from database to warehouse. In production it's schema drift, replication slot management, and full resyncs every time DDL runs. The honest guide.

Data Pipeline Observability: Monitoring Airflow + dbt Without Drowning in Alerts
200 tests, 80 weekly alerts, and a Slack channel nobody reads. That's alert debt. The monitoring stack that catches real failures without burying your team.