ivinco
Data Quality at Scale: Great Expectations vs. Soda vs. dbt Tests

Data Quality at Scale: Great Expectations vs. Soda vs. dbt Tests

Ivinco Team·

The most common data quality failure we see in client pipelines is not a missing test. It's a test in the wrong place. dbt tests pass every morning. The revenue dashboard is still wrong. The source API renamed a column three weeks ago, and dbt never saw it — because dbt doesn't look at sources.

Gartner estimates that poor data quality costs organizations $12.9 million per year on average. Most of that cost isn't from missing tests. It's from tests that run after the damage is done.

dbt tests run after transformation. Great Expectations validates before loading. Soda monitors across the full pipeline. Each tool catches a different category of failure. Choosing one doesn't mean you don't need the others.

The question isn't which tool is best. It's which combination covers your actual failure modes.

The Trust Gradient

Data quality tools sit on a spectrum. We call it the Trust Gradient.

At one end: deterministic tests. You write order_id must be unique and not null. If it fails, you know exactly what's wrong. dbt tests and Great Expectations live here — catching known failure modes with certainty.

At the other end: statistical monitoring. The system learns baseline distributions and alerts when something deviates. Monte Carlo and Elementary live here — catching unknown failure modes, but with less precision and more noise.

Most teams cluster at one end and ignore the other. dbt-heavy teams have comprehensive transformation tests but zero monitoring at the source boundary. Teams that buy Monte Carlo get observability dashboards but skip the deterministic tests that would catch schema corruption before it enters the pipeline.

Both ends matter. Our standard practice for client pipelines is to cover the full Trust Gradient — deterministic tests at the boundary, statistical monitoring in the warehouse. Here's how each tool fits.

dbt Tests: The Transformation Layer

dbt's built-in testing framework is the most widely adopted data quality tool in the modern data stack. According to Astronomer's State of Airflow 2026, dbt pairs with Airflow at a 44% adoption rate. The testing framework is part of why.

What dbt tests catch

dbt ships four generic tests out of the box:

models:
  - name: orders
    columns:
      - name: order_id
        tests:
          - unique
          - not_null
      - name: status
        tests:
          - accepted_values:
              values: ['pending', 'shipped', 'delivered', 'cancelled']
      - name: customer_id
        tests:
          - relationships:
              to: ref('customers')
              field: customer_id

These four — unique, not_null, accepted_values, and relationships — cover referential integrity and basic constraint enforcement. The dbt_utils and dbt_expectations packages extend this significantly: range checks, regex matching, recency validation, row count comparisons, and cross-database assertions.

Where dbt tests fail

dbt tests run inside your warehouse, on transformed data, after your models execute. This means:

They can't catch source-side corruption. If your extraction layer loads truncated data or maps a renamed column to NULL, the bad data is already in your warehouse by the time dbt tests run. A not_null test on customer_id will catch the NULL values — but only after they've propagated through every model that depends on the staging table.

They don't learn. dbt tests are deterministic: they pass or fail based on rules you wrote. If revenue_per_order gradually drifts from a mean of $85 to $42 over three weeks because of a timezone bug in the source, no dbt test will catch it unless you wrote a specific range check. And you probably didn't, because you didn't know the normal range before the drift started.

They run at build time, not continuously. dbt tests execute when dbt test runs — typically once or twice a day as part of a scheduled pipeline. If data quality degrades between runs, the next test catches it, but hours of bad data may have already reached dashboards.

Cost

dbt Core is open source. dbt Cloud starts at $100/seat/month for the Starter plan (up to 5 seats, 15,000 models/month). Enterprise contracts for 10-25 seats typically run $30,000-$90,000/year. The testing framework is fully available in dbt Core — you don't need dbt Cloud for tests.

Great Expectations: The Source Boundary

Great Expectations (GX) validates data at the point of extraction — before it enters your warehouse. This is the architectural distinction that matters most.

What GX catches

GX defines validation rules as "Expectations" — Python assertions against data:

validator.expect_column_to_exist("customer_id")
validator.expect_column_values_to_be_of_type("customer_id", "INTEGER")
validator.expect_column_values_to_not_be_null("customer_id")
validator.expect_column_values_to_be_between(
    "order_total", min_value=0, max_value=100000
)
validator.expect_table_row_count_to_be_between(
    min_value=1000, max_value=50000
)

Expectations are grouped into suites, and suites run through Checkpoints that integrate into your pipeline. If a Checkpoint fails, the pipeline halts before bad data reaches the warehouse. GX also generates Data Docs — human-readable HTML reports showing validation results — which makes it easier to communicate data quality status to non-technical stakeholders.

Where GX wins over dbt tests

Schema validation at the boundary. When a source API renames customer_id to cust_id, a GX Expectation like expect_column_to_exist("customer_id") catches it before the data is loaded. The same problem in dbt requires the bad data to be loaded and transformed first.

Row count and freshness gates. GX can check that a source delivered the expected number of records before any transformation happens. If your API source returns 800 records when it normally returns 10,000, GX blocks the load. dbt can check row counts in the warehouse, but only after the partial load has already landed.

Non-warehouse data sources. GX validates data in pandas DataFrames, Spark DataFrames, and SQL databases directly. For teams processing files, APIs, or streaming data before it reaches a warehouse, GX provides validation that dbt structurally cannot.

Where GX falls short

Python-only, code-heavy. Every Expectation suite is Python code. For analytics engineers who think in SQL, the learning curve is steeper than dbt's YAML-based test definitions. This matters for team adoption — the tool the team actually uses beats the technically superior tool they don't.

No anomaly detection. GX is purely deterministic. It catches what you test for and nothing else. If you don't know to check for timezone drift, GX won't find it.

Operational overhead. GX requires its own execution environment — a Python process running Checkpoints, managing Data Contexts, connecting to data sources. This is additional infrastructure to maintain alongside your orchestrator and warehouse.

Cost

GX Core is open source (Apache 2.0). GX Cloud is a managed SaaS option that simplifies deployment and collaboration. Pricing for GX Cloud is not publicly listed — contact sales.

Soda: The Full-Stack Monitor

Soda sits between dbt's warehouse-native testing and GX's source-boundary validation. SodaCL — the Soda Checks Language — is a YAML-based DSL with 50+ built-in metrics that can run against any data source.

What Soda catches

checks for stg_orders:
  - row_count > 0
  - freshness(loaded_at) < 2h
  - duplicate_count(order_id) = 0
  - missing_percent(customer_id) < 1%
  - avg(order_total) between 50 and 150
  - schema:
      fail:
        when required column missing: [order_id, customer_id, total]
        when wrong type:
          order_id: integer

SodaCL's syntax is readable enough that non-engineers can write and review checks. The freshness, schema, and anomaly_score metrics address monitoring gaps that neither dbt tests nor GX cover natively.

Soda's architectural advantage

Soda integrates with Airflow, Dagster, Prefect, and Azure Data Factory as a pipeline step. But it also runs independently on a schedule — checking data quality between pipeline runs, not just during them. This means Soda can catch degradation that occurs between dbt builds.

Soda Cloud adds anomaly detection on top of SodaCL checks: automatic baseline learning, threshold suggestions, and alerting when distributions shift. This puts Soda at both ends of the Trust Gradient — deterministic checks via SodaCL and statistical monitoring via Soda Cloud.

Where Soda falls short

Less depth than GX at the source boundary. GX's Expectation library is larger and more customizable than SodaCL's built-in metrics. For complex validation logic — custom regex patterns, multi-column constraints, conditional expectations — GX provides more expressive power.

Less depth than dbt in the warehouse. dbt tests are tightly integrated with the dbt project: they understand model dependencies, run within the same build process, and benefit from dbt's materializations and macros. Soda runs against tables independently — it doesn't understand the DAG.

Pricing scales with data. Soda Core is open source. Soda Cloud's free tier covers 3 datasets. The Team plan is $8/dataset/month. At 100 datasets, that's $800/month — not expensive, but it's a cost that dbt tests (free in dbt Core) and GX Core (free) don't carry.

The Enterprise Tier: Monte Carlo and Elementary

Two additional tools deserve mention for teams at scale.

Monte Carlo provides ML-based anomaly detection across five observability pillars: freshness, quality, volume, schema, and lineage. It monitors without requiring you to write checks — the system learns baseline patterns and alerts on deviation. Enterprise pricing starts above $100K/year. Monte Carlo's 150+ enterprise customers include CNN, JetBlue, and Toast. For most mid-market teams we work with, the ROI calculation only works when undetected data quality incidents have demonstrable business cost exceeding the tool's price. Below that threshold, open-source tools cover the Trust Gradient adequately.

Elementary is an open-source dbt-native observability tool. It reads dbt artifacts — test results, model run times, row counts — and generates reports with anomaly detection. For teams already using dbt that want monitoring without adding another vendor, Elementary fills the gap between dbt tests and a full observability platform.

The Decision Framework

| | dbt Tests | Great Expectations | Soda | Monte Carlo | |---|---|---|---|---| | Where it runs | Warehouse (post-transform) | Source boundary (pre-load) | Any data source | Any data source | | Test definition | YAML + SQL | Python | SodaCL (YAML) | Auto-learned | | Catches known issues | Yes | Yes | Yes | Limited | | Catches unknown issues | No | No | Soda Cloud only | Yes | | Freshness monitoring | Via dbt source freshness | Custom expectations | Built-in freshness() | Automatic | | Schema validation | on_schema_change='fail' | expect_column_to_exist | Built-in schema check | Automatic | | Cost | Free (Core) / $100+/seat/mo | Free (Core) / Contact sales | Free (Core) / $8/dataset/mo | $100K+/year | | Best for | Teams already on dbt | Pre-warehouse validation | Cross-stack monitoring | Enterprise observability |

For teams with 1-5 data engineers

Start with dbt tests. They're free, they're integrated into your build process, and they catch the most common transformation issues. Add dbt_expectations for range checks and freshness. You don't need another tool yet.

For teams with 5-15 data engineers

Add Soda Core for source-boundary validation and freshness monitoring. SodaCL's YAML syntax is approachable for the whole team, and running Soda checks before dbt models catches source corruption before it propagates. Total cost: $0 (Soda Core + dbt Core).

For teams with 15+ data engineers or regulated industries

Add Great Expectations at the ingestion layer for complex validation logic, and Soda Cloud or Elementary for continuous monitoring between pipeline runs. Evaluate Monte Carlo if undetected data quality issues have demonstrable business cost exceeding $100K/year.

Where This Framework Stops

This covers structured, tabular data in a warehouse-centric architecture. We haven't tested these patterns enough against streaming pipelines where validation happens at event speed, ML feature stores where "quality" means statistical stability across training and serving distributions, or unstructured data where the concept of a schema barely applies. Those are different problems requiring different tools — and we don't have strong opinions on them yet.

Data contracts — formal agreements between producers and consumers about schema, SLAs, and quality — are a related but separate category. Tools like Gable and Datafold address contracts specifically. Worth evaluating, but orthogonal to the Trust Gradient.

No tool substitutes for a team that treats data quality as an engineering discipline, not a checkbox.


Need help building data quality into your pipelines? Talk to an engineer — we'll tell you honestly if we can help.

Frequently Asked Questions

What is the difference between Great Expectations and dbt tests?

Great Expectations validates raw source data at the point of extraction — before data enters your warehouse. dbt tests validate transformed data after models run inside the warehouse. They cover different pipeline stages: Great Expectations catches schema drift and source corruption, while dbt tests catch transformation logic errors and referential integrity failures. Most production pipelines need both.

Is Soda free?

Soda Core is open-source and free. Soda Cloud offers a free tier covering 3 datasets with basic checks. The Team plan costs $8/dataset/month, scaling with the number of datasets monitored. At 100 datasets, that's $800/month. For comparison, dbt Core tests and Great Expectations Core are completely free with no dataset limits.

How do you choose a data quality tool?

Match the tool to the pipeline stage where you have the most failures. If source data is unreliable, start with Great Expectations or Soda at the ingestion boundary. If transformation logic is your main source of bugs, dbt tests cover that layer. If you need anomaly detection for problems you can't anticipate, Soda Cloud or Monte Carlo provide statistical monitoring. Most mature teams use at least two tools covering different stages.

What is a data quality testing framework?

A data quality testing framework is a system for defining, running, and reporting on checks that validate data correctness. Frameworks range from deterministic tests (dbt's unique and not_null checks that pass or fail based on explicit rules) to statistical monitors (Monte Carlo's ML-based anomaly detection that learns baseline patterns). The Trust Gradient describes this spectrum — from explicit checks you write to anomaly detection the system learns.

How much does enterprise data quality monitoring cost?

Open-source options (dbt Core, Great Expectations Core, Soda Core, Elementary) are free. Mid-market tools like Soda Cloud run $8/dataset/month. Enterprise platforms like Monte Carlo start above $100K/year for ML-based observability across all data assets. dbt Cloud adds team collaboration at $100/seat/month. Total cost depends on pipeline complexity, team size, and whether undetected data issues have measurable business impact.