
Building a Reverse-ETL Pipeline: Hightouch, Census, and When It's Worth It
The warehouse knows which 500 users are about to churn. The sales team does not. A propensity score sits in Snowflake with a daily-refresh dbt model behind it, and the Salesforce contact records those users own have no field for it. No rep has seen the list. No automation references it. Friday goes by. Monday goes by. The users churn.
This is the single most expensive gap in most data stacks: analytical conclusions that never reach the operational systems where action happens. Reverse ETL is the infrastructure pattern built to close it. Instead of ingesting data into the warehouse (ETL), you push warehouse outputs back out to Salesforce, HubSpot, Braze, Marketo, Iterable, Google Ads, Zendesk — the CRM and marketing platforms where someone clicks "send" or "call." The warehouse becomes a source of record for customer segments, propensity scores, and enriched attributes. Not a dead end.
This is the Operational Gap: the distance between what the data team can compute and what the business can act on. A churn model predicting 500 at-risk accounts this week has no business value until an account executive sees those names in their pipeline view. The model is not the product. The sync is.
Reverse ETL is the category of tool built to close that gap at scale. The two names worth knowing are Hightouch and Census, and the category consolidated meaningfully in 2025 when Fivetran acquired Census.
What Reverse ETL Actually Does
A reverse ETL pipeline does four things:
- Reads from the warehouse. SQL or dbt model against Snowflake, BigQuery, Databricks, Redshift, or a lakehouse like Iceberg/Delta.
- Models records for destinations. Maps warehouse columns to the object model of the destination API (Salesforce Contact, HubSpot Company, Braze User Profile, etc.).
- Syncs to the destination. Pushes the data via the destination's API, handling auth, rate limiting, pagination, and retries.
- Tracks sync state. Records what was synced, when, what changed, and what failed — typically in a state table back in the warehouse.
The value proposition is that you don't build this yourself. Each destination has its own API quirks — Salesforce's 15-character vs 18-character record IDs, HubSpot's batch size limits, Braze's identifier hierarchy, Marketo's 50,000 record/day export cap. A reverse ETL tool abstracts those over a common sync interface.
Sync types matter. Two patterns dominate:
- Full sync. Recompute the entire destination state on every run. Simple, expensive, and produces lots of unnecessary writes against destination APIs that count toward rate limits.
- Incremental sync. Only push records that changed since the last sync, identified by a primary key plus a change detection column (usually
updated_at). This is the default for anything beyond small datasets.
The rare third pattern is mirror sync, where the reverse ETL tool observes the destination and keeps warehouse state in sync — this is an edge case most teams don't need.
Hightouch vs. Census: The 2026 State
As of 2025, the reverse ETL market consolidated meaningfully.
Fivetran acquired Census in May 2025 and rebranded it as Fivetran Activations. Census's product continues to operate as the activations layer inside Fivetran's platform, with migration guidance published for existing Census users. The pitch is end-to-end: Fivetran pulls data in (ELT), dbt or native Fivetran transforms model it, and Activations pushes it out (reverse ETL), all under one pricing and support surface.
Hightouch remains independent and partnered with Fivetran for data delivery. Hightouch's pitch runs opposite to Fivetran's consolidation: a focused activation layer that plugs into whatever warehouse, ELT stack, or CDP strategy the customer already has. Hightouch ships 300+ destination integrations across marketing, sales, customer success, advertising, and support. The company has also pushed into AI Decisioning — agents that choose message, channel, and timing based on warehouse-resident signals.
What this means for buyers:
- Already running Fivetran for ingestion: Fivetran Activations is the low-friction path. Same contract, same logins, same audit log. Integration overhead is zero, even if the standalone product comparison doesn't always favor it.
- Need the widest destination coverage: Hightouch's 300+ integrations lead the category, and the product roadmap is activation-focused rather than shared across an ELT product line.
- No existing vendor lock-in: Three factors drive the decision — which destinations matter, what pricing looks like at projected sync volume, and whether vendor independence is worth paying for.
We do not have firm data yet on how Census's roadmap will evolve under Fivetran versus Hightouch's independent direction. In mid-2026 this is worth re-evaluating, particularly if Fivetran shifts Activations pricing to the new per-MAR model that hit Fivetran ingestion customers with 40-70% increases in March 2025.
Other Options Worth Naming
The category isn't just Hightouch and Census-now-Fivetran:
- Polytomic — open-source-friendly, bi-directional sync, smaller but growing
- Grouparoo — Airbyte acquired the project; continuity is uncertain
- RudderStack Reverse ETL — for teams already running RudderStack as a CDP
- Build your own with dbt + custom sync jobs — still the right choice for narrow use cases with one or two destinations
The build-vs-buy decision lives later in this post. For most teams with three or more destinations and volumes above 100K records/day, a tool pays for itself.
When Reverse ETL Is Worth Building
Reverse ETL is worth it when two conditions hold:
- The data has a consumer. A named person or automated workflow that changes what they do based on the synced data. The CS team calls at-risk accounts. The ads team excludes high-value customers from acquisition campaigns. The product team enrolls users in an onboarding experiment.
- The alternative is manual export. Without reverse ETL, someone exports a CSV from the warehouse, uploads it to the destination, and repeats that every week. The manual process is both error-prone and a bottleneck on how often the data refreshes.
Concrete use cases where the math routinely works:
- Customer segmentation to marketing automation. Braze, Iterable, Marketo, Customer.io, Klaviyo all benefit from warehouse-resident segments with propensity scores and lifetime value.
- Lead enrichment to CRM. Pushing account-level firmographics, product usage scores, and intent signals into Salesforce or HubSpot.
- Ad audience sync. Pushing custom audiences to Google Ads, Meta, LinkedIn, and TikTok. The platforms' native audience tools have narrower targeting than warehouse-derived segments.
- Personalization signals to product surfaces. Feature flags (LaunchDarkly, Statsig), A/B test assignments, and user attributes pushed into real-time decision systems.
- Support context to help desks. Enriching Zendesk and Intercom tickets with user lifetime value, recent purchase history, and at-risk flags.
These workflows fail without reverse ETL because the warehouse sits outside the operational tool's native integration graph. A Salesforce admin cannot natively query Snowflake for account enrichment; a marketing automation tool cannot natively read dbt marts.
When Reverse ETL Is a Distraction
The category markets itself as essential infrastructure. It isn't — not for every stack.
- Single-destination syncs to one tool. A daily export to HubSpot is a scheduled Airflow task wrapping the HubSpot Python SDK in roughly 20 lines. Don't buy Hightouch for this.
- Sub-minute activation latency. Reverse ETL tools run on batch schedules (typically 15 minutes to 24 hours). Sub-second push is streaming territory — Kafka consumers hitting destination APIs directly, not reverse ETL.
- Low data volumes with few destinations. A startup syncing 5,000 records to two destinations weekly is paying $500+/month for tooling that a Python script and a cron job cover at zero.
- Source-of-truth conflicts. When the destination owns a field (Salesforce-owned account status, for example), pushing warehouse-derived values into it creates write conflicts. Reverse ETL assumes the warehouse is authoritative or the destination accepts overrides cleanly — verify which applies per field.
- Compliance-restricted data. PII and regulated health/finance data often cannot leave specific geographic boundaries. Reverse ETL to a US-based marketing tool from an EU warehouse requires explicit DPA work that teams routinely underestimate.
Architecture Patterns That Hold Up
The reverse ETL architectures that work in production share a small number of properties:
Identity resolution in the warehouse, not the tool
Your warehouse should hold the canonical identity graph — user_id, email, customer_id, org_id — and reverse ETL pushes records keyed by those identifiers. Trying to do identity resolution in the reverse ETL tool itself (matching records across destinations) is a rabbit hole. Keep the reverse ETL tool as a pipe, not a logic layer.
Rate-limit-aware scheduling
Destination APIs have hard daily, hourly, or per-minute limits. Salesforce Enterprise: 100K API calls/day by default. HubSpot: tiered by plan. Google Ads: batch upload API with complex quota logic. Schedule syncs so that concurrent pushes don't exhaust quota. Most reverse ETL tools handle this, but only if you give them accurate quota metadata.
Change detection done right
updated_at timestamps on warehouse models are the cheapest reliable change detection. dbt makes this trivial: add a current_timestamp() column to every staging and marts model. Without it, incremental sync falls back to hashing every row, which is slow and error-prone.
State tables back in the warehouse
Every reverse ETL tool writes sync state (what was pushed, when, with what result) into a sync log. Configure it to land back in the warehouse so analysts can see sync lag, failure rates, and per-destination health in the same place they build models. Tools that hide this state inside their own UI become operational blind spots.
Idempotency at the destination
Some destination APIs are idempotent on the same record id (upsert semantics); some are not. Reverse ETL tools handle this abstraction, but the failure mode when they get it wrong is creating duplicate records in Salesforce or Marketo. Test first syncs against a sandbox, not production.
Common Failure Modes
Three patterns show up repeatedly when reverse ETL goes wrong:
- No consumer for the synced data. A reverse ETL pipeline runs beautifully for six months, syncing propensity scores into Salesforce fields that no sales rep has mapped to their pipeline view. Nobody takes action. The sync is infrastructure without purpose. Check for this before building — ask who opens the destination tool and what they do with the data.
- Mapped fields drift out of sync. The warehouse field is
lifetime_value_usd; the Salesforce field isLTV_USD__c. Someone renames the warehouse column. The reverse ETL sync silently writes NULL into the destination for two weeks before anyone notices. Field mapping documentation and alerting on unexpected NULL rates prevents this. - Destination rate limit exhaustion. An analyst adds a new sync without checking the destination quota. The daily run pushes a million records to Marketo, hits the 50K/day export limit, and blocks every other sync to that destination for the rest of the day. Quota budgeting is operational hygiene that gets skipped until it hurts.
Build vs. Buy: The Break-Even
For 1-2 destinations with low volume, build. A Python script using the destination's SDK, scheduled via Airflow, with logging to your warehouse. Under 200 lines of code.
For 3-5 destinations with meaningful volume, buy. The tool's abstraction over API quirks, rate limits, and schema mapping pays for itself quickly. Hightouch self-serve tiers start at a few hundred dollars a month; Fivetran Activations pricing is bundled with Fivetran's ingestion billing.
Above 5 destinations or in any organization with multiple data-producing teams, buy without hesitation.
The build-yourself failure mode is consistent: a scrappy sync job to one destination grows into a library of 10 scripts, each with different logging, different retry semantics, and different owners. By year two, nobody remembers who wrote the original Braze sync. That's when "why didn't it run last night" becomes everyone's problem.
What We Don't Know Yet
The major open question for 2026 is how the Fivetran-Census integration evolves. Census customers who liked the standalone product now depend on Fivetran's roadmap prioritization, and Fivetran's ingestion pricing history suggests that bundling can come with cost surprises. Hightouch's independent position is a hedge, but only if the market keeps valuing composability over consolidation — and Hightouch's pricing is not cheap.
The tool choice matters less than the discipline around what gets synced and why. Teams that treat reverse ETL as plumbing with a named business consumer get value from either vendor. Teams that build generic "activate all the data" pipelines do not.
Reverse ETL is worth building only when someone in the business takes action on what syncs.
Need help evaluating reverse ETL tools or designing a sync architecture? Talk to an engineer — we'll tell you honestly if we can help.
Frequently Asked Questions
What is reverse ETL and how is it different from ETL?
ETL moves data from operational systems into the warehouse. Reverse ETL moves it back out — from the warehouse to CRMs, marketing tools, ad platforms, and support systems where business teams take action. ETL feeds dashboards. Reverse ETL feeds the tools where someone clicks "send."
Should I use Hightouch or Fivetran Activations (formerly Census)?
Fivetran Activations is the low-friction choice if you already run Fivetran for ingestion — same contract, same audit log. Hightouch has roughly 300 destination integrations and a more active activation-focused roadmap. If vendor independence matters or you need the widest destination coverage, Hightouch wins. If you want a single-vendor data movement stack, Fivetran's bundle is simpler.
When is reverse ETL not worth the cost?
A single-destination sync with low volumes is cheaper as a Python script using the destination's SDK on an Airflow schedule. Sub-second activation latency belongs to streaming, not reverse ETL. And any pipeline without a named business consumer for the synced data is pure overhead — infrastructure without a user is infrastructure without value.
How does reverse ETL handle API rate limits?
Reverse ETL tools schedule syncs against destination API quotas — Salesforce's 100K/day limit, HubSpot's plan-tiered limits, Google Ads batch upload quotas, Marketo's 50K/day export cap. Most tools track quota usage and throttle or retry automatically, but only if you provide correct quota metadata. Incremental sync (only changed records) reduces API volume compared to full refresh.
What is a composable CDP and how does reverse ETL fit?
A composable CDP uses your warehouse as the customer database instead of buying a packaged CDP product (Segment, mParticle, Amplitude). Reverse ETL is the activation layer that pushes warehouse-resident segments and traits to operational tools, replacing the packaged CDP's built-in connectors. Fivetran, Snowplow, and Hightouch have all marketed warehouse-first composable CDP stacks since 2022.
Related posts

Airflow vs. Dagster vs. Prefect: A 2026 Orchestration Comparison
Airflow 3 shipped breaking changes, Dagster bet on assets over DAGs, Prefect cut 90% of runtime overhead. A decision framework based on team size, not feature lists.

CDC at Scale: Debezium, Flink CDC, and the Real-Time Replication Problem
CDC looks simple — stream row-level changes from database to warehouse. In production it's schema drift, replication slot management, and full resyncs every time DDL runs. The honest guide.

Data Pipeline Observability: Monitoring Airflow + dbt Without Drowning in Alerts
200 tests, 80 weekly alerts, and a Slack channel nobody reads. That's alert debt. The monitoring stack that catches real failures without burying your team.