Data quality on Databricks, layer by layer

kept in this browsersaved to your account

Constraints, expectations, DQX, data profiling and anomaly detection each catch a different failure. What each layer sees, what it costs, and how to choose.

What it is

“Data quality” on Databricks is not a single product you turn on. It is four different mechanisms, each watching a different moment in the life of a row:

LayerWhere it runsCatchesReaction
Delta constraintson the table, for every writernulls and predicate violationsthe write fails
Pipeline expectationsinside a declarative pipelinerows that break a rulewarn, drop, or fail the update
DQXany PySpark job or streamthe same rules, outside a pipelineannotate or quarantine
Data quality monitoringafter the write, on a scheduledrift, staleness, missing volumemetrics, dashboards, alerts

The first three act on data in transit. The last one watches data at rest and tells you that something changed even when every rule still passes.

Why it exists

Each layer is blind to what the others see. A CHECK constraint cannot tell you that the row count dropped by 90%, because every surviving row is valid. An expectation cannot protect a table that someone writes to from a notebook. Profiling cannot stop anything, it can only report afterwards. Choosing one and calling it “data quality” is how a pipeline ends up green while the dashboard is wrong.

How it works

Delta constraints: the floor

NOT NULL and CHECK live in the table definition, so they apply to every writer, in every language, forever. Violating one fails the transaction. They are cheap and absolute, which is exactly why they should hold only the rules that are true by definition: a primary key is not null, an amount is not negative. See Data quality: expectations and constraints for the syntax and for what happens to an existing table when you add one.

Expectations: the pipeline’s own rules

Inside a Lakeflow pipeline (the product formerly called Delta Live Tables), an expectation is a named boolean condition with an action: keep the row and count the violation, drop the row, or fail the update. Results land in the event log, so “how many rows failed valid_amount last week” is a query, not an archaeology project. This is the default choice for anything that already runs as a pipeline.

DQX: the same discipline for everything else

DQX: data quality checks for PySpark applies named rules to any PySpark DataFrame or table, batch or streaming, and splits valid rows from quarantined ones. Use it when the data never touches a declarative pipeline, or when the same rule set has to be shared by several jobs and owned as configuration rather than code.

Data quality monitoring: the trend

Unity Catalog groups two features under data quality monitoring, and they are not the same maturity.

Data profiling is the feature formerly called Lakehouse Monitoring. Attach a monitor to a table and it computes summary statistics on a schedule, writing two Delta tables: a profile metrics table with the statistics and a drift metrics table comparing each window with the previous one and with a baseline. Three monitor types cover the cases: time series for timestamped data, inference for model request logs, and snapshot for everything else. Because the output is a table, alerts and dashboards are ordinary queries over it. It is generally available, though not in every region.

Anomaly detection works at the schema level rather than per table and is in Public Preview as of September 2026. It learns two things from history: freshness, how recently a table is usually updated, and completeness, how many rows normally arrive in a day. It then flags tables that went quiet or arrived thin. This is the layer that catches an upstream job that silently stopped, which no row-level rule can see. Both are billed as serverless compute, so a monitor on every table is a cost decision, not a free win.

Outside the platform

Several mature open-source frameworks solve the same problem, and are worth knowing if the team already uses one or if the rules have to run somewhere other than Databricks.

FrameworkShapeWhy you would pick it over DQX
Great Expectations (GX Core)Expectation suites plus generated documentation, many backendsThe team already has suites, or you want the data docs as an artefact. Note that stewardship moved to Fivetran in May 2026
Soda CoreChecks in YAML, positioned around data contractsYou want contracts between teams, with a commercial cloud for the reporting side
spark-expectations (Nike)In-process Spark rules with quarantine and statistics tablesThe closest thing to DQX outside Labs: rules live in a table, alerting goes to Kafka or email
Deequ and PyDeequ (AWS)Scala-first “unit tests for data”, with a metrics repositoryYou want constraint suggestion and anomaly detection over a history of metrics, on any Spark, not only Databricks
PanderaSchema and statistical typing for pandas, Polars and PySparkThe rules are really a schema contract in code, checked in CI as well as in the job

None of them know about Unity Catalog, Lakeflow pipelines or workspace deployment, which is the one thing DQX gets for free.

Choosing

  • The rule is true by definition and must hold for every writer: a Delta constraint.
  • The rule belongs to one pipeline and you want it in the event log: an expectation.
  • The rule has to run outside a pipeline, or be shared across jobs, or produce a quarantine table: DQX: data quality checks for PySpark.
  • You want to know when the data changes shape rather than breaks: data profiling.
  • You want to be told when a table goes quiet without writing any rule: anomaly detection.

Most teams end up with a constraint layer of five rules, expectations or DQX at the bronze-to-silver boundary, and profiling on the gold tables the business actually reads.

Common mistakes

  • Only checking at the end. A quality rule on gold tells you the number is wrong. A rule at the silver boundary tells you which source row made it wrong.
  • Rules with no owner. Every rule needs a name, a severity and someone who is expected to look when it fires. A rule that fires weekly and is ignored is worse than no rule, because it trains everybody to ignore the alert channel.
  • Confusing profiling with enforcement. Profiling never blocks a write. If the requirement is “this must not be possible”, it is a constraint.
  • Building a bespoke framework. Between expectations, DQX and profiling, the interesting work left is the rules themselves, not the runner.

Where this sits

Resources

10All resources
Report a problem with this page
What kind of problem?

Reports about "Data quality on Databricks, layer by layer" go to the maintainer, not to a public thread.

Suggest a resource
What kind?

Nothing appears on the site automatically. A person reads every suggestion, checks the link and writes the note that goes with it.