Production monitoring for GenAI apps
Runs the same evaluation scorers against a sample of live traces and attaches the results to each trace.
Why it is worth watching
The only honest way to know whether a deployed agent is still good.
What Beta promises
Open to most customers but not for production: support runs through engineering, and the shape of it can still move. On by default on Premium, off on Enterprise.
| Who can use it | Most customers |
|---|---|
| Production use | No |
| Support | Engineering only |
| On by default | Premium yes, Enterprise no |
The definitions are Databricks' own, on its release types page. The label on this page is the one the feature's documentation states, last checked on 13 Sept 2026.
Where to read more
We wrote it up: Production monitoring for GenAI apps — Registered scorers run continuously against a sampled fraction of live traces and attach their verdicts to each trace as feedback, so quality drift shows up without a scheduled evaluation. That page carries the Beta label while this entry does.
The Databricks documentation for Production monitoring for GenAI apps