Explore
The reference, by area. The groups are the domains the documentation is organised by; the areas carry the names the workspace gives them, so what you see on screen is what you find here. For the order to learn in, take a path; for how it all connects, open the graph.
Foundations2 areas
The SQL fundamentals that actually matter on Databricks: Spark SQL, differences from Postgres, types, functions, and the patterns that keep showing up in exams.
Python and PySpark in the Databricks context: the DataFrame API, differences from pandas, notebooks, and transformation patterns.
Platform & governance6 areas
The Databricks workspace: platform architecture, notebooks, files, Git folders, bundles, and the development workflow.
Unity Catalog: metastore, catalogs, schemas, managed and external tables, privileges, row filters, column masks, ABAC, lineage.
Orchestration with Lakeflow Jobs (tasks, dependencies, triggers, parameters, repair) and declarative pipelines with Lakeflow pipelines (formerly Delta Live Tables).
Clusters, serverless, SQL warehouses, policies, pools, runtimes, Photon: how to choose compute and how to diagnose problems.
Databricks Marketplace and the Discover section: datasets, models, and apps shared through OpenSharing (formerly Delta Sharing). Covered only to the extent the exams require.
Databricks Apps, the serverless home for data and AI applications, and the applications Databricks itself ships on it.
Data engineering · Lakeflow7 areas
Monitoring job and pipeline runs: states, run history, duration, failure rate, spotting upstream blockers.
Getting data into the lakehouse: batch, streaming, and incremental patterns; COPY INTO, Auto Loader, Lakeflow Connect, JDBC/REST, semi-structured data.
Delta Lake as the lakehouse storage format: transactions, time travel, medallion architecture, gold-layer objects, liquid clustering, predictive optimization.
Structured Streaming on Databricks: triggers, checkpoints, watermarks, stateful operations, real-time mode.
Keeping bad rows out of silver and gold: pipeline expectations, Delta constraints, the DQX framework, and the monitoring that tells you when quality drifts.
The open-source projects around the platform: Databricks Labs, the tools that migrate, test and validate, and what is safe to depend on.
Lakebase, the managed Postgres inside Databricks: projects and branches, computes that scale to zero, connecting and authenticating, synced tables from Unity Catalog, changes back into Delta, and apps built on it.
Data warehousing & BI6 areas
The workspace SQL editor: saved queries, snippets, parameters, running on SQL warehouses.
AI/BI Dashboards: datasets, visualizations, filters, publishing, and dashboard tasks in jobs.
The Genie family: Genie Agents for questions over curated data, Genie One for business users, Genie Code for whoever is building, and the Genie Ontology they share.
Alerts on SQL queries: conditions, notification destinations, scheduling, and integration with jobs.
Query History and Query Profile: reading the plan, spotting bottlenecks, optimizing queries on warehouses.
SQL warehouses: types (classic, pro, serverless), sizing, scaling, costs, and when to use one instead of a cluster.
AI & ML8 areas
AI Playground: trying out foundation models and tool calling from the workspace before writing code.
Agent Framework and Agent Bricks: building, evaluating, and deploying agents on governed data.
Unity Gateway (formerly Mosaic AI Gateway): governance, rate limiting, logging, and fallback for models, MCP servers and agents.
MLflow Experiments: tracking runs, parameters, metrics, and artifacts.
Feature engineering and the Feature Store in Unity Catalog: feature tables, lookups, point-in-time joins.
Models in Unity Catalog: versions, aliases, promotion across environments.
Model Serving: endpoints for custom models, foundation models, and agents; traffic, scaling, monitoring.
Databricks AI Search (formerly Vector Search): indexes, syncing from Delta tables, hybrid and full-text search for RAG.