Python in notebooks: dbutils, widgets, modules

kept in this browsersaved to your account

How dbutils, widgets, and %pip fit into a Databricks notebook, and how to structure code so it survives the move to a proper module.

What it is

A Databricks notebook is a Python (or SQL, Scala, R) session attached to a cluster, split into cells, plus a set of platform-specific helpers that don’t exist in a plain .py script: dbutils, widgets, and magic commands like %pip and %run. None of this is Spark itself — it’s the layer that makes a notebook a convenient place to develop, separate from the DataFrame API you use once code is running.

Why it exists

A notebook needs to do things a script rarely does interactively: install a library for this session only, expose parameters so the same notebook runs for different dates or tables without editing code, move files around before Spark reads them, chain notebooks together, and read secrets without hardcoding credentials. dbutils bundles those as one namespace instead of scattering them across separate libraries.

How it works

dbutils submodules

SubmodulePurposeExample
dbutils.fsList, copy, move files on cluster-attached storagedbutils.fs.ls("/Volumes/main/default/raw")
dbutils.widgetsDefine and read notebook parametersdbutils.widgets.text("run_date", "2026-09-10")
dbutils.notebookRun another notebook and get its return valuedbutils.notebook.run("clean_orders", timeout_seconds=600)
dbutils.jobs.taskValuesPass small values between tasks in a jobdbutils.jobs.taskValues.set("row_count", 1200)
dbutils.secretsRead credentials from a secret scopedbutils.secrets.get(scope="etl", key="api_token")

dbutils.fs overlaps with plain Python’s os/pathlib, but talks to cloud storage and volumes uniformly across backends — it isn’t a replacement for os on the driver’s local disk.

Widgets

A widget (dbutils.widgets.text/dropdown/combobox/multiselect) creates a control at the top of the notebook and a value you read back with dbutils.widgets.get("name"). This is what turns a hardcoded notebook into one a Lakeflow Job can call with different parameters per run (see Job and task parameters, dynamic values, and task values) instead of maintaining a copy per environment.

Libraries: %pip install and restarting Python

%pip install some-package installs a library scoped to the current session only — it doesn’t touch the cluster for anyone else, and disappears when the session detaches. Because Python loads a module once per process, upgrading a package mid-session usually needs dbutils.library.restartPython() afterward; this clears local Python state but leaves Spark and table state untouched.

%run versus dbutils.notebook.run

Both let one notebook use another, but they’re not interchangeable:

%run ./helpersdbutils.notebook.run("helpers", 60)
Executes inSame session, same variablesSeparate, isolated job run
Shares variables/functions backYesNo — only a string return value
Accepts parametersNoYes, as a dict
Use caseShared functions, constants used inlineA step that should run independently, possibly retried on its own

From notebook to module

A notebook full of dbutils.widgets calls and top-level code is hard to unit test and can’t be imported. Isolating logic into plain functions in a .py file — imported once that file sits in the same Repo/Git folder, or installed as a package — keeps the notebook itself a thin entry point: read widgets, call functions, write output. That split is what lets the same logic later ship as a wheel dependency for a job instead of being copy-pasted between notebooks.

display() versus show()

df.show() is plain Spark: a fixed-width text dump to the console. display() is a Databricks notebook feature: a rich, sortable table with one-click charts — but it only exists inside a notebook, so code meant to run as a plain script or job task shouldn’t depend on it.

Example

dbutils.widgets.text("run_date", "2026-09-10", "Run date")
run_date = dbutils.widgets.get("run_date")

orders = spark.table("shop.silver.orders").filter(f"order_date = '{run_date}'")
display(orders)          # rich, interactive — notebook only
orders.show(5)            # plain text — works anywhere Spark runs

row_count = orders.count()
dbutils.jobs.taskValues.set(key="row_count", value=row_count)
# helpers.py, imported like a normal module from a Git folder.
def clean_orders(df):
    return df.dropna(subset=["order_id"]).dropDuplicates(["order_id"])
# In the notebook: reusable logic stays testable outside the notebook.
from helpers import clean_orders
silver_orders = clean_orders(spark.table("shop.bronze.orders"))

Common mistakes

  • Using %run when the goal is an isolated, parameterized, independently retriable step — that’s what dbutils.notebook.run or a separate job task (see Lakeflow Jobs, what a job is) is for.
  • Forgetting dbutils.library.restartPython() after %pip install, then wondering why the old version of a package is still active.
  • Reading a widget value without a default and without checking it’s been created first, which raises on a fresh session.
  • Leaving business logic as top-level notebook code instead of functions in a module — untestable, and it can’t be reused when the pipeline moves to a wheel-based job.
  • Calling display() inside code meant to run outside a notebook UI (a wheel task, a plain script) — it isn’t defined there.

Where this sits

Resources

1All resources
Report a problem with this page
What kind of problem?

Reports about "Python in notebooks: dbutils, widgets, modules" go to the maintainer, not to a public thread.

Suggest a resource
What kind?

Nothing appears on the site automatically. A person reads every suggestion, checks the link and writes the note that goes with it.