Skip to content

Execution & Compute

Skill: databricks-execution-compute

You can execute code on Databricks directly from your local editor — no browser, no notebook UI. Three execution modes cover every workload: Databricks Connect for local Spark development, serverless jobs for fire-and-forget heavy processing, and interactive clusters for stateful multi-step workflows. Your AI coding assistant picks the right mode based on what you’re doing, manages dependencies, and captures output cleanly.

“Run this ETL script on serverless compute. The script reads a table, aggregates events, and writes the summary back.”

/tmp/event_summary.py
from databricks.connect import DatabricksSession
import pyspark.sql.functions as F
import json
spark = DatabricksSession.builder.serverless(True).getOrCreate()
df = (
spark.read.table("catalog.schema.raw_events")
.filter(F.col("event_date") >= "2026-01-01")
.groupBy("user_id", "event_type")
.agg(F.count("*").alias("event_count"))
)
df.write.mode("overwrite").saveAsTable("catalog.schema.user_event_summary")
dbutils.notebook.exit(json.dumps({"rows_written": df.count()}))

Then submit it via MCP:

execute_code(
file_path="/tmp/event_summary.py",
compute_type="serverless"
)

Key decisions:

  • compute_type="serverless" — no cluster to provision. Serverless cold-starts in 25-50 seconds and tears down after execution. Best for Python and SQL workloads that do not need persistent state across calls.
  • file_path over inline code — non-trivial scripts belong in files. You can edit, version, and re-run them. Inline code is for one-liners.
  • dbutils.notebook.exit() for output — on serverless, print() output is unreliable. Use dbutils.notebook.exit(json.dumps(result)) to return structured results your assistant can read after the run.
  • Databricks Connect inside the script — the script itself uses DatabricksSession.builder.serverless(True) so the Spark calls execute on serverless compute. This is the canonical pattern.

“Set up an interactive session on my cluster. Run some setup code, then query the results in a follow-up.”

# First call -- creates an execution context
result = execute_code(
code="""
import pandas as pd
df = pd.DataFrame({
"region": ["US", "EU", "APAC", "US", "EU"],
"revenue": [1200, 950, 800, 1100, 1050],
})
spark_df = spark.createDataFrame(df)
spark_df.createOrReplaceTempView("regional_revenue")
print("View created")
""",
compute_type="cluster",
)
# Second call -- reuses the same context, variables persist
execute_code(
code="spark.sql('SELECT region, SUM(revenue) FROM regional_revenue GROUP BY region').show()",
context_id=result["context_id"],
cluster_id=result["cluster_id"],
)

The context_id preserves variables, temp views, and imports between calls. This is the cluster equivalent of running cells in a notebook. Drop the context when done by passing destroy_context_on_completion=True on the last call.

“Execute my local transform.py file on the dev cluster.”

execute_code(
file_path="/Users/me/project/src/transform.py",
compute_type="cluster",
)

The tool detects the language from the file extension (.py, .scala, .sql, .r) and uploads it for execution. This is the fastest path from local development to remote testing — no manual upload, no workspace notebook creation.

“Spin up a 4-worker autoscaling cluster with ML Runtime for model training.”

# Create an autoscaling cluster
manage_cluster(
action="create",
name="ml-training-cluster",
spark_version="15.4.x-ml-scala2.12",
node_type_id="i3.xlarge",
autoscale_min_workers=2,
autoscale_max_workers=8,
autotermination_minutes=60,
)
# Later: terminate when done (does not delete)
manage_cluster(action="terminate", cluster_id="0123-456789-abcdef")
# Check status while it stops
list_compute(resource="clusters", cluster_id="0123-456789-abcdef")

terminate stops the cluster but preserves its configuration for restarting later. delete is permanent and irreversible. Always confirm before deleting.

  • print() on serverless is unreliable — serverless compute does not guarantee stdout capture. Use dbutils.notebook.exit(json.dumps(result)) to return data from serverless runs. This catches everyone at least once.
  • Scala and R require interactive clusters — serverless only supports Python and SQL. For Scala or R, fall back to an interactive cluster with compute_type="cluster".
  • Never start a cluster without asking — manage_cluster(action="start") takes 3-8 minutes and costs money. Always run list_compute(resource="clusters") first. If nothing is running, present the user a choice: start a cluster, or switch to serverless for instant execution.
  • Context leaks on clusters — execution contexts consume memory. If you create many contexts without destroying them, the cluster can run out of memory. Pass destroy_context_on_completion=True when you finish iterating.