Jupyter is where most people first touch Spark, and it is also where a surprising share of Spark incidents start: a kernel that holds forty executors over a weekend, a spark.jars.packages setting that silently does nothing, an executor that cannot connect back to a laptop, a toPandas() call that kills the kernel. Almost all of these come from one question that the notebook interface hides: where does the Spark driver run?
This article answers that question for the three placements you will meet in practice. In the first, the notebook kernel is the driver. In the second, the driver runs on a server behind a REST gateway such as Livy. In the third, the kernel is a thin Spark Connect client and the driver lives in a Connect server. For each one it covers what you install, how you configure it, what breaks, and what it costs. The worked example runs a client-mode driver inside a Kubernetes notebook pod, which is the setup that needs the most care to get right.
Three places the driver can live
Why the driver's location matters
A Spark application has exactly one driver. The driver turns your DataFrame code into a plan, splits it into stages and tasks, hands those tasks to executors, tracks shuffle locations and receives every row you collect. Executors are long-lived processes that must be able to open connections back to the driver. Everything else follows from that design.
A notebook kernel is a long-lived process that a human drives interactively, with pauses of minutes or days between cells. A Spark driver was designed for batch jobs that start, run and exit. Put the two together and you get a driver whose lifetime is set by a browser tab. Idle time, kernel restarts, laptop sleep and Wi-Fi changes now affect a distributed system, and the driver keeps holding its executors through all of them unless you tell it not to.
The three placements trade off how much of the driver lives in the kernel. The more of it lives there, the simpler the setup and the more fragile the result.
Placement A: the kernel is the driver
Placement A is the default when you pip install pyspark and open a notebook. Creating a SparkSession starts a JVM next to the Python kernel, connected through Py4J, and that JVM is the driver. With master('local[*]') the executors are threads in the same JVM, so it is a single-machine Spark that is ideal for learning, unit-sized data and testing transformations.
from pyspark.sql import SparkSession
spark = (
SparkSession.builder
.appName("nb-exploration")
.master("local[4]")
# Must be set BEFORE the first getOrCreate(); the JVM reads them once.
.config("spark.jars.packages", "org.postgresql:postgresql:42.7.4")
.config("spark.driver.memory", "4g")
.config("spark.sql.execution.arrow.pyspark.enabled", "true")
.config("spark.sql.repl.eagerEval.enabled", "true")
.getOrCreate()
)
print(spark.version, spark.sparkContext.uiWebUrl)Two lines in that block cause most of the confusion in this placement. First, getOrCreate() returns the existing session if one is already running in the kernel, and options that only apply when the JVM starts, such as packages, driver memory and extra class path, are ignored on the second call. If you edit the cell and re-run it, nothing changes. Restart the kernel, or call spark.stop() and then rebuild the session. Second, spark.driver.memory is the heap of the driver JVM, and it cannot grow after start. Every collect() and toPandas() lands there first and then in Python, so the kernel needs room for two copies of the result.
Turning on Arrow makes toPandas() and createDataFrame(pandas_df) move data in columnar batches instead of pickled rows. That is far faster, but it does not make a large collect safe. Eager evaluation renders a DataFrame as an HTML table when it is the last expression in a cell, capped by spark.sql.repl.eagerEval.maxNumRows (default 20). That is pleasant for exploration and it also triggers a job on every display, which is easy to forget when the DataFrame reads a large table.
Placement A can also point at a real cluster: set the master to yarn, a k8s:// URL or a standalone spark:// URL, and the kernel becomes a client-mode driver. The executors run on the cluster but connect back to the kernel's machine. That works from an edge node or a pod inside the cluster network. From a laptop on home Wi-Fi behind NAT it does not, because the executors cannot route to you.
Placement B: the driver behind Livy
In placement B the kernel is not a Spark driver at all. The Livy gateway runs the driver on the cluster and the notebook's sparkmagic kernel ships each cell's code to it as text over REST. The kernel can crash or restart and the session survives until Livy's idle timeout. Dependencies, cluster credentials and the YARN queue are all managed on the server, so analysts never need cluster network access.
The cost is a split runtime. Cell code runs remotely, so local Python variables are not visible to it. Results come back as text or through sparkmagic's %%sql -o and %%send_to_spark magics, which serialise data through the REST channel and are only suitable for small results. Errors arrive as strings, not tracebacks you can inspect. The Livy article covers its configuration, state machine and security in depth, so this page does not repeat them.
Placement C: Spark Connect
In placement C the kernel runs a Spark Connect client. The DataFrame API in the notebook builds unresolved logical plans, sends them over gRPC to a Connect server, and the driver inside that server analyses, optimises and executes them. Results stream back as Arrow batches. The client has no JVM, so the notebook image does not need Java and a client crash cannot take down the driver.
# Spark 3.4+ client: pip install "pyspark[connect]"
# Spark 4.0+ client: pip install pyspark-client (pure Python, no JARs or JRE)
from pyspark.sql import SparkSession
spark = SparkSession.builder.remote("sc://spark-connect.data.svc:15002").getOrCreate()
orders = spark.read.table("sales.orders")
daily = orders.groupBy("order_date").sum("amount")
pdf = daily.orderBy("order_date").limit(1000).toPandas() # bounded resultUnlike placement B, this is the normal PySpark DataFrame API, so local Python code, pandas and plotting work as they do in placement A. The limits are the parts of PySpark that need direct JVM access. RDD APIs and sparkContext are not available over Connect, and libraries that reach into _jvm fail. Check your library list before you standardise on it. Also note that by default a Connect server's sessions share one Spark application, so one user's heavy job competes with everybody else's for the same executors. Tenant isolation is something you design, for example one server per team or per user.
Worked example: a client-mode driver in a Kubernetes notebook pod
A platform team runs JupyterHub on Kubernetes. Each user gets a notebook pod in the notebooks namespace and wants PySpark with full API access, so the team chooses placement A in client mode against the same cluster, as described in Spark on Kubernetes. Executor pods must dial back into the notebook pod, and four settings make that work.
import os, socket
from pyspark.sql import SparkSession
pod_ip = os.environ["POD_IP"] # injected via the downward API
pod_name = os.environ["HOSTNAME"]
spark = (
SparkSession.builder
.master("k8s://https://kubernetes.default.svc:443")
.config("spark.kubernetes.namespace", "notebooks")
.config("spark.kubernetes.container.image", "registry.local/spark-py:3.5.3")
# The notebook pod's service account must be allowed to create and delete pods.
# 1. Executors reach the driver by pod IP on fixed ports.
.config("spark.driver.host", pod_ip)
.config("spark.driver.bindAddress", "0.0.0.0")
.config("spark.driver.port", "29413")
.config("spark.blockManager.port", "29414")
# 2. Executors are owned by this pod, so they die with it.
.config("spark.kubernetes.driver.pod.name", pod_name)
# 3. Idle notebooks give executors back.
.config("spark.dynamicAllocation.enabled", "true")
.config("spark.dynamicAllocation.shuffleTracking.enabled", "true")
.config("spark.dynamicAllocation.minExecutors", "0")
.config("spark.dynamicAllocation.maxExecutors", "8")
.config("spark.dynamicAllocation.executorIdleTimeout", "120s")
# 4. The UI is reachable through the Hub's proxy.
.config("spark.ui.proxyBase", f"/user/{os.environ['JUPYTERHUB_USER']}/proxy/4040")
.getOrCreate()
)Fixed driver and block manager ports let the NetworkPolicy for the namespace allow exactly those two ports from executor pods to notebook pods, rather than opening every port. Pinning assumes one Spark session per notebook server. A second kernel in the same pod would retry onto the next ports, which the policy blocks, and its UI would move to 4041. Enforce one session with a startup hook, or allow a small port range sized to spark.port.maxRetries. Setting spark.kubernetes.driver.pod.name gives every executor pod an owner reference to the notebook pod, so Kubernetes garbage-collects the executors when the Hub culls the notebook. Without it, a culled notebook leaves orphaned executors running until someone notices them.
Dynamic allocation with shuffle tracking (described in dynamic allocation) lets the session drop to zero executors two minutes after the last task, while keeping executors that still hold shuffle data that later stages need. Kubernetes has no external shuffle service by default, so shuffle tracking is what makes dynamic allocation safe here. The Spark UI binds port 4040, or 4041 and higher if 4040 is taken, and jupyter-server-proxy exposes it under the Hub URL. spark.ui.proxyBase makes the UI's links include that prefix.
The numbers show why this matters. Before the change, the team's 30 users each held a fixed four executors of 4 cores and 16 GB from the moment they opened a notebook until it was culled, so the cluster carried 480 cores for an interactive load whose busy-hour peak, read from the Spark UI event logs, was closer to 100. After the change the same users rarely held more than 120 cores at once, and executors disappeared within minutes of a notebook going quiet.
Notebook hygiene that prevents incidents
Several habits keep notebooks from turning into production incidents, whichever placement you use.
- Bound every result that crosses to Python. Use
limit()beforetoPandas(), sample before plotting, and write large results to a table instead of collecting them. Raisingspark.driver.maxResultSizeonly moves the out-of-memory error. - Match Python versions. In placement A against a cluster, the executor image's Python must match the kernel's minor version, or Python UDFs fail with a version mismatch. Set
PYSPARK_PYTHONexplicitly and build the notebook and executor images from one base. - Configure before the first session. If you set options through environment variables,
PYSPARK_SUBMIT_ARGSmust end withpyspark-shelland must be set before the firstSparkSessionis created. - Stop sessions you are done with. End analysis notebooks with
spark.stop()and configure the server's idle kernel culling, for exampleMappingKernelManager.cull_idle_timeout, so abandoned kernels release their drivers. - Promote notebooks deliberately. Tools such as papermill run a notebook headlessly with injected parameters, which is fine for reports. Pipelines that need retries, tests and review belong in a module submitted as a job.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Executors start, then the job hangs and tasks never run | Executors cannot reach the client-mode driver (NAT, firewall, wrong host) | Set spark.driver.host to a routable address, pin ports, allow them in the network policy |
| New package or memory setting has no effect | getOrCreate() returned the existing session | Restart the kernel or stop the session first |
| Kernel dies on toPandas() | Result exceeds driver heap or kernel memory | limit() or aggregate first; write big results to storage |
| Executors keep running after the notebook closes | No owner reference or no idle culling | Set the driver pod name; enable kernel and Hub culling |
| Python UDF fails with a version mismatch | Kernel and executor Python differ | One base image; set PYSPARK_PYTHON |
| Spark UI links return 404 behind the Hub | UI unaware of the proxy prefix | Set spark.ui.proxyBase to the proxied path |
| Library fails under Spark Connect | It needs sparkContext or the JVM gateway | Use placement A for that work or replace the library |
Trade-offs
| A: kernel is driver | B: Livy / sparkmagic | C: Spark Connect | |
|---|---|---|---|
| Setup effort | Lowest locally, highest in client mode | Gateway to run and secure | Server to run; small client |
| API coverage | Full, including RDDs | Full, but remote and text-based | DataFrame and SQL; no RDDs |
| Kernel crash | Kills the application | Session survives | Session survives |
| Network from notebook | Executors must reach it | HTTPS to the gateway only | gRPC to the server only |
| Local pandas and plotting | Native | Through magics; small data | Native |
| Isolation between users | One application per user | One session per user | Shared by default; design it |
A sensible default is placement C for analysts and placement A in client mode for engineers who need the full API. Use placement B when you already run Livy or need its REST batch interface. Placement A in local mode remains the right tool for learning and for testing transformations on small samples.
What to do next
- Decide which placement each group of users gets, and write the reason down next to the notebook image definition.
- Put all JVM-start options (packages, driver memory, ports) in one session-builder cell at the top of the notebook, or in the image's
spark-defaults.conf. - Turn on Arrow, and add a lint or review rule against unbounded
collect()andtoPandas()calls. - In client mode, set the driver host and pin the driver and block manager ports, then write the NetworkPolicy that allows only those ports.
- Enable dynamic allocation with shuffle tracking and a minimum of zero, and configure idle culling in both Jupyter and the Hub.
- Try a Spark Connect server in a test namespace, run your three most common notebooks against it, and list the calls that fail.
- Move any notebook that runs on a schedule into a parameterised job with tests.