Jupyter is where most people first touch Spark, and it is also where a surprising share of Spark incidents start: a kernel that holds forty executors over a weekend, a spark.jars.packages setting that silently does nothing, an executor that cannot connect back to a laptop, a toPandas() call that kills the kernel. Almost all of these come from one question that the notebook interface hides: where does the Spark driver run?

This article answers that question for the three placements you will meet in practice. In the first, the notebook kernel is the driver. In the second, the driver runs on a server behind a REST gateway such as Livy. In the third, the kernel is a thin Spark Connect client and the driver lives in a Connect server. For each one it covers what you install, how you configure it, what breaks, and what it costs. The worked example runs a client-mode driver inside a Kubernetes notebook pod, which is the setup that needs the most care to get right.

Three places the driver can live

A: kernel is the driverpyspark in the kernel processB: driver behind a gatewaysparkmagic to Livy over RESTC: Spark Connect clientgRPC to a Connect serverKernel = JVM driverplans, schedules, collectsLivy sessiondriver on the clusterConnect serverdriver on the clustersame processcode as textunresolved plansExecutors (YARN, Kubernetes or standalone)must reach the driver's RPC and block manager portsexecutors dial backA: easiest start, hardest networking and resource hygiene.B and C: the kernel can restart without killing the job; dependencies live server-side.Results only cross back to the notebook when you collect, show or convert to pandas.
The notebook kernel can be the driver, a client of a REST gateway that runs the driver, or a Spark Connect client of a server that runs it. Executors always connect to the driver, wherever it is.

Why the driver's location matters

A Spark application has exactly one driver. The driver turns your DataFrame code into a plan, splits it into stages and tasks, hands those tasks to executors, tracks shuffle locations and receives every row you collect. Executors are long-lived processes that must be able to open connections back to the driver. Everything else follows from that design.

A notebook kernel is a long-lived process that a human drives interactively, with pauses of minutes or days between cells. A Spark driver was designed for batch jobs that start, run and exit. Put the two together and you get a driver whose lifetime is set by a browser tab. Idle time, kernel restarts, laptop sleep and Wi-Fi changes now affect a distributed system, and the driver keeps holding its executors through all of them unless you tell it not to.

The three placements trade off how much of the driver lives in the kernel. The more of it lives there, the simpler the setup and the more fragile the result.

Placement A: the kernel is the driver

Placement A is the default when you pip install pyspark and open a notebook. Creating a SparkSession starts a JVM next to the Python kernel, connected through Py4J, and that JVM is the driver. With master('local[*]') the executors are threads in the same JVM, so it is a single-machine Spark that is ideal for learning, unit-sized data and testing transformations.

from pyspark.sql import SparkSession

spark = (
    SparkSession.builder
    .appName("nb-exploration")
    .master("local[4]")
    # Must be set BEFORE the first getOrCreate(); the JVM reads them once.
    .config("spark.jars.packages", "org.postgresql:postgresql:42.7.4")
    .config("spark.driver.memory", "4g")
    .config("spark.sql.execution.arrow.pyspark.enabled", "true")
    .config("spark.sql.repl.eagerEval.enabled", "true")
    .getOrCreate()
)
print(spark.version, spark.sparkContext.uiWebUrl)

Two lines in that block cause most of the confusion in this placement. First, getOrCreate() returns the existing session if one is already running in the kernel, and options that only apply when the JVM starts, such as packages, driver memory and extra class path, are ignored on the second call. If you edit the cell and re-run it, nothing changes. Restart the kernel, or call spark.stop() and then rebuild the session. Second, spark.driver.memory is the heap of the driver JVM, and it cannot grow after start. Every collect() and toPandas() lands there first and then in Python, so the kernel needs room for two copies of the result.

Turning on Arrow makes toPandas() and createDataFrame(pandas_df) move data in columnar batches instead of pickled rows. That is far faster, but it does not make a large collect safe. Eager evaluation renders a DataFrame as an HTML table when it is the last expression in a cell, capped by spark.sql.repl.eagerEval.maxNumRows (default 20). That is pleasant for exploration and it also triggers a job on every display, which is easy to forget when the DataFrame reads a large table.

Placement A can also point at a real cluster: set the master to yarn, a k8s:// URL or a standalone spark:// URL, and the kernel becomes a client-mode driver. The executors run on the cluster but connect back to the kernel's machine. That works from an edge node or a pod inside the cluster network. From a laptop on home Wi-Fi behind NAT it does not, because the executors cannot route to you.

Placement B: the driver behind Livy

In placement B the kernel is not a Spark driver at all. The Livy gateway runs the driver on the cluster and the notebook's sparkmagic kernel ships each cell's code to it as text over REST. The kernel can crash or restart and the session survives until Livy's idle timeout. Dependencies, cluster credentials and the YARN queue are all managed on the server, so analysts never need cluster network access.

The cost is a split runtime. Cell code runs remotely, so local Python variables are not visible to it. Results come back as text or through sparkmagic's %%sql -o and %%send_to_spark magics, which serialise data through the REST channel and are only suitable for small results. Errors arrive as strings, not tracebacks you can inspect. The Livy article covers its configuration, state machine and security in depth, so this page does not repeat them.

Placement C: Spark Connect

In placement C the kernel runs a Spark Connect client. The DataFrame API in the notebook builds unresolved logical plans, sends them over gRPC to a Connect server, and the driver inside that server analyses, optimises and executes them. Results stream back as Arrow batches. The client has no JVM, so the notebook image does not need Java and a client crash cannot take down the driver.

# Spark 3.4+ client:   pip install "pyspark[connect]"
# Spark 4.0+ client:   pip install pyspark-client   (pure Python, no JARs or JRE)
from pyspark.sql import SparkSession

spark = SparkSession.builder.remote("sc://spark-connect.data.svc:15002").getOrCreate()

orders = spark.read.table("sales.orders")
daily = orders.groupBy("order_date").sum("amount")
pdf = daily.orderBy("order_date").limit(1000).toPandas()   # bounded result

Unlike placement B, this is the normal PySpark DataFrame API, so local Python code, pandas and plotting work as they do in placement A. The limits are the parts of PySpark that need direct JVM access. RDD APIs and sparkContext are not available over Connect, and libraries that reach into _jvm fail. Check your library list before you standardise on it. Also note that by default a Connect server's sessions share one Spark application, so one user's heavy job competes with everybody else's for the same executors. Tenant isolation is something you design, for example one server per team or per user.

Worked example: a client-mode driver in a Kubernetes notebook pod

A platform team runs JupyterHub on Kubernetes. Each user gets a notebook pod in the notebooks namespace and wants PySpark with full API access, so the team chooses placement A in client mode against the same cluster, as described in Spark on Kubernetes. Executor pods must dial back into the notebook pod, and four settings make that work.

import os, socket
from pyspark.sql import SparkSession

pod_ip   = os.environ["POD_IP"]          # injected via the downward API
pod_name = os.environ["HOSTNAME"]

spark = (
    SparkSession.builder
    .master("k8s://https://kubernetes.default.svc:443")
    .config("spark.kubernetes.namespace", "notebooks")
    .config("spark.kubernetes.container.image", "registry.local/spark-py:3.5.3")
    # The notebook pod's service account must be allowed to create and delete pods.
    # 1. Executors reach the driver by pod IP on fixed ports.
    .config("spark.driver.host", pod_ip)
    .config("spark.driver.bindAddress", "0.0.0.0")
    .config("spark.driver.port", "29413")
    .config("spark.blockManager.port", "29414")
    # 2. Executors are owned by this pod, so they die with it.
    .config("spark.kubernetes.driver.pod.name", pod_name)
    # 3. Idle notebooks give executors back.
    .config("spark.dynamicAllocation.enabled", "true")
    .config("spark.dynamicAllocation.shuffleTracking.enabled", "true")
    .config("spark.dynamicAllocation.minExecutors", "0")
    .config("spark.dynamicAllocation.maxExecutors", "8")
    .config("spark.dynamicAllocation.executorIdleTimeout", "120s")
    # 4. The UI is reachable through the Hub's proxy.
    .config("spark.ui.proxyBase", f"/user/{os.environ['JUPYTERHUB_USER']}/proxy/4040")
    .getOrCreate()
)

Fixed driver and block manager ports let the NetworkPolicy for the namespace allow exactly those two ports from executor pods to notebook pods, rather than opening every port. Pinning assumes one Spark session per notebook server. A second kernel in the same pod would retry onto the next ports, which the policy blocks, and its UI would move to 4041. Enforce one session with a startup hook, or allow a small port range sized to spark.port.maxRetries. Setting spark.kubernetes.driver.pod.name gives every executor pod an owner reference to the notebook pod, so Kubernetes garbage-collects the executors when the Hub culls the notebook. Without it, a culled notebook leaves orphaned executors running until someone notices them.

Dynamic allocation with shuffle tracking (described in dynamic allocation) lets the session drop to zero executors two minutes after the last task, while keeping executors that still hold shuffle data that later stages need. Kubernetes has no external shuffle service by default, so shuffle tracking is what makes dynamic allocation safe here. The Spark UI binds port 4040, or 4041 and higher if 4040 is taken, and jupyter-server-proxy exposes it under the Hub URL. spark.ui.proxyBase makes the UI's links include that prefix.

The numbers show why this matters. Before the change, the team's 30 users each held a fixed four executors of 4 cores and 16 GB from the moment they opened a notebook until it was culled, so the cluster carried 480 cores for an interactive load whose busy-hour peak, read from the Spark UI event logs, was closer to 100. After the change the same users rarely held more than 120 cores at once, and executors disappeared within minutes of a notebook going quiet.

Notebook hygiene that prevents incidents

Several habits keep notebooks from turning into production incidents, whichever placement you use.

  • Bound every result that crosses to Python. Use limit() before toPandas(), sample before plotting, and write large results to a table instead of collecting them. Raising spark.driver.maxResultSize only moves the out-of-memory error.
  • Match Python versions. In placement A against a cluster, the executor image's Python must match the kernel's minor version, or Python UDFs fail with a version mismatch. Set PYSPARK_PYTHON explicitly and build the notebook and executor images from one base.
  • Configure before the first session. If you set options through environment variables, PYSPARK_SUBMIT_ARGS must end with pyspark-shell and must be set before the first SparkSession is created.
  • Stop sessions you are done with. End analysis notebooks with spark.stop() and configure the server's idle kernel culling, for example MappingKernelManager.cull_idle_timeout, so abandoned kernels release their drivers.
  • Promote notebooks deliberately. Tools such as papermill run a notebook headlessly with injected parameters, which is fine for reports. Pipelines that need retries, tests and review belong in a module submitted as a job.

Failure modes

SymptomCauseFix
Executors start, then the job hangs and tasks never runExecutors cannot reach the client-mode driver (NAT, firewall, wrong host)Set spark.driver.host to a routable address, pin ports, allow them in the network policy
New package or memory setting has no effectgetOrCreate() returned the existing sessionRestart the kernel or stop the session first
Kernel dies on toPandas()Result exceeds driver heap or kernel memorylimit() or aggregate first; write big results to storage
Executors keep running after the notebook closesNo owner reference or no idle cullingSet the driver pod name; enable kernel and Hub culling
Python UDF fails with a version mismatchKernel and executor Python differOne base image; set PYSPARK_PYTHON
Spark UI links return 404 behind the HubUI unaware of the proxy prefixSet spark.ui.proxyBase to the proxied path
Library fails under Spark ConnectIt needs sparkContext or the JVM gatewayUse placement A for that work or replace the library

Trade-offs

A: kernel is driverB: Livy / sparkmagicC: Spark Connect
Setup effortLowest locally, highest in client modeGateway to run and secureServer to run; small client
API coverageFull, including RDDsFull, but remote and text-basedDataFrame and SQL; no RDDs
Kernel crashKills the applicationSession survivesSession survives
Network from notebookExecutors must reach itHTTPS to the gateway onlygRPC to the server only
Local pandas and plottingNativeThrough magics; small dataNative
Isolation between usersOne application per userOne session per userShared by default; design it

A sensible default is placement C for analysts and placement A in client mode for engineers who need the full API. Use placement B when you already run Livy or need its REST batch interface. Placement A in local mode remains the right tool for learning and for testing transformations on small samples.

What to do next

  1. Decide which placement each group of users gets, and write the reason down next to the notebook image definition.
  2. Put all JVM-start options (packages, driver memory, ports) in one session-builder cell at the top of the notebook, or in the image's spark-defaults.conf.
  3. Turn on Arrow, and add a lint or review rule against unbounded collect() and toPandas() calls.
  4. In client mode, set the driver host and pin the driver and block manager ports, then write the NetworkPolicy that allows only those ports.
  5. Enable dynamic allocation with shuffle tracking and a minimum of zero, and configure idle culling in both Jupyter and the Hub.
  6. Try a Spark Connect server in a test namespace, run your three most common notebooks against it, and list the calls that fail.
  7. Move any notebook that runs on a schedule into a parameterised job with tests.
Key takeaway: Every Spark-in-Jupyter problem starts with where the driver runs. If the kernel is the driver, configure everything before the first getOrCreate, bound every result that crosses into Python, and in client mode make the driver routable on pinned ports with executors owned by the notebook pod and released by dynamic allocation. If the driver sits behind Livy or a Spark Connect server, the kernel can restart freely and needs no cluster network access, at the cost of a remote, text-based workflow with Livy or a DataFrame-only API with Connect. Pick the placement per user group on purpose, and move scheduled notebooks into tested jobs.