Real data is rarely flat. An order has line items, an event has a map of attributes, a sensor reading carries a nested location. Spark DataFrames model this with three complex types: StructType for a fixed set of named fields, ArrayType for an ordered list of one element type, and MapType for key-value pairs. Used well, they let you store and process nested records without flattening them into many tables and joining them back. Used badly, they produce row explosions, unreadable column expressions and queries that read every nested field from storage.

This article explains how each type is represented, how to build, read and transform them with built-in functions, why higher-order functions usually beat explode, how to change one field deep inside a struct, how JSON and Spark 4's VARIANT type fit in, and what the optimizer can and cannot do for nested data. It assumes you know DataFrame basics; examples are PySpark, and the same functions exist in Scala and SQL.

The three complex types and how they are stored

A schema is a tree. Each column has a data type, and complex types nest other types, so array<struct<sku:string,qty:int,price:decimal(10,2)>> is a perfectly ordinary column. Inside the engine a row is stored in Tungsten's binary row format; a struct is an embedded row, and arrays and maps are variable-length regions with a header holding the element count and null bits, followed by the elements (a map is a pair of arrays, keys and values). This means reading one element does not deserialise the rest, and functions over arrays run on compact memory rather than on Python or JVM objects. Tungsten describes the layout in more depth.

TypeHoldsAccessTypical use
structFixed named fields of mixed typescol('a.b'), getFieldA nested record such as an address
arrayOrdered elements of one typecol[i] (0-based), element_at (1-based)Line items, tags, time series
mapKeys of one type to values of one typecol[key], element_at, map_keysSparse attributes, counters

Structs are the cheapest: field names live in the schema, not in each row, and Parquet stores each leaf as its own column. Maps are the most flexible and the least optimizable, because keys are data, so the engine cannot prune a single key at read time. Prefer a struct whenever the set of keys is known.

Building and reading nested values

Build complex values with struct, array and create_map or map_from_arrays, and read them with dot paths and element access.

from pyspark.sql import SparkSession, functions as F

spark = SparkSession.builder.getOrCreate()

orders = spark.createDataFrame(
    [("o1", "c9", [("A1", 2, 9.5), ("B7", 1, 40.0)], {"channel": "web", "coupon": "FALL"}),
     ("o2", "c3", [("A1", 5, 9.5)], {"channel": "store"})],
    "order_id string, customer string, "
    "items array<struct<sku:string,qty:int,price:double>>, attrs map<string,string>")

orders.select(
    "order_id",
    F.col("items")[0]["sku"].alias("first_sku"),         # 0-based; raises under ANSI if out of range
    F.element_at("items", -1)["sku"].alias("last_sku"),  # 1-based, negative from end
    F.col("attrs")["coupon"].alias("coupon"),             # null when the key is absent
    F.size("items").alias("n_items"),
    F.struct("order_id", "customer").alias("key"),
).show()

Two indexing rules cause most bugs. Bracket access on arrays is 0-based while element_at is 1-based and accepts negative indexes from the end. And out-of-range behaviour depends on ANSI mode, which is on by default from Spark 4.0: under ANSI both brackets and element_at raise an error for an index past the end, where non-ANSI mode returns null. A missing map key still returns null. try_element_at always returns null instead of failing, which is what you usually want for optional positions.

Selecting items.sku on an array of structs returns an array of the field, ['A1', 'B7'], without exploding anything; that projection is often all a report needs.

The explode family, and when not to use it

explode turns each array element (or map entry) into its own row; posexplode adds the position; inline explodes an array of structs straight into columns; the _outer variants keep rows whose collection is null or empty, producing nulls instead of dropping the parent. That last point is a frequent silent data loss: an order with no items vanishes from an explode result.

lines = (orders
         .select("order_id", F.posexplode_outer("items").alias("pos", "item"))
         .select("order_id", "pos", "item.*"))       # sku, qty, price as columns

attrs = orders.select("order_id", F.explode("attrs").alias("k", "v"))  # maps give key, value

Explode is the right tool when each element genuinely becomes a row of the result: loading line items into a fact table, joining items to a product dimension, counting SKUs across orders. It is the wrong tool for computing something per parent row, because to get back to one row per order you must groupBy again, usually with a shuffle. With a million orders averaging 50 items, that is 50 million intermediate rows to compute a million totals.

Higher-order functions

Two ways to compute per-row totals over an array columnorders1 row, items: 3 structsexplode + groupByexplode(items)3 rowsshufflegroupBy order_idsum1 rowhigher-order functionaggregate(items, 0, ...)inside the row, no shuffletotal1 rowRule of thumbkeep work inside the rowwhen the answer is per rowexplode multiplies rows and often forces a shuffle to put them back together;array functions keep the computation local to each row.
Figure 3. Exploding and regrouping multiplies rows and shuffles; a higher-order function computes the same per-row answer in place.

Higher-order functions take a lambda and apply it inside each collection, entirely in the engine. They are available in SQL since Spark 2.4 and in the PySpark function API since 3.1.

enriched = orders.select(
    "order_id",
    F.aggregate("items", F.lit(0.0),
                lambda acc, x: acc + x["qty"] * x["price"]).alias("total"),
    F.filter("items", lambda x: x["qty"] >= 2).alias("bulk_items"),
    F.transform("items", lambda x: x.withField("line", x["qty"] * x["price"])).alias("items"),
    F.exists("items", lambda x: x["sku"].startswith("B")).alias("has_b"),
    F.forall("items", lambda x: x["price"] > 0).alias("all_priced"),
    F.map_filter("attrs", lambda k, v: k != "coupon").alias("public_attrs"),
    F.transform_values("attrs", lambda k, v: F.upper(v)).alias("attrs_uc"),
)

The equivalent SQL uses arrow lambdas: aggregate(items, 0D, (acc, x) -> acc + x.qty * x.price). Watch the accumulator type: with F.lit(0) the accumulator is an integer and adding a double fails analysis, so start from the type you want to finish with. Other useful members are zip_with for element-wise work on two arrays, array_sort with a comparator, transform_keys and map_zip_with. Set functions (array_distinct, array_union, array_intersect, array_except) and flatten and sequence cover most remaining needs without a UDF.

Changing fields inside structs

Before Spark 3.1, changing one field inside a struct meant rebuilding the entire struct by listing every field. Column.withField and Column.dropFields fix that, and they accept dotted paths for nested structs.

customers = customers.withColumn(
    "profile",
    F.col("profile")
     .withField("address.country", F.upper("profile.address.country"))
     .dropFields("legacy_id", "address.fax"))

For arrays of structs, combine with transform as shown above. Positional union matches nested struct fields by position, so two structs with the same fields in a different order can union into the wrong fields or fail. Prefer unionByName and test nested reordering on your Spark version.

Nullability is part of the schema at every level. A struct field, an array element (containsNull) and a map value (valueContainsNull) each carry their own nullable flag. Writing data whose nested nullability differs from a table's declared schema can fail or force a cast, and an array declared with nullable elements forces every lambda over it to handle null. Declare nullability deliberately when you define table schemas, rather than accepting whatever inference produced.

JSON strings and the VARIANT type

Nested data often arrives as JSON strings. from_json parses a string column with an explicit schema, to_json goes the other way, and schema_of_json infers a DDL schema from a sample literal, which is handy for writing the schema once rather than at runtime.

schema = "struct<user:struct<id:string,tier:string>,events:array<struct<t:long,kind:string>>>"
parsed = raw.select("payload", F.from_json("payload", schema, {"mode": "PERMISSIVE"}).alias("p"))
# Since 3.0, PERMISSIVE turns a malformed payload into a struct of nulls, not a null struct
bad = parsed.where(F.col("payload").isNotNull()
                   & F.col("p.user").isNull() & F.col("p.events").isNull())

Always pass an explicit schema in production. Schema inference on a sample is a time bomb: the first record with a new field or a number in a string field changes the inferred type and every downstream job. Use FAILFAST mode when a bad record should stop the job, and in PERMISSIVE mode count rows whose payload is present but whose parsed fields are all null, because since Spark 3.0 a malformed record yields a struct of null fields rather than a null struct.

Spark 4.0 adds a VARIANT type for semi-structured data whose shape is genuinely unknown or changing. parse_json produces a variant in a binary encoding that is faster to query than re-parsing strings, and variant_get(v, '$.user.id', 'string') extracts a typed value by path. Use variant for the long tail of fields you cannot schema up front, and promote stable, frequently queried fields into real struct columns so the optimizer can prune them.

Performance: pruning, pushdown and UDFs

Nested schema pruning. With Parquet and ORC, Spark reads only the leaf columns a query touches, even inside structs; spark.sql.optimizer.nestedSchemaPruning.enabled is on by default since 3.0. Selecting profile.address.country reads one leaf, not the whole profile. Pruning works through projections and many filters but is defeated by anything that needs the whole struct, such as to_json(profile), a Python UDF taking the struct, or a select('profile') early in the pipeline. Check the ReadSchema in the scan node of the physical plan; if it lists every field, pruning did not happen.

Predicate pushdown. Filters on nested struct fields can be pushed into Parquet and ORC scans, so row-group statistics skip data. Filters on array elements or map values cannot use those statistics, because there is no per-element min and max. If you filter on a map key in every query, that key belongs in a top-level or struct column. Parquet reading and writing covers the file-level side.

Avoid Python for nested logic. A Python UDF over an array of structs serialises each row to Python and back. Higher-order functions run in the JVM with code generation. If you truly need Python, pandas UDFs with Arrow handle nested types far better than row-at-a-time UDFs.

Huge collections. An array with millions of elements lives in one row, on one task, in memory. Very large per-row collections cause executor out-of-memory errors that no partitioning fixes; explode them or model them as rows from the start.

Worked example: sessions without a shuffle

Bring it together. A clickstream table has one row per session: session_id, user (a struct), events (an array of structs with timestamp, kind and a map of properties). Product wants, per session, the number of add-to-cart events, the time from first view to first purchase, and the session row with personally identifying fields removed.

s = sessions.select(
    "session_id",
    F.col("user").dropFields("email", "phone").alias("user"),
    F.size(F.filter("events", lambda e: e["kind"] == "add_to_cart")).alias("carts"),
    F.array_min(F.transform(F.filter("events", lambda e: e["kind"] == "view"),
                            lambda e: e["t"])).alias("first_view"),
    F.array_min(F.transform(F.filter("events", lambda e: e["kind"] == "purchase"),
                            lambda e: e["t"])).alias("first_buy"),
).withColumn("ms_to_buy", F.col("first_buy") - F.col("first_view"))

There is no explode and no shuffle; each session row produces one output row. The scan reads events.kind, events.t and the remaining user fields, but not the event property maps, which pruning skips. A session with no purchase has a null first_buy and therefore a null ms_to_buy, which is the correct answer rather than a dropped row.

Failure modes

  • Rows disappearing after explode because collections were null or empty. Use the _outer variants.
  • Off-by-one from mixing 0-based brackets with 1-based element_at.
  • Jobs failing after a Spark 4 upgrade because ANSI mode now raises on out-of-range access or invalid casts inside nested expressions. Use the try_ functions where nulls are acceptable.
  • Duplicate map keys. Building a map with repeated keys raises an error by default; spark.sql.mapKeyDedupPolicy set to LAST_WIN keeps the last value instead. Decide deliberately.
  • Schema drift in JSON silently turning fields null under PERMISSIVE parsing. Monitor the null rate of parsed columns.
  • Full-struct reads from an early select('*') or a UDF, visible as a wide ReadSchema.
  • Maps in comparisons. Maps are not orderable, and support for maps in grouping and set operations has varied by version; check your release before relying on it, or convert with map_entries and array_sort.

What to do next

  1. Print the schema of your three largest nested tables and mark which maps have a known key set; plan to turn those into structs.
  2. Search your jobs for explode followed by groupBy on the parent key, and rewrite one with aggregate or filter.
  3. Replace explode with explode_outer wherever parent rows must survive, and add a row-count check.
  4. Check ReadSchema in explain() for a hot query and remove whatever stops nested pruning.
  5. Give every from_json an explicit schema and alert on the rate of rows whose parsed fields are all null.
  6. Before moving to Spark 4, run your nested pipelines with ANSI on and switch to try_element_at where nulls are intended.
  7. Replace one row-at-a-time Python UDF over nested data with higher-order functions and compare runtimes.
Key takeaway: Spark's struct, array and map types let you keep nested records intact, and built-in functions process them without flattening. Prefer structs over maps when keys are known, use higher-order functions instead of explode-then-groupBy for per-row answers, use the outer explode variants when parents must survive, give JSON an explicit schema, and check ReadSchema to make sure nested pruning is reading only the fields you use.