Task decomposition means turning one request into several smaller ones whose answers combine into the result. Almost every serious language model system does it somewhere: a pipeline of prompts, a planner handing work to tools, an agent writing itself a to-do list. The technique is easy to apply and hard to apply well, because every cut you make is a bet. It buys focus, checkability and parallelism, and it costs calls, latency and new places for information to get lost.
This page is about choosing the decomposition: where to cut, who draws the plan, how to represent and validate it, how to execute and repair it, and how to tell whether it helped. How to wire fixed sequences of prompts, with contracts between steps, is covered in Prompt chaining. Least-to-most and decomposed prompting, which order subproblems from easy to hard, are covered in Least-to-most prompting.
Why decomposing helps, and what it costs
A single prompt asked to retrieve data, compute figures, interpret them and write a memo fails in characteristic ways. It recalls numbers instead of querying them, it drops one of six requirements, and when the result is wrong you cannot tell which part went wrong. Splitting fixes three of these at once. Each step gets only the context it needs, so attention is not spread over irrelevant material. Each step has an output small enough to check. And steps that need different capabilities, a SQL engine, a search index, a model, can each use the right one.
The costs are just as concrete. Every step is a call with its own latency and tokens. Every boundary is a summary that can lose detail the next step needed. Errors compound across steps. Whether the gains outweigh those losses is an empirical question to measure, not assume.
Where to cut: five seam tests
Good boundaries are not arbitrary. A candidate cut is worth making when it passes at least one of these tests, and better when it passes several.
| Test | Question | Example |
|---|---|---|
| Checkable output | Can a program or reviewer verify this piece alone? | A SQL result with known row count and value ranges |
| Different capability | Does this piece need a tool, model or permission the rest does not? | Querying the warehouse versus writing prose |
| Independence | Can it run without the others' outputs? | Fetching survey data while computing churn |
| Context isolation | Does it need a large input the rest does not? | Reading a 40-page policy to extract three dates |
| Failure isolation | Would a failure here be cheaper to retry alone? | A flaky external search call |
The anti-seam matters as much. Do not cut through tightly coupled reasoning, where each part constantly needs the other's intermediate state; splitting a proof or a tricky refactor forces that state across boundaries, and the next model reconstructs it badly. If you cannot write down what passes between two steps in a few fields, they probably belong together.
Static plans versus model-generated plans
Decompositions sit on a spectrum. At one end, a designer writes the steps once and the code runs them for every input: fast, testable and identical every time, ideal when the task shape is stable. At the other, the model writes a new plan per request, which handles varied tasks at the cost of a planning call and a plan that can itself be wrong. Between them sit templates: a fixed skeleton whose branches the model selects or fills in.
The research literature explores the model-generated end. Plan-and-Solve prompting asks the model to devise a plan and then carry it out, in one response. Self-Ask has the model pose and answer explicit follow-up questions. ReWOO separates a planner, which writes the whole tool plan up front with placeholders for results, from workers that fill them in and a solver that combines them, so the planner does not reread every observation. LLMCompiler goes further and has the planner emit a dependency graph of function calls, so independent calls run in parallel. The shared idea worth keeping is that the plan becomes an explicit object, not a paragraph of reasoning.
The plan as data
Ask the planner for structured output with a fixed shape: an id, a one-sentence goal, the tool, the ids it needs and a completion check. Each node declares its dependencies, so the plan is a directed acyclic graph rather than a list, and independence is visible to the scheduler instead of buried in prose.
You are planning, not answering. Break the task into steps.
Return JSON only, matching this shape:
{"nodes": [{"id": "n1", "goal": "...", "tool": "sql|search_docs|llm",
"needs": ["ids of nodes whose output this step reads"],
"done_when": "a check a program or reviewer can apply to the output"}]}
Rules:
- Each goal is one sentence with one verb and one output.
- A step needs another only if it reads that step's output. Independent steps must not
depend on each other, so they can run in parallel.
- Use sql for numbers that exist in the warehouse; never ask llm to recall a number.
- At most 10 nodes. If the task fits in one step, return one node.
- The last node produces the final answer and needs every node it summarises.
Task: {task}
Available tables: {schema_summary}Two rules in that prompt carry most of the value. Routing numbers to SQL stops the model from recalling figures it should compute. And the done_when field forces the planner to say what success looks like for each step, which you will need for checking anyway.
Validate before anything runs
A generated plan is untrusted input. Validate it with code before spending a single worker call: known tools only, no missing dependencies, no cycles, a bounded node count and exactly one final node. Return the error list to the planner and ask for a corrected plan; most malformed plans are fixed in one round, and if two rounds fail, fall back to a static plan or to a single prompt.
from graphlib import TopologicalSorter, CycleError
TOOLS = {"sql", "search_docs", "llm"}
def validate(plan, max_nodes=10):
nodes = {n["id"]: n for n in plan.get("nodes", [])}
errors = []
if not 1 <= len(nodes) <= max_nodes:
errors.append(f"plan has {len(nodes)} nodes, allowed 1..{max_nodes}")
for n in nodes.values():
if n.get("tool") not in TOOLS:
errors.append(f"{n['id']}: unknown tool {n.get('tool')!r}")
if not n.get("done_when"):
errors.append(f"{n['id']}: missing done_when")
for dep in n.get("needs", []):
if dep not in nodes:
errors.append(f"{n['id']}: needs unknown node {dep}")
try:
graph = {k: set(v.get("needs", [])) & nodes.keys() for k, v in nodes.items()}
list(TopologicalSorter(graph).static_order())
except CycleError as e:
errors.append(f"dependency cycle: {e.args[1]}")
used = {d for n in nodes.values() for d in n.get("needs", [])}
sinks = [k for k in nodes if k not in used]
if len(sinks) != 1:
errors.append(f"expected one final node, found {sinks}")
return errorsPython's standard graphlib does the cycle detection. The single-sink rule catches a common planner error: steps whose output nothing consumes, which means work done and then ignored.
Execute in parallel, check every node
The executor walks the graph, runs every ready node at once, checks each output against its done_when and passes each node only the outputs it declared. That last point is deliberate: a node that silently reads everything finished so far defeats context isolation and hides missing dependencies.
import asyncio
from graphlib import TopologicalSorter
class NodeFailed(Exception):
pass
async def run_node(node, results, workers, check, attempts=2):
inputs = {d: results[d] for d in node["needs"]} # only declared dependencies
for _ in range(attempts):
out = await workers[node["tool"]](node["goal"], inputs)
ok, reason = await check(node["done_when"], out)
if ok:
return out
raise NodeFailed(f"{node['id']}: {reason}")
async def run_plan(plan, workers, check, results=None):
results = dict(results or {}) # finished nodes survive a replan
spec = {n["id"]: n for n in plan["nodes"]}
ts = TopologicalSorter({k: set(v["needs"]) for k, v in spec.items()})
ts.prepare()
while ts.is_active():
ready = [i for i in ts.get_ready()]
todo = [i for i in ready if i not in results]
outs = await asyncio.gather(*(run_node(spec[i], results, workers, check) for i in todo))
results.update(zip(todo, outs))
for i in ready:
ts.done(i)
return resultsThis version waits for a whole wave before releasing the next, which is simple and slightly slower than releasing each node the moment its own dependencies finish. The check can be code, a schema, a row count or a small judge model; cheap deterministic checks first. A node that fails twice raises, and the caller replans with the error and the finished results.
Worked example: a churn memo
Request: compare Q3 churn by segment with the pricing change and write a memo for leadership. A single prompt tends to run one query, guess the rest and write confident prose. The planner produced this:
{"nodes": [
{"id": "n1", "goal": "Compute monthly churn by segment for Q2 and Q3", "tool": "sql",
"needs": [], "done_when": "one row per segment-month, 6 months, churn between 0 and 1"},
{"id": "n2", "goal": "Retrieve the pricing change announcement and its effective date", "tool": "search_docs",
"needs": [], "done_when": "contains an effective date and the list of changed plans"},
{"id": "n3", "goal": "Compute churn by segment for the 60 days before and after the effective date", "tool": "sql",
"needs": ["n2"], "done_when": "two numbers per segment with customer counts"},
{"id": "n4", "goal": "Retrieve exit-survey reasons for churned customers in Q3", "tool": "sql",
"needs": [], "done_when": "reason counts per segment"},
{"id": "n5", "goal": "Identify segments where churn rose after the change and the top stated reasons", "tool": "llm",
"needs": ["n1", "n3", "n4"], "done_when": "every claim cites a number from n1, n3 or n4"},
{"id": "n6", "goal": "Write a one-page memo with findings, caveats and a recommendation", "tool": "llm",
"needs": ["n2", "n5"], "done_when": "under 500 words, includes a caveats section"}
]}Nodes n1, n2 and n4 have no dependencies and run together in the first wave. n3 waits for the effective date from n2. n5 interprets the numbers, with a check that every claim cites one, and n6 writes the memo. When n3 failed its check because the sales calendar used a different date format, only n3 was retried; the three finished nodes were kept. The plan also made a gap visible during review: nothing checks whether churn was already rising before the change, so a reviewer added a node comparing the same window a year earlier.
Granularity
Too coarse and a step is the original problem again. Too fine and you pay a call per trivial operation and multiply the boundaries where context is lost. A workable heuristic: one goal sentence, one verb, one output that a check can judge, and an output that a downstream step actually consumes. Merge adjacent steps that use the same tool and context, and split a step whose check keeps failing for different reasons, a sign that it hides two jobs.
Put a hard cap on node count. Plans that grow past it are usually decomposing the explanation rather than the work.
Replanning without thrashing
Retry first, replan second. A node that fails repeatedly suggests the plan is wrong: a missing input, a wrong tool, an impossible goal. Then send the planner the original task, the current plan, the finished results and the failure, and ask for a revised plan that reuses finished nodes by id. Cap replans at one or two per request, log the plan diffs and alert when the replan rate climbs; a planner that replans often is telling you the plan space needs a static template. Longer-running agents that replan as they act are covered in Agentic prompting.
Evaluating the decomposition itself
Score the final answers, but also score the plans, because a plan can be wrong in ways the final answer hides. Useful plan-level metrics: the share of plans that validate first time, node count distribution, nodes whose output nothing uses, requirements from the task that no node covers, and the replan rate. Build a set of fifty or so representative tasks with reference answers.
Then run the ablation that justifies the whole exercise: the same tasks through a single well-written prompt, through a static plan and through generated plans, comparing quality, cost and latency. Decomposition often wins on tasks mixing retrieval, computation and writing, and often loses on short reasoning tasks a strong model handles in one pass, especially with native reasoning, as discussed in Chain-of-thought prompting.
Failure modes
- Plan theatre: nodes that restate the task in smaller words without changing who does the work or what is checked.
- Lost constraints: a requirement in the original task appears in no node goal and silently disappears.
- Recall instead of retrieval: numbers produced by an llm node rather than a sql node.
- Hidden dependencies: a node that needs another's output but did not declare it, run in parallel on missing data.
- Synthesis overload: a final node fed every intermediate output, rebuilding the context problem decomposition was meant to solve.
- Replan loops: a planner that keeps proposing the same failing step; cap replans and fall back.
Trade-offs and when not to decompose
Do not decompose when a single prompt already meets the quality bar, when latency matters more than marginal accuracy, or when the work is tightly coupled reasoning. Prefer a static plan when inputs share a shape, because it is cheaper, testable and debuggable. Use generated plans when tasks vary and the plan can be validated. In every case the plan should make work checkable; if it does not, it is only adding calls.
What to do next
- Pick one multi-part task you run today and list its parts against the five seam tests.
- Write the plan as JSON by hand first, with done_when checks, and execute it as a static plan.
- Measure it against a single-prompt baseline on fifty tasks: quality, tokens and latency.
- Only if inputs vary enough, add a planner with the schema above and the validator in front of execution.
- Pass each node only its declared inputs, and make the cheapest deterministic checks run first.
- Cap node count and replans, and log every plan and plan diff.
- Track plan validity rate, unused nodes and uncovered requirements alongside answer quality.