Federated fine-tuning adapts a shared base model using data that stays in several places: hospitals, banks, regional subsidiaries, or user devices. Each participant trains locally and sends only an update; a server combines the updates into the next model. The basic protocol, its threat model, LoRA-factor averaging, secure aggregation and poisoning defences are covered in federated learning for LLMs.
This article picks up where that one stops, at the point where a basic FedAvg loop meets real silos. Their data is not identically distributed, their hardware and adapter ranks differ, and some of them are offline for a given round. The standard fixes for those problems, server optimizers, FedProx, SCAFFOLD, heterogeneous-rank LoRA and asynchronous aggregation, each change what is sent, what is stored and what an observer can learn. A team that picks an optimizer for accuracy can silently break the privacy guarantee it built the federation to provide. The goal here is to make those interactions explicit and choose a combination that trains well and keeps its guarantees.
Client drift: why FedAvg stalls on real silos
FedAvg runs K local optimizer steps on each client and averages the resulting models. With identically distributed data, local steps all move roughly toward the same optimum and averaging is close to running K steps of large-batch training. With non-IID data, which is the reason federation exists in the first place, each client's local optimum is different, and the more local steps it takes the further it walks toward its own optimum. The average of those endpoints is not the optimum of the average objective. This is client drift.
A one-dimensional example shows it. Client 1 minimises (w - 0)^2 and client 2 minimises 3 * (w - 4)^2. The joint objective, their sum, has its optimum at w = 3. Starting from w = 0 with step size 0.1 and many local steps, client 1 stays at 0 and client 2 converges to 4, so the average is 2, and every following round returns to 2. One local step per round would converge to 3, at the cost of many more rounds. The gap between 2 and 3 is drift, and in LLM fine-tuning it shows up as a global adapter that underperforms on every silo.
The knobs are the number of local steps, the local learning rate, and the corrections below. Fewer local steps cost communication rounds, and every round is another upload that must be clipped, noised and accounted for, so drift control is also a privacy-budget question.
Server-side optimizers
Adaptive federated optimization, from Reddi and colleagues, treats the averaged client change as a pseudo-gradient and feeds it to a server-side optimizer. With FedAdam, the server keeps first and second moment estimates m and v and updates the global weights from them, exactly as Adam does for a gradient. Server momentum smooths noisy rounds, which matters when only a fraction of silos report each time or when differential privacy noise is large.
def server_round(x, deltas, m, v, lr=1e-3, b1=0.9, b2=0.99, tau=1e-3):
# deltas: list of (x_i - x) from participating clients, already clipped
d = sum(deltas) / len(deltas) # with secure aggregation: sum arrives masked-off
m = b1 * m + (1 - b1) * d
v = b2 * v + (1 - b2) * d * d
x = x + lr * m / (v.sqrt() + tau)
return x, m, vThe security reading is mostly good news. The server optimizer consumes only the aggregate, so it is compatible with secure aggregation and with central differential privacy noise added to the sum; running Adam on a noised sum is post-processing and costs no extra privacy budget. The caveat is persistence. With momentum 0.9, a single poisoned round keeps influencing updates for roughly ten rounds, so robust aggregation must sit before the optimizer, not after it.
Client-side corrections: FedProx and SCAFFOLD
Two client-side corrections address drift directly. FedProx, from Li and colleagues, adds a proximal term to each client's local loss, F_i(w) + (mu/2) * ||w - x||^2, which pulls local training back toward the global model x. It changes nothing in what is uploaded, so it has no privacy cost; its price is one more hyperparameter and some loss of local progress.
SCAFFOLD, from Karimireddy and colleagues, is more powerful and more intrusive. The server keeps a global control variate c and each client keeps its own c_i, estimates of the global and local gradient directions. Each local step corrects the gradient by c - c_i, which cancels the drift term. After local training, the client updates c_i and uploads the change in c_i alongside its model change.
def scaffold_client(x, c, c_i, data, K, lr):
y = x.clone()
for _ in range(K):
g = grad(loss(y, next(data)))
y = y - lr * (g - c_i + c) # drift-corrected step
c_i_new = c_i - c + (x - y) / (K * lr) # option II update from the paper
return y - x, c_i_new - c_i, c_i_new # upload both deltas; keep c_i_new locallyThe second upload matters for security. The control-variate delta is a function of the client's private data, roughly its average gradient, and it is exactly the kind of signal gradient-inversion and membership inference attacks exploit. It must be clipped, included in the secure sum and covered by the privacy accounting, which doubles upload size and halves the noise budget per quantity if the two are noised separately; the alternative is to clip and noise the concatenated pair jointly. The persistent c_i on each client is also sensitive state: if a silo is compromised later, its c_i summarises its data across every round it joined.
What leaves each silo
The diagram is the checklist for any optimizer change. List every value that leaves a silo, and confirm each one is clipped, is inside the secure sum and is counted by the privacy accountant. Anything else that leaves, including metadata such as rank, sample counts or timing, is visible to the server in the clear.
LoRA choices that change the privacy story
Averaging the A and B factors of LoRA separately does not average the product B A, a problem the overview article covers. Two refinements matter here because of how they interact with privacy. FFA-LoRA, from Sun and colleagues at ICLR 2024, freezes the randomly initialised A on every client and trains only B. Because A is shared and fixed, averaging B averages the product exactly, uploads are half the size, and differential privacy noise is added to B alone instead of to two factors whose product would compound it.
Heterogeneous ranks arise when silos have different GPU budgets. Two families of fix exist. Zero-padding every client's factors to the maximum rank keeps every upload the same shape, so it fits secure aggregation, where masks cancel only on equal-length vectors. Stacking, as in FLoRA, concatenates each client's A and B so that the product equals the exact sum of client products. That is mathematically clean, but the server needs each client's factors individually, which is incompatible with secure aggregation by construction, and the concatenated rank grows with the number of clients. Under a threat model that does not trust the server, padding is the only option of the two.
For a sense of scale, LoRA at rank 16 on the query and value projections of a 32-layer model with hidden size 4096 has 32 * 2 * 16 * (4096 + 4096) = 8,388,608 trainable parameters, 33.5 MB in fp32. FFA-LoRA halves that to about 16.8 MB, and SCAFFOLD doubles whichever applies.
Partial participation and asynchrony
Waiting for the slowest silo every round wastes everyone's GPUs. FedBuff, from Nguyen and colleagues, lets clients train asynchronously; the server buffers updates and applies them once it has a fixed number, discounting stale ones. That fixes throughput and creates two security problems.
First, secure aggregation needs a cohort: a set of clients whose masks cancel together. Asynchronous arrivals have to be grouped into buffers that act as cohorts, and a small buffer weakens the protection, since the sum of three updates reveals more about each than the sum of thirty. Set a minimum buffer size and refuse to aggregate below it.
Second, differential privacy accounting for federated learning usually relies on privacy amplification by sampling, which assumes each client participates with a known probability chosen by the server. In cross-silo deployments, participation is driven by availability, which an adversary may influence, and the assumption fails. Without it, use the accounting for full participation, which spends far more budget per round, or have the server sample from available clients with a fixed probability and log the coin flips. Track the result with the methods in privacy budgets.
Worked example: eight banks, one complaint assistant
Eight regional banks want a shared assistant for classifying and drafting responses to customer complaints, on a 7B base model, without pooling complaint text. Each bank's complaints skew toward its own products, so data is strongly non-IID. Two banks have one GPU each; six have eight. The threat model assumes an honest-but-curious coordinator and allows for one malicious bank.
| Decision | Choice | Reason |
|---|---|---|
| Adapter | FFA-LoRA, rank 16 on q and v | Exact aggregation, 16.8 MB uploads, noise on B only |
| Heterogeneous hardware | Same rank everywhere; small banks run fewer local steps | Keeps uploads equal-shape for secure aggregation |
| Drift | FedProx plus FedAdam on the server | No extra uploads, unlike SCAFFOLD |
| Privacy | Record-level DP-SGD inside each bank, plus secure aggregation | Protects each customer; the coordinator sees only the sum |
| Participation | Synchronous rounds, all eight banks, with a timeout | Full-participation accounting is honest for eight silos |
| Robustness | Norm bound plus a held-out probe set per round | Secure sum hides individual updates from inspection |
The privacy row reflects a choice of unit. Client-level differential privacy, noise on the aggregate that hides any one bank's whole contribution, needs hundreds or thousands of participants before the noise stops swamping the signal; with eight silos and millions of adapter coordinates it would destroy the model. What the banks actually need to protect is each customer's complaint, so each bank runs DP-SGD locally with its own record-level budget, and secure aggregation keeps individual bank updates confidential from the coordinator.
SCAFFOLD would likely give the best accuracy on data this skewed, and the team considered it. They rejected it because it doubles the uploads that must be noised and leaves a persistent per-bank summary of complaint gradients on each bank's systems. Instead they tuned the FedProx coefficient on a public complaints dataset split by product to imitate the skew, and measured accuracy per bank, not just on the pooled average, because a global number can hide a bank whose products the federated model serves badly. A bank that benefits little can fine-tune a small private adapter on top of the shared one, locally and outside the federation.
Failure modes
- Drift mistaken for poor data. A global model worse than every local one usually means too many local steps, not bad silos. Reduce K before blaming participants.
- Side-channel uploads. Control variates, sample counts used for weighting, per-client loss values and ranks all leak if sent outside the secure sum. Weight by a clipped, rounded count or not at all.
- Stacked heterogeneous LoRA behind a privacy claim. The server sees every factor, so any secure-aggregation claim is false.
- Assumed sampling amplification. Reporting an epsilon that assumes random client sampling when participation was availability-driven understates the real privacy loss.
- Small async buffers. Aggregating two or three updates reveals nearly as much as individual ones.
- Momentum carrying poison. Robust aggregation after the server optimizer is too late; a bad round lingers in the momentum for many rounds. Pair with the ingestion controls in data poisoning.
- Global-only evaluation. Average accuracy improves while a minority silo regresses. Report per-silo metrics every round.
Trade-offs
| Technique | Accuracy on non-IID data | Extra uploads | Compatible with secure sum and DP |
|---|---|---|---|
| FedAvg | Baseline; drifts | None | Yes |
| FedAdam or other server optimizer | Better, smoother | None | Yes, post-processing |
| FedProx | Better | None | Yes |
| SCAFFOLD | Often best | Doubles | Yes, if both deltas are clipped, summed and noised |
| FFA-LoRA | Close to LoRA | Halves | Yes, and DP noise is better behaved |
| Padded heterogeneous ranks | Good | Up to max rank | Yes |
| Stacked heterogeneous ranks | Exact sum | Grows with clients | No |
| FedBuff asynchronous | Faster wall-clock | None | Only with minimum buffer size and honest accounting |
What to do next
- Write down every value that leaves a silo, and check each is clipped, inside the secure sum and accounted for.
- Measure per-silo accuracy against a locally fine-tuned baseline to detect drift before tuning anything else.
- Start with FedProx and a server optimizer; add SCAFFOLD only if accuracy demands it and you can noise both uploads.
- Prefer FFA-LoRA or another exact-aggregation adapter scheme when differential privacy is required.
- Keep adapter shapes identical across silos; pad ranks rather than stack them.
- Base the privacy accountant on how participation really happens, and set a minimum cohort size.
- Place robust aggregation before the server optimizer, and review the full protocol in secure aggregation before launch.