NeMo-DCR: stop shipping whole checkpoints between RL steps
The paper, submitted October 6, 2026 by Aalto University and NVIDIA researchers, starts from a workflow problem: "Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout clusters before the next batch." When those clusters sit in different regions, that is slow: "Transferring a full 1T checkpoint ... takes 87.5 min between two AWS regions."
- Key observation. "Only 0.6–1.2% of training-side source elements change their bfloat16 (BF16) values per step," so NeMo-DCR sends compressed deltas instead of full weights, and "receivers obtain the same parameter and buffer bits as a dense refit."
- Results. At a 3% change rate it takes "22.6 s instead of 750 s at 120B and 150 s instead of 87.5 min at 1T." Even at 3% and 5% stress rates, refits of 30B-1T models are "12-40× faster than a transport-only full-checkpoint reference." The test bed paired 32 GB300 GPUs in one AWS region with 64 H100s in another.
- Remaining bottleneck. "The transport lower bound accounts for 77–94% of refit latency," so the network, not encoding, now dominates.
- Availability. "NeMo-DCR is about 7K lines of Python," built on Megatron Bridge and vLLM's loader, and is open source in NeMo RL; the implementation pull request was merged in July, before the paper.
Note the baselines are transport-only references, not full end-to-end runs. Analysis: for teams running RL with rollout capacity rented wherever GPUs are available, this makes cross-region setups practical -- and turns refit cost into a networking question.
Byteification: retrofitting tokenizer-free models
"Retrofitting language models to operate over bytes," published in Nature on October 7, 2026 (first author Benjamin Minixhofer, with authors from Ai2, Cambridge, the University of Washington, Imperial College London and LMU Munich), argues that subword tokenization hides detail that matters for code and biological sequences, but training byte-level models from scratch is expensive.
- Method. "We use a two-stage conversion procedure to retrofit existing subword-based models into byte-level models with minimal extra training," using "less than 1% of a typical pretraining budget (49.1B tokens in total."
- Models. Bolmo 7B and Bolmo 1B from Olmo 3 7B and OLMo 2 1B, Bwen 8B from Qwen3 8B Base, and Blama 8B from Llama 3 8B.
- Result. "Bolmo 7B achieved a +16.5% absolute improvement in STEM tasks over BLT 7B, which was trained from random initialization." Against their subword parents, the paper says the models "come close to matching" them, including in inference speed.
- Open. Training code is on GitHub (allenai/bolmo-core) and the paper is open access.
Analysis: byteification makes tokenizer-free models a conversion job rather than a pretraining project, which matters most for domains where tokenizers fail -- code, DNA, rare languages. The paper also notes the two-stage training "adds implementation complexity," and it reports results "close to" the parent models, not better across the board.
NeMo-DCR makes cross-region RL refits of trillion-parameter models take minutes instead of over an hour by sending bit-exact weight deltas, leaving the network as the bottleneck; byteification shows existing subword LLMs can be converted to byte-level models with under 1% of a pretraining budget, with code released.