Miles v0.1 deep dive - Fully Async RL

0. Intro

Goal: verified, clean, and customizable

Key features

  • Rollout engine: SGLang
  • Trainer: Megatron or Pytorch FSDP
  • 3 weight-sync methods
  • Full-param RL, LoRA RL support
  • OPD, SFT support
  • True-on-policy rollout-training alignment

Changellges in RL

  • multi-turn rollouts
  • tool use; complicated external env
  • O(T) param MoE models
  • → needs a verified, clean, and customizable RL framework

Limitations of Miles

  • need better supports for some precision formats
  • some weight-transfer paths only cover only certain models
  • some benchmarks are for specific configs / setup

1. The Miles RL Loop

image.png

  • Rollout: generate trajectories (seq of state and actions) and produces rewards
  • Training: consume trajactories groups, compute loss, and update the policy (model weights)
  • Weight update: sync new weights with the Rollout engine

Fully async mode: Rollout continues generates while Trainer is working

2. Rollout

What matters during Rollout?

  • Throughput
  • Fidelity: the trainer’s view of a trajectory can silently diverge from what the policy actually sampled

Key terms

  • Prompt: one task drawn from the dataset
  • Trajectory: one attempt at that task by the current policy (can include tool calls, and env replies)
  • Group: trajectories from the same prompt
  • Session: what the serving layer sees, namely the ordered set of requests belonging to a single episode
    • may produce one or more trajectories
    • multiple env. each has a trajectory

2.1 Multi-turn request routing

Request routing larges determines the throughput.

For multi-turn RL, route the request to the engine of the previous turns (to reuse KV cache)

Two-level affinity

Implementation

  • key-based routing: attaches a stable routing key to every request in a session, and maps to the same engine / DP rank

Load balancing: A new session goes to the engine with least number of active requests

2.2 Fully async RL

  • Fully async: rollout and training can progress concurrently
    • The trainer consumes whichever trajectories when it needs a batch
    • Rollouts runs continiously
  • The rollout engine and the trainer are disaggregated on different GPUs
    • Miles refuses fully async mode when they are colocated
  • Pros: don’t have to wait for long and tail trajectories

2.2.1 The replacement rule: keep the Rollout busy

  • We want to keep the Rollout Engines busy
  • Replacement rule: a background worker checks the current rollouts and decide when to fill new trajectories. This “when to fill” is the replacement rule.
  • We can choose one of the two rules:
    • Group granularity: waits until every trajectory in a group finishes before starting a replacement
    • Sample granularity: fill a new trajectory as soon as a current one finishes. This is the default.

2.2.2 Data Buffer

  • Finished rollout groups are placed in the data buffer, and the trainer cosumes them
  • Role of Data buffer: decouple rollout engines from the trainer, absorbing diff in their rates
  • supported ops (only 3): put a trajectory group in, take a batch out, and report metrics
    • easy to customize (e.g., change the data selector when taking a batch out)
  • 3 rules to drop a group:
    • The first two are the properties of groups, so they can be checked when arriving
    • The 3rd (staleness) depends on how long a group waits, so it’s checked when it’s collected for the trainer

image.png

  • Data Buffer has a bounded capacity (a multiple of the training batch size). It blocks adding groups until it has free capacity.

2.2.3 Observability

  • In fully async RL, rollout and training can be at different rates.
  • So we need to monitor them and avoid the rates drifting too much.
  • Miles report Buffer metric on every training step:

image.png

  • Queue size is the quickest signal for identifying which stage limits progress → can be a singal for scaling rollout / training

2.2.4 Async Eval

  • Eval can compete with trajectory generation
  • Sync RL: the rollout engine is idle when training is in progress → eval using the rollout engine is free
  • Async RL: eval displaces generation
  • Async eval modes:

image.png

  • Snapshot-based modes keep evaluation off the critical path after exporting the snapshot. But exporting a fresh snapshot imposes a pause → it’s a collective op on the trainer
  • Reusing a periodically saved checkpoint can avoid such pause.

image.png

2.3 Async Env

  • Agentic RL accommodates multiple isolated envs of different scopes.
  • Each trajectory comes from an env
  • For example, a coding-agent env has a sandbox to run cmds, edit files, and get test results
  • Different scopes
    • Some envs only own the episode loop
    • Other envs controls batching, rewards, and token recording
  • Miles organizes env integration as three nested plug-in layers:

image.png

  • An env reaches Miles through a plug-in point.
  • An env can take over as much or as little of the rollout as it needs.
  • Supported envs (experimental)
    • Agent function: Harbor (terminal), NeMo Gym (a lib for some simple enviroments), and OpenEnv (Huggingface lib of envs)
    • Generate function: HUD, Strands Agents, and $\tau$-bench
    • Rollout function: Prime Intellect Verifiers. It brings its own taskset, groups the episodes, and computes per-rollout and group rewards before returning completed traces
  • The sandboxes run on any backend of your choosing. Miles (experimentally) supports AgentENV, Daytona, E2B, and Modal
  • nixos is another popularEnv
  • CPU usage for containers: Env is often encapsuled in a container. Let’s say we have 64 prompts, each with 8 trajectories (under GRPO), each is 4GB CPU mem for container, then total is 2TB → higher demand for memory

2.4 Token-In-Token-Out (TITO)

  • In multi-turn RL, the trainer may see modified token sequences that are not exactly the same as what the rollout engine produces.
  • Reason
    • there can be multiple states of each turn: e.g., message parsing, tool execution, and chat-template rendering
    • each stage may change tokenization, prune historical reasoning, or reserialize tool calls
  • The TITO session server adresses this by letting the server, rather than the harness, control
    tokenization, preserving the exact token IDs generated by the model.
  • https://github.com/THUDM/slime/issues/2285

2.5 Rollout Routing Replay (R3)

  • In MoE layers, the rollout engine and the trainer may choose different experts due to numeric diff.
  • R3 treats the expert choices as part of the rollout data
  • Recorded assignments are cheap to replay but expensive to carry
    • Each routing tensor holds (tokens − 1) × layers × k 32-bit integers.
    • For a 32K-token sequence over 60 layers at k = 8, that tensor occupies roughly 60 MB per trajectory
    • For long-context RL, this can be severe. Section 3.2 has mitigations.
  • Limitation: in async RL, many other factors contributes to train-rollout mismatch, so enabling R3 could have limited effects.

3. Training

3.1 Low-precision training

  • Low-recision training can lead to training-rollout mismatch.
  • Solution: training and rollouts must quantize the same weights in the same way
  • Miles supports the following recipes, which are verified with low KL divergence between SGLang and Megatron-LM.

image.png

3.2 Memory Efficiency & Disk Offload

Two offloading mechanisms that act at different time

  • Offloading the paused actor entirely (when actor is idle)
  • Streaming the optimizer state (during training)
  • These two can work together

3.2.1 Offloading the Paused Actor

  • In a colocated run, the default is to offload the trainer’s weights, gradient buffers, and optimizer state during rollout
  • Megatron backend transfers the actor’s state at the allocator level rather than tensor by tensor
  • Offload to local DRAM if possible, otherwise to local disk
  • In disagg run, offloading is off by default.
  • PPO is a special case: it needs an actor and a critic, which are always colocated. Offloading is on by default.

3.2.2 Streaming the Optimizer State

  • Optimizer states dominates the HBM usage, including FP32 master weights, AdamW momentums an variances, and optional FP32 master gradients
  • Streaming: split params into buckets, one file for each bucket’s optmizer states
  • Streaming saves HBM but make the training slower. We could we FP16 states to speed up the transfer.

Skip 3.3 (two training backends Megatron and FSDP)

3.4 The Objective and Rollout–Training Corrections

The Objective

  • Miles defines a training objective by
    • an advantage estimator
    • a typed loss interface
  • Five advantage estimators:
    • GRPO and GSPO
    • REINFORCE++ [26] in plain and baseline-relative forms
    • PPOwith a learned value function
  • The typed loss interface
    • provides one protocol with policy, value, and supervised variants, plus a hook for a user-supplied one

Rollout–Training Corrections

  • Training and rollout use differnet engines, so they have numeric differences.
  • importance ratio $r = \frac{\pi_\text{train}}{\pi_\text{rollout}}$
  • Miles offers two corrections that act on the importance ratio differently
    • Truncated importance sampling (TIS): clamps the ratio to a configured interval. The default interval is [0, 2]
    • Clip-or-pop: sets the per-token weight to zero for any token whose ratio falls outside the same interval (can drop outlier tokens)

4. Weight update

  • When training and rollout occupy different GPUs, the weight transfer can become a major pipeline bottleneck, and at frontier scale it can dominate the step
    • a full NCCL broadcast of Kimi K2 1T-A32B takes almost a minute
  • How does a prepared bucket of weights reach the rollout engines? → Miles provides three weight transports: NCCL broadcast, P2P, and disk-delta

image.png

  • Broadcast is the default. P2P is useful only when the trainer fleet and the rollout fleet each span more than one node.

4.1 The Shared Bucketed Pipeline

  • NCCL Broadcast and P2P use one preparation pipeline: the pipeline prepares each bucket once, then hands it to the selected transport.
  • Megatron owns the preparation pipeline and both transports.

Megatron preparation pipeline

  • AG weights: for each pipeline stage, Miles all-gathers Mrgatron’s TP shards
  • weight conversion: convert weights to Hugging Face names & layout that SGlang expects
  • buffering: append the weights to a 512MB buffer
  • flush: Miles hands the whobucket to the selected transport (broadcast or P2P)
  • transfer: either one batch of async NCCL broadcast, or RDMA P2P to rollout-rank mem
  • expert pass: (optional) a second pass of expert router decisions

4.2 P2P tansfer

  • With P2P weight transfer, training ranks re-shard each weight bucket for the target SGLang layout and write only the required shards directly into rollout-rank memory over RDMA.
  • P2P treats each weight update as direct writes from training ranks (sources) to rollout ranks (targets)
  • Transfer plan: assign each training rank to its destination rollout ranks
  • CPU-resident model replica: Miles builds SGLang model in CPU memory, so we can call it and transfer the weights to the expected sharding. This way, the sender doesn’t have to re-implement SGLang’s sharding rules
  • Single shared pinned buffer: Miles register the buffer for RDMA once, and every rollout engine can resue it.

4.3 Disk-Delta Updates

  • Consecutive RL steps change only a small fraction of the model’s bytes, so sending just the changed bytes costs far less than sending all of them.
  • With disk-delta updates, every rollout host starts from a shared base checkpoint
  • At each update the trainer writes the changed bytes to a shared filesystem, together with a reference to the base they apply to, and each rollout host reads them and patches its own local copy of the checkpoint.

4.4 Pausing Generation, and Checking the Result

  • Rollout engines apply and verify the delta-weights while generation continues, then pause only to load the combined weights into SGLang.
  • What happens to requests in flight when the weights change? 3 options:
    • abort a request outright
    • leave it in place while applying the update around it (may apply TIS)
    • or retract it so that it rolls back and resumes against the new weights. retraction is the default.
  • How does a run confirm that the engines received the trainer’s weights?
    • Miles ships an opt-in check that confirms the rollout engines hold the weights the trainer sent. The check runs once, at the start of training, and compares each engine’s weights against the trainer’s own.
    • Miles first fills the engine’s tensors with random values and only then runs the first weight update, so any tensor the transport failed to write still holds that noise when the check reaches

Skip the rest of sections


Miles v0.1 deep dive - Fully Async RL
https://gdymind.github.io/2026/09/27/Miles-deep-dive-fully-async-RL/
Author
gdymind
Posted on
September 27, 2026
Licensed under