Experiments run on Lambda Cloud 8×H100 instances using SkyRL.
| Model | Colocated | Disagg Async | Split (Train, Generate) | power utilization | wall clock time reduction |
|---|---|---|---|---|---|
| Qwen3-14B (Thinking) | 34.6% | 51.4% | 4+4 | +16.8 % | -16.1% |
In single-turn RL (e.g., math, code generation), each trajectory consists of one generation call per sample. In agentic RL, each trajectory spans multiple turns of interleaved LLM inference and tool execution — the model generates an action, calls an external tool (web search, browser, code interpreter), waits for the result, and repeats. For detailed trace examples, see SkyRL-Agent and this architecture walkthrough.
In single-turn RL, trajectory completion times within a batch are relatively uniform. In agentic RL, two sources of variance compound:
The figure below illustrates this for a single batch (Qwen3-14B, batch size 24). Each bar represents one trajectory, sorted by completion time. Trajectories with more tool calls (red) take longer; In synchronous colocated mode, training cannot begin until the last trajectory completes — the gap between the median and the straggler represents GPU time spent waiting.

Colocated RL training alternates between generation and training within each step. The time breakdown from the Qwen3-14B colocated baseline (8×H100, averaged over 8 steps):
| Phase | Avg Time | % of Step |
|---|---|---|
| Generate (inference + tools) | ~385s | ~85% |
| Forward (logprobs + values) | ~24s | ~5% |
| Train (policy update) | ~33s | ~7% |
| Weight sync + other | ~9s | ~2% |
| Total step | ~451s | 100% |
Compute-intensive phases (forward + train) account for ~12% of wall clock time. The remaining ~85% is generation, during which GPUs alternate between short inference bursts and idle waits for tool execution. Generation time is determined by the slowest trajectory in the batch, not the average, further reducing effective utilization.