Experiments run on Lambda Cloud 8×H100 instances using SkyRL.

Key Takeaways

Model Colocated Disagg Async Split (Train, Generate) power utilization wall clock time reduction
Qwen3-14B (Thinking) 34.6% 51.4% 4+4 +16.8 % -16.1%

Why Agent RL Training Underutilizes GPUs

Multi-Turn Trajectories and Tool Latency

In single-turn RL (e.g., math, code generation), each trajectory consists of one generation call per sample. In agentic RL, each trajectory spans multiple turns of interleaved LLM inference and tool execution — the model generates an action, calls an external tool (web search, browser, code interpreter), waits for the result, and repeats. For detailed trace examples, see SkyRL-Agent and this architecture walkthrough.

The Straggler Effect

In single-turn RL, trajectory completion times within a batch are relatively uniform. In agentic RL, two sources of variance compound:

The figure below illustrates this for a single batch (Qwen3-14B, batch size 24). Each bar represents one trajectory, sorted by completion time. Trajectories with more tool calls (red) take longer; In synchronous colocated mode, training cannot begin until the last trajectory completes — the gap between the median and the straggler represents GPU time spent waiting.

turn_distribution_straggler.png

GPU Utilization Breakdown

Colocated RL training alternates between generation and training within each step. The time breakdown from the Qwen3-14B colocated baseline (8×H100, averaged over 8 steps):

Phase Avg Time % of Step
Generate (inference + tools) ~385s ~85%
Forward (logprobs + values) ~24s ~5%
Train (policy update) ~33s ~7%
Weight sync + other ~9s ~2%
Total step ~451s 100%

Compute-intensive phases (forward + train) account for ~12% of wall clock time. The remaining ~85% is generation, during which GPUs alternate between short inference bursts and idle waits for tool execution. Generation time is determined by the slowest trajectory in the batch, not the average, further reducing effective utilization.