RL Is Bottlenecked by Inference. Scale It Independently

(skypilot.ai)

10 points | by alex000kim 5 hours ago ago

3 comments

  • 5 hours ago
    [deleted]
  • efiop 5 hours ago

    what did gpu hours look like here? with 3 replicas for a 1.8x speedup, the cost tradeoff isn’t obvious.

    • efiop 4 hours ago

      ah, nevermind. 3 engines seem cheaper overall too: 7x661s vs 5x1200s of allocated H100 time per step. Nice.