Rust SIMD on the GPU

(vectorware.com)

52 points | by sagacity 3 hours ago ago

20 comments

  • O3marchnative 9 minutes ago

    The author mentions Rust's portable SIMD library [0]. The only issue with portable SIMD is it's only available on nightly. I used it in my FFT crate, but we had to switch to the fearless_simd crate in order to get a portable SIMD solution that works on stable [1].

    [0] https://doc.rust-lang.org/std/simd/index.html [1] https://github.com/linebender/fearless_simd

  • 6r17 an hour ago

    My heard hurts - i was stupid enough to think that SIMD was a CPU only thing - I don't understand why it would be ported to GPU - huge kudos to managing to surprise me

    • chlorion 24 minutes ago

      GPUs work on vectors and matrices very often, that's what they are good at, so it makes a lot of sense that they can operate with SIMD I think!

    • hingler36 38 minutes ago

      Welcome to the lucky 10,000! SIMD is actually a pretty integral part of how GPUs are able to work efficiently, it's part of why there's such a strong focus on branchless programming in the field.

  • nynx 7 minutes ago

    Do you have examples of complex algorithms running on the gpu with rust with competative performance? Radix sort might be a good one to start with

  • LegNeato an hour ago

    Author here, AMA.

    • lbhdc 39 minutes ago

      What is vectorware's business model? Are you planning to sell support/consulting to companies using your stack? Or are you looking to sell licenses to your tool? Or something else?

      • LegNeato 18 minutes ago

        The tentative plan is to open source all the compiler and `std` bits with our products built on top (compilers are not good businesses). More about our products coming in the next couple of months!

    • jcranmer an hour ago

      The post is kind of vague on the IR you're targeting. Can you give some examples of what the SIMD-ized IR looks like, and how it maps to the target PTX?

      • the__alchemist an hour ago

        I'm confused too. How does this fit between these approaches for paraellization:

          - CUDA kernels and Tiles (e.g. Cudarc, cuda-oxide, rust-gpu etc) - SIMD on the GPU. (E.g. as in the title...)
          - CPU SIMD using avx or SSE instructions (And probably thin wrappers for vectors so you can have sane syntax). Or the maybe-upcoming core simd which should abstract over architecture-specific instructions. Magic floats etc which do 4-16 computations at once, but are a bit clumsy to work with
          - Rayon thread pools - arbitrary parallel computations, including SIMD, one per CPU core.
        
        It looks like from the code samples like maybe a cleaner syntax for writing code on the GPU than CUDA kernels? E.g. without mucking with serialization, host and device by abstracting over it? And inspired by core::simd. (Good choice if so, in the interest of standardizing on syntax; I did this for my x86 SIMD vector/quaternion lib as well)
      • LegNeato 20 minutes ago

        Didn't want to go into crazy detail in the post.

        Each family of operations is a trait parameterized by the operation itself:

          pub trait EvaluateReduction<Operation, T>: LaneEvaluator {
              /// Reduce one distributed definition to an ordinary uniform scalar.
              fn evaluate_reduction(&self, value: LaneValue<Self, role::Distributed, T>) -> T;
          }
        
        
        Call sites name the operation:

          let one   = evaluator.splat::<Splat, _>(1_u32);
          let two   = evaluator.splat::<Splat, _>(2_u32);
          let three = evaluator.binary::<Add, _>(one, two);
        
          let total   = evaluator.reduce::<Sum, u32>(three);   // a uniform u32
          let running = <Executor as EvaluateScan<Scan<Sum, Exclusive>, u32>>::scan(&evaluator, three);
        
        
        Operations like Sum, Max, ReduceXor, Inclusive, and Exclusive are all distinct types.

        As mentioned in the post, execution shape is typed too. A static shuffle takes its control as a type-level constant, and the shuffle mode constrains which controls are expressible:

          // Shift down one lane, keeping our own value where the source is inactive.
          let down  = <Executor as EvaluateShuffle<Shuffle<Down>, DownOrSelf<1>, u32>>::shuffle(&ev, v);
          // Broadcast from lane zero.
          let bcast = <Executor as EvaluateShuffle<Shuffle<Broadcast>, WarpLane<0>, u32>>::shuffle(&ev, down);
          // Butterfly exchange with the neighbor one bit away.
          let bfly  = <Executor as EvaluateShuffle<Shuffle<Xor>, Butterfly<1>, u32>>::shuffle(&ev, bcast);
        
        
        For an example of errors caught, a warp-scoped executor for a device-scoped barrier is a compile error:

          <ScopedWarpExecutor<'_, WarpUniform> as EvaluateBarrier<Barrier<Device>>>::barrier(evaluator)
          // error[E0277]: the trait bound `Device: NvptxBarrierScope` is not satisfied
          //               help: the trait `NvptxBarrierScope` is implemented for `Warp`
        
        
        Strip mining is typed on the amount of work and the lane capacity, and it hands back one chunk at a time along with the predicate saying which lanes live in that chunk:

          // Six work items across four active lanes: two chunks, based at 0 and 4.
          <Executor as EvaluateStripMine<StripMine, (WorkItems, ActiveLanes<StripMined<4>>), i32>>::
              for_each_strip_mined(
                  &evaluator,
                  (WorkItems::new(6)?, ActiveLanes::new(4)?),
                  |index, active| {
                   // ...
                  },
              );
        
        
        Hopefully that gives the flavor of it.
    • Eridrus 17 minutes ago

      Given the massive demand for GPUs for LLMs, what sorts of work do you expect to economically benefit from utilizing GPUs more?

      • LegNeato 13 minutes ago

        Part of our thesis is that decent GPUs are in every shipping device and most software doesn't use them and should.

    • lbhdc 43 minutes ago

      This is really cool! It sounds like y'all have a compiler fork that you are using to make this work. I wanna tinker with this, is your compiler available?

      • LegNeato 6 minutes ago

        It is not currently available but we intend to make it available after we launch our products.

  • efnx an hour ago

    Congrats to the Rust-GPU folks! Nice to see the good work flowing.

  • rust-lang 36 minutes ago

    Good job!

  • the__alchemist 44 minutes ago

    Hey - this is probably off-topic/meta, but what is going on with the comments here? Is it bots?

    • dev_l1x_be 39 minutes ago

      No idea, but it seems HN needs POW challenges.

      • lukan 16 minutes ago

        Could also just be trolls attracted by the Rust topic.