This is why I love Go. Nobody was asking for this, but they took the time to do it right and continue to Push go as a memory safe, high-level systems language.
This feature opens many doors for optimizing low-level performance in Go projects, that are already running multicore. IIRC there aren’t a lot of languages with built-in std lib support for SIMD and variants. Love the way Go is trying new stuff lately.
Besides the usual C and C++, we have Java, .NET, D, Zig, Julia, Swift, Rust.
So yeah, also appreciate having Go in the group instead of manually having to write Assembly.
However not many languages adopt ways to manually write SIMD, because most of us have no idea how to write good SIMD code in first place, I surely don't.
Even with languages that adopt ways to manually write SIMD, it’s mostly left to library maintainers rather than application developers.
I work for a C++ timeseries database startup that leverages SIMD about as much as we possibly can, and except for some extremely rare places we just use libraries.
Yeah, that is what I have heard from some NVidia folks as well, like Bryce Adelstein, use the libraries as much as possible, and leave the kernels for experts.
However even then, it depends on how the libraries API surface looks like.
But it’s not necessary at all, the whole point is that these utility libraries bring you more elegant code that work on all platforms without having to pollute your codebase with SIMD intrinsics.
Unless this was tongue in cheek, because this is in fact a problem with AI that it degrades your codebase in these types of ways.
It's hard to make predictions with an open source project, but my personal guess is some flavor of it will land (including it is already demonstrating good results without an enormous level of code complexity in the compiler and without overly slowing down compile speeds), but I guess we'll see.
It's being driven by an external contributor who has landed some good changes in the past to the Go compiler. (I think the autovectorization work might be part of their PhD or other academic research, but not sure.)
As a first step, it might be possible to write a linter rule that rewrites suitable numeric loops to SIMD. There are already rules to rewrite several loop types, so that should be doable.
The problem with Go isn't performance but with the C/C++ interop overhead, even with the "30% less overhead" from a few updates ago which isnt true for 99% of cases, it isnt enough
Already using this for foreground estimation of cutouts in my project, around 30% speedup over non-SIMD, but the algorithm is probably not very optimised yet.
This is why I love Go. Nobody was asking for this, but they took the time to do it right and continue to Push go as a memory safe, high-level systems language.
This feature opens many doors for optimizing low-level performance in Go projects, that are already running multicore. IIRC there aren’t a lot of languages with built-in std lib support for SIMD and variants. Love the way Go is trying new stuff lately.
Besides the usual C and C++, we have Java, .NET, D, Zig, Julia, Swift, Rust.
So yeah, also appreciate having Go in the group instead of manually having to write Assembly.
However not many languages adopt ways to manually write SIMD, because most of us have no idea how to write good SIMD code in first place, I surely don't.
Even with languages that adopt ways to manually write SIMD, it’s mostly left to library maintainers rather than application developers.
I work for a C++ timeseries database startup that leverages SIMD about as much as we possibly can, and except for some extremely rare places we just use libraries.
Yeah, that is what I have heard from some NVidia folks as well, like Bryce Adelstein, use the libraries as much as possible, and leave the kernels for experts.
However even then, it depends on how the libraries API surface looks like.
With AI I'm pretty sure SIMD will be easier to integrate when necessary.
But it’s not necessary at all, the whole point is that these utility libraries bring you more elegant code that work on all platforms without having to pollute your codebase with SIMD intrinsics.
Unless this was tongue in cheek, because this is in fact a problem with AI that it degrades your codebase in these types of ways.
With AI, I expect it to eventually be good enough for us to finally have 5 GLs, so it won't really matter.
"CGO 2022 Keynote: Compiler 2.0"
https://www.youtube.com/watch?v=w_sX9aZoZxg
Vectorizing computations has been Matlabs secret sauce.
Does matlab these days do stuff like JIT operator fusing to avoid memory roundtrips and take advantage of FMAs?
Yes: https://www.mathworks.com/help/fixedpoint/ref/half.fma.html
I'm grateful that Go a non-proprietary language offers these features.
Julia does that too.
Oh this is great, it was one of my biggest bugbears about Go since you almost always have to link C/C++ code to get the appropriate performance.
The one negative I'd say is that often autovectorisation is 'good enough' and this doesn't really tackle that gap.
FWIW, there is some pretty substantial autovectorization work that is already in-flight for the Go compiler.
There's a CL stack here:
https://go.dev/cl/791740
It's hard to make predictions with an open source project, but my personal guess is some flavor of it will land (including it is already demonstrating good results without an enormous level of code complexity in the compiler and without overly slowing down compile speeds), but I guess we'll see.
It's being driven by an external contributor who has landed some good changes in the past to the Go compiler. (I think the autovectorization work might be part of their PhD or other academic research, but not sure.)
As a first step, it might be possible to write a linter rule that rewrites suitable numeric loops to SIMD. There are already rules to rewrite several loop types, so that should be doable.
The poor Assembler and the unsafe package forgotten in the corner.
While reaching out to CGO is the easier way, it doesn't mean it is the only tool available in Go.
The problem with Go isn't performance but with the C/C++ interop overhead, even with the "30% less overhead" from a few updates ago which isnt true for 99% of cases, it isnt enough
Already using this for foreground estimation of cutouts in my project, around 30% speedup over non-SIMD, but the algorithm is probably not very optimised yet.