logoalt Hacker News

raphlinus • today at 9:31 PM • 1 reply • view on HN

You've got a point but are overstating it considerably. There is a big gap between just autovectorization and the portable primitives a library like Highway or Fearless SIMD will give you. For example, I haven't seen autovectorization do select or swizzle.

But there's another point in the tradeoff space. One of the explicit design decisions in Fearless SIMD is to support "downcasting," or specialization to a specific microarchitecture. At least for the kind of problems I've worked on, even when you're doing something fancy with arch-specific permutations or what not, the majority of the operations will be pretty vanilla, and can be expressed well in the portable subset.

So you can think of a library like Fearless SIMD as enabling your extreme optimization use case, just more ergonomically.

Of course, this depends on LLVM compiling intrinsics to assembly efficiently. That hasn't always been the case, and is not perfect now (a number of issues have been filed against rustc and LLVM while developing Fearless SIMD), but is pretty good.

As always, though, you do have to measure performance, and I frequently look at the assembler output to double-check that it's doing the right thing. The day of "fire and forget" portable SIMD has not yet arrived.


Replies

Asmod4n • today at 9:57 PM

The issue with libraries which offer you portable simd is that the auto vectorizer of the compiler will likely generate faster code.

➕ show 1 reply