Fun update to this: Daniel Lemire added another optimization to make this even faster. https://github.com/jadidbourbaki/llama.cpp/pull/12
I’ll benchmark his change and add it to the article, crediting him for this improvement.