Just hard fork the project. Frankly, llama.cpp is so badly written that these kind of speedups are trivial, and a hard fork (or a total rewrite) has been needed for the longest time.