Just run /goal to optimise it and you should be good in less than an hour. Also best to use models that support speculative decoding.
Optimize llama.cpp? Hmm.
WRT speculative decode, basically zero finetunes keep it. I'm testing with some ridiculous abliterated amalgamation so spec decode has been gone for most of its ancestry.
Optimize llama.cpp? Hmm.
WRT speculative decode, basically zero finetunes keep it. I'm testing with some ridiculous abliterated amalgamation so spec decode has been gone for most of its ancestry.