logoalt Hacker News

mirekrusintoday at 10:57 AM1 replyview on HN

Just run /goal to optimise it and you should be good in less than an hour. Also best to use models that support speculative decoding.


Replies

LoganDarktoday at 11:51 AM

Optimize llama.cpp? Hmm.

WRT speculative decode, basically zero finetunes keep it. I'm testing with some ridiculous abliterated amalgamation so spec decode has been gone for most of its ancestry.