it "fits" in 64 ram with mmap. granted it runs at 15 tk/s with a 9070xt but it runs
Right, I meant "fits" in the sense of I can load the whole thing into some combination of system RAM and GPU at llama-server launch.
15 tk/s isn't useless if you can give it big tasks to do overnight, or like ask it to do something and check back 3-4 hours later.
Right, I meant "fits" in the sense of I can load the whole thing into some combination of system RAM and GPU at llama-server launch.
15 tk/s isn't useless if you can give it big tasks to do overnight, or like ask it to do something and check back 3-4 hours later.