Hopefully it will be open weights and have the same architecture and size as the current v4 flash vision, which is probably the best LLM that can be run on 128G devices.
Interesting, I had assumed it'd be too large to fit. What quant and context size are you running?
Interesting, I had assumed it'd be too large to fit. What quant and context size are you running?