512 option isnt worth it imo, you get severe slowdowns when weights are that large. 256 is the sweet spot, you can run large open weight models at decent speeds for full private inference.
> 512 option isnt worth it imo, you get severe slowdowns when weights are that large.
I think most people are getting 512 for running Chrome with a bunch of tabs open. /s
a) we don't actually know what the prices will look like yet, b) what about same weights + huge context? or, same weights that you'd run on 128gb/256gb, but multiple models running for different tasks?