Well, yes and no.
The author has an M3.
Here's reality, MLX on the software layer will not magically place hardware matrix multiplication units in your GPU cores.
Newer Macs are always just gonna smoke anything earlier than an M5.
Not even sure why this person is trying to get this stuff to run on hardware that wasn't designed for AI?
Everyone is pointing him to newer hardware precisely because you need the newer stuff to get models to be performant. You can go with AMD, NVidia or Apple, but you're gonna be using stuff designed well after the M3 if you want to push >100tok/s.
> Not even sure why this person is trying to get this stuff to run on hardware that wasn't designed for AI?
Posturing/overclaiming like this shades rather than illuminates, there are no worlds in which the "M3...wasn't designed for AI". My M4 Max 64 GB gets the same speed.