> Mythic M1 stores up to 80 million neural network weight parameters directly on-chip
Which means connecting ~350 chiplets to run a Qwen 3.8 27b and over 30000 chiplets to run Qwen3.8-2.4T-A95B. Cost? Space? Feasibility?
Edit: wrong values, lost a zero...
Edit: seemingly, the M1 is only part of the whole need. With the M1, you would run a feedforward pass of the NN but use the rest of the Von Neumann architecture to manage the data. The pass in the M1 will be lightning fast, the rest still a bottleneck. The M1 is almost explicitly not for LLMs.
[dead]
I think we are at the apex of Von Neumann machines. Once there is enough money to make alternatives, they will thrive. And AI is that catalyst.