this is cool but like, are we just vibe coding NAND burners at this point? these decode times don't really tell the whole story, because prefill becomes the bottleneck.
half an hour to process 10k tokens on an M5 seems... not great
Not great for coding, or realtime agent interactions. But for background processing tasks overnight? Seems like it’d work pretty well
Not great for coding, or realtime agent interactions. But for background processing tasks overnight? Seems like it’d work pretty well