Interesting. Is the speedup from specializing for the shape of Llama 3.1 or are they (contra my mental model) actually winning on burning in the weights?
The weights are in SRAM, so the LLM architecture is burned in but the weights can be updated.
The weights are in SRAM, so the LLM architecture is burned in but the weights can be updated.