Super awesome. Wish they would release the paper about what they did to achieve this. I remember nous released the token superposition paper which improved pretraining FLOPs some, but not 50x: https://nousresearch.com/token-superposition. Wondering if they also found some cool tokenization strategiesa
"Write a paper" < "Sell to a big AI lab for $$$"
Going to be interesting to see what happens to discoveries like this in the future.