you may notice Kimi, GLM have also started telling how their model is able to optimise it's own inference pipeline
https://www.kimi.com/blog/kimi-k3