logoalt Hacker News

cubefoxtoday at 2:30 PM1 replyview on HN

I would expect this only to be true for linear architectures like Mamba or Gated DeltaNet. Transformers and hybrid architectures do not have constant compute cost per token.


Replies

monkpittoday at 3:50 PM

Performant could certainly mean “higher performing” and not “quicker”.