logoalt Hacker News

DrJokeputoday at 4:46 AM1 replyview on HN

> It sounds to be like 5.5 and below may have been using a conventional attention scheme where a cached KV sequence can be easily used to restore a prefix of itself, but perhaps 5.6 is using linear attention or LSTM or another recurrent scheme where you cannot rewind the model state by just truncating it.

I feel like this is the kind of substantial change to your product that you would need to tell your customers about. It would be simply disrespectful to your customers to not disclose this upfront.


Replies

zx8080today at 5:05 AM

It depends on who they consider the customers. Shareholders and govt are the customers, not users.

Users is the product.