logoalt Hacker News

VulgarExigencytoday at 4:06 PM1 replyview on HN

It was being served for free. They were almost certainly being overloaded.


Replies

Aurornistoday at 6:16 PM

Presumably the efficiency numbers they're quoting are for the high concurrency state they were serving.

RAM was probably the bottleneck for the amount of context they were offering.

I assume it would run a little faster with lower concurrency but "RIP nVidia" is a little premature. The cutting edge inference hardware is amazingly powerful