We're serving it around 150-200tok/s (uses our new speculative decoding implementation on a DFlash2 draft model).
https://mixlayer.com, LAUNCH-Q38-27B gets you $5 in credits if you want to kick the tires.
I don't see any kind of input cache discount listed on your pricing page. Do you offer that, or is all input priced the same?
I tried in your playground and got 14.2 tok/s?
This feels great
I don't see any kind of input cache discount listed on your pricing page. Do you offer that, or is all input priced the same?