I think 300 tps might be roughly what an astra sized model would be at if running on Cerebras hardware no?
Or are they simply lower the batch sizes on whatever hardware they are currently running?
Yeah sorry I misunderstood you. And yeah I bet that it is the cerebras collaboration although they didn't explicitly say this, maybe their own chip is fast enough for this inference speed? But I could only take a guess if its this or that
Yeah sorry I misunderstood you. And yeah I bet that it is the cerebras collaboration although they didn't explicitly say this, maybe their own chip is fast enough for this inference speed? But I could only take a guess if its this or that