Scaleway have two separate (one fully managed one a bit less) services for that:
https://www.scaleway.com/en/generative-apis/
https://www.scaleway.com/en/inference/
yea, but no prompt caching right? This makes it unusable for my usecase at least, the cost would be insane
yea, but no prompt caching right? This makes it unusable for my usecase at least, the cost would be insane