Does 3.6x faster also = 3.6x cheaper?
Yes, 3.6x cheaper when it comes to training!
When it comes to inference it should also be cheaper (since we’ve cut attention sequence length by 4x), but I don’t have a hard number for you how much cheaper for inference
Yes, 3.6x cheaper when it comes to training!
When it comes to inference it should also be cheaper (since we’ve cut attention sequence length by 4x), but I don’t have a hard number for you how much cheaper for inference