Can we assume that model performance at 90% of the 256k limit != 90% of 1M token limit?
Is this the exact same model just with less VRAM allocated for context window?