I think it's safe to start the conversation as about bad as the jump from 256 to 512 on the M3, which was a little more than double base to 256. If it's surprisingly different at launch then it can be a party, but there is no sense getting your hopes up for that at the moment.
Longer context also slows token prediction proportional to the context size. If it wasn't regularly referenced then there would be no need to keep it in RAM.
Usually the pitch for more memory is "I can run a massive model/context and get my answer in a while instead of next weekend from disk".