Related to this idea, there are interesting papers that explicitly examine the adversarial case. Basically, besides the provider hiding watermarks, one could also think of an adversary training a model to exhibit this behaviour depending on the Input of the prompt. So if you use the manipulated model, not only information about the author that the platform knows is encoded, but also, e.g. one-time tokens from your email. This works surprisingly well (albeit with the naive approach still noticeable in most cases).
TrojanStego: https://arxiv.org/abs/2505.20118 Improvement: https://arxiv.org/abs/2606.09411