I think what you're seeing is the probability of a particular signature or style of signature appearing on a particular style of cartoon, not an intent to sign.
This. It's also while you'll sometimes get a mangled Getty Images watermark on some image generations, or a logo in the bottom left corner. If it's a prominent feature in the training dataset it'll show up, exactly how these models are supposed to work.
The 'bug' here is whatever post-processing step or system prompt is in place to steer the model away from doing this.
Yes, but that's just the pixel generation layer. The harness or whatever infrastructure around could nudge it towards common sense.
I do many voice transcriptions. Many times I've had empty silence at the end of recordings transcribed as "Thank you" or even "Like and subscribe".