Not really, though. Sequences of tokens are predicted as {spam, not spam} via training on pair sets. It's like a token predictor where the output vocabulary is two tokens.
The complete inability to use it to generate spam doesn't make it different to you?
The complete inability to use it to generate spam doesn't make it different to you?