It's a simple binary classification. AI-generated proofs can't be "honest", and the only other possibility is "malicious".