To train a large model to do what? Break AES? How would that work?
Train on plaintext, ciphertext -> key.
Train on plaintext, ciphertext -> key.