It’s always been interesting to see how similar output is across ostensibly very different models. I remember testing short story writing in the early days and having the models all choose the same niche topic across e.g GPT, Llama, Phi, Claude, etc