I recently was testing something, I asked some models to provide me a single random word:
claude-opus-5: Lantern
claude-opus-5-5: Lantern
claude-fable-5-1: Lantern
claude-fable-5: Lantern
gemini-3.8-flash: Zephyr
gemini: Petrichor
qwen3.5-dashscope: Zephyr
glm-5.1: Lantern
gpt-6-astra: Lantern
grok-4: octopus
mimo-v2.5-pro: Breeze
minimax-m2.5: serendipity
kimi2.6-or: Gossamer
grok-4.20: luminescent
deepseek-v4-flash: serendipity
deepseek-v4-pro: Endurance
deepseek-chat: Serendipity
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out.That is a cool idea. That astra gave the same word as claude is highly unexpected.
Just tried Mistral Large 4: Serendipity.
I pointed something similar out on a related question several weeks ago - absent strong direction, LLM output regresses toward the mean.
The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel.
I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”)
Cool idea! I won't paste my prompt here to avoid letting LLMs train on it but here's my attempt: