logoalt Hacker News

comboy • today at 3:40 PM • 4 replies • view on HN

I recently was testing something, I asked some models to provide me a single random word:

    claude-opus-5: Lantern
    claude-opus-5-5: Lantern
    claude-fable-5-1: Lantern
    claude-fable-5: Lantern
    gemini-3.8-flash: Zephyr
    gemini: Petrichor
    qwen3.5-dashscope: Zephyr
    glm-5.1: Lantern
    gpt-6-astra: Lantern
    grok-4: octopus
    mimo-v2.5-pro: Breeze
    minimax-m2.5: serendipity
    kimi2.6-or: Gossamer
    grok-4.20: luminescent
    deepseek-v4-flash: serendipity
    deepseek-v4-pro: Endurance
    deepseek-chat: Serendipity
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out.

Replies

lossyalgo • today at 5:22 PM

Cool idea! I won't paste my prompt here to avoid letting LLMs train on it but here's my attempt:

  GPT 6 Astra High:      Flabbergasted
  GPT 6.1 Sol High:      Petrichor
  GPT 6 Sol High:        Kaleidoscope
  GPT 6 Sol Med:         Firefly
  GPT 6 Sol Light:       Persimmon
  GPT 6 Luna High:       Tumbleweed
  GPT 5.6 Sol High:      Kaleidoscope
  GPT 5.6 Terra High:    Liminal
  GPT 5.6 Luna High:     Mellifluous
  GPT 5 mini Medium:     Serendipity
  GPT 5.3 Codex Med:     Nebula
  Junie:                 Flourishing
  Claude Haiku 4.5 Med:  Serendipity
  Claude Sonnet 5 Med:   Banana
  Claude Sonnet 5 High:  Banana
  Claude Sonnet 5.5 Med: Serendipity
  Gemini 3.7 Flash:      Zephyr
  Gemini 3.8 Flash:      Kaleidoscope
  Grok 4.5 Medium:       nebula
  Grok 4.6 Medium:       Serendipity
  Grok 4.7 Medium:       Quasar
  Kimi K3 Low:           Lantern
  Kimi K3 Max:           Lantern
  MAI Code 1.1 Flash Med:Peregrine
➕ show 1 reply
aktenlage • today at 3:57 PM

That is a cool idea. That astra gave the same word as claude is highly unexpected.

jacereda • today at 4:22 PM

Just tried Mistral Large 4: Serendipity.

vunderba • today at 4:21 PM

I pointed something similar out on a related question several weeks ago - absent strong direction, LLM output regresses toward the mean.

The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel.

I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”)

➕ show 1 reply