On a related note, I was just noting to my co-founder, as we struggle to write good case studies for our website, that I find LLMs are astoundingly bad at writing good prose.
We all know the "AI-tics" that give away a sloppily AI-written piece, but even if you steer them, they still struggle to write consistently high-quality prose.
Somehow I feel that the work of a good copywriter has never been more noticeable.
Random observation: Google's Gemma 4 models write so much nicer prose than ChatGPT or Claude.
Though this might be me as a British reader, simply preferring a rather less American turn of phrase.
I reckon the more transatlantic, english-as-international language DeepMind team have had a subliminal (or maybe deliberate) impact on the way it chooses to write.
Or perhaps small open weights models simply aren't under the same commercial pressure to be engaging and sycophantic and are therefore less likely to adopt the samey overly casual, upbeat, Californian sales assistant manner. (Don't get me wrong, I like this from real human Californians just fine!)
Either way, the default tone is much less showy. I would be interested to find out if you agree.
I am very much an LLM cynic. I am engaging because I must, and trying to learn fundamentals, but I would not say I am overly excited by any of this, just glad that small open weights models exist as a counterpoint.
I loathe the way ChatGPT writes, and the Claude-isms that are everywhere; it is actually quite enraging, especially when you start seeing it in internet comments from people who used to try to write out their own thoughts.
But in my experiments with open weights models I have found I am much less aggravated by summaries and outlines written by Gemma 4, so much that I am happy enough to read them, because they have fewer irritants that take me out of the reading flow.
Though this evening it told me very kindly that my photography is a bit "safe". How very dare it… understand me that well.
A lot of what would be the top "reference" works aren't even that good either, they were a successful marketing phenomenon or had cultural or social relevance at their time. So, you can get a lot of bad prose going by a number of somewhat logical, externally measurable parameters.
I gave Claude Fable $25 in Pangram API credits and, after hundreds of attempts, it was unable to produce a single readable original piece of writing that was not immediately identified as AI.
This seems to be a hard problem for LLMs, as passing would probably require good self-perception ("oh no, I am writing like an AI!") and fine-grained control over its own output ("let's write like a human instead!").
I've not really enjoyed finding out lately just how few people seem to notice what ought to be unmissable.
More noticeable to me is the lack of the work of a good copy editor which, sadly, we haven't had for a really long time. At least, not on the interwebs. Even the news sites reduced where their print copies were known for rigorous editing saw obvious issues with the various corporate overlords doing serious headcount reductions. The rush to be first to publish reduced even further the time any editors might have had, and then the wide spread use of CMS style articles that slammed output together with something as unintelligent as 'cat segmentFromAuthor1 segmentFromAuthor2 segmentFromAuthor3 > article' where you can tell where each segment started over again with the same basic information as if it was content meant to stand on its own.
Of course, the amount of self published work has also helped make the lack of a good copy editor noticeable. I can excuse self published blogs though. But the stuff released "professionally" has really become farcical.
I can sniff out AI writing immediately but from what I hear AI writing is more popular than ever
I recently realized this as well and I think what I’ve discovered is that AI just produces mediocre content in all realms, but you don’t really notice it except in the realms where you have real expertise. With a lot of harness and prompting you can have it pump out something that’s pretty good but by default the next best token rarely produces anything of quality it seems like and if you think it does, perhaps you may want to recheck your assumptions on your expertise of the topic at hand