Most of my and my peers PortCos run their own eval and benchmark sets, simply because they know what they need best.
The reality is, capabilities have largely converged across foundation models over the last 18 months, and much of the value add is coming from the harness layer itself now.
This has been the operating assumption for me and my peers, and has largely played out that way.
That said, this has always been an issue with benchmarking since the very beginning. DB Benchmarks, compute benchmarks, and others that were external facing were always inherently a content and product marketing tool. The actual internal benchmarking used to model, understand, and enhance your product was always a closely held secret.
Most of these conversations are happening, but largely in person and not on HN.
> capabilities have largely converged across foundation models over the last 18 months
For reference, in March '25 the models du jour were Sonnet 3.7, gpt o4 and gemini 2.5 pro. GPT5 was in august '25.
It's been a while since we've heard the old "models have stagnated". Oh well.
[flagged]
> The reality is, capabilities have largely converged across foundation models over the last 18 months, and much of the value add is coming from the harness layer itself now.
Can you please elaborate on what are your thoughts on open weights models (GLM 5.3, Kimi K3, deepseek etc.)
and if the value add is coming from the harness layer itself, then thoughts on open source harnesses (there are so many harnesses but to name a few: opencode, pi [omp as well], maki, codex is OSS as well, fx.sh) and you can always combine them with skills (Obra/superpowers, matt pocock skills plus using these skills and others to create some other custom skills tailored to your use case as well)
And what about the combination of both now with this cheap open weights models + open source harnesses and other things to compete over the closed garden ecosystems?
How does that comparison follow in reality
Could you in theory use these methods to save on the massively expensive $$$ token spending on Anthropic/OAI?
(Personal anecdote but I have GLM 5.3 + maki [sometimes omp/opencode but mostly maki] and its good enough for most use cases out there that I have and I dont know of too many use cases outside of say recreation of games for examples maybe that I would prefer complete SOTA models. I would also love to know where you believe that SOTA models absolutely do still make the difference discounting the benefits provided by the harness.)