The entire safety evals industry is essentially funded and controlled by OpenAI/Anthropic. Notice that on recent models, they exclusively use internal testing or black box external vendors (e.g., Gray Swan) whose entire business is to serve OpenAI/Anthropic. And all these companies just share the same pool of researchers back and forth.
That doesn't sound like it describes SecureBio to me?
(Disclosure: I work at SecureBio, but not on the biological evals side.)
Yes this should be immediately replaced by a federal agency, like we do for other kinds of potentially harmful products.
Anyone who calls it “safety” probably has a certain world view and is more aligned with the big 2 (and stuck in 2023).
There is a growing industry of commercially focused risk evals that has a broader customer base.
that's pretty damn smart if this was a long-term plan to block competitors
So, it's a cottage industry.
Who gets to decide what is safety?
I expect some of those tests (prolly not public) will basically be "wokeness" tests or "PC correctness" tests or "western media filter" tests.
China has different objectives. Sure.
I'm not sure one is safer than the other; I would know which one to go to if I want to research on topic that are viewed very different on both sides of this "new iron curtain".
The USG has a safety organization (CAISI), but it has been neutered by the current administration (with the recent stop-work order etc.). Perhaps UK AISI would be closest to what you are looking for? See their recent work on Kimi K3 cyber (which was declared safe) [1].
It's tricky because a lot of the safety researchers have ties to the labs since those were the only companies training LLMs >5 years ago.
[1]: https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-...