logoalt Hacker News

Catloafdevyesterday at 11:37 PM1 replyview on HN

It's absolutely reasonable to have safeguards on sufficiently dangerous models being released - if you disagree, can you explain your perspective?

I think it's wildly irresponsible to release models that are extremely capable at things like bio-weapons. Do you really think information anarchy is the answer?

The problem with open models compared to closed models is not about protecting profit - it's about protecting capability. Any open model can be retrained or fine-tuned for anything. There's no such thing as an open model that is both capable _and_ permanently safe when it comes to certain dangerous topics. It's not possible to prevent 'uncensoring' a model.


Replies

philipkglasstoday at 12:21 AM

The biggest problem is the infectious nature of restrictions that start narrowly. Fable is too touchy about helping people with biology and chemistry problems. It was initially released with an even more insidious safety mandate:

https://simonwillison.net/2026/Jun/10/if-claude-fable-stops-...

In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design).

...

Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT).

(And although the "silent" downgrade part was quickly dropped, Fable still won't help you here.)

Anthropic won't teach you how to build bioweapons, or enable you to make your own software infrastructure so that you can train your own biology model. That's where lawmakers may arrive too if they buy Anthropic-style safety arguments. It's too dangerous to publish models that understand biology. It's too dangerous to publish training software. It's too dangerous to publish tools that allow you to build training software.

If you keep following the implications of their safety argument, it's as broad an assault on the distribution of software and computing as has ever been proposed. Worse than the Clipper Chip proposal of the 1990s era Crypto Wars. I have seen how "children must be protected online" has in practice turned into an attack on adult privacy affecting a wide swath of services and devices. I'm taking a maximalist position on openness now because I think that I can anticipate the next steps on the safety side, and I reject those steps.