AI News
Anthropic updates Fable 5 biology safeguards to cut benign-query fallbacks
Anthropic says a refined classifier reduces biology-related fallbacks by about 85% while keeping dual-use research requests behind existing safeguards.
Anthropic has updated the biology safeguards attached to Fable 5, saying the revised system will send far fewer ordinary biology questions to a less capable model. The company reported that biology-related fallbacks fell by about 85% in its testing across product surfaces. The update is intended to make the model more useful for everyday health and educational questions while preserving restrictions for requests the company classifies as dual use, including areas such as virology, toxicology, molecular design and professional drug-development work. The change is a refinement of an existing safety boundary, not a decision to open unrestricted frontier biology assistance.
Fable 5 originally used a broad classifier that was designed to err on the side of caution. When the classifier detected a potentially safeguarded biology request, the system rerouted the user to Opus 5, a model Anthropic describes as capable but less biologically capable than Fable 5. That route is what users experienced as a fallback. The approach limited exposure to difficult dual-use questions, but it also blocked benign requests that could arise in a classroom, a patient’s effort to understand lab results or a general discussion of symptoms. Anthropic now says its new classifier separates those categories more precisely.
The company’s explanation highlights the central difficulty of biology safeguards: beneficial and harmful uses can be technically close. Researching a disease, designing a treatment or studying a pathogen can overlap with knowledge that might be misused. Anthropic said Fable 5 can outperform experts on some complex biological tasks and can provide operational support on others, which is why it initially placed a wide range of queries behind a safeguard. It argues that the same capabilities that can help a legitimate researcher can also provide material assistance to a malicious actor, so a model’s usefulness cannot be assessed only by whether a question sounds educational on its surface.
To make the boundary narrower, Anthropic said it rewrote the classifier’s constitution, a collection of rules used to distinguish allowed material from safeguarded material. It then developed updated training data, retrained the classifier and sought feedback from internal and external experts. The company says the resulting system still triggers for harmful and dual-use research biology content while allowing more benign requests. It also says the classifier must be robust against jailbreak attempts, a requirement that complicates the effort to reduce false positives. The reported 85% figure is Anthropic’s own testing result, not an independent measurement.
The expected user experience varies by product surface. Anthropic said total fallbacks, including those prompted by other causes, should decline by roughly 67% on Claude.ai, 55% on Cowork, 17% on Claude Code and 7% on the Claude Platform. Those numbers should not be interpreted as a guarantee that every biology question will now be answered by Fable 5. The company said it will continue to route dual-use professional biology and drug-development requests away from the model. Its stated longer-term goal is to create trusted access paths for researchers who need stronger capabilities under more controlled conditions.
The release matters because it illustrates a common trade-off in frontier-model deployment. A safety system that blocks too broadly can frustrate users and weaken the usefulness of a product in legitimate settings. A system that blocks too narrowly can create risks that are harder to reverse after a model is widely available. Anthropic’s approach is to change the classifier rather than remove the boundary, using narrower rules and new training data to move more clearly benign questions to the allowed side. Whether that balance holds will depend on real-world performance, adversarial testing and the company’s willingness to adjust again when errors emerge.
For clinicians, educators and ordinary users, the practical message is measured. More everyday biology and health questions may receive direct assistance instead of a fallback, but the product is not being presented as a substitute for medical judgment or a universal research tool. For researchers working in controlled, dual-use areas, the restriction remains material. The announcement shows that model safety is increasingly an operational product question: users see it in which model responds, how often a request is rerouted and whether a helpful answer is available. The next evidence will come from how the refined classifier performs outside Anthropic’s internal tests.