Anthropic released its AI model Fable on Tuesday as a public and limited version of its cybersecurity-focused Mythos model, initially launched in April under Project Glasswing to select companies [1].

Fable includes strict guardrails that block or alter responses to queries related to cybersecurity, biology, chemistry, and distillation attempts. When these guardrails activate, the system rejects requests or falls back to Anthropic’s earlier Claude Opus 4.8 AI model [1, 2]. Anthropic designed the restrictions to prevent malicious uses such as malware creation or biological weapons development [1].

However, cybersecurity researchers criticized the guardrails as overly broad and inconsistent, hindering even benign tasks. Valentina “Chompie” Palmiotti of IBM X-Force noted that Fable "rejects any request that could be tangentially cyber related. Even innocuous tasks like reading a blog post" were blocked [1]. Veteran cybersecurity expert Matt Suiche said, "If you ask it to write secure code, it assumes it is cybersecurity related work instead of software engineering best practices, and you get downgraded." He added, "It’s better to catch more people than not enough. But it is understandable as we are still in the early days and they are still adapting their guardrails" [1].

Anthropic acknowledged criticism about stealthily applying invisible guardrails, especially those filtering distillation attempts, without notifying users that responses were degraded or changed [2]. The company said it will make restrictions more transparent by showing fallback notifications whenever guardrails trigger. In a statement to The Verge, Anthropic said, "Visible safeguards can be probed, so they have to be robust, which takes time to get right. Invisible safeguards can be targeted more narrowly, allowing us to ship quickly with very few false positives" [2].

Since April, Anthropic has expanded Mythos access to hundreds of organizations in 15 countries [1]. Fable is positioned as a safer, publicly accessible subset of Mythos’s powerful capabilities.

Anthropic’s next steps include implementing visible notifications for guardrail triggers on Fable to improve user awareness about restrictions and fallbacks to older models [2].