Anthropic is leaning hard on scientific research as a Fable 5.1 proof point. On the new Terminal-Bench-Science 0.1 benchmark, the company reports 52.6% for Fable 5.1 versus much lower scores for Fable 5, Opus 5, and GPT-5.6 Sol in its harness — a jump Simon Willison highlighted as the standout chart of launch day. The bench is built for CLI-style scientific problem solving, not trivia chat.
Anthropic also says partners saw long unattended runs, rare crash root-cause finds, and research directions other models missed. Treat vendor benches carefully, but the directional claim is clear: Fable 5.1 is being sold as a model that can stay coherent on multi-hour scientific and engineering loops, not just autocomplete a function.
Source: https://simonwillison.net/2026/Sep/1/claude-fable-5-1/