Meta AI's EvoHarness-RL framework, which trains agents to manage their execution environments through a unified Belief, Progress, Experience interface, achieved a score of 96.9 per cent on the ALFWorld agentic benchmark using a Qwen3 8B parameter model — edging past Claude Opus 4.5's 96.4 per cent and representing a forty-nine-point improvement over the base model's unassisted performance. A "harness annealing" mechanism reduces latency and compute costs as agents internalise routine execution patterns over time. The framework demonstrates that frontier-level agentic performance on structured task environments is achievable with open-weight models at substantially lower inference costs than frontier model APIs — a finding with direct implications for the enterprise AI pricing dynamics that have allowed the major laboratories to command significant premiums on agentic workloads.
LLMs
Claude Fable 5.1 Cuts Cache Pricing by 75% — and the Real Story Is What Anthropic Is Building Toward
Anthropic's Fable 5.1 and Mythos 5.1 carry the same underlying model but two different deployment re…