Tencent released Hy4-Preview on Hugging Face and ModelScope on Thursday — a 770 billion parameter Mixture-of-Experts model activating 49 billion parameters per token across 78 layers with 256 routed experts. The model's custom attention mechanism, Gated DeepSeek Sparse Attention with IndexCache, enables cross-layer sparse index reuse and represents a notable architectural departure from standard transformer attention. Benchmark scores — 92.3 on GPQA Diamond, 65.7 on SWE Bench Pro, 64.3 on Deep SWE — place it at the competitive open-source frontier tier, particularly for software engineering tasks. One million token context window, Apache 2.0 licence, available in standard and FP8-quantized versions. Tencent explicitly describes the release as a preview with "real headroom left in both pre-training and post-training." The timing — released into a week in which the primary open-weight hosting and model development infrastructure is being absorbed by Nvidia and Stripe via the Hugging Face and Poolside acquisitions — is worth noting. Hy4-Preview lands on a commons that may, within months, no longer be a commons.
LLMs
Claude Fable 5.1 Cuts Cache Pricing by 75% — and the Real Story Is What Anthropic Is Building Toward
Anthropic's Fable 5.1 and Mythos 5.1 carry the same underlying model but two different deployment re…