FrontierThe story, in brief

Tencent Open-Sources AngelSpec: A Unified Training Framework for MTP and Block-Parallel Speculative Decoding on Hy3 Models

Tencent open-sources AngelSpec: a unified framework for speculative decoding that hits 2.4× speedup on 295B models. Here's why inference speed just became a competitive lever.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Speculative decoding is moving from research to production tooling. Tencent's open-source framework and DFly drafter architecture give practitioners a path to 2–2.4× faster inference on large models without retraining — a capability gap closer for teams without frontier-lab resources.

The key facts

6 to know
  1. AngelSpec: open-source torch-native framework for speculative-decoding draft models

  2. Supports six architectures (HY3-295B-A21B confirmed)

  3. DFly drafter: block-diffusion with hybrid target conditioning + hidden-correction autoregressive head

  4. D-cut: runtime-adaptive verification budgeting for inference optimization

  5. 1.98–2.40× speedup vs autoregressive decoding on HY3-295B at TP=8, concurrency 4–64

  6. Open-source release enables broader adoption of speculative decoding

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Tencent has released AngelSpec, an open-source torch-native framework for training speculative-decoding draft models across six architectures. It introduces DFly, a block-diffusion drafter with hybrid target conditioning and a hidden-correction autoregressive head, and integrates D-cut for…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier