architectures · tencent
Deep sparse mixture-of-experts transformer — 80 layers with 192 experts and 8 routed per token — using grouped-query attention over 64 query heads against 8 key/value heads. Declared as model_type `hy_v3`.
| Publisher | tencent |
|---|---|
| Origin | China |
| Licence | Tencent-Hunyuan |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.