architectures · qwen
Alibaba Qwen3.5 extended MoE architecture. Builds on Qwen3 with wider expert banks (up to 397B total / 17B active), Grouped-Query Attention, and dynamic expert dropout during training. Supports 262K context via dual ROPE scaling. Used by Qwen3.5-27B through Qwen3.5-397B-A17B leaderboard models.
| Weight file SHA-256 | bdb2545734738da2dcd1b68c0ae9a896b4823206efe129fb48500eefe53dbd35 |
|---|
| Publisher | qwen |
|---|---|
| Origin | Recorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which. |
| Licence | Apache-2.0 |
| Pulls recorded here | 92,466 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.