architectures · minimax

MiniMax Lightning Attention (Hybrid)

MiniMax hybrid architecture combining linear Lightning Attention with standard softmax attention in alternating layers. Enables O(1) per-token inference cost at 1M context without KV-cache growth. Used in MiniMax M2 through M2.7. Achieves 400+ tok/s throughput at 230B scale.

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Weight file SHA-256c321fb89f473be5e518d8ca15faa86a0f6b39797977b0ac55ac28ad9dfb7bae5

What is recorded

Publisherminimax
OriginRecorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which.
LicenceMiniMax-Open
Pulls recorded here51,918

Elsewhere in the registry

Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.