architectures · liquidai
Hybrid backbone that interleaves short-range convolution blocks with full-attention blocks rather than using attention at every layer (config declares layer_types of `conv` and `full_attention`, with a 3-step convolution cache). Only the attention layers contribute to the KV cache, so cache growth with context is a fraction of an all-attention stack of the same depth. Grouped-query attention on the attention layers; 128K context. Declared as model_type `lfm2`.
| Publisher | liquidai |
|---|---|
| Origin | United States |
| Licence | Apache-2.0 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.