architectures · baidu
Compact vision-language architecture for document OCR: a vision encoder paired with a small sparse-expert language decoder (12 layers, 64 routed experts with 6 active per token). Task-specific rather than general-purpose, which is why the parameter count is small relative to its capability on document images. Declared as model_type `unlimited-ocr`.
| Publisher | baidu |
|---|---|
| Origin | China |
| Licence | Apache-2.0 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.