Tech & AI News
Hacker News

Muse Glimmer is a memory hierarchy disguised as a 30B Transformer

Muse Glimmer is a 30-billion-parameter decoder-only multimodal model. It uses a memory hierarchy with 52 text layers that alternate between local 2,048-token RoPE-bounded attention and global full-context attention. The model has 32 query heads, 2 KV cache heads, and a 4-bit compressed language core under 20 GB. A Vision Transformer processes images once, compresses patches four-to-one, and feeds tokens to the language decoder, allowing long-term context and tool usage on 24–32 GB hardware.