Meta
Meta releases open-weight Muse Glimmer for single-GPU agents
The 30-billion-parameter model, distilled from Meta's closed Muse Spark 1.2, is Meta's first fully open release since it retired Llama.
The answer
Meta open-sourced Muse Glimmer, a 30B model that runs agents on one GPU.
Meta released the weights of Muse Glimmer on Hugging Face on 10 August 2026, a 30-billion-parameter open model licensed under Apache 2.0 for commercial use, modification and redistribution. Meta said the model is designed to run AI agents locally on a single consumer GPU.
Meta Superintelligence Labs built the model by distilling it from Muse Spark 1.2, Meta's closed flagship model launched on 5 August, according to gHacks.
Muse Glimmer is Meta's first fully open model since it retired the Llama line in favour of the proprietary Muse Spark earlier in 2026. Meta chief executive Mark Zuckerberg said: "Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally."
The full BF16 weights total 59.6 GB, but a 4-bit version comes in under 20 GB and runs on a 24 or 32 GB GPU or a Mac, with under 1% accuracy loss reported, Meta said.
The model ships with a small drafter called DFlash for speculative decoding, which Meta said speeds up inference by 3.1 times on an RTX 5090, 1.5 times on an M4 Max and 1.8 times on an M5 Max.
Muse Glimmer handles text and images through an 1.8-billion-parameter vision encoder, and Meta said it was trained for multi-step tool use, function calling and recovering when an API call fails. It supports a context window of more than 131,072 tokens.
On the MCP Atlas benchmark, Muse Glimmer scored 75.5, ahead of Gemma 4 31B at 54.2 and Qwen 3.6 27B at 62.5. It scored 35 on the Artificial Analysis Intelligence Index.
Zuckerberg said Meta will also open-source Muse Spark 1.2, the closed model from which Muse Glimmer was distilled.