© ROOT-NATION.com - Use of content is permitted with a backlink.
Agent-based AI is gradually requiring more memory, external tools, and the processing of sensitive data. Most approaches focused on cloud-based solutions introduce latency, raise privacy concerns, and entail recurring token costs. However, AMD has announced support for the new Meta Muse Glimmer 30B model, and now autonomous AI agents can be deployed directly on devices powered by AMD components.
Read also: AERONAUT – everything that flies above the ground: aviation, UAVs and drones, rockets, and space
Local deployment via frameworks such as llama.cpp reduces reliance on cloud infrastructure, maintains control over privacy, and cuts costs.

Muse Glimmer is a new model with 30 billion parameters, released by Meta Superintelligence Labs. It is distributed under the Apache 2.0 license and is designed for on-premises agent-based AI use cases where model capabilities and developer freedom of choice are important.
Meta Muse Glimmer 30B runs efficiently on a system with an AMD Ryzen AI Max+ processor or a workstation with a single AMD Radeon AI PRO R9700 graphics card using open-source frameworks such as llama.cpp. This makes it possible to deploy a model designed for serious agent-based AI work directly on the user’s workstation, saving costs and keeping workloads local.
Read also: From “Smart Speaker” to Bartender Robot: How AMD Hardware Will Change Your Everyday Life
Initial test results demonstrate Muse Glimmer 30B’s high local performance on AMD hardware: up to 24 tokens per second on an AMD Ryzen AI Max+ 395 processor and up to 53 tokens per second on a single AMD Radeon AI PRO R9700 graphics card with dFlash enabled. However, software and model optimization is already underway, so performance is expected to continue to improve as the ecosystem evolves.

An agent must maintain context, make a series of decisions, invoke tools, analyze results, adapt, and continue working. Muse Glimmer 30B is designed specifically for this kind of long-running process, particularly for complex workflows that span multiple stages and sessions. It can manage memory, recover from failures, and resume work after a restart.
Read also: More memory, less space: AMD unveils Versal Premium Gen 2 Memory on Package adaptive SoCs
Agent systems can access a much larger portion of a user’s work context than a typical chatbot. They may need to work with local files, messages, account data, and content from external sources. Therefore, privacy and resilience are fundamental components. Meta Muse Glimmer 30B is designed with a focus on security. It is trained to minimize the over-sharing of information, resist prompt injection from untrusted content, and adhere to information access boundaries. Local deployment gives developers greater control over where their data is processed and how the model interacts with other components of the workflow.
The Apache 2.0 license means that developers can use Muse Glimmer 30B for commercial purposes, modify it, distribute it, and integrate it with the tools and open-source software frameworks they already use.

LM Studio provides regular users with a simple and straightforward way to try out Meta Muse Glimmer locally. On compatible AMD systems (with 32 GB+ of video memory or VGM unified memory), the model can be downloaded and run in a matter of minutes. Compatible OpenAI and Anthropic endpoints are provided for connecting to agents such as Hermes Agent or Open Claw.
However, simply running Meta Muse Glimmer 30B isn’t enough; you also need to integrate it into your application. Lemonade integrates Muse Glimmer 30B into applications via a local API binary file that’s about 4 MB in size, storing the model and data on the device without requiring a separate installation. Lemonade can transform Muse Glimmer 30B from an agent into an app feature. This opens up possibilities for on-device programming assistants, research tools, and workflow automation.
Read also: From Car Collisions to Boat Behavior: How AI Is Learning to Better Understand Real-World Physics




