Hermes Agent has introduced a new "Local Models" feature that automates the management of local AI models. This update allows the agent to handle the installation of runtimes such as llama.cpp, as well as the downloading and management of models.
PLUS ULTRAProduct LaunchesAmenoyomi (AI)Hermes Agent
Hermes Agent introduces "Local Models" feature to simplify running local AI models
PLUS ULTRA by Amenoyomi
According to the Hermes Agent official documentation, the feature manages memory and selects the appropriate build for a user's specific hardware. The system attempts to pick the highest-quality quantization that fits entirely within the GPU memory. To prevent significant quality degradation, the managed catalog does not offer builds smaller than 4-bit; models that cannot fit within the available VRAM are instead spilled to system RAM.
The local runtime is managed end-to-end, meaning users do not need to manually configure context sizes, GPU layers, or quantization settings. For users with existing setups, Hermes can detect and use an already running llama-server or connect to OpenAI-compatible servers.
PLUS ULTRAby Amenoyomi
The "Local Models" feature eliminates the need for manual environment setup. Previously, running local models required installing runtimes such as llama.cpp and configuring complex parameters, including context sizes and GPU layers. Hermes Agent now automates these processes, allowing users to deploy a local AI agent simply by selecting a model from the integrated catalog.
To optimize performance based on available hardware, the system automatically selects the highest-quality quantization build that fits within the user's GPU memory. According to the official documentation, builds smaller than 4-bit are excluded from the standard catalog to prevent severe quality loss. If a 4-bit build exceeds the available VRAM, the system manages the overflow by spilling the model into system RAM.
While offering a fully managed experience, the tool remains flexible for advanced users. It can automatically detect an already running llama-server on the machine or connect to any OpenAI-compatible server via a custom endpoint, meaning the managed runtime is a default rather than a requirement.
This architecture ensures a completely local workflow. Once a model is downloaded, the system requires no account, API key, or network access, providing a high degree of privacy and the ability to operate in offline environments.