🎙️hermes-gemini-live❖ Community
Realtime duplex voice with the Gemini Live API, built around handing real work to Hermes. A single control in the composer (not a floating panel): tap, speak, hear the model, and when the ask needs files, a terminal, the web or memory the voice model delegates the job to the api server and keeps talking, reading the outcome aloud when it lands. Three verbs rather than thirty schemas — hermes_task starts, hermes_tasks reads the call's task board, and hermes_task_update denies, steers or stops a task. A task that stops to ask permission reaches the user as a spoken prompt and a "Hermes needs you" indicator with Approve and Deny buttons; only the user's press approves it, never the voice model. Multi-task is one Live session plus a board, not a session per task. Mute drops the uplink without hanging up, and the call survives composer remounts, so opening a session mid-sentence does not kill it. Full-proxy architecture: the plugin backend owns the upstream Live socket and the key, and the renderer is never given either, because the Live API carries the credential in the socket URL. Ships no Python dependencies of its own, which keeps it out of the uv workspace and lets one clone serve several profiles. Tested model behaviour is published in the README, including that gemini-3.8-live-extended-thinking accepts a function declaration and then delegates only about 2 times in 5, which is why gemini-3.8-live is the default. Disclosure — pressing Start captures microphone audio and streams it through the local backend to generativelanguage.googleapis.com using GEMINI_API_KEY from the Hermes home's .env; delegated work runs as a Hermes run in the serving profile's own memory with the profile's full toolset via the local api_server and API_SERVER_KEY, and those runs' results (up to 4000 characters each), the task board and approval prompts are sent back to Gemini as text. Nothing is sent anywhere else and no telemetry is emitted.
Voice