Skip to main content

fish-audio

❖ Communityv1.3.1

The official Fish Audio plugin, maintained by Electric Sheep: expressive text-to-speech in 80+ languages, speech-to-text, Fish Audio's community voice library, voice cloning and voice design, with a Voices page in Hermes Desktop. Fish Audio becomes a Hermes speech and transcription provider for voice replies, read-aloud, voice notes and Desktop dictation, with sentence-by-sentence streaming in CLI voice mode and Desktop. Adds the fish_speak, fish_voices and fish_transcribe tools, /fish and hermes fish commands, and three skills.

Open in Hermes Desktop
hermes plugins install fish-audio

Screenshots

What it adds

Tools 3

fish_speakfish_voicesfish_transcribe

Hooks 2

transform_llm_outputpre_tool_call

Environment variables it needs 1

FISH_API_KEY

README

From the reviewed commit 6f2497e ↗; it updates when the author re-pins.

Fish Audio for Hermes

Fish Audio for Hermes

The official Fish Audio plugin for Hermes Agent, maintained by Electric Sheep.

Give your agent a voice: expressive speech in 80+ languages, Fish Audio's library of community voices, your own cloned or designed voices, and accurate speech-to-text. It works in the Hermes CLI, every messaging app Hermes supports and the Desktop app.

hermes plugins install https://github.com/100yenadmin/fish-audio-hermes --enable
hermes fish login        # paste your key from https://fish.audio/app/api-keys

Hermes 0.21.5 installs the plugin's Python dependencies automatically. Newer Hermes builds ask first; add --yes-deps there to answer yes. hermes fish login checks the key against your Fish account, saves it to the active profile. It makes Fish Audio the speech and transcription provider where none is set, and replaces another provider only if you confirm at the prompt or pass --yes. Once Fish Audio is the provider, voice replies, read-aloud and voice notes go through it.

What you get

Expressive speech Fish Audio S2 models follow emotion and delivery cues such as [excited], [whispering] and [laughing]. Multi-speaker dialogue and custom pronunciations are supported too.
Streaming voice CLI voice mode and Desktop voice replies stream each sentence as it is ready. In our tests the first audio bytes arrived 0.24–0.38 s after a sentence was sent (see Streaming voice for how this was measured).
Voice library Search Fish Audio's community voices from chat (/fish voices narrator) and switch with /fish use <id>.
Your own voices Clone a voice from a short sample, or describe one in words and pick from the candidates. Cloning and deleting from chat go through Hermes's approval gate, so by default Hermes asks you first.
Speech-to-text Voice notes and Desktop dictation are transcribed by transcribe-1-pro. fish_transcribe adds speaker turns, timestamps and SRT subtitles.
Desktop Voices page In Hermes Desktop: browse and preview voices, pick one for the agent, clone or design your own, and check your credit.
Account at a glance /fish balance shows your API credit and plan. When credits run out, tools and /fish commands reply with a top-up link; for voice replies, /fish status shows the last error.

Setup

  1. Get a key at https://fish.audio/app/api-keys. New accounts can start on the free s2.1-pro-free model.
  2. Add it with hermes fish login on the machine running Hermes, or in Desktop ▸ Settings ▸ Plugins ▸ Fish Audio (Capabilities ▸ Plugins on older Desktop), or in hermes tools ▸ Text-to-Speech ▸ Fish Audio. login checks the key against your Fish wallet and saves it to the active profile. It makes Fish Audio the speech and transcription provider where none is set; it replaces another provider only if you confirm at the prompt or pass --yes. --key-stdin reads the key from standard input.
  3. Check it with hermes fish doctor, which makes one short billed synthesis (--no-synth skips it). Restart a running gateway or the Desktop app to pick up a new key.

Never paste API keys into a chat. /fish refuses them and tells you to rotate the key.

hermes fish doctor: key, wallet, providers, Hermes version, a billed TTS round trip and STT all pass

Use

In conversation: just ask.

  • "Read that back to me in an excited voice."
  • "Find me a deep British narrator voice and use it."
  • "Design a voice for our support agent and let me pick."
  • "Transcribe meeting.m4a with speakers and give me subtitles."

Commands

Command What it does
/fish status Key, providers, voice, model and why it was chosen, credit, last error
/fish voices [query] Search the voice library
/fish use <voice-id> Use a voice for this profile
/fish model <id> Pick a model (s2.1-pro, s2.1-pro-free, s2-pro, s1, drama-3-preview)
/fish preview <voice-id> [text] Hear a voice (a billed synthesis)
/fish balance API credit and plan
hermes fish login [--key-stdin] [--yes] Save and check a key, offer to switch providers
hermes fish status · hermes fish use <voice-id> Status and voice choice from your terminal
hermes fish doctor [--no-synth] Check the setup with one billed synthesis, or without it

Model tools (toolset fish_audio; calls are billed to your Fish Audio account)

Tool What it does
fish_speak Expressive or multi-speaker speech; with timestamps: true, word timings plus SRT and WebVTT subtitles
fish_voices Search, list and inspect voices; clone, design, save, update and delete them
fish_transcribe Rich transcription with speakers, timestamps and SRT

Plain read-aloud uses Hermes's own text_to_speech tool, with Fish Audio as the provider. Include the media_tag that fish_speak returns in the reply to deliver its audio; the plugin also appends missing audio tags through its output hook. Plugin skills: fish-audio:fish-audio-setup, fish-audio:fish-audio-expressive-speech and fish-audio:fish-audio-voice-studio.

Hermes Desktop

The Voices page in Hermes Desktop: the Fish Audio voice library searched for "narrator", with Preview, Use and favourite on each voice

The Voices page follows Hermes Desktop's display language: English, Simplified Chinese (zh), Traditional Chinese (zh-hant), Japanese (ja), Arabic (ar), Russian (ru), French (fr), German (de) and Spanish (es). Changing the display language updates the page, sidebar, status bar and command palette without a reload. The translations are machine-authored and checked against Hermes Desktop's own wording; native-speaker corrections are welcome as pull requests to src/desktop/locales/.

The plugin adds a Voices page to Hermes Desktop for the selected agent:

  • Library: search Fish Audio's voices by name and language, preview them (a short billed sample), star favourites, and choose Use to make a voice the agent's voice.
  • My voices: the voices you cloned or designed; delete one by typing its name.
  • Create: clone a voice from 1–3 recordings (up to 10 MB each, with the speaker's permission), or describe a voice and save the candidate you like.
  • Account: API credit, plan credits and top-up links. The status bar shows your API credit, in orange when it runs low.
Create Account
Create: design a voice from a description and save a candidate, or clone one from recordings with the speaker's consent Account: API credit, plan credits, top-up and plan links

The page has two halves in one package: routes that run on the agent's gateway and screens that run in Hermes Desktop. To install from Desktop:

  1. Open Capabilities → Plugins → Install from Git, enter https://github.com/100yenadmin/fish-audio-hermes, choose Review repository, then Install. Desktop installs the agent half into the connected agent and the desktop half on this computer.
  2. In Capabilities → Plugins, open Fish Audio and switch on Desktop. Desktop plugins stay off until you turn them on.
  3. Restart the agent's gateway once. Hermes mounts plugin routes only at startup.

If you installed the plugin with hermes plugins install on the gateway machine, do steps 1 and 2 on each computer that runs Desktop, and restart the gateway once after upgrading to 0.3.0. The Voices row appears only for agents whose gateway has the plugin; an agent without a Fish Audio key shows a card that links to the key page and the plugin's settings.

Configuration

All settings are per profile.

tts:
  provider: fish-audio
  fish-audio:
    voice: 933563129e564b19a115bedd57b7406a   # any Fish voice id (this one: "Sarah", a conversational narrator)
    model: s2.1-pro                           # optional; see "Which model?"
    latency: balanced                         # normal | balanced | low
    temperature: 0.7
stt:
  provider: fish-audio
  fish-audio:
    model: transcribe-1-pro
plugins:
  entries:
    fish-audio:
      settings:
        streaming: auto        # auto | "off"
        allow_free_model: true # false = never use s2.1-pro-free

The plugins.entries.fish-audio.settings keys (base_url, streaming, transport, allow_free_model, operator_account) and the API key also appear in Desktop ▸ Settings ▸ Plugins ▸ Fish Audio (Capabilities ▸ Plugins on older Desktop). Set the voice and model from the Voices page, /fish use or /fish model.

Which model? A valid tts.fish-audio.model is used as set (for fish_speak, a model named in the call comes first); invalid ids are ignored. Otherwise the plugin uses s2.1-pro, and picks s2.1-pro-free only when Fish reports a wallet with no API credit and no past top-ups that doesn't carry Fish's free-credit flag. If the wallet can't be read, it uses s2.1-pro. When it picks the free model it tells you once that the model is free until 30 November 2026 and that Fish Audio may use free-tier requests to improve its models. allow_free_model: false replaces s2.1-pro-free with s2.1-pro everywhere, even when you name the free model yourself.

Streaming voice. With tts.provider: fish-audio, on a Hermes build that has the plugin streaming hook (Hermes main since 2026-10-06; the next release), the plugin streams 24 kHz PCM from POST /v1/tts to CLI and TUI voice playback, Desktop read-aloud and gateway streaming. Older Hermes builds, including 0.21.5, use whole-file synthesis instead: sentence by sentence in CLI voice mode, the whole reply in Desktop read-aloud and gateway voice. Each streamed sentence is one billed request and is not retried. Measured first audio bytes: Fish HTTP API p50 243 ms (3 samples, 238–356 ms), CLI voice mode 373–381 ms, Desktop speak-stream handler 310 ms (in-process, not through the Electron UI); these exclude the model's own reply time. Set streaming: "off" for whole-file speech; transport: ws switches to Fish's live WebSocket for diagnosis.

Under host plugin isolation (plugins.isolation: host, newer Hermes builds), the hermes fish terminal command is unavailable and voice replies use whole-file speech instead of streaming; /fish, the tools, the hooks and both providers still work. Config changes apply to the active profile; managed installs may refuse writes.

Operator-managed keys

Set plugins.entries.fish-audio.settings.operator_account: true (Desktop: Operator-managed account) when the Fish account belongs to the agent operator. This hides balance, plan, billing and API-key links in chat and Desktop, and directs setup or credit problems to the operator. Because one operator account can serve many agents, the account's own voices aren't listed, edited or deleted from an agent (no My voices tab; fish_voices refuses mine, update and delete); cloning and designing still work, and a voice created on the Voices page goes into that agent's Favourites. It defaults to false. Tool descriptions update after a gateway restart. Operators should also pin tts.fish-audio.model or set allow_free_model: false, so the model never depends on the wallet that users can't see.

Privacy and security

  • No telemetry. The plugin talks only to Fish Audio's API (api.fish.audio, or the base_url you configure), and only for the work you ask for. See SECURITY.md for exactly what each request carries. Links to Fish Audio are plain links, with no tracking parameters.
  • Keys come from the active profile's secret scope. When one Hermes process serves several profiles, each profile uses its own key; a single-profile install may also read FISH_API_KEY from the environment. Keys are never logged or put in error messages.
  • Voice fields: search accepts author_id, title_language and licensed, and returns boolean has_more. Design accepts num_step (1–128), guidance_scale and instruct_guidance_scale (finite, ≥ 0). Clone/update accept private/unlist visibility and cover_image_path (PNG/JPEG/WebP, ≤ 5 MiB); clone also accepts boolean generate_sample. Transcription returns language when Fish supplies a string.
  • Files the model tools upload (fish_voices clone samples, fish_transcribe recordings) must be regular audio files (covers must be images). They refuse symlinks, Hermes config and secret files, and SSH or cloud credential paths.
  • Approvals: cloning and deleting voices from chat or the model tools go through Hermes's approval gate and follow your Hermes approval settings. By default Hermes asks you, and refuses when no one is there to answer. The Desktop Voices page confirms them on the page instead (see Disclosure).
  • Report vulnerabilities privately; see SECURITY.md.

Disclosure

  • Synthesis, transcription, cloning and voice design are paid Fish Audio requests billed to your Fish account (the free s2.1-pro-free model excepted).
  • Fish Audio may use free-model requests to improve its models.
  • The Desktop Voices page adds gateway routes under /api/plugins/fish-audio/, behind the Hermes dashboard's existing authentication. On that page, previews, clones and voice designs are billed; cloning a voice or saving a designed one creates a voice in your Fish account; cloning and deleting are confirmed on the page (a speaker-consent box, a typed voice name) instead of through Hermes's approval gate; and Use sets the profile's Fish voice, and its speech provider when none is set. Preview and design audio files are deleted from the gateway once returned; a design's audio is held in gateway memory so you can save it, until it is saved, a later design or save finds it over an hour old, or the gateway stops; clone uploads are deleted after the clone, or by the next upload once idle for 15 minutes. With a key set, Desktop reads your Fish wallet when an agent is selected and about every five minutes, for the status-bar credit; favourites stay in Desktop's plugin storage on that computer, and fetched previews stay in the window's memory until it closes.

Compatibility

Hermes Agent 0.21.5 or newer, Python 3.11+, macOS and Linux. CI tests it against the latest Hermes release, Hermes main and the Electric Sheep fork. Windows is untested.

Troubleshooting

Message Fix
"Fish Audio isn't set up for this profile" hermes fish login
"Top up Fish Audio API credits" Top up at https://fish.audio/app/developers/billing. Plan credits and API credits are separate.
"Fish Audio concurrency limit reached" Fish allows 5 concurrent requests below $100 of lifetime top-ups, 15 from $100, and 50 from $1,000.
"The voice id was not found" The id is wrong or the voice is private to another account. Try /fish voices.
Voice replies stay silent Run /fish status; it shows the last Fish error, with its Fish trace id when Fish sent one.
Anything else Run hermes fish doctor, and quote the Fish trace id from the error when you contact Fish Audio support.

Known limitations

The auto-append fallback is per session. Mid-turn context compression can drop that fallback; the model normally includes the returned media_tag itself.

Core gives plugins a .mp3 path for plain text_to_speech and re-encodes it to Opus for voice bubbles (upstream #133133). Core's text normaliser also alters <|speaker:N|> markup on that path: use fish_speak for multi-speaker speech (upstream #133131).

License

Apache-2.0. Maintained by Electric Sheep with the Fish Audio team. Fish Audio is a trademark of Hanabi AI Inc.

boardstate❖ Community★ 3

A live board the agent builds for you: a Board tab in the Hermes dashboard and a Desktop page where the agent lays out tabs and widgets (notes, markdown, tables, charts, KPI cards, live Hermes data, approvals) through 19 native boardstate_* tools. Runs a local Node >= 20 sidecar on loopback with a per-spawn nonce; no network access unless you opt in to operator-gated connectors. Disclosure — spawns a local Node.js ≥20 sidecar (from PATH/HERMES_NODE_BIN, nothing downloaded) on 127.0.0.1 on first use; state under HERMES_HOME/boardstate-state/. Reads the dashboard's loopback session token from Hermes internals (web_server._SESSION_TOKEN) and passes it to the sidecar for live usage/sessions/cron widgets; mounts one token-gated static route for approved widget assets. No outbound network unless you author boardstate.connectors.json (stdio/HTTP MCP connectors); connector actions need operator approval.

Desktop
plan-mode❖ Community★ 1

Claude Code-style plan mode on every interactive Hermes surface: /plan or /planmode on blocks mutating Hermes tools dispatched through pre_tool_call (writes outside .hermes/plans, terminal, code execution, delegation, MCP) while the agent explores, asks clarifying questions and writes a plan. The agent submits the plan with plan_mode(submit) and you approve it through Hermes' own approval prompt (gateway buttons, the TUI/Desktop card, the CLI panel or /approve); it then implements in the same turn and tracks todo_list progress. Deny keeps planning and the agent revises. When yolo or approvals.mode: off would auto-approve, it asks one clarify question instead (Approve plan rev N / Keep planning), and only that answer approves. By default plans are compact and decision-complete and the agent does not commit unless asked (plan_style, allow_commits and plan_skill settings; plan_style: core keeps the zero-context /plan craft). Classic CLI enforced on every supported Hermes; gateway, TUI and Desktop from Hermes v2026.9.24 (0.21.5), and refused on older builds. Disclosure — plan-mode does not enforce Codex app-server exec/applyPatch or TUI /background and btw side agents. It reads seven internal Hermes seams for session identity, workspace, profile and skills.inline_shell, each guarded with a documented fallback. It creates .hermes/plans under the session working directory and keeps per-session state in plugin state. It adds a turn note while planning or executing, a ≤200-char system-prompt hint, and a one-line footer to short replies on chat platforms. When you pick "Always" on a plan approval, Hermes core writes a plugin_rule:plan-mode entry to command_allowlist in config.yaml. With allow_gateway_injection: true it queues one message to start work after a typed /planmode approve. No network calls, no subprocesses, no self-updater.

Tools
hermes-cloud-file-manager❖ Community★ 0

Cloud Files for Hermes Desktop — browse, search, create folders and bulk-upload files and whole folders to the machine your agent runs on, read and edit its Markdown and text files (view first; nothing is written until you review the line diff and press Save), import files from the agent's own Google Drive, and hand the agent file locations from the chat + menu without re-uploading. Built for remote gateways (token or sign-in); works locally too. Disclosure — adds a file API to the gateway (dashboard plugin routes under the gateway's existing auth) over the agent's working folders, by default terminal.cwd or else the gateway user's home with its dot-entries hidden. It lists and searches those folders, reads Markdown and text files, creates folders and writes uploads and Drive imports (staged as hidden .cfm-* temp files in the destination folder; a name clash keeps both files). It overwrites only an existing .md/.markdown/.txt file you open and explicitly Save in its editor (refused if the file changed since you opened it). It never deletes your files (it removes only its own stale or aborted .cfm-* temp files), never follows links out of those folders, and never lists, reads or writes the agent's Hermes home apart from a working folder set inside it (roots, terminal.cwd, or a workspace/ folder it creates when the working folder would be the home itself). Its only network use is Google Drive, read-only, through the agent's own google-workspace skill, whose scripts (from skills/productivity/google-workspace in the agent's Hermes home, else the copy bundled with Hermes) it runs with the gateway's Python and a minimal environment without API keys; the plugin never reads the Google token. While Desktop is open it re-checks Drive sign-in about once a minute; otherwise it makes no network calls.

Desktop
hermes-teammates❖ Community★ 0

Named subagent teammates for Hermes: the operator defines a roster (instructions, toolsets, model), and the agent assigns them work, checks results, steers a running teammate, follows up with a digest of earlier results, and can work a Kanban teammates lane with a review handoff. Launches subagents only through the public ctx.subagent_lifecycle API. Stock Hermes; no patched core. Disclosure — adds five tools and a /teammates command; the model picks only a teammate name, and the operator's config decides what it means. Teammates run as in-process Hermes subagents on the parent's provider credentials; toolsets narrow the parent's tools at toolset granularity only (the file toolset includes write_file and patch). Steering dispatches delegate_task's steer action with the turn-bound parent, so under plugins.isolation: host every tool except the roster refuses. It keeps a per-profile SQLite run ledger in its plugin-data folder (goals, statuses, clipped summaries; no transcripts or credentials). For a card assigned to its non-profile lane, it claims the card through the Kanban DB, heartbeats the claim, then moves the card to review without a reviewer, or blocks it on failure; it never completes a card. It runs one in-process monitor thread per lane run to keep the claim alive; no network calls, child processes or self-updates of its own.

Automation

← Back to the catalog · catalog built Oct 7, 2026