跳到主要内容

groq-provider

❖ Communityv1.0.0★ 0

Groq as a Hermes model provider (`--provider groq`, or Groq in `hermes model`): reasoning effort is sent in the form each Groq model accepts, so gpt-oss-120b and gpt-oss-20b work (through a custom endpoint, Hermes 0.21.5 sends them a value Groq rejects with HTTP 400), and the model list leaves out Groq's speech-to-text, text-to-speech, prompt-guard and 4K-context models that Hermes cannot run. Not affiliated with Groq.

Open in Hermes Desktop
hermes plugins install groq-provider

What it adds

Environment variables it needs 1

GROQ_API_KEY

README

From the reviewed commit b5be038 ↗; it updates when the author re-pins.

Groq provider for Hermes

Use Groq's models in Hermes Agent as a first-class provider: pick Groq in hermes model, or run hermes chat --provider groq -m openai/gpt-oss-120b.

An independent community plugin, not affiliated with Groq.

Why not a custom endpoint?

Hermes can already point a custom OpenAI-compatible endpoint at Groq. Measured on 2026-09-29 against Groq's live API with Hermes v0.21.5 and main:

Custom endpoint This plugin
openai/gpt-oss-120b, openai/gpt-oss-20b Every turn fails, whatever the reasoning setting: Hermes sends reasoning_effort: "default" ("none" when reasoning is off) and Groq answers HTTP 400 (reasoning_effort must be one of low, medium, or high) Works, and your reasoning setting is sent in the form the model accepts
Model list All 11 models Groq returns, including speech-to-text, text-to-speech, prompt-guard and 4K-context models that Hermes cannot run The 4 models Hermes can run

An open Hermes pull request, #124103, would send gpt-oss-120b and gpt-oss-20b a graded level on custom endpoints. The plugin does not depend on it and works the same with or without it.

Install

hermes plugins install groq-provider

Get an API key at console.groq.com/keys and add it to ~/.hermes/.env:

GROQ_API_KEY=gsk_...

Then choose Groq in hermes model, or set it in ~/.hermes/config.yaml:

model:
  provider: groq
  default: openai/gpt-oss-120b

Models and reasoning effort

The model list comes from Groq's API, minus the models Hermes cannot run: whisper (speech-to-text), orpheus (text-to-speech), prompt-guard (a 512-token classifier) and allam-2-7b (a 4,096-token window; Hermes needs at least 64K). On 2026-09-29 that left:

Model Reasoning effort Hermes sends
openai/gpt-oss-120b, openai/gpt-oss-20b, openai/gpt-oss-safeguard-20b low, medium or high. These models always reason, so none (off) and minimal become low; xhigh, max and ultra become high
qwen/qwen3.8-27b none, low, medium or high. minimal becomes low; anything above high becomes high
Any other model Nothing: Groq rejects the field on models without reasoning

With no reasoning effort configured, nothing is sent and Groq uses the model's default. A level is only ever lowered to the nearest one the model accepts, never raised.

Groq's free plan

Groq's free plan allows 8,000 tokens per minute per model (6,000 for allam-2-7b), as reported in its rate-limit headers on 2026-09-29. One Hermes request with tools is larger than that: about 18.8K tokens with the default toolsets, and 10.7K to 11.7K with a single toolset. On the free plan Groq therefore rejects each agent turn with HTTP 413 (Request too large ... on tokens per minute (TPM)), which Hermes reports as the conversation having grown too large to send. A turn without tools fits. For agent work you need Groq's paid Developer plan. This is Groq's limit; the plugin does not change the size of a request.

Security and footprint

  • Registers one model provider, groq, when Hermes loads it. No tools, hooks, commands or background threads.
  • Network: chat requests and the model list go to https://api.groq.com/openai/v1 (or the model.base_url you set) through Hermes' own client, carrying GROQ_API_KEY. The plugin itself opens no connections.
  • Reads no files, writes no files, starts no processes and has no dependencies beyond Hermes. Hermes, not the plugin, reads GROQ_API_KEY.

Compatibility

Hermes 0.21.5 or newer. CI runs the tests against Hermes v2026.9.24 (0.21.5) and main every day.

Changelog

1.0.0

  • The groq model provider: per-model reasoning effort, and a model list without the models Hermes cannot run.

License

MIT

cloudflare-workers-ai-provider❖ Community★ 0

Cloudflare Workers AI as a Hermes model provider (`--provider workers-ai`, or Cloudflare Workers AI in `hermes model`, which asks for the token and the account's base URL): the 16 Workers AI models with function calling and 64K+ context, Cloudflare's context windows (a custom endpoint lists no models and assumes 256K for the 128K gpt-oss), reasoning effort in the form each model accepts, and a clear message when a free-plan account picks a paid-plan model. Not affiliated with Cloudflare. Disclosure — registers one model provider at import and nothing else (no tools, hooks, commands or threads); the model list and details are built in, so nothing is fetched to show them; chat requests go to the URL in CLOUDFLARE_WORKERS_AI_BASE_URL through Hermes' own client with CLOUDFLARE_WORKERS_AI_API_TOKEN, both read by Hermes per profile; the plugin itself opens no connections, reads and writes no files and starts no processes.

Models
memory-rewind❖ Community★ 0

Automatic version history for the built-in memory (MEMORY.md, USER.md), SOUL.md and skills/ (including curator-archived skills): browse, diff and restore any past version with `hermes memory-rewind`, or view it from any chat with the read-only `/memory-history`. Works alongside the built-in memory; not a memory provider. Disclosure — runs the local git binary (isolated from the user's git config, hooks and signing) to keep a bare repository in $HERMES_HOME/plugin-data/memory-rewind/; reads only the tracked files (never .env, auth.json, databases or skills/.hub/); rewrites tracked files only when the user runs `restore`; no network.

General
ntfy-approval❖ Community★ 0

Answer Hermes' approval prompts from your phone: once selected with security.approval.transport: ntfy, each dangerous-command prompt arrives as an ntfy push notification with Approve once, Approve for session and Deny buttons, and no answer still means deny. Disclosure — once selected, anyone who can read the ntfy topic can approve dangerous commands (the one-time answer codes are in the notification), so on the public ntfy.sh default the topic name is the only secret and NTFY_APPROVAL_TOKEN is optional; each pending approval publishes Hermes' reason, the redacted command (unless send_command is false), the non-default profile name and one-time answer codes to the configured ntfy server, then reads <topic>-reply and deletes the notification; any failure or timeout is a deny; reads NTFY_APPROVAL_TOPIC/NTFY_APPROVAL_TOKEN from the profile .env, never puts the token in a notification, follows no redirects, writes no files and starts no processes.

Automation
reaction-feedback❖ Community★ 0

Lets the agent see the emoji reactions you leave on its Telegram messages: react 👍 or 👎 to a reply and the agent's next turn in that private chat starts with a short note naming the reaction and the message it followed. Disclosure — registers pre_gateway_dispatch (always returns None, so dispatch is never changed), pre_llm_call and gateway_platform_event hooks; keeps a bounded in-memory record of private-chat message ids, 60-character excerpts and reactions; writes nothing to disk, makes no network requests, starts no processes and reads no secrets.

General

← Back to the catalog · catalog built Oct 3, 2026