跳到主要内容

cloudflare-workers-ai-provider

❖ Communityv1.0.0★ 0

Cloudflare Workers AI as a Hermes model provider (`--provider workers-ai`, or Cloudflare Workers AI in `hermes model`, which asks for the token and the account's base URL): the 16 Workers AI models with function calling and 64K+ context, Cloudflare's context windows (a custom endpoint lists no models and assumes 256K for the 128K gpt-oss), reasoning effort in the form each model accepts, and a clear message when a free-plan account picks a paid-plan model. Not affiliated with Cloudflare.

Open in Hermes Desktop
hermes plugins install cloudflare-workers-ai-provider

What it adds

Environment variables it needs 2

CLOUDFLARE_WORKERS_AI_API_TOKENCLOUDFLARE_WORKERS_AI_BASE_URL

README

From the reviewed commit f6821d0 ↗; it updates when the author re-pins.

Cloudflare Workers AI provider for Hermes

Use Cloudflare Workers AI models in Hermes Agent as a first-class provider: pick Cloudflare Workers AI in hermes model, or run hermes chat --provider workers-ai -m @cf/openai/gpt-oss-120b.

An independent community plugin, not affiliated with Cloudflare.

Why a plugin?

Measured on 2026-09-29 against Cloudflare's live API on the Workers Free plan, with Hermes v0.21.5 and main:

Without the plugin With it
Choosing Workers AI as a provider Unknown provider 'cloudflare-workers-ai' Listed in hermes model
Model list, through a custom endpoint Empty: Cloudflare's OpenAI-compatible /models answers HTTP 405 The 16 Workers AI models that have function calling and at least 64K of context
Context window, through a custom endpoint gpt-oss-120b is assumed to have 256K (it has 128K), so Hermes compresses too late Cloudflare's own values, 128K to 1.3M
qwen3.8-27b with reasoning high, through a custom endpoint The turn fails: HTTP 400 Unexpected reasoning effort high. Supported types are xhigh (default), medium, and low Sent as medium
A paid-plan model on the free plan "custom rejected your API key … Update it" "Cloudflare Workers AI rejected the request and retrying won't help. Pick another model", followed by Cloudflare's "not available on the Workers Free plan"

Install

hermes plugins install cloudflare-workers-ai-provider

In the Cloudflare dashboard, open AI → Workers AI → Use REST API. It shows your Account ID and a Create a Workers AI API Token button. Then choose Cloudflare Workers AI in hermes model: it asks for the token and then for the base URL, which is https://api.cloudflare.com/client/v4/accounts/<your account ID>/ai/v1, and saves both to ~/.hermes/.env. You can also add them there yourself, with the account ID written out (Hermes does not expand ${...} in .env):

CLOUDFLARE_WORKERS_AI_API_TOKEN=<your token>
CLOUDFLARE_WORKERS_AI_BASE_URL=https://api.cloudflare.com/client/v4/accounts/<your account ID>/ai/v1

The plugin uses its own variable names so that a Cloudflare Pages token in CLOUDFLARE_API_TOKEN (which Hermes' publish-site skills use) is never sent to Workers AI. To pick the model in ~/.hermes/config.yaml instead of hermes model:

model:
  provider: workers-ai
  default: "@cf/openai/gpt-oss-120b"

Models

From Cloudflare's model catalog on 2026-09-29: the models with function calling and at least 64K of context. The Workers Free plan includes 10,000 Neurons a day; the last seven models need Workers Paid.

Model Context Vision Reasoning effort sent Plan
@cf/openai/gpt-oss-120b, @cf/openai/gpt-oss-20b 128K low, medium, high Free
@cf/qwen/qwen3.8-27b 262K ✓ low, medium, xhigh Free
@cf/zai-org/glm-4.7-flash 131K — Free
@cf/google/gemma-4-26b-a4b-it 256K ✓ — Free
@cf/nvidia/nemotron-3-120b-a12b 256K — Free
@cf/meta/llama-4-scout-17b-16e-instruct 131K ✓ — Free
@cf/mistralai/mistral-small-3.1-24b-instruct 128K — Free
@cf/ibm-granite/granite-4.0-h-micro 131K — Free
@cf/deepseek-ai/deepseek-v4-flash-0731 1.3M none, low, high, max Paid
@cf/deepseek-ai/deepseek-v4-pro-0813 1M none, low, high, max Paid
@cf/moonshotai/kimi-k2.6 262K ✓ none, high Paid
@cf/moonshotai/kimi-k2.7-code 262K ✓ — Paid
@cf/zai-org/glm-5.2 262K none, high, max Paid
@cf/zai-org/glm-5.3, @cf/zai-org/glm-5.3-flash (vision) 1.3M low, high, max Paid

Reasoning effort follows Hermes' own rule. A level the model does not list goes to the nearest weaker level it does list. A level below the model's lowest goes to that lowest level. With reasoning turned off, a model that lists none gets none, and any other model gets its lowest level: gpt-oss always reasons, qwen3.8 would otherwise run at its xhigh default, and Cloudflare's catalog says glm-5.3 turns none into max. Models marked — get no reasoning_effort field. With no reasoning effort configured, nothing is sent and the model uses its default. The free-plan values were checked against the live API; the paid-plan values come from Cloudflare's catalog.

Security and footprint

  • Registers one model provider, workers-ai, when Hermes loads it. No tools, hooks, commands or background threads.
  • Network: chat requests go to the URL in CLOUDFLARE_WORKERS_AI_BASE_URL through Hermes' own client, carrying CLOUDFLARE_WORKERS_AI_API_TOKEN. The plugin itself opens no connections. The model list and model details are built into the plugin, so nothing is fetched to show them.
  • Reads no files, writes no files, starts no processes and has no dependencies beyond Hermes. Hermes, not the plugin, reads the two variables, per profile.

Compatibility

Hermes 0.21.5 or newer. CI runs the tests against Hermes v2026.9.24 (0.21.5) and main every day.

Changelog

1.0.0

  • The workers-ai model provider: the Workers AI models Hermes can run, with Cloudflare's context windows, per-model reasoning effort, and a clear message when a free-plan account picks a paid-plan model.

License

MIT

groq-provider❖ Community★ 0

Groq as a Hermes model provider (`--provider groq`, or Groq in `hermes model`): reasoning effort is sent in the form each Groq model accepts, so gpt-oss-120b and gpt-oss-20b work (through a custom endpoint, Hermes 0.21.5 sends them a value Groq rejects with HTTP 400), and the model list leaves out Groq's speech-to-text, text-to-speech, prompt-guard and 4K-context models that Hermes cannot run. Not affiliated with Groq. Disclosure — registers one model provider at import and nothing else (no tools, hooks, commands or threads); chat requests and the model list go to https://api.groq.com/openai/v1 (or the configured model.base_url) through Hermes' own client with GROQ_API_KEY, which Hermes reads; the plugin itself opens no connections, reads and writes no files and starts no processes.

Models
memory-rewind❖ Community★ 0

Automatic version history for the built-in memory (MEMORY.md, USER.md), SOUL.md and skills/ (including curator-archived skills): browse, diff and restore any past version with `hermes memory-rewind`, or view it from any chat with the read-only `/memory-history`. Works alongside the built-in memory; not a memory provider. Disclosure — runs the local git binary (isolated from the user's git config, hooks and signing) to keep a bare repository in $HERMES_HOME/plugin-data/memory-rewind/; reads only the tracked files (never .env, auth.json, databases or skills/.hub/); rewrites tracked files only when the user runs `restore`; no network.

General
ntfy-approval❖ Community★ 0

Answer Hermes' approval prompts from your phone: once selected with security.approval.transport: ntfy, each dangerous-command prompt arrives as an ntfy push notification with Approve once, Approve for session and Deny buttons, and no answer still means deny. Disclosure — once selected, anyone who can read the ntfy topic can approve dangerous commands (the one-time answer codes are in the notification), so on the public ntfy.sh default the topic name is the only secret and NTFY_APPROVAL_TOKEN is optional; each pending approval publishes Hermes' reason, the redacted command (unless send_command is false), the non-default profile name and one-time answer codes to the configured ntfy server, then reads <topic>-reply and deletes the notification; any failure or timeout is a deny; reads NTFY_APPROVAL_TOPIC/NTFY_APPROVAL_TOKEN from the profile .env, never puts the token in a notification, follows no redirects, writes no files and starts no processes.

Automation
reaction-feedback❖ Community★ 0

Lets the agent see the emoji reactions you leave on its Telegram messages: react 👍 or 👎 to a reply and the agent's next turn in that private chat starts with a short note naming the reaction and the message it followed. Disclosure — registers pre_gateway_dispatch (always returns None, so dispatch is never changed), pre_llm_call and gateway_platform_event hooks; keeps a bounded in-memory record of private-chat message ids, 60-character excerpts and reactions; writes nothing to disk, makes no network requests, starts no processes and reads no secrets.

General

← Back to the catalog · catalog built Oct 3, 2026