Qdrant memory provider for Hermes
Self-hostable vector memory for Hermes Agent, backed by Qdrant. No cloud dependency, no embedding API key, no per-query cost.
Replaces Hermes' built-in memory with a Qdrant collection: turns are embedded locally and stored as points, and recall is a dense vector search filtered by session.
What it does
- Dense semantic search over a named
densevector. - Session-scoped recall through a
session_idpayload index. - Verbatim turn storage — each turn is embedded as one point
(
user: … assistant: …). There is no LLM summarization step, so recall is cheap and quote-faithful. - Circuit breaker — 5 consecutive I/O failures open a 120s cooldown. A Qdrant outage degrades to "no new memories", not a crash.
- Honest availability —
is_available()/check_backend()/unavailable_reason()distinguish a bad config from a dead server, sohermes memory statuscan tell you which one you have. - 5 agent tools:
qdrant_search,qdrant_upsert,qdrant_recall,qdrant_collect(read-only),qdrant_prepare(reports the model's dimensions and cache location before you commit to a collection). - Progress display — optional status events during memory operations.
When enabled, the CLI/TUI shows
💾 qdrant — stored (127,778 points)after each turn and💾 qdrant — recalled 3 memoriesafter each retrieval. Controlled by theprogressconfig key:off,minimal(default),verbose(start + completion events).
Requirements
-
Hermes
>=0.21.4 -
Python
>=3.11 -
A reachable Qdrant. Quickest local option:
docker run -p 6333:6333 qdrant/qdrantOr Qdrant Cloud (set
api_key). -
Dependencies are managed by the repo's
pyproject.toml:package constraint why qdrant-client>=1.10.0,<21.10.0 is the oldest floor with query_points/Prefetch/FusionQueryfastembed>=0.4.0,<1default embedder — ONNX, no torch, ~287 MB peak RSS sentence-transformers>=2.7.0,<7opt-in GPU backend — heavy, see below
torch comes in transitively and is intentionally not pinned here.
Size warning.
sentence-transformersis the heavy dependency: it pullstorchplus the CUDA wheels — measured at 1.2 GB oftorch, 3.2 GB ofnvidia-*and 895 MB oftritonin a CUDA build, ~5.3 GB total. The defaultfastembedbackend needs none of it. The wide>=2.7.0,<7span is deliberate — a tight pin would freeze users onto an old dependency and miss upstream security fixes. If you only intend to use a remote Qdrant endpoint, you can skip the local embedder entirely; see theembedderconfig key above.
If you uninstall torch, it can come back silently. Removing
torch(or thenvidia-*/tritonwheels) is safe while the defaultfastembedbackend is in use — nothing in the runtime imports torch. But if any other package later pulls in a torch dependency, pip/uv will reinstall the full CUDA build and give back all ~5.3 GB without saying why. Selectingembedder: sentence-transformersafter such a removal fails loudly with aModuleNotFoundErrornaming torch; the fix is to reinstall it (uv pip install torch, ortorch --index-url .../cu124for the CUDA build) rather than to debug the plugin.
Keep torch in your development/test venv — do not remove it. A runtime install is well served by dropping torch entirely (the default
fastembedbackend never imports it, and neither does the plugin at import time). A development venv is not: the test suite exercises thesentence-transformersbackend on purpose, both to prove vector parity with the stored collection and to prove that selecting it fails loudly when torch is absent. Without torch in the dev venv those tests either skip or pass vacuously, and the guarantee that a backend swap needs no re-embedding loses its only guard.Practically: keep torch in whichever venv runs
pytest(here the 3.11 dev venv, with a CPU-only build), and omit it from the venv the gateway serves from. The plugin's own imports are backend-agnostic — neithertransformersnortorchis loaded byimporting the plugin or by taking thefastembedpath, so a torch-free runtime produces no warnings on its own.
Setup
Start a server first — the provider is REST-only and always needs a reachable Qdrant (there is no embedded mode):
docker run -p 6333:6333 qdrant/qdrant
Then install and configure in one step:
hermes plugins install qdrant
hermes memory setup # pick "qdrant" in the picker
The picker walks the config fields, probes the server, and sets
memory.provider: qdrant. To skip the picker — the provider name is a
positional argument, there is no --provider flag:
hermes memory setup qdrant
That path activates the provider and installs its dependencies, but it does
not prompt for settings. To configure by hand, create <HERMES_HOME>/qdrant.json
(created 0600) — ~/.hermes/qdrant.json for the default profile,
~/.hermes/profiles/<name>/qdrant.json for a profile:
{
"url": "http://localhost:6333",
"collection": "hermes_memories",
"vector_size": 384,
"distance": "Cosine"
}
This file holds no credentials. api_key is deliberately stripped in
save_config() (and any key an older version wrote is scrubbed on the next
save), so the API key reaches the provider only through QDRANT_API_KEY in
.env, read via Hermes' scoped-secret path.
Do not put it in the plugin directory. That directory is a build input: Hermes hashes every file in it into its dependency stamp, so any state written there re-syncs the venv on every launch.
Verify with hermes memory status. QDRANT_URL and QDRANT_API_KEY are
optional environment overrides — a local server on the default URL needs
neither.
There is no
hermes qdrantCLI. This plugin registers no CLI commands; everything is reached through the five agent tools andhermes memory status.
Configuration
| key | type | default | meaning |
|---|---|---|---|
url |
string | http://localhost:6333 |
Qdrant server (REST) |
api_key |
secret | "" |
Qdrant Cloud only |
collection |
string | hermes_memories |
collection name |
vector_size |
integer | 384 |
must match your embedder |
distance |
select | Cosine |
Cosine, Dot, Euclid |
embedder |
select | fastembed |
fastembed (default, no torch) or sentence-transformers (GPU, needs torch) |
model |
string | sentence-transformers/all-MiniLM-L6-v2 |
must match the backend's catalog; changing it changes vector space |
device |
select | cpu |
auto, cpu — cuda requires the sentence-transformers backend |
Resolution order, lowest to highest: built-in defaults → config.yaml's
memory.qdrant → <HERMES_HOME>/qdrant.json → QDRANT_URL / QDRANT_API_KEY
from the environment. Secrets are read through Hermes' scoped-secret path and
are never written into config.yaml.
Embeddings
Local, at 384 dimensions, via one of two backends:
| backend | requires | when to pick it |
|---|---|---|
fastembed (default) |
fastembed only — ONNX, no torch |
CPU hosts; ~287 MB peak RSS, no torch install |
sentence-transformers |
torch + CUDA wheels (~5.3 GB) |
GPU hosts needing CUDA, or models outside fastembed's catalog |
Both are pinned to sentence-transformers/all-MiniLM-L6-v2 by default, and
they produce identical vectors for that model (measured cosine 1.0000),
so switching backends does not require re-embedding an existing collection.
The first call downloads the model (~90 MB via the Hugging Face Hub for the
sentence-transformers backend; fastembed keeps an ONNX export in
<hermes home>/state/qdrant/model_cache, shared by every profile,
FASTEMBED_CACHE_PATH overrides it) — that cold-start download is the warning
you will see. The path is pinned deliberately: fastembed's own default is
$TMPDIR/fastembed_cache, and Hermes reaps scratch entries idle for 24 hours,
which re-downloaded the weights on every prune, once per distinct $TMPDIR.
qdrant_prepare prints the directory it will load from, so a cache miss is
reported before it costs you a download. Nothing leaves your machine.
Choosing a different model changes the vector space and invalidates
every existing point; re-embed or start a new collection. qdrant_prepare
reports the model's dimensions and cache location before you commit to that.
If you change vector_size away from 384 you must also supply a matching
embedding model; the mismatch surfaces as a write error, not a config error.
Back up, restore, move
Memories live in the Qdrant collection — the plugin keeps no local state
besides <HERMES_HOME>/qdrant.json (and the qdrant-status.json bookkeeping
next to it), so moving or backing up means moving the server's data:
- Snapshot (recommended):
POST /collections/<name>/snapshots(dashboard UI does the same) writes a snapshot file including vectors — restoring it needs no re-embedding. Measured: ~408 MB for a 127k-point collection. - Docker volume: the default
qdrant/qdrantimage stores everything under/qdrant/storage— copy the volume, ordocker cpit out, and the whole memory moves with it. hermes backupdoes not carry your points. This provider reports nobackup_paths(), sohermes backupcaptures config only; back the server up separately (snapshot or volume).
To verify a restore: qdrant_collect(action="info") shows the point count,
and a session recall should return the same lines as before the move.
Not implemented
Documented here so nobody has to read the source to find out:
- Lifecycle hooks — none. The provider takes
session_idon every call (sync_turn(..., session_id=...), andsession_idis a required parameter ofqdrant_recall), so there is no session state to rebind and nothing to flush at a session boundary.on_pre_compresswould be a real addition — a digest point written before the transcript is discarded — and is not built. - Hybrid dense+sparse RRF search and INT8 scalar quantization exist in
_backend.pybut are not wired into the provider — the provider talks toQdrantClientdirectly. Unreachable today. - gRPC transport is not supported; the client is REST.
- No collection pruning. No TTL, no dedup, no pruning job — a long-lived collection grows unbounded.
- No LLM extraction. Memories are verbatim turns, not summaries.
Troubleshooting
hermes update says qdrant is "configured but not installed" / "not in catalog" — known cosmetic false negative: the catalog check runs before user plugins register. Do NOT reinstall or switch memory.provider. hermes update touches the core venv and checkout only; it does NOT touch ~/.hermes/plugins/qdrant/, <HERMES_HOME>/qdrant.json, or the Qdrant server data. Confirm health with:
hermes plugins list # qdrant shows enabled
hermes memory status
unavailable_reason() mentions a missing dependency — install the
packages above in the same interpreter Hermes runs from. The default backend
needs only fastembed; a ModuleNotFoundError for torch means you selected
sentence-transformers on a host where torch was removed (see the size note).
It says available but recall returns nothing — the server is up but the
collection is empty or empty-filtered; check qdrant_collect(action="info").
Write errors after changing vector_size — embedder dimension mismatch.
Leave it at 384 unless you have changed the embedding model too.
First recall is slow — that is the one-time model download plus load, not the query path.
Development
uv venv && uv sync # dependencies only — this is a directory plugin
uv run pytest -q # live-server tests self-skip without a Qdrant on :6333
The repo root is the plugin directory, so tests/conftest.py imports the
shipped files as plugins.memory.qdrant and stubs the three Hermes core symbols
the provider needs. Point HERMES_SOURCE at a checkout of
hermes-agent to run the same
tests against the real core instead of the stubs.
Platform support
Linux and macOS on x86_64 and arm64. Windows x86_64 works; Windows on ARM
does not — grpcio (pulled in by qdrant-client) has no win_arm64 wheel.
License
MIT
