Back to home

chenkezhen480

dsh-semantic-memory

为deepseek-harness添加向量化跨会话记忆插件

Stars
1
Language
TypeScript
Created
Aug 15, 2026
Updated
Aug 15, 2026

Introduction

dsh-plugin-semantic-memory

中文README.zh.md,推荐) | English

Semantic long-term memory for DeepSeek Harness.

A dsh-plugin (Cordis plugin) that gives the model a persistent, embedding-based memory across sessions — unlike the harness's built-in session_query (literal FTS5), this store retrieves by meaning.

Zero-config out of the box: install, restart, and start a new session. The local embedding model downloads itself on first use (~100 MB); the four tools, per-question recall, and 5-turn auto-summarization all work with defaults. You only configure when you want something different (see Configuration).

Features

  • Cross-session semantic memory — facts, decisions, preferences, and notes persisted as JSONL under $DSH_HOME/memories/memories.jsonl.
  • Embedding retrieval — cosine similarity over normalized vectors; provider is pluggable:
    • local (default): ONNX inference via @huggingface/transformers with Xenova/bge-small-zh-v1.5 (offline, ~100 MB model, cached in ~/.cache/huggingface).
    • api: any OpenAI-compatible /embeddings endpoint (e.g. SiliconFlow, Zhipu, DashScope).
  • Memory decay & strengthening — each entry's effective strength halves over the configured half-life since its last access; searching an entry refreshes it. Importance (1–5) sets the base strength.
  • Model-facing tools:
    • memory_write — persist a fact / decision / preference / note (content hash dedup, repeats update in place).
    • memory_search — semantic top-k recall with kind/tag/workspace filters.
    • memory_forget — delete by id.
    • memory_stats — store summary.
  • Automatic injection — the plugin watches the session event stream: every new user message is embedded and searched asynchronously, and the freshest per-session recall is rendered into the system prompt before the turn's prompt assembly (question-aware). With no fresh recall yet, the strongest resident memories are injected as a fixed-size fallback.
  • Proactive writing guidance — the injected prompt tells the model to call memory_write on its own when the user states a durable preference, an established fact, or an explicit decision (no need to say "remember").
  • Auto-summarization — every N user messages (default 5), the plugin asks the harness LLM to distill the recent transcript into memory entries and writes them (tagged auto). One in-flight summary per session; silent on failure; only active when llm and agentDefaultModel services exist.
  • Workspace tagging — entries record the caller session's cwd; search scopes to that workspace by default and can opt into cross-workspace recall.

Install into a DSH profile

Official one-liner (the package ships an in-package cordis.patch.yml declared via dsh.bundle.patch, so dsh plugin add mounts it automatically — no manual profile edits):

dsh plugin --profile web add dsh-plugin-semantic-memory
# or from a local checkout / tarball:
dsh plugin --profile web add file:C:/path/to/dsh-embedding

Then restart dsh web and start a new session. All knobs have schema defaults; the in-package cordis.patch.yml is the deployment config source — in-package config overrides outer layers (settings.yaml and user patch rows only fill keys the package does not declare, they do not override it). With a file: linked install the loader reads that file live on every boot: edit it and restart dsh web, no reinstall needed.

Manual equivalent (for older installs): add the dependency to the profile's package.json, insert a mount row — new entries must be inserted (a bare - id: row only overrides an existing bundle id and is silently ignored):

- insert:
    - id: semantic-memory
      name: 'dsh-plugin-semantic-memory'

Leave mode/provider unset unless you need an explicit switch: selection is automatic (see below).

Usage

Provider selection

The embedding provider is chosen by mode (explicit deployment switch), falling back to the automatic selection:

ConfigurationProvider
mode: 'cloud'API (OpenAI-compatible /embeddings endpoint); requires apiKey
mode: 'local'local (ONNX via @huggingface/transformers, offline), even with an apiKey set
no mode, apiKey present (non-empty)API
no mode, no apiKeylocal
provider: 'local' (explicit)local, even with an apiKey set
provider: 'api' (explicit)API; requires apiKey

Switching deployment mode is editing mode in the in-package cordis.patch.ymlin-package config overrides outer layers (settings.yaml or user profile patch rows only fill keys the package does not declare; they do not override it). A restart is needed after patch-file changes; the settings document (~/.dsh/settings.yaml, semantic-memory: section) hot-reloads for the keys it is allowed to supply. The first local embed downloads the model (~100 MB, cached in ~/.cache/huggingface; use remoteHost for a mirror).

Verify the plugin is live

Open a new session (existing sessions keep their original tool set) and ask the model: "Do you have memory_ tools?"* — it should list memory_write, memory_search, memory_forget, and memory_stats. The system prompt also carries a ## Long-term memory section once memories exist.

What the model can do

  • Persist on its own — state a durable preference, fact, or decision; the injected guidance makes the model call memory_write without being asked.
  • Ask it to remember"记住:我在用硅基流动的 API"memory_write.
  • Recall"我之前对回答风格有什么偏好?" → the per-turn semantic recall surfaces relevant memories automatically; memory_search digs deeper (supports kind, tags, workspace, limit, min_score).
  • Managememory_forget <id> deletes; memory_stats summarizes the store.

Automatic behaviors

TriggerBehavior
Every user messageAsynchronous embedding + search; the freshest per-session hits are injected into the next prompt assembly (## Long-term memory (recalled for your current question))
Every N user messages (default 5)The harness LLM distills only the messages since the last summary (per-session seq cursor — no re-digesting, nothing skipped) into memory entries, written with the auto tag; the cadence can be set with the DSH_SEMANTIC_MEMORY_SUMMARIZE_EVERY environment variable (0 disables, overrides the config document)
Prompt assembly, no fresh recallStrongest memories (importance × recency × access) injected as fallback

Where the data lives

  • Store: $DSH_HOME/memories/memories.jsonl (one JSON line per entry, vectors included; edit/backup freely).
  • Settings: in-package cordis.patch.yml (deployment source of truth — package config overrides outer layers); ~/.dsh/settings.yaml under semantic-memory: only fills keys the package does not declare (hot-reloaded).

Troubleshooting

  • No memory_ tools in a session* — the session predates the plugin; start a new one.
  • First local embed is slow / fails — the model downloads on first use; set remoteHost: https://hf-mirror.com in restricted networks.
  • api provider errors — confirm mode/apiKey are set and apiBase points at an OpenAI-compatible endpoint (a /v1 base gets /embeddings appended).
  • Auto-summary never fires — it needs the llm and agentDefaultModel services (present in the standard web profile) and autoSummarizeEvery > 0.

Configuration

KeyDefaultMeaning
mode(unset)Deployment switch: local forces the local model, cloud forces the API (requires apiKey). Unset keeps the automatic selection.
providerautoauto selects by apiKey (non-empty → api, else local); explicit local/api overrides. An explicit mode overrides both.
localModelXenova/bge-small-zh-v1.5Local transformer model id.
remoteHosthttps://huggingface.coModel download host; set https://hf-mirror.com in restricted networks.
apiBasehttps://api.siliconflow.cn/v1API base URL (an /embeddings route is appended).
apiKey''API key. When non-empty and provider is not explicitly local, the API provider is used.
apiModelBAAI/bge-m3API embedding model name.
memoryPath$DSH_HOME/memories/memories.jsonlStore file path.
promptTopK3Memories injected per system-prompt assembly (0 disables).
maxSearchResults10Default memory_search hit cap.
minScore0.35Default minimum relevance for search hits.
halfLifeMs30 daysMemory strength half-life.
autoSummarizeEvery5Auto-summarize every N user messages (0 disables; needs llm + agentDefaultModel). The DSH_SEMANTIC_MEMORY_SUMMARIZE_EVERY env var overrides this (0..100).
summarizeWindow12Most recent messages included in one auto-summary.
summarizeMaxTokens800Token budget for the summary call.
summarizeTemperature0.2Sampling temperature for the summary call.

Memory model

interface MemoryEntry {
  id: string            // sha1(kind + content), 16 hex chars — upsert key
  kind: 'fact' | 'decision' | 'preference' | 'note'
  content: string       // one-sentence, self-contained text
  tags: string[]
  workspace?: string    // caller session cwd at write time
  source?: { sessionId: string; seq: number }
  importance: number    // 1..5
  embedding: number[]   // normalized vector
  createdAt: number
  updatedAt: number
  accessCount: number
  lastAccessAt: number
}

Effective strength = importance / 5 × 0.5^(age / halfLife); search rank = cosine(query, entry) × strength.

Known Limitations

  • Recall is best-effort and async — the user-message listener embeds in the background; on a cold start (model still downloading) or with a slow API the first recall may arrive one step late, and the strength-ranked fallback covers that turn. Recall caches are per-session and stale after 60 s.
  • Sync prompt injection — the injected section renders from resident data only; the store is loaded lazily on first tool call, so a brand-new process may start with an empty injection for the first assembly.
  • No embedding persistence cache — vectors are stored inside each entry, so no separate index file is needed, but full re-embedding never happens either (entries keep their vectors forever).
  • Brute-force search — O(n) cosine over all entries per query; fine for personal-scale stores (thousands), not for millions of entries.