dsh-jev-prune
Jev-judged context compaction for DeepSeek Harness: semantic tool-result pruning + deterministic receipt compaction
- Stars
- 4
- Language
- JavaScript
- Created
- Sep 21, 2026
- Updated
- Oct 6, 2026
Introduction
dsh-jev-prune

Jev-judged context compaction for DeepSeek Harness. Structured judgments from TypeSafe Jev drive DSH's two-layer context compaction. The compaction algorithms are untouched; the judgment backend is pluggable (Jev / rules / a self-hosted model).
English · 简体中文
The problem it solves
DSH's built-in context reclamation is purely volumetric. Once a tool result crosses a size threshold, its middle is chopped out and the head and tail are kept; region compaction, meanwhile, has the model write a summary to stand in for old history. The first approach cannot tell "this result is large but I still need it" from "this one is spent", and the second one invites summary hallucination.
This plugin replaces the decision in both places with Jev's structured output (noul / choice, returning calibrated probabilities), under one design rule:
What should not be generated by a model is not generated by a model. Trimming only ever decides keep or discard; the original text is preserved verbatim. Region compaction injects a deterministic receipt produced by code, containing no model inference at all.
The two layers

| Layer | Interception point | DSH default | This plugin |
|---|---|---|---|
| 1 · Result trimming | ctx.toolResultPruner.pruneSession | Chops the middle once thresholdChars is exceeded | Jev decides, per tool result, whether it will still be needed. Needed ones are never trimmed, however large; stale ones are trimmed however small (unless shorter than minCharsToPrune); with no judgment available it falls back to DSH's original behaviour |
| 2 · Receipt compaction | ctx.compaction.summarize + compactRegion, or a single-result surface replacement for mixed batches | The model reads the raw history and writes a summary | Moves fully eligible read-only steps out of the surface. In a parallel batch where only some results qualify, it keeps every call/result envelope and replaces only the eligible result bodies with deterministic receipts. Tool name, command, path, character count and seq are all computed by code |
A layer-2 receipt looks like this:
[已压缩 · 确定性回执] 原历史 s25–s27 是 1 次工具调用(共约 16489 字符输出),
为释放上下文已移出。以下为事实清单(工具名/入参/字符数/seq 由代码算出;模型原话逐字引用,不含任何推断):
· s25 模型原话(逐字引用):identifiers:把版本字符串拆成标识符。Next inc.js.
· s27 read:C:\Users\you\project\src\state.js → 16489 字符输出
原始事件仍完整保存在会话日志中(seqs 25–27)。需要内容时重跑相同命令/读取相同文件即可;本回执不含对内容的解释。
Parallel tool batches are evaluated per call/result pair. DSH 0.1.5 can replace one surface node or one balanced contiguous region, but cannot remove five pairs from a six-call assistant message in one atomic operation. For a mixed batch, the plugin therefore replaces each eligible tool/result body with a short receipt through DSH's native single-node replace protocol; rejected and recent results remain byte-for-byte unchanged, the assistant head stays intact, and tool pairing remains valid after every replacement. Results are matched to calls by callId, so completion order may differ from declaration order. A fully eligible batch still uses compactRegion and removes the whole balanced step.
The receipt body is emitted by the plugin's own JavaScript, so its wording is Chinese today — as is the
/jevstatus output. Thejev_*tool descriptions are already English. Localizing the runtime strings is a separate change.
Measured results
Numbers below come from a three-arm measurement of a 37-step "read every file in a directory, strictly one at a time" long task over semver@6e05b76: A = vanilla DSH (baseline), B = the plugin with compactReceipts: false (layer 1 only), C = the plugin with defaults (both layers). Every number is taken from the usage object the API actually returned in the DSH session logs — no estimator is involved in scoring.
Layer 1 acts on every long run; the baseline never does. Layer 1 trimmed stale results in every plugin-arm run (4–7 nodes per long run; arm A trims 0 by definition), and answers on the short/medium/long tasks were correct in every arm.
Long tasks run on a fraction of the full-price input. Uncached ("full-price") input tokens per completed 37-step run: baseline 75,781–87,872 vs full plugin 25,516–45,071 — roughly a third of the baseline at equal task scope. Report token classes separately: the cache-hit share of input is high (78–94%) and hit/miss prices differ by ~50×, so total-token comparisons mislead by design.
Fewer destructive compactions — some of them receipts instead of summaries. Per completed long run, DSH's own model-written summary compactions dropped from 31–40 (baseline) to 15–23 with the full plugin, of which 6–8 were replaced by deterministic receipts rendered by code. Layer 2 stays deliberately silent on short/medium tasks — its gates require a contiguous run of eligible read-only steps — so on those workloads the effect comes from layer 1. Even when layer 2 fires, DSH-initiated compactions still exist and fall back to model summaries: the plugin reduces them, it does not eliminate them.
A failure mode we found and fixed. Receipts replace a whole "call + result" step, including the assistant message that carried it; in one long run the model's intermediate notes were erased step by step (12 of 13) and it stopped issuing tool calls mid-task. Receipts now carry each step's assistant-visible text verbatim (text blocks only, reasoning drafts excluded, zero model generation; bounded by receiptTextChars, default 400, 0 disables). After the fix, the same long task completed in every run with correct answers.
Gating (layer 2)
Moving a whole pair out of the surface is destructive, so the default is deliberately conservative. Every one of the following must hold:
- Intersection of two axes:
result(is the content still needed) andeffect(did the call change state outside the session) must each fall inside this session's trailingcompactQuantile - The tool is not in
neverCompactTools(write-type calls are excluded by a hard rule, never by a probability) - Evidence guard: results matching
error/assert/fail/todoand friends are never moved out. If layer 1 has already trimmed a result, the guard followssourceEventSeqsand scans the original event too - Steps whose assistant text exceeds
maxStepTextChars, or whosereasoningexceedsmaxStepReasoningChars, are never moved out. The two are measured separately on purpose: longtextmeans the step is delivering a conclusion worth keeping, while longreasoningis just scratch work — merging them into one budget let reasoning length alone silently shut layer 2 off - Anything within the most recent
compactPreserveRecentnodes is skipped by layer 2 (layer 1 usespreserveRecent) - Both ends of the range must satisfy DSH's tool-pairing balance; the span must save at least
compactMinCharscharacters; and the receipt must stay belowreceiptMaxRatioof the original content's tokens
Probabilities are consumed as relative quantiles, never as a fixed threshold: the output distribution of a small judge model is narrow, and only the relative ordering within one session carries stable information.
Degradation on small populations. Read-only tools are often a minority in write/execute-heavy sessions (measured: 1 in 6), which can leave a quantile population of only two or three items — too few for ordering to mean anything. Rather than giving up, the mode degrades to an absolute floor: both axes must fall below floorThreshold (default 0.2, materially stricter than compactThreshold, compensating for the missing relative information). If the population is below minCandidatesForFloor (default 2), nothing is moved out — a single sample is not a distribution. The default was lowered from 3 to 2 so that the two-candidate populations that batch-read sessions really produce are not skipped outright; one sample still never acts. Degradation is always reported in the report and the heartbeat; it never happens silently.
Install and quick start
Version note: use 0.1.1 for DSH 0.2.0-rc.2. It supports both root and preset-local services and passed actual CLI installation/activation; older 0.1.0 is rejected by this host. Compatibility evidence and rolling verification.
Requires Node ^22.19.0 || >=24.0.0. Regression baseline: DSH 0.1.5-rc.2; real CLI install/preset activation verified on 0.2.0-rc.2. The profile needs pruner/compaction and tokenMeter services. Set TYPESAFE_API_KEY before starting DSH.
Install a pinned version:
dsh plugin --profile web add dsh-jev-prune@0.1.1
# GitHub alternative:
dsh plugin --profile web add github:yangyu666/dsh-jev-prune#v0.1.1
Local development:
git clone https://github.com/yangyu666/dsh-jev-prune.git
cd dsh-jev-prune
git checkout v0.1.1
npm ci
npm run check
npm run smoke
dsh plugin --profile web add link:/absolute/path/to/dsh-jev-prune
Without pnpm, run node scripts/wire_profile.mjs <DSH_HOME> <profile-name>. After host upgrades, call jev_probe_shapes to verify the event shapes.
Configuration
Start with a dry run in the profile's plugin entry:
- id: jev-prune
config:
dryRun: true
judgeOn: pressure
softLimit: "55%"
compactReceipts: true
compactOn: pressure
compactSoftLimit: "70%"
After inspecting /jev, set dryRun: false to apply reductions.
| Option | Default | Purpose |
|---|---|---|
keepMode | budget | Rank candidates; trimming budget follows context pressure. |
preserveRecent / compactPreserveRecent | 4 / 1 | Independent recent-node protection. |
headChars / tailChars | 600 / 200 | Verbatim content retained by layer 1. |
compactQuantile | 0.34 | Tail fraction on each judgment axis; use the intersection. |
compactTools | Read-only allow-list | Limits which tool pairs may be compacted. |
receiptTextChars | 400 | Bounds verbatim assistant-visible notes in receipts. |
dryRun | false | Judge and report without applying changes. |
Full configuration and pressure-gate details · Configuration example
In-session usage
| Entry point | Purpose |
|---|---|
/jev, jev_prune_status | Ledger for both layers: cached judgments, cumulative savings, takeover state, tool-name index |
jev_prune_now | Force one layer-1 trimming pass |
jev_compact_now (supports dryRun) | Force one layer-2 receipt compaction and list every gate's exclusion counts; in dryRun it also prints the full receipt text |
jev_restore | Safety valve: fetch back the original text that a checkpoint moved out (read-only) |
jev_probe_shapes | Print the real event shapes and the resolved tool names, for adapting to a different DSH version |
In normal operation both layers are driven automatically by context pressure; no manual step is needed.
Demo

The recording runs the real plugin apply(), registered tools and replacement paths against a simulated DSH host with fixed judge probabilities. It shows trimming, preservation of a high-scored result, receipt compaction, and retrieval of a region checkpoint's original text. It does not measure live Jev quality or prove live-host compatibility.
Replay, assertions and regeneration · Machine-readable evidence
Limits and data handling
- Live judging sends history, paths, snippets and command output to TypeSafe and adds inference cost and latency.
- Automated CI covers helpers, simulated-host behavior and loading against a locked host dependency tree. It does not exercise a live DSH session end to end.
- On current source, receipts are scoped to an asynchronous transaction and validated against the selected replay messages; foreign, canceled or mismatched summaries fall back to the host. The published 0.1.0 still has the older session-scoped behavior. See architecture.
- Current source invalidates result judgments when the last three user instructions change, retaining effect judgments and discarding in-flight answers for older goals. Small-population layer-1 fallback can still trim at zero pressure ratio.
jev_restorecurrently finds region summary checkpoints, truncates each returned event to 4,000 UTF-16 units, and does not directly resolve layer-1 or partial-result replacement IDs. Original events remain in the log.
Layout and development
index.js Plugin entry (unchanged installation contract)
src/ Judge, pruning, receipts and event projection
test/ Helper and simulated-host checks
scripts/ Profile wiring and offline session inspection
docs/ Architecture, configuration, porting and releases
demo/ Reproducible behavior demonstration
assets/ Diagrams and recorded demonstration
examples/ Profile configuration examples
npm run check
npm run smoke
npm run coverage
npm run demo
Contributing · Architecture · Porting · Releasing · Changelog