← Back to home@yangyu666

dsh-jev-prune

Jev-judged context compaction for DeepSeek Harness: semantic tool-result pruning + deterministic receipt compaction

Stars
4
Language
JavaScript
Created
Sep 21, 2026
Updated
Oct 6, 2026
GitHub repo

Introduction

dsh-jev-prune

dsh-jev-prune — Jev-judged context compaction for DeepSeek Harness

Jev-judged context compaction for DeepSeek Harness. Structured judgments from TypeSafe Jev drive DSH's two-layer context compaction. The compaction algorithms are untouched; the judgment backend is pluggable (Jev / rules / a self-hosted model).

English · 简体中文

license node dsh CI smoke checks Listed on dsh-plugin.org

The problem it solves

DSH's built-in context reclamation is purely volumetric. Once a tool result crosses a size threshold, its middle is chopped out and the head and tail are kept; region compaction, meanwhile, has the model write a summary to stand in for old history. The first approach cannot tell "this result is large but I still need it" from "this one is spent", and the second one invites summary hallucination.

This plugin replaces the decision in both places with Jev's structured output (noul / choice, returning calibrated probabilities), under one design rule:

What should not be generated by a model is not generated by a model. Trimming only ever decides keep or discard; the original text is preserved verbatim. Region compaction injects a deterministic receipt produced by code, containing no model inference at all.

The two layers

The two layers: result trimming and receipt compaction

LayerInterception pointDSH defaultThis plugin
1 · Result trimmingctx.toolResultPruner.pruneSessionChops the middle once thresholdChars is exceededJev decides, per tool result, whether it will still be needed. Needed ones are never trimmed, however large; stale ones are trimmed however small (unless shorter than minCharsToPrune); with no judgment available it falls back to DSH's original behaviour
2 · Receipt compactionctx.compaction.summarize + compactRegion, or a single-result surface replacement for mixed batchesThe model reads the raw history and writes a summaryMoves fully eligible read-only steps out of the surface. In a parallel batch where only some results qualify, it keeps every call/result envelope and replaces only the eligible result bodies with deterministic receipts. Tool name, command, path, character count and seq are all computed by code

A layer-2 receipt looks like this:

[已压缩 · 确定性回执] 原历史 s25–s27 是 1 次工具调用(共约 16489 字符输出),
为释放上下文已移出。以下为事实清单(工具名/入参/字符数/seq 由代码算出;模型原话逐字引用,不含任何推断):
· s25 模型原话(逐字引用):identifiers:把版本字符串拆成标识符。Next inc.js.
· s27 read:C:\Users\you\project\src\state.js → 16489 字符输出
原始事件仍完整保存在会话日志中(seqs 25–27)。需要内容时重跑相同命令/读取相同文件即可;本回执不含对内容的解释。

Parallel tool batches are evaluated per call/result pair. DSH 0.1.5 can replace one surface node or one balanced contiguous region, but cannot remove five pairs from a six-call assistant message in one atomic operation. For a mixed batch, the plugin therefore replaces each eligible tool/result body with a short receipt through DSH's native single-node replace protocol; rejected and recent results remain byte-for-byte unchanged, the assistant head stays intact, and tool pairing remains valid after every replacement. Results are matched to calls by callId, so completion order may differ from declaration order. A fully eligible batch still uses compactRegion and removes the whole balanced step.

The receipt body is emitted by the plugin's own JavaScript, so its wording is Chinese today — as is the /jev status output. The jev_* tool descriptions are already English. Localizing the runtime strings is a separate change.

Measured results

Numbers below come from a three-arm measurement of a 37-step "read every file in a directory, strictly one at a time" long task over semver@6e05b76: A = vanilla DSH (baseline), B = the plugin with compactReceipts: false (layer 1 only), C = the plugin with defaults (both layers). Every number is taken from the usage object the API actually returned in the DSH session logs — no estimator is involved in scoring.

Layer 1 acts on every long run; the baseline never does. Layer 1 trimmed stale results in every plugin-arm run (4–7 nodes per long run; arm A trims 0 by definition), and answers on the short/medium/long tasks were correct in every arm.

Long tasks run on a fraction of the full-price input. Uncached ("full-price") input tokens per completed 37-step run: baseline 75,781–87,872 vs full plugin 25,516–45,071 — roughly a third of the baseline at equal task scope. Report token classes separately: the cache-hit share of input is high (78–94%) and hit/miss prices differ by ~50×, so total-token comparisons mislead by design.

Fewer destructive compactions — some of them receipts instead of summaries. Per completed long run, DSH's own model-written summary compactions dropped from 31–40 (baseline) to 15–23 with the full plugin, of which 6–8 were replaced by deterministic receipts rendered by code. Layer 2 stays deliberately silent on short/medium tasks — its gates require a contiguous run of eligible read-only steps — so on those workloads the effect comes from layer 1. Even when layer 2 fires, DSH-initiated compactions still exist and fall back to model summaries: the plugin reduces them, it does not eliminate them.

A failure mode we found and fixed. Receipts replace a whole "call + result" step, including the assistant message that carried it; in one long run the model's intermediate notes were erased step by step (12 of 13) and it stopped issuing tool calls mid-task. Receipts now carry each step's assistant-visible text verbatim (text blocks only, reasoning drafts excluded, zero model generation; bounded by receiptTextChars, default 400, 0 disables). After the fix, the same long task completed in every run with correct answers.

Measurement details and scope

Gating (layer 2)

Moving a whole pair out of the surface is destructive, so the default is deliberately conservative. Every one of the following must hold:

  • Intersection of two axes: result (is the content still needed) and effect (did the call change state outside the session) must each fall inside this session's trailing compactQuantile
  • The tool is not in neverCompactTools (write-type calls are excluded by a hard rule, never by a probability)
  • Evidence guard: results matching error / assert / fail / todo and friends are never moved out. If layer 1 has already trimmed a result, the guard follows sourceEventSeqs and scans the original event too
  • Steps whose assistant text exceeds maxStepTextChars, or whose reasoning exceeds maxStepReasoningChars, are never moved out. The two are measured separately on purpose: long text means the step is delivering a conclusion worth keeping, while long reasoning is just scratch work — merging them into one budget let reasoning length alone silently shut layer 2 off
  • Anything within the most recent compactPreserveRecent nodes is skipped by layer 2 (layer 1 uses preserveRecent)
  • Both ends of the range must satisfy DSH's tool-pairing balance; the span must save at least compactMinChars characters; and the receipt must stay below receiptMaxRatio of the original content's tokens

Probabilities are consumed as relative quantiles, never as a fixed threshold: the output distribution of a small judge model is narrow, and only the relative ordering within one session carries stable information.

Degradation on small populations. Read-only tools are often a minority in write/execute-heavy sessions (measured: 1 in 6), which can leave a quantile population of only two or three items — too few for ordering to mean anything. Rather than giving up, the mode degrades to an absolute floor: both axes must fall below floorThreshold (default 0.2, materially stricter than compactThreshold, compensating for the missing relative information). If the population is below minCandidatesForFloor (default 2), nothing is moved out — a single sample is not a distribution. The default was lowered from 3 to 2 so that the two-candidate populations that batch-read sessions really produce are not skipped outright; one sample still never acts. Degradation is always reported in the report and the heartbeat; it never happens silently.

Install and quick start

Version note: use 0.1.1 for DSH 0.2.0-rc.2. It supports both root and preset-local services and passed actual CLI installation/activation; older 0.1.0 is rejected by this host. Compatibility evidence and rolling verification.

Requires Node ^22.19.0 || >=24.0.0. Regression baseline: DSH 0.1.5-rc.2; real CLI install/preset activation verified on 0.2.0-rc.2. The profile needs pruner/compaction and tokenMeter services. Set TYPESAFE_API_KEY before starting DSH.

Install a pinned version:

dsh plugin --profile web add dsh-jev-prune@0.1.1
# GitHub alternative:
dsh plugin --profile web add github:yangyu666/dsh-jev-prune#v0.1.1

Local development:

git clone https://github.com/yangyu666/dsh-jev-prune.git
cd dsh-jev-prune
git checkout v0.1.1
npm ci
npm run check
npm run smoke
dsh plugin --profile web add link:/absolute/path/to/dsh-jev-prune

Without pnpm, run node scripts/wire_profile.mjs <DSH_HOME> <profile-name>. After host upgrades, call jev_probe_shapes to verify the event shapes.

Configuration

Start with a dry run in the profile's plugin entry:

- id: jev-prune
  config:
    dryRun: true
    judgeOn: pressure
    softLimit: "55%"
    compactReceipts: true
    compactOn: pressure
    compactSoftLimit: "70%"

After inspecting /jev, set dryRun: false to apply reductions.

OptionDefaultPurpose
keepModebudgetRank candidates; trimming budget follows context pressure.
preserveRecent / compactPreserveRecent4 / 1Independent recent-node protection.
headChars / tailChars600 / 200Verbatim content retained by layer 1.
compactQuantile0.34Tail fraction on each judgment axis; use the intersection.
compactToolsRead-only allow-listLimits which tool pairs may be compacted.
receiptTextChars400Bounds verbatim assistant-visible notes in receipts.
dryRunfalseJudge and report without applying changes.

Full configuration and pressure-gate details · Configuration example

In-session usage

Entry pointPurpose
/jev, jev_prune_statusLedger for both layers: cached judgments, cumulative savings, takeover state, tool-name index
jev_prune_nowForce one layer-1 trimming pass
jev_compact_now (supports dryRun)Force one layer-2 receipt compaction and list every gate's exclusion counts; in dryRun it also prints the full receipt text
jev_restoreSafety valve: fetch back the original text that a checkpoint moved out (read-only)
jev_probe_shapesPrint the real event shapes and the resolved tool names, for adapting to a different DSH version

In normal operation both layers are driven automatically by context pressure; no manual step is needed.

Demo

Recorded deterministic demonstration

The recording runs the real plugin apply(), registered tools and replacement paths against a simulated DSH host with fixed judge probabilities. It shows trimming, preservation of a high-scored result, receipt compaction, and retrieval of a region checkpoint's original text. It does not measure live Jev quality or prove live-host compatibility.

Replay, assertions and regeneration · Machine-readable evidence

Limits and data handling

  • Live judging sends history, paths, snippets and command output to TypeSafe and adds inference cost and latency.
  • Automated CI covers helpers, simulated-host behavior and loading against a locked host dependency tree. It does not exercise a live DSH session end to end.
  • On current source, receipts are scoped to an asynchronous transaction and validated against the selected replay messages; foreign, canceled or mismatched summaries fall back to the host. The published 0.1.0 still has the older session-scoped behavior. See architecture.
  • Current source invalidates result judgments when the last three user instructions change, retaining effect judgments and discarding in-flight answers for older goals. Small-population layer-1 fallback can still trim at zero pressure ratio.
  • jev_restore currently finds region summary checkpoints, truncates each returned event to 4,000 UTF-16 units, and does not directly resolve layer-1 or partial-result replacement IDs. Original events remain in the log.

Layout and development

index.js              Plugin entry (unchanged installation contract)
src/                  Judge, pruning, receipts and event projection
test/                 Helper and simulated-host checks
scripts/              Profile wiring and offline session inspection
docs/                 Architecture, configuration, porting and releases
demo/                 Reproducible behavior demonstration
assets/               Diagrams and recorded demonstration
examples/             Profile configuration examples
npm run check
npm run smoke
npm run coverage
npm run demo

Contributing · Architecture · Porting · Releasing · Changelog

MIT