visail
dsh-vision-tool
Paste an image into the chat box and text-only DSH models can "see" it — auto-rewrite of pasted images + analyze_image tool routed to a Kimi vision model.
- Stars
- 0
- Language
- JavaScript
- Created
- Aug 14, 2026
- Updated
- Aug 14, 2026
Introduction
dsh-vision-tool
Image routing for text-only models in DeepSeek Harness (DSH).
Text-only models (e.g. deepseek-v4-flash) cannot see images. This bundle gives
them eyes in two coordinated steps:
vision-prompt— shadowsPOST /api/session.prompt. When the active session model does not support image input, pasted images are persisted as content-addressed attachments and rewritten in place into text prompts that carry the full attachment reference JSON. Any other request (no image, or a model that already supports images) is forwarded unchanged, with the official/apitrust fence (DNS rebinding / cross-site defense) reimplemented.vision-tool— registers a globalanalyze_imagetool. The model calls it with the attachment reference (or a local filepath); the tool routes the image to a vision model and returns its text answer.
Verified end-to-end: paste an image → attachment persisted → model calls
analyze_image → vision model describes the image.
Install
Requires the dsh CLI (see package and install a plugin).
# from npm (once published)
dsh plugin --profile <name> add dsh-vision-tool
# straight from a git host (pin a commit; no build step needed — pure ESM)
dsh plugin --profile <name> add github:<you>/dsh-vision-tool#<sha>
# or from a local tarball
pnpm pack
dsh plugin --profile <name> add ./dsh-vision-tool-0.1.0.tgz
The bundle's cordis.patch.yml inserts two rows: vision-tool and
vision-prompt. Restart the profile afterwards:
dsh --profile <name>
Verify the layer landed without booting:
dsh --profile <name> --dump-config # look for the "# == dsh-vision-tool" layer
Configuration
Credential (required)
The tool resolves KIMI_CODE_API_KEY — from $DSH_HOME/.credentials.yaml or an
environment variable of the same name. Get the key from your
Kimi Code subscription page (sk-kimi- prefix).
# $DSH_HOME/.credentials.yaml
KIMI_CODE_API_KEY: sk-kimi-...
Switching vision models
The defaults target kimi-for-coding at https://api.kimi.com/coding/v1.
Override the row in your profile's cordis.patch.yml (a patch replaces the
whole config, so restate every key you keep):
- id: vision-tool
name: dsh-vision-tool
config:
baseURL: https://api.kimi.com/coding/v1
model: kimi-for-coding
apiKeyEnv: KIMI_CODE_API_KEY
maxImageBytes: 20971520
timeoutMs: 120000
Note:
kimi-for-codingonly acceptstemperature: 1(anything else is rejected with HTTP 400). The tool hard-codestemperature: 1as its default and is not configurable for this model. Other OpenAI-compatible vision endpoints generally work as long as they accepttemperature: 1.
Supported inputs
attachment— full reference JSON injected by the paste-rewrite mechanism ({"attachmentId":"sha256:...","mediaType":...,"bytes":N,"width":N,"height":N}). Pass it verbatim; do not strip fields.path— local image file (absolute, or relative to the session cwd).- Formats:
png/jpg/jpeg/webp/gif. Local files up tomaxImageBytes(default 20 MB). Attachments are bounded by the harness attachment store limits.
Security
vision-promptreimplements the official/apitrust fence: loopback /trustedHostshost check,sec-fetch-siteandOriginchecks.- Request bodies are capped at 160 MB (413 otherwise), matching the harness http-bridge default.
- Any failure degrades to passthrough — the original request is forwarded unchanged, never swallowed or mangled.
- The tool only reads the attachment you reference and your configured credential; it never stores prompt or image content beyond the attachment store the harness itself maintains.
Diagnostics
Both plugins append to $DSH_HOME/vision-trace.log:
handle: rewrite result = REWRITTEN
vision-tool: execute: resolve KIMI_CODE_API_KEY -> source=file len=72