Back to home

Zhangbo-cn

dsh-vision-plugin

No description

Stars
0
Language
TypeScript
Created
Aug 15, 2026
Updated
Aug 15, 2026

Introduction

dsh-vision-plugin

Vision capability for DeepSeek Harness: lets a text-only model "understand" an image by routing to an external OpenAI-compatible multimodal API. Built as a standalone dsh-plugin from the official vision capability seam proposal.

Packages

PackageRole
@zhangbo-cn/dsh-visionService Definition: ctx.vision (registerAdapter, describe, listProviders)
@zhangbo-cn/dsh-vision-openai-compatibleProvider: OpenAI-compatible chat-completions adapter
@zhangbo-cn/dsh-tool-visionConsumer: view_image tool

Install

pnpm add @zhangbo-cn/dsh-vision @zhangbo-cn/dsh-vision-openai-compatible @zhangbo-cn/dsh-tool-vision

Mount in your cordis.yml:

- id: vision
  name: '@zhangbo-cn/dsh-vision'

- id: vision-openai-compatible
  name: '@zhangbo-cn/dsh-vision-openai-compatible'
  config:
    baseURL: 'https://api.example.com/v1'   # required at request time
    model: 'gpt-4o'                          # required at request time
    apiKeyEnv: 'OPENAI_API_KEY'              # env var holding the key

- id: tool-vision
  name: '@zhangbo-cn/dsh-tool-vision'

Then ask the model: "use view_image to look at ./screenshot.png" — it reads the file, commits the bytes through the attachment seam, and returns a text description from your configured vision model.

How it works

view_image(file_path, prompt)
  → ctx.fs reads the image bytes
  → attachments.saveImage (durable, content-addressed)
  → ctx.vision.describe({ ref, prompt })
      → vision provider posts a data:image/...;base64 image_url to /chat/completions
      → returns text

Image input reuses the durable ImageAttachmentRef from the attachment seam; output is text (no ImageBlock), so it is independent of whether the harness LLM route itself accepts images.

Requirements

  • DeepSeek Harness with an attachment store (dsh-attachment-local) and filesystem (dsh-fs-local).
  • A configured OpenAI-compatible multimodal endpoint (any OpenAI-chat-completions-compatible vision model).

Development

npm install
npm run build -ws
npx vitest run

Tests include a real Loader composition booting the plugin with a fake in-memory vision provider (52 tests).

License

MIT