akqwpeter-prog
dsh-media-skills
Free image reading & generation for DeepSeek Harness — paste an image into any chat, even text-only sessions. 免费读图·生图 · 9 种语言 · 无 Key 入库
- Stars
- 1
- Language
- Python
- Created
- Aug 14, 2026
- Updated
- Aug 14, 2026
Introduction
🎨 dsh-media-skills
Give DeepSeek Harness eyes — and a brush. Read images in any chat, generate new ones, all with free models.
DeepSeek Harness is brilliant at reasoning — but a text-only model can't see the image you just dragged into the chat. This bundle fixes that with two free skills and a free vision model route:
- 📎 Paste to read — paste, drag, or pick an image in any session; the free vision model turns it into text your current model understands.
- 👁️
vision-review— analyze images and screenshots, catch UI visual bugs, detect watermarks, turn images into text. - 🎨
media-tools— generate illustrations, avatars, backgrounds and banners with a free, watermark-free model.
No hardcoded keys, no paid API, no file saving, no session switching.
Why · Quick start · See it in action · Usage · Keys & privacy · FAQ · Examples
English · 简体中文 · 繁體中文 · 日本語 · 한국어 · Español · Deutsch · Português · Русский
🤔 Why
Most DSH vision plugins only read images — and many push you through a shared third-party endpoint. dsh-media-skills takes a different stance:
| This bundle | Typical vision-only plugin | |
|---|---|---|
| Read images for free | ✅ Zhipu GLM-4V-Flash | ✅ |
| Generate images for free | ✅ SiliconFlow Kolors | ❌ usually absent |
| Auto model route in the picker | ✅ installed automatically | sometimes |
| Keys committed to the repo | ❌ never — keys stay local | ⚠️ often required |
| Docs in multiple languages | ✅ 9 languages | ❌ usually English only |
| Privacy | ✅ you choose the provider; images only go to your provider | shared free endpoints can see your images |
Why bring your own free key instead of a built-in anonymous endpoint? Privacy and reliability. Your images go only to the provider you choose, under your account and your rate limits — no shared third-party service in the middle.
✨ What you get
| Capability | What it does | Model | Cost |
|---|---|---|---|
| 📎 Paste-image reading | In a text-only session, the input bar gains an “Add image” button (paperclip); pasted images are auto-described by the vision model and handed to the current model as text | Zhipu GLM-4V-Flash | Free |
| 🧠 Vision model route | 「智谱 GLM-4V-Flash(视觉)」 appears in the model selector automatically — pick it for a new conversation and talk about images directly | Zhipu GLM-4V-Flash | Free |
👁️ vision-review | Analyze / recognize / describe images & screenshots; catch UI visual bugs (overlap, overflow, misalignment); detect watermarks/logos; turn images into text | Zhipu GLM-4V-Flash | Free |
🎨 media-tools | Generate images, illustrations, avatars, backgrounds, banners | SiliconFlow Kolors | Free, no watermark |
⚡ Quick start
dsh plugin --profile <name> add github:akqwpeter-prog/dsh-media-skills
-
Get two free keys (~2 minutes, no payment):
- Zhipu — open.bigmodel.cn → API Keys (
glm-4v-flashis free) - SiliconFlow — siliconflow.cn → API Keys (Kolors is free)
- Zhipu — open.bigmodel.cn → API Keys (
-
Add them in the Web GUI (Settings → Models → the zhipu-vision provider's API Key field), or use the credentials file:
# ~/.dsh/.credentials.yaml (chmod 600) GLM_API_KEY: <your key> -
Restart
dsh web, then hard-refresh (Cmd+Shift+R).
Verify: the model selector shows 智谱 GLM-4V-Flash(视觉). If your Harness build supports paste-image reading, the input bar also has a 📎 Add image button — paste an image in any session and it arrives as a text description.
Full walkthrough and troubleshooting: docs/SETUP_VISION_EN.md.
📸 See it in action
Paste an image in a text-only session → the free vision model describes it → your model answers. The same bundle also generates new images on demand.
How it works in one picture:
🚀 Usage
Three ways to read images:
| Way | How | When |
|---|---|---|
| A. Paste directly (recommended) | In any session, click the 📎 button / drag / paste an image and send | Everyday image questions — no file saving, no model switching |
| B. Vision model session | New conversation, pick 智谱 GLM-4V-Flash(视觉), paste images and chat | Multi-turn image conversations, native read_image |
| C. Files + skill | Put the image in the workspace and say “read this image with vision-review” | Batch review, scripted workflows |
Descriptions follow your message language (Chinese message → Chinese description; English message → English description; no text → Chinese).
Also just say:
- “Look at this image / check this screenshot for visual bugs” →
vision-review - “Generate an image of …” →
media-tools
🔑 Keys & privacy
Keys are never stored in this repo. Skill scripts read, in order: environment variables → ~/.dsh/secrets/media-tools.env → ~/.codex/secrets/media-tools.env (legacy fallback). The vision model route reads GLM_API_KEY from DSH's credential store.
# ~/.dsh/secrets/media-tools.env (chmod 600, one KEY=value per line)
GLM_API_KEY=...
SILICONFLOW_API_KEY=...
Your images are sent only to the provider you configure — never to this repo, never to a shared anonymous endpoint.
❓ FAQ
Does paste-image reading require a DeepSeek Harness core patch?
The auto-describe pipeline lives in the Harness core (api-proxy image-admission logic; see docs/HARNESS_PATCH_EN.md). This bundle ships the model route + skills: the vision model works on any DSH build, but paste-image reading requires a Harness build with that core support — see FAQ Q1 in docs/SETUP_VISION_EN.md.
Why not just use a built-in free endpoint with no key at all? We prefer to let you own the route: your images go to the provider you pick, under your rate limits, with no shared middleman. The keys are free and take about two minutes to create.
Is media-tools really free?
Yes — SiliconFlow Kolors is free and watermark-free. If a model is temporarily disabled, the skill lists available models and you can switch.
🎁 Examples
Sample material to try instantly — 6 AI-generated images with their prompts, plus a purpose-built vision test card (title, buttons, bar-chart values) for checking reading accuracy:

🗺️ Layout
dsh-media-skills/
├── package.json # dsh.bundle manifest
├── cordis.patch.yml # plugin layer
├── index.js # registers skills + seeds the zhipu-vision model route
├── skills/
│ ├── vision-review/ # image reading
│ └── media-tools/ # image generation
├── examples/ # sample images + vision test card
├── docs/
│ ├── screenshots/ # demo mockup & how-it-works diagram
│ ├── SETUP_VISION_EN.md # detailed setup guide (English)
│ ├── SETUP_VISION.md # 详细配置指南(中文)
│ ├── HARNESS_PATCH_EN.md# core patch notes (English)
│ ├── HARNESS_PATCH.md # 本体补丁说明(中文)
│ └── lang/ # READMEs in 9 languages
├── scripts/make-banner.py # regenerates docs/social-preview.png
└── docs/social-preview.png
🤝 Join the DSH plugin ecosystem
DeepSeek Harness developer preview is still in its testing phase for Harness developers; core plugins and base APIs will keep iterating. We look forward to exploring the upper limits of intelligence together with developers worldwide, on top of open-source, open, reusable, and composable infrastructure.
Tag this repo with
dsh-pluginso others can discover it. PRs, issues and translations are welcome.