← Back to home@jiumengya

dsh-computer-use

DeepSeek Harness desktop-control plugin (dsh-computer-use): virtual mouse/keyboard, screen capture, and driver-level input

Stars
0
Language
JavaScript
Created
Sep 12, 2026
Updated
Sep 12, 2026

Introduction

dsh-computer-use

A computer-use plugin for DeepSeek Harness, adapted from the OpenAI Codex computer-use action vocabulary: the agent "sees" the real Windows desktop through screenshots and "operates" it through a virtual mouse. Each action is delivered directly to the window at the target coordinates, so the user's physical mouse and keyboard are completely unaffected — the user can keep using their own computer while the agent works (Codex virtual-mouse semantics). Zero npm dependencies, structurally identical to dsh-browser-use / dsh-wallpaper.

Tool surface (7 tools)

ToolPurpose
computer_screenshotCaptures the full desktop (multi-monitor), JPEG downscaled to ≤2000px, returned as a model-visible image
computer_clickClick / double-click / triple-click with left/right/middle button
computer_moveMoves the virtual mouse without clicking (hover events reach the window below)
computer_dragDrags along a path
computer_typeTypes text into the window at the virtual cursor (supports Chinese, bypasses IME)
computer_keyPresses a key/combination into the window at the virtual cursor (ctrl+a, alt+f4, win, ...)
computer_scrollVertical/horizontal scroll (Playwright semantics: positive y = down, positive x = right)

Work loop

Codex's computer-use loop: screenshot → model reads the image and decides → action → screenshot to verify. Coordinates are pixel coordinates in the screenshot image (origin at top-left); the tool internally converts them back to real screen coordinates using the scale factor of the most recent screenshot (the attachment store caps a single image at ≤3.5MB / longest edge ≤2000px, so a 4K screen is downsampled, and the coordinate back-mapping is done by the driver layer). Every action returns the cursor position (in screenshot coordinates) and the foreground window (process name + title), so the agent can confirm focus before typing.

Implementation notes

  • Virtual input (default): win-vinput.ps1 uses PostMessage to deliver mouse/keyboard messages directly to the window at the action coordinates (WindowFromPoint + ChildWindowFromPointEx recursively drills down to the top-most control; coordinates are converted to client-area coordinates). The physical mouse/keyboard are never touched — user and agent each use their own. The driver layer tracks the virtual cursor position: actions without coordinates (type/key) land at the current virtual cursor position (screen center before the first action); the system prompt tells the agent to "click the input box first, then type".
  • Physical input (optional): with inputMode: real, win-input.ps1 injects via SendInput, which really moves the physical cursor; UIPI interception (when the foreground is an elevated window) reports an explicit error.
  • Screenshot: GDI+ CopyFromScreen of the whole virtual desktop (VirtualScreen, including negative-coordinate secondary monitors); SetProcessDPIAware ensures true pixels. JPEG uses the GDI+ default quality; when over 3.5MB it downscales step by step.
  • Script split: the three driver scripts are split into win-input.ps1 (physical input) / win-vinput.ps1 (virtual input) / win-shot.ps1 (screenshot) — merging input injection and screenshots into a single script gets blocked by Windows Defender AMSI (ScriptContainedMaliciousContent). Likewise, the screenshot script cannot use the JPEG quality encoder (EncoderParameters); it must save with default quality only.
  • Screenshot into the model: goes through the attachment service (saveImage) → render returns an image block, the same pipeline as read_image; the screenshot tool is gated by model capability (models that do not declare image input are rejected outright).
  • Safety: actions are serialized (isConcurrencySafe: false); in virtual mode the process blacklist is rejected on the driver side before delivery (the JS side adds a second check); the system prompt requires the agent to confirm destructive actions first.
  • Driver scripts are pure ASCII (PowerShell 5.1 parses UTF-8-without-BOM as GBK).

Configuration (desktop.yml)

- id: computer-use
  name: '@deepseek-ai/dsh-computer-use'
  config:
    inputMode: virtual       # virtual (default, virtual mouse) / real (physical injection)
    powershellPath: ''       # defaults to powershell.exe from PATH
    maxImageDimension: 2000  # screenshot longest edge in pixels (do not exceed the attachment cap)
    blockedProcesses: []     # target process blacklist (lowercase, without .exe), e.g. ['cmd','regedit']

Install / uninstall

install.ps1 / uninstall.ps1 (the same three steps as dsh-wallpaper: copy the package → farm junction → data-dir overlay). After installing, restart DeepSeek Harness and tell the agent "take a screenshot of my screen".

Known limitations

  • Windows only (PowerShell 5.1+ / .NET GDI+).
  • Dynamic screen content (video/games) is captured as a static frame; that is expected behavior.
  • Virtual mode delivers synthetic window messages: a few applications ignore them (games, raw-input programs, some global hotkeys); the system prompt already requires the agent to report honestly rather than blindly retry when an action has no effect.
  • In virtual mode, dialog keys such as Enter are routed according to the application's own focus semantics (e.g. activating the button that has focus), consistent with physical key behavior; that is expected behavior.
  • In physical mode (real), input on elevated windows (UAC high integrity) is blocked by UIPI and reported as an error.
  • Cannot inject on the lock screen / secure desktop (login screen).

Verification

Four scripts under test\ (all runnable directly, no API key required):

  • verify-load.mjs: plugin load validation (4 checks) — registers all 7 tools using the real dsh-tools compiler from the runtime, specifically guarding against "schema violations are only discovered after install, killing the backend".
  • verify-driver.mjs: screenshot / cursor / coordinate mapping / physical key basic path (7 checks).
  • verify-vinput.mjs: virtual input end-to-end (14 checks) — real listening windows receive messages, exact coordinates, Enter dialog-key routing, the physical cursor is never moved, and the blacklist rejects before delivery.
  • verify-amsi.mjs: all three driver scripts pass Windows Defender AMSI.

Schema hard constraint (from the 2026-08-24 backend-startup incident, guarded by verify-load.mjs): dsh-tools' value-schema DSL requires every type: object node (at any depth) to explicitly declare the boolean additionalProperties; in the parameter table, required may only be true or omitted (optional parameters must not write required: false). A violation throws UNSUPPORTED_SCHEMA at tool registration and kills the backend process.