voice-input
ποΈ Voice Input plugin for DeepSeek Harness β record, pause, playback review, and transcribe speech directly into message drafts with custom OpenAI-compatible models.
- Stars
- 1
- Language
- JavaScript
- Created
- Sep 27, 2026
- Updated
- Oct 1, 2026
Introduction
Voice Input for DeepSeek Harness
A feature-rich, high-performance voice recording and transcription plugin for DeepSeek Harness (DSH). It places a clean microphone icon directly beside the model selector in your composer bar, allowing you to record, pause, listen back to your audio, and transcribe it into your message draft using any OpenAI-compatible API or multimodal model (such as gemini-3.8-flash-high).
β¨ Features
- Toolbar Integration: Docked right beside the model selector in the composer toolbar without jumping across the input row.
- Multiple AI Engines / Providers:
- Save and manage multiple transcription backends (e.g. Gemini 3.8 Flash Proxy, OpenAI Whisper, Groq, local vLLM).
- Each engine has its own Display Name, Provider API Base URL, API Key, and Model ID.
- Switch active engines instantly via the quick-switch chip on the toolbar or inside the Settings modal.
- Add, edit, or delete provider engines with ease.
- Permanent Host Disk Persistence: Credentials and provider profiles are stored on the local disk (
%APPDATA%\dsh-desktop\voice-input-config.json) and synchronized with the browser, remaining preserved across restarts and port changes. - One-Click Recording: Click the microphone icon to begin recording audio immediately.
- Real-time Recording Timer: Live elapsed time counter (
MM:SS) with a visual recording indicator. - Pause & Resume: Pause recording anytime, and resume right from where you stopped.
- Audio Playback & Review: Listen to what you recorded before sending. Includes Play, Pause Playback, and Stop Playback controls.
- Discard / Cancel: Discard the recording at any stage with a single click.
- Direct Send: Click Send while recording or during review. Audio is converted to standard 16kHz mono WAV in the browser and forwarded to your active provider engine.
- Auto Draft Insertion: Transcribed text is automatically inserted directly into your conversation draft.
- Fail-Safe "Crashed Recording" Protection:
- If a transcription request fails (due to HTTP 429 quota exhaustion, network issues, or proxy timeouts), the raw audio is automatically saved to the
Crashed Recording/directory inside the project folder so your speech is never lost. - If transcription succeeds, the audio is automatically discarded, keeping your disk clean and clutter-free.
- In addition, an explicit Download button (β¬) is always available on the toolbar during review mode.
- If a transcription request fails (due to HTTP 429 quota exhaustion, network issues, or proxy timeouts), the raw audio is automatically saved to the
- Interactive Settings Modal: Clean, centered modal with Show/Hide API key toggle, provider management, and Escape key dismissal.
π¦ What Is Included vs. What Is Not
Included:
- Client Web UI Module (
client.js): Pure JavaScript component that mounts into theconversation.input.activityslot, managing recording, timers, playback, engine switching, and the settings modal. - Host Backend Service (
index.js): Fast Node.js service that hosts/api/voice-input/transcribeand/api/voice-input/config, managing disk persistence and forwarding audio to your OpenAI-compatible endpoint. - Multi-Modal & Whisper Fallback: Supports both OpenAI Chat Completions with
input_audio(e.g. Gemini 3.8 / GPT-4o multimodal models) and traditional/v1/audio/transcriptions(Whisper endpoints). - Desktop Microphone Enabler Utility (
enable-desktop-mic.cjs): Automates patching the DeepSeek Harness Desktop Electron app on Windows to grant media/microphone access.
NOT Included:
- No API Keys or Private URLs: By default, no API keys or backend URLs are bundled. You must supply your own provider endpoint and credentials in the Settings modal.
- No Heavy Native Dependencies: Uses native browser Web Audio API (
AudioContext,MediaRecorder) and Node.js standard built-ins (fetch,Buffer,Blob,fs).
π Transcription System Instruction
When transcribing speech through multimodal LLMs (e.g. gemini-3.8-flash-high), the backend applies this transcription directive:
"You are an expert transcriber and translator. Translate or transcribe the exact meaning of the audio into clean English. Remove any spoken filler words, stutters, repetitions, and hesitation marks (like 'um', 'uh', 'you know'). Do NOT add any summaries, conversational responses, or formatting. Output ONLY the raw, clean translated text."
π DeepSeek Harness Desktop App: Microphone Permission Fix
If you are using the DSH Desktop (Windows Electron desktop app), you might encounter:
Microphone access denied or unavailable: Permission denied
Why This Happens:
DeepSeek Harness Desktop is an Electron application. In its main process (out/main/index.js), it enforces a strict permission check:
function canGrantWindowPermission(permission, requestingUrl, isMainFrame) {
return (permission === "clipboard-sanitized-write" || permission === "notifications") && ...
}
Because the media (microphone) permission was omitted from this whitelist, Electron automatically denies all microphone requests before Windows even sees them. Consequently, DeepSeek Desktop never prompts for microphone access and does not show up in Windows Privacy Settings.
How to Fix It (Automated):
This repository provides an automated patch script: enable-desktop-mic.cjs.
- Close DSH Desktop.
- Run the patch script from your terminal:
The script automatically creates a backup (node enable-desktop-mic.cjsapp.asar.bak) and patchesapp.asarto includemediain the permission check. - Ensure Windows Desktop microphone access is enabled:
- Open Windows Settings (
Win + I) β Privacy & security β Microphone. - Make sure Microphone access is ON.
- Make sure Let apps access your microphone is ON.
- Scroll to the bottom and ensure Let desktop apps access your microphone is ON.
- Open Windows Settings (
- Relaunch DSH Desktop.
π Web Browser Access
You can also use DeepSeek Harness directly through any web browser (Google Chrome, Microsoft Edge, Brave, Firefox):
- Find your local Harness port (displayed when starting DSH, typically
http://127.0.0.1:43129or similar). - Open that URL in your browser.
- Click the microphone icon β your browser will display the standard permission prompt:
"127.0.0.1 wants to use your microphone: [Allow] [Block]"
- Click Allow.
π Installation into DeepSeek Harness
-
Clone or copy this repository into your workspace or plugins directory:
git clone https://github.com/likhonmain/voice-input.git -
Install the bundle using the Harness Plugin Manager or CLI:
- From DSH chat: ask the agent to install bundle
dsh-voice-inputfrom this folder path. - Or install via profile manifest.
- From DSH chat: ask the agent to install bundle
-
Open Settings by clicking the β (gear) icon next to the microphone icon in the composer bar:
- Configure one or more AI Engines with their Provider API Base URL, API Key, and Model ID.
- Choose which engine is Active.
-
Click Save All Engines and start speaking!
π License
MIT License. Feel free to use, modify, and distribute.