At a glance
What it does
Voice dictation and push-to-talk transcription for the DeepSeek Harness web UI.
Web Profile
Not declared in supplied evidence
Evidence-verified
Checked Sep 6, 2026, 1:52 PM UTC
Code-evidenced contributions
What it adds to DSH
Adds browser dictation and push-to-talk voice-message controls to the DeepSeek Harness web composer.
Mechanism evidence ↗Registers a transcribe_audio tool for agents to transcribe audio files from disk.
Mechanism evidence ↗Before you choose it
dsh-voice adds pause-segmented dictation, voice messages, and mouse or keyboard push-to-talk to a DeepSeek Harness web profile. It supports browser speech recognition, cloud transcription providers, and optional local speech-to-text fallback chains; it also exposes an agent audio-transcription tool.
Best for
DeepSeek Harness web users who want to compose prompts by voice or let agents transcribe local audio files.
Common tasks
- Dictate text into the web composer, split at natural pauses.
- Record a voice message and cancel it during the configured send window.
- Configure cloud and local transcription providers as fallbacks.
- Let an agent transcribe an audio file using transcribe_audio.
Permissions and data
Requires browser microphone access for voice capture. Audio handling depends on the selected provider chain.
Permissions- Browser microphone access for dictation and voice messages.
- Access to configured local audio paths when an agent invokes transcribe_audio.
- Optional execution of local whisper.cpp and ffmpeg for local transcription.
- Browser recognition is described as local in Chrome.
- Cloud-provider modes may submit recorded audio to the configured provider.
- Credential references are described as resolved on the host rather than sent to browser clients.
- Optional Deepgram, Groq, Hugging Face, OpenAI-compatible, and other configured transcription services.
- Optional local whisper.cpp or SenseVoice/Sherpa-ONNX service.
- Selected cloud providers require their corresponding configured API-key credential references; browser and documented local modes do not require one.
Limitations
- Installation and runtime behavior were not executed in the supplied evidence.
- No DeepSeek Harness version range is declared in the supplied evidence.
- Local whisper configuration requires user-provided absolute paths to the whisper.cpp server and model.
- Provider availability, accuracy, privacy terms, and fallback behavior depend on the user’s configuration and were not independently verified.
What DSHub checked
- Pinned Git source and DSH bundle structure were verified.
- Package version 0.8.17, MIT license, web client platform, and declared peer dependencies were captured.
- Registry package identity was verified, but package contents were not audited.
What DSHub did not check
- Successful installation, browser microphone behavior, transcription quality, provider failover, and local-process execution were not tested.
Pinned install
Install dsh-voice
This plugin bundle does not have a DSH Plugin install action. Use its source documentation for the delivery method.
Maintainer source
Project README
📦 @goodandready/dsh-voice
<div align="center"><h3>Zero-Latency Streaming Dictation & Multi-Provider Voice Input for DeepSeek Harness</h3><p align="center"> <a href="https://www.npmjs.com/package/@goodandready/dsh-voice"><img src="https://img.shields.io/npm/v/@goodandready/dsh-voice.svg?style=for-the-badge&color=6366f1&labelColor=1e1b4b" alt="npm version"></a> <a href="LICENSE"><img src="https://img.shields.io/badge/License-MIT-10b981.svg?style=for-the-badge&color=10b981&labelColor=064e3b" alt="license"></a> <a href="https://github.com/topics/dsh-plugin"><img src="https://img.shields.io/badge/DSH-Plugin-8b5cf6.svg?style=for-the-badge&labelColor=2e1065" alt="DSH Plugin"></a> <a href="https://nodejs.org"><img src="https://img.shields.io/badge/Node-20%2B-f59e0b.svg?style=for-the-badge&labelColor=451a03" alt="Node version"></a> </p><p align="center"> <a href="https://goodandready.app/"><img src="https://img.shields.io/badge/All_Author_Projects-goodandready.app-ff4500.svg?style=for-the-badge&logo=rocket&logoColor=white&labelColor=1a1a2e" alt="All Author Projects"></a> </p><p align="center"> <a href="README.md"><b>🇬🇧 English</b></a> • <a href="README.ru.md"><b>🇷🇺 Русский</b></a> • <a href="README.zh.md"><b>🇨🇳 中文说明</b></a> </p></div>⚡ Overview
dsh-voice brings voice superpowers to the DeepSeek Harness Web UI. Whether you need hands-free real-time streaming dictation segmented on natural breath pauses or crisp voice notes with keyboard/mouse Push-to-Talk gestures, dsh-voice ensures your audio is never lost thanks to automatic multi-provider fallback chains.
graph LR
subgraph Client [Browser Web UI]
Mic[🎙️ Dictation Mic] -->|VAD Cut on Pause| Stream[Audio Chunks]
Wave[🌊 Voice Message] -->|Hold / Release| PTT[Push-to-Talk]
end
subgraph Host [DSH Host Backend]
Stream --> FFMPEG[ffmpeg 16kHz Transcoder]
PTT --> FFMPEG
FFMPEG --> Chain{Fallback Chain}
Chain -->|1st Priority| P1[Deepgram / Nova-2]
Chain -.->|On Rate Limit / 429| P2[Groq / Whisper Turbo]
Chain -.->|On Failure| P3[Local whisper.cpp / Offline]
end
subgraph Output [Target]
P1 --> Composer[💬 Web Composer / Chat]
P2 --> Composer
P3 --> Composer
end
style Client fill:#1e1e2e,stroke:#89b4fa,stroke-width:2px,color:#cdd6f4
style Host fill:#181825,stroke:#cba6f7,stroke-width:2px,color:#cdd6f4
style Output fill:#11111b,stroke:#a6e3a1,stroke-width:2px,color:#cdd6f4
✨ Key Features
- 🎙️ Streaming Dictation with VAD: Speech is automatically sliced at natural pauses (
vadSilenceMs, default 700ms) and typed into the composer in real time. - 🌊 Voice Notes with Cancel Window: Record your thought and have it automatically dispatched to the agent after a safety countdown (
autoSendMs, default 4000ms). - 🎮 Tactile Push-to-Talk:
- Mouse: Hold the wave button — releasing sends the message; dragging pointer away discards.
- Keyboard: Hold <kbd>Ctrl</kbd> (or custom hotkey) for hands-free speaking; press <kbd>Esc</kbd> to cancel.
- ⚡ Zero-Latency In-Browser Captions (
browser): Chrome Web Speech API recognition runs 100% locally with live floating captions as you speak. - 🛡️ Ironclad Multi-Provider Fallbacks: If your primary cloud provider runs out of credits or hits a 429 rate limit, requests seamlessly fail over down the chain.
- 🧠 Context Glossary Injection: Automatically extracts code variables and identifiers from your composer draft to steer STT model accuracy on technical jargon.
- 🎵 Embedded Audio Player: Preview, scrubber, and playback of your recorded voice message directly in chat and the composer dock.
- 🔇 Hardware Noise Suppression Toggle: Configurable in settings to toggle browser-level noise suppression, echo cancellation, and auto gain control.
- 📊 Provider Latency & Health Dashboard: Live visual telemetry of provider latency (ms), success rates, and errors directly within the settings UI.
- 🔒 Zero API Key Leakage: Keys are resolved on the host via
ctx.credentials(credentialRef) and never transmitted to browser clients. - 🖥️ Offline Local Whisper Server: Automatically boots and manages whisper.cpp (
whisper-server) with on-the-flyffmpegtranscode. - ⚡ SenseVoice-ONNX / Sherpa-ONNX (0.8.11): Ultra-fast (~50–100ms) non-autoregressive local STT engine with automatic emotion/event tag stripping. Supports both Sherpa-ONNX HTTP and OpenAI-compatible endpoints.
- 🌐 Realtime Audio Streaming (0.8.11): Low-latency WebSocket bridge (
/dsh-voice/realtime) for OpenAI Realtime API or local Sherpa-ONNX streaming. API keys stay securely on the host. - 🌊 Liquid Wave & Dynamic Orb Visualizer (0.8.12): Smooth animated audio visualization in the recording pill with real-time mic volume reactivity. Switch between organic multi-layer liquid waves, pulsating radiant orb, classic bars, or off.
🎮 Four Ways to Speak
| Mode | Gesture / Trigger | Behavior |
|---|---|---|
| Dictation | Click <kbd>🎙️ Mic</kbd> | Speech is sliced on pauses (vadSilenceMs) and typed live into composer |
| Voice Message | Click <kbd>🌊 Wave</kbd> | Records until stopped, then sends after cancel window (autoSendMs) |
| Mouse PTT | Hold <kbd>🌊 Wave</kbd> | Records while held; release sends message, drag off button to discard |
| Keyboard PTT | Hold <kbd>Ctrl</kbd> | Hands-free recording; release sends message, press <kbd>Esc</kbd> to discard |
[!TIP] You can customize the keyboard modifier in settings (
hotkey:Control,Alt,Shift, or anyKeyboardEvent.code).
🛠️ Supported Providers Matrix
| Provider Key | Service Backend | Default Model | Credential Ref | Features & Notes |
|---|---|---|---|---|
browser |
Web Speech API | Native Browser | None | Zero latency, floating live captions in Chrome |
deepgram |
Deepgram API | nova-2 |
DEEPGRAM_API_KEY |
Ultra-fast cloud transcription |
groq |
Groq Whisper | whisper-large-v3-turbo |
GROQ_API_KEY |
Near-instant inference speed |
hf |
HuggingFace Inference | openai/whisper-large-v3 |
HF_TOKEN |
High-accuracy open Whisper |
local-whisper |
Local whisper.cpp | Server defined | None | 100% private, offline, no internet needed |
sensevoice |
SenseVoice-ONNX / Sherpa-ONNX | SenseVoiceSmall |
None | Ultra-fast (~50ms) local non-autoregressive STT |
🚀 Ready-Made Presets (Plug & Play)
Just specify the name in your fallback chain and add the corresponding API key:
openai(whisper-1) →OPENAI_API_KEYsiliconflow(SenseVoiceSmall) →SILICONFLOW_API_KEYmistral(voxtral-mini-latest) →MISTRAL_API_KEYopenrouter(google/gemini-2.5-flash) →OPENROUTER_API_KEYdeepinfra(whisper-large-v3-turbo) →DEEPINFRA_API_KEYfireworks(whisper-v3-turbo) →FIREWORKS_API_KEY
📦 Quick Installation
dsh plugin --profile web add @goodandready/dsh-voice
[!IMPORTANT] Restart DSH Web UI after installation (
systemctl --user restart dsh-web) and refresh your browser tab.
⚙️ Configuration
Open Settings → Plugins → Plugin settings → Voice in the Web UI:
- id: dsh-voice
config:
dictation:
language: ru
vadSilenceMs: 700
chain:
- provider: deepgram
- provider: groq
- provider: local-whisper
message:
language: ru
autoSendMs: 4000
chain:
- provider: openai
- provider: local-whisper
hotkey: Control
autoStart: true
whisperModel: /models/ggml-medium-q8_0.bin
🤖 Agent Tool & HTTP API
Agent Tool (transcribe_audio)
Registers transcribe_audio(file_path, language?) in ctx.tools, allowing agents to analyze audio files, interview recordings, and voice notes directly from disk.
Internal HTTP Endpoints
POST /dsh-voice/transcribe—{ dataBase64, mimeType, mode }→{ ok, text, provider, tookMs }POST /dsh-voice/polish—{ text }→{ ok, text }GET /dsh-voice/status— Returns daemon status, active chains, SenseVoice and realtime config.GET /dsh-voice/realtime— WebSocket upgrade for low-latency audio streaming (OpenAI Realtime API / Sherpa-ONNX). Accepts binary audio chunks, returns JSON text deltas.
📄 License
MIT © GooDAnDReaDY
Operate deliberately
Install and manage
Prerequisites and target Profile
Target: Web Profile
Delivery: Dsh Bundle Git — GooDAnDReaDY/dsh-voice#febfe7b42e8400084acfe8d030c6cea9695731aa。
Verify, update, and remove
Show lifecycle commands
dsh plugin --profile web listCompatibility and access
DeepSeek Harness web profile with declared peer dependencies: Not declared in supplied evidence。
Review compatibility evidence ↗
Risk facts
Uses browser microphone input and may send audio to the transcription provider selected in its fallback chain.
Evidence ↗Optional local-whisper mode is documented to run whisper.cpp and transcode audio with ffmpeg.
Evidence ↗Evidence and editorial reviewManifest, Bundle patch, distribution and freshness
Immutable evidence
Review status and source activity
The README describes extensive UI and transcription capabilities; treat those runtime claims as publisher documentation until independently tested.
AI reviewed Sep 10, 2026, 11:36 AM UTC。GitHub facts last checked Sep 10, 2026, 11:36 AM UTC。
No material source change has been recorded since this evidence baseline.