Evidence snapshot reviewed Sep 10, 2026GitHub checked Aug 21, 2026
Evidence-verifiedPlugin BundleSearch, Vision & DataWeb Profile

dsh-voice

Voice dictation and push-to-talk transcription for the DeepSeek Harness web UI.

At a glance

What it does

Voice dictation and push-to-talk transcription for the DeepSeek Harness web UI.

Use cases
Search, Vision & DataUIMedia CaptureIntegrations
Works with
Deepseek HarnessWebSpeech To Text
Compatibility

Web Profile
Not declared in supplied evidence

Trust & status

Evidence-verified
Checked Sep 6, 2026, 1:52 PM UTC

Code-evidenced contributions

What it adds to DSH

Web UIVoice input controls

Adds browser dictation and push-to-talk voice-message controls to the DeepSeek Harness web composer.

Mechanism evidence
Model ToolsAudio transcription tool

Registers a transcribe_audio tool for agents to transcribe audio files from disk.

Mechanism evidence

Before you choose it

dsh-voice adds pause-segmented dictation, voice messages, and mouse or keyboard push-to-talk to a DeepSeek Harness web profile. It supports browser speech recognition, cloud transcription providers, and optional local speech-to-text fallback chains; it also exposes an agent audio-transcription tool.

Best for

DeepSeek Harness web users who want to compose prompts by voice or let agents transcribe local audio files.

Common tasks

  • Dictate text into the web composer, split at natural pauses.
  • Record a voice message and cancel it during the configured send window.
  • Configure cloud and local transcription providers as fallbacks.
  • Let an agent transcribe an audio file using transcribe_audio.

Permissions and data

Requires browser microphone access for voice capture. Audio handling depends on the selected provider chain.

Permissions
  • Browser microphone access for dictation and voice messages.
  • Access to configured local audio paths when an agent invokes transcribe_audio.
  • Optional execution of local whisper.cpp and ffmpeg for local transcription.
Data handling
  • Browser recognition is described as local in Chrome.
  • Cloud-provider modes may submit recorded audio to the configured provider.
  • Credential references are described as resolved on the host rather than sent to browser clients.
External services
  • Optional Deepgram, Groq, Hugging Face, OpenAI-compatible, and other configured transcription services.
  • Optional local whisper.cpp or SenseVoice/Sherpa-ONNX service.
Credentials
  • Selected cloud providers require their corresponding configured API-key credential references; browser and documented local modes do not require one.

Limitations

  • Installation and runtime behavior were not executed in the supplied evidence.
  • No DeepSeek Harness version range is declared in the supplied evidence.
  • Local whisper configuration requires user-provided absolute paths to the whisper.cpp server and model.
  • Provider availability, accuracy, privacy terms, and fallback behavior depend on the user’s configuration and were not independently verified.

What DSHub checked

  • Pinned Git source and DSH bundle structure were verified.
  • Package version 0.8.17, MIT license, web client platform, and declared peer dependencies were captured.
  • Registry package identity was verified, but package contents were not audited.

What DSHub did not check

  • Successful installation, browser microphone behavior, transcription quality, provider failover, and local-process execution were not tested.

Pinned install

Install dsh-voice

This plugin bundle does not have a DSH Plugin install action. Use its source documentation for the delivery method.

Visit the source project

Maintainer source

Project README

View at commit febfe7b
Maintainer-authored contentCaptured from README.md on Sep 6, 2026. The text and repository-relative media are fixed to commit febfe7b42e84 with content hash 01132861bb03; provider-hosted badges may update independently. README commands are upstream documentation; the DSHub copy action above is the verified, version-pinned install.

📦 @goodandready/dsh-voice

<div align="center"><h3>Zero-Latency Streaming Dictation & Multi-Provider Voice Input for DeepSeek Harness</h3><p align="center"> <a href="https://www.npmjs.com/package/@goodandready/dsh-voice"><img src="https://img.shields.io/npm/v/@goodandready/dsh-voice.svg?style=for-the-badge&color=6366f1&labelColor=1e1b4b" alt="npm version"></a> <a href="LICENSE"><img src="https://img.shields.io/badge/License-MIT-10b981.svg?style=for-the-badge&color=10b981&labelColor=064e3b" alt="license"></a> <a href="https://github.com/topics/dsh-plugin"><img src="https://img.shields.io/badge/DSH-Plugin-8b5cf6.svg?style=for-the-badge&labelColor=2e1065" alt="DSH Plugin"></a> <a href="https://nodejs.org"><img src="https://img.shields.io/badge/Node-20%2B-f59e0b.svg?style=for-the-badge&labelColor=451a03" alt="Node version"></a> </p><p align="center"> <a href="https://goodandready.app/"><img src="https://img.shields.io/badge/All_Author_Projects-goodandready.app-ff4500.svg?style=for-the-badge&logo=rocket&logoColor=white&labelColor=1a1a2e" alt="All Author Projects"></a> </p><p align="center"> <a href="README.md"><b>🇬🇧 English</b></a> • <a href="README.ru.md"><b>🇷🇺 Русский</b></a> • <a href="README.zh.md"><b>🇨🇳 中文说明</b></a> </p></div>

⚡ Overview

dsh-voice brings voice superpowers to the DeepSeek Harness Web UI. Whether you need hands-free real-time streaming dictation segmented on natural breath pauses or crisp voice notes with keyboard/mouse Push-to-Talk gestures, dsh-voice ensures your audio is never lost thanks to automatic multi-provider fallback chains.

graph LR
    subgraph Client [Browser Web UI]
        Mic[🎙️ Dictation Mic] -->|VAD Cut on Pause| Stream[Audio Chunks]
        Wave[🌊 Voice Message] -->|Hold / Release| PTT[Push-to-Talk]
    end

    subgraph Host [DSH Host Backend]
        Stream --> FFMPEG[ffmpeg 16kHz Transcoder]
        PTT --> FFMPEG
        FFMPEG --> Chain{Fallback Chain}
        
        Chain -->|1st Priority| P1[Deepgram / Nova-2]
        Chain -.->|On Rate Limit / 429| P2[Groq / Whisper Turbo]
        Chain -.->|On Failure| P3[Local whisper.cpp / Offline]
    end

    subgraph Output [Target]
        P1 --> Composer[💬 Web Composer / Chat]
        P2 --> Composer
        P3 --> Composer
    end

    style Client fill:#1e1e2e,stroke:#89b4fa,stroke-width:2px,color:#cdd6f4
    style Host fill:#181825,stroke:#cba6f7,stroke-width:2px,color:#cdd6f4
    style Output fill:#11111b,stroke:#a6e3a1,stroke-width:2px,color:#cdd6f4

✨ Key Features

  • 🎙️ Streaming Dictation with VAD: Speech is automatically sliced at natural pauses (vadSilenceMs, default 700ms) and typed into the composer in real time.
  • 🌊 Voice Notes with Cancel Window: Record your thought and have it automatically dispatched to the agent after a safety countdown (autoSendMs, default 4000ms).
  • 🎮 Tactile Push-to-Talk:
    • Mouse: Hold the wave button — releasing sends the message; dragging pointer away discards.
    • Keyboard: Hold <kbd>Ctrl</kbd> (or custom hotkey) for hands-free speaking; press <kbd>Esc</kbd> to cancel.
  • Zero-Latency In-Browser Captions (browser): Chrome Web Speech API recognition runs 100% locally with live floating captions as you speak.
  • 🛡️ Ironclad Multi-Provider Fallbacks: If your primary cloud provider runs out of credits or hits a 429 rate limit, requests seamlessly fail over down the chain.
  • 🧠 Context Glossary Injection: Automatically extracts code variables and identifiers from your composer draft to steer STT model accuracy on technical jargon.
  • 🎵 Embedded Audio Player: Preview, scrubber, and playback of your recorded voice message directly in chat and the composer dock.
  • 🔇 Hardware Noise Suppression Toggle: Configurable in settings to toggle browser-level noise suppression, echo cancellation, and auto gain control.
  • 📊 Provider Latency & Health Dashboard: Live visual telemetry of provider latency (ms), success rates, and errors directly within the settings UI.
  • 🔒 Zero API Key Leakage: Keys are resolved on the host via ctx.credentials (credentialRef) and never transmitted to browser clients.
  • 🖥️ Offline Local Whisper Server: Automatically boots and manages whisper.cpp (whisper-server) with on-the-fly ffmpeg transcode.
  • SenseVoice-ONNX / Sherpa-ONNX (0.8.11): Ultra-fast (~50–100ms) non-autoregressive local STT engine with automatic emotion/event tag stripping. Supports both Sherpa-ONNX HTTP and OpenAI-compatible endpoints.
  • 🌐 Realtime Audio Streaming (0.8.11): Low-latency WebSocket bridge (/dsh-voice/realtime) for OpenAI Realtime API or local Sherpa-ONNX streaming. API keys stay securely on the host.
  • 🌊 Liquid Wave & Dynamic Orb Visualizer (0.8.12): Smooth animated audio visualization in the recording pill with real-time mic volume reactivity. Switch between organic multi-layer liquid waves, pulsating radiant orb, classic bars, or off.

🎮 Four Ways to Speak

Mode Gesture / Trigger Behavior
Dictation Click <kbd>🎙️ Mic</kbd> Speech is sliced on pauses (vadSilenceMs) and typed live into composer
Voice Message Click <kbd>🌊 Wave</kbd> Records until stopped, then sends after cancel window (autoSendMs)
Mouse PTT Hold <kbd>🌊 Wave</kbd> Records while held; release sends message, drag off button to discard
Keyboard PTT Hold <kbd>Ctrl</kbd> Hands-free recording; release sends message, press <kbd>Esc</kbd> to discard

[!TIP] You can customize the keyboard modifier in settings (hotkey: Control, Alt, Shift, or any KeyboardEvent.code).


🛠️ Supported Providers Matrix

Provider Key Service Backend Default Model Credential Ref Features & Notes
browser Web Speech API Native Browser None Zero latency, floating live captions in Chrome
deepgram Deepgram API nova-2 DEEPGRAM_API_KEY Ultra-fast cloud transcription
groq Groq Whisper whisper-large-v3-turbo GROQ_API_KEY Near-instant inference speed
hf HuggingFace Inference openai/whisper-large-v3 HF_TOKEN High-accuracy open Whisper
local-whisper Local whisper.cpp Server defined None 100% private, offline, no internet needed
sensevoice SenseVoice-ONNX / Sherpa-ONNX SenseVoiceSmall None Ultra-fast (~50ms) local non-autoregressive STT

🚀 Ready-Made Presets (Plug & Play)

Just specify the name in your fallback chain and add the corresponding API key:

  • openai (whisper-1) → OPENAI_API_KEY
  • siliconflow (SenseVoiceSmall) → SILICONFLOW_API_KEY
  • mistral (voxtral-mini-latest) → MISTRAL_API_KEY
  • openrouter (google/gemini-2.5-flash) → OPENROUTER_API_KEY
  • deepinfra (whisper-large-v3-turbo) → DEEPINFRA_API_KEY
  • fireworks (whisper-v3-turbo) → FIREWORKS_API_KEY

📦 Quick Installation

dsh plugin --profile web add @goodandready/dsh-voice

[!IMPORTANT] Restart DSH Web UI after installation (systemctl --user restart dsh-web) and refresh your browser tab.


⚙️ Configuration

Open Settings → Plugins → Plugin settings → Voice in the Web UI:

- id: dsh-voice
  config:
    dictation:
      language: ru
      vadSilenceMs: 700
      chain:
        - provider: deepgram
        - provider: groq
        - provider: local-whisper
    message:
      language: ru
      autoSendMs: 4000
      chain:
        - provider: openai
        - provider: local-whisper
    hotkey: Control
    autoStart: true
    whisperModel: /models/ggml-medium-q8_0.bin

🤖 Agent Tool & HTTP API

Agent Tool (transcribe_audio)

Registers transcribe_audio(file_path, language?) in ctx.tools, allowing agents to analyze audio files, interview recordings, and voice notes directly from disk.

Internal HTTP Endpoints

  • POST /dsh-voice/transcribe{ dataBase64, mimeType, mode }{ ok, text, provider, tookMs }
  • POST /dsh-voice/polish{ text }{ ok, text }
  • GET /dsh-voice/status — Returns daemon status, active chains, SenseVoice and realtime config.
  • GET /dsh-voice/realtimeWebSocket upgrade for low-latency audio streaming (OpenAI Realtime API / Sherpa-ONNX). Accepts binary audio chunks, returns JSON text deltas.

📄 License

MIT © GooDAnDReaDY

Operate deliberately

Install and manage

Prerequisites and target Profile

Target Web Profile

Delivery Dsh Bundle Git — GooDAnDReaDY/dsh-voice#febfe7b42e8400084acfe8d030c6cea9695731aa

Verify, update, and remove

Show lifecycle commands
Verify
dsh plugin --profile web list

Compatibility and access

DeepSeek Harness web profile with declared peer dependencies Not declared in supplied evidence

Review compatibility evidence

Risk facts

Microphone And Audio

Uses browser microphone input and may send audio to the transcription provider selected in its fallback chain.

Evidence
Local Processes

Optional local-whisper mode is documented to run whisper.cpp and transcode audio with ffmpeg.

Evidence
Evidence and editorial reviewManifest, Bundle patch, distribution and freshness

Immutable evidence

Review status and source activity

AI reviewed

The README describes extensive UI and transcription capabilities; treat those runtime claims as publisher documentation until independently tested.

AI reviewed Sep 10, 2026, 11:36 AM UTCGitHub facts last checked Sep 10, 2026, 11:36 AM UTC

No material source change has been recorded since this evidence baseline.

Next step

Follow the Plugin installation workflow

Subscribe to material changes for dsh-voice