Back to Browse

Docker Talkies MCP Server

by Psyb0t
Developer ToolsUse Caution4.2MCP RegistryLocal
Free

Server data from the Official MCP Registry

Self-hosted MCP server for speech: ASR transcription, TTS synthesis, and file staging tools.

About

Self-hosted MCP server for speech: ASR transcription, TTS synthesis, and file staging tools.

Security Report

4.2
Use Caution4.2High Risk

This MCP server is a speech API bridge with reasonable security practices overall. Authentication is optional but properly implemented via bearer tokens. The main concerns are broad exception handling that could mask errors, potential information disclosure through debug logging of user input, and the inherent risks of accepting arbitrary file uploads and executing external commands (ffmpeg) on untrusted input. Permissions align well with the stated purpose. Supply chain analysis found 8 known vulnerabilities in dependencies (0 critical, 5 high severity).

4 files analyzed · 17 issues found

Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.

Permissions Required

This plugin requests these system permissions. Most are normal for its category.

File System Read

Reads files on your machine. Normal for tools that analyze or process local data.

File System Write

Writes or modifies files on your machine. Check that this is expected for the tool.

HTTP Network Access

Connects to external APIs or services over the internet.

env_vars

Check that this permission is expected for this type of plugin.

process_spawn

Check that this permission is expected for this type of plugin.

system_info

Check that this permission is expected for this type of plugin.

How to Install

Add this to your MCP configuration file:

{
  "mcpServers": {
    "io-github-psyb0t-talkies": {
      "args": [
        "-y",
        "@psyb0t/talkies"
      ],
      "command": "npx"
    }
  }
}

Documentation

View on GitHub

From the project's GitHub README.

talkies

CI version license Docker Pulls

Self-hosted speech services in one Docker image: OpenAI-compatible file transcription and text-to-speech, Talkies live ASR over WebSocket, file staging, model lifecycle controls, and an MCP endpoint for ASR workflows.

Contents

Start here

Restrict the first boot to the models you need; otherwise the entrypoint downloads every model in the bundled registry.

docker run --rm -it --name talkies \
  -p 127.0.0.1:8000:8000 \
  -v "$PWD/talkies-data:/data" \
  -e TALKIES_ENABLED_MODELS=whisper-large-v3-turbo,kokoro-82m \
  psyb0t/talkies:latest

curl -s http://127.0.0.1:8000/healthz
curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
  -F "file=@/path/to/clip.wav" \
  -F "model=whisper-large-v3-turbo"

For CUDA-only models — Parakeet-TDT, the larger Canary models, Qwen3 TTS and Chatterbox Turbo — use psyb0t/talkies:latest-cuda with --gpus all. The loopback port mapping keeps the service local; see Getting started for first boot and authentication.

What it provides

SurfacePurposeReference
POST /v1/audio/transcriptionsFile transcription and subtitlesHTTP API
WS /v1/audio/transcriptions/streamLive 16 kHz PCM ASRStreaming
POST /v1/audio/speechSpeech synthesis in six formatsHTTP API
GET /v1/modelsEnabled slugs and their modalityHTTP API
GET /v1/audio/voicesPer-model voice catalog with origin tagsModels
GET/PUT/DELETE /v1/files/*Server-side file stagingHTTP API
/api/ps, /unloadModel inspection and evictionOperations
/v1/mcpStreamable HTTP MCP with ASR/file toolsHTTP API
GET /healthzLiveness probe; the only unauthenticated routeOperations

The HTTP transcription and speech routes use the corresponding OpenAI wire shapes where those contracts overlap. Streaming ASR, files, lifecycle controls, and MCP are Talkies extensions.

Models at a glance

  • CPU: two Whisper models, Canary-180M-Flash, Nemotron ASR via parakeet.cpp, four English Sherpa-ONNX Zipformer choices, Vosk small English, two phoneme recognizers, and two Kokoro TTS backends.
  • CUDA: the CPU set plus Parakeet-TDT, Canary 1B/Qwen ASR, five Qwen3 TTS variants, and Chatterbox Turbo.
  • Live ASR: bundled Nemotron, Sherpa-ONNX, and Vosk are native; bundled Whisper is a bounded rolling decoder. Sherpa and Vosk also work through the OpenAI-compatible file-transcription endpoint.
  • Phoneme recognition: wav2vec2-xlsr-53-espeak and zipa-ipa return the IPA phones that were spoken, not words, with no language model correcting them toward the nearest dictionary entry. Same transcription endpoint and timestamp options as the other ASR models; see Phoneme recognition.
  • Per-model concurrency limits cover WebSocket, HTTP, MCP, ASR, and TTS; the bundled Nemotron CPU and CUDA entries admit two requests.
  • Streaming TTS: Qwen3 returns incremental raw PCM for response_format="pcm"; other TTS formats and Kokoro are buffered.
  • Expressive TTS: Chatterbox Turbo (English) takes 19 inline tags such as [sigh], [whispering] and [laugh] directly in the input text. Its output carries a neural watermark by default; set TALKIES_CHATTERBOX_WATERMARK to false to emit unmarked audio.
  • Voice cloning: drop a .wav into /data/custom-voices and it appears on GET /v1/audio/voices. Qwen3 pairs it with an optional sibling .txt transcript; Chatterbox needs only the clip, longer than five seconds.

Exact slugs, executors, tag list, and registry format: Models and registries.

Reading phonemes

wav2vec2-xlsr-53-espeak and zipa-ipa use the same transcription call as every other ASR slug; only the model changes. text comes back as a space-separated IPA phone stream rather than words, and no language model corrects a mispronunciation toward a real word.

curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
  -F "file=@/path/to/clip.wav" \
  -F "model=zipa-ipa"
# {"text": "a ɪ m k ə n f j u z ...", ...}

Add -F "response_format=verbose_json" (or timestamp_granularities[]=word) to get each phone as a words entry with start and end in seconds.

Prompting Chatterbox with emotion

Tags go inline in input, in square brackets, lowercase. They are real tokens in the model's tokenizer, so only these 19 do anything — any other bracketed word is spoken as literal text:

[angry] [fear] [surprised] [whispering] [advertisement] [dramatic] [narration]
[crying] [happy] [sarcastic] [clear throat] [sigh] [shush] [cough] [groan]
[sniff] [gasp] [chuckle] [laugh]
curl -s http://127.0.0.1:8000/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{
        "model": "chatterbox-turbo",
        "voice": "builtin",
        "input": "Oh, that is hilarious. [chuckle] Anyway [sigh] back to work.",
        "response_format": "mp3"
      }' --output out.mp3

Swap "voice" for the name of any .wav you dropped in /data/custom-voices (extension stripped) to speak the same line in a cloned voice.

Documentation

GuideContents
Getting startedRun CPU/CUDA, persist data, authenticate, verify
Models and registriesBundled slugs, image availability, custom registries
ArchitectureRequest flow, backend selection, on-disk layout
HTTP APIRequests, responses, files, lifecycle, MCP
StreamingLive ASR protocol, streaming backends, PCM TTS
ConfigurationSupported environment variables and limits
Operations and securityExposure, model memory, data retention, logs
DevelopmentMake targets, test suites, image builds

Agent integrations

The Talkies skill teaches agents to use the HTTP, WebSocket, and MCP surfaces. Install it through the shared psyb0t marketplace or let Codex discover it directly from this checkout.

Claude Code

claude plugin marketplace add psyb0t/agents
claude plugin install talkies@psyb0t

Claude Code prompts for the Talkies URL and, when enabled, the bearer token; the sensitive token is stored through the client's protected configuration.

Codex

codex plugin marketplace add psyb0t/agents
codex plugin add talkies@psyb0t

A marketplace install invokes the skill as $talkies:talkies. Codex also discovers .agents/skills/talkies directly in this repository, where it is invoked as $talkies without installation.

OpenClaw

The skill and MCP bridge are published through ClawHub:

openclaw skills install @psyb0t/talkies
openclaw plugins install clawhub:@psyb0t/talkies

The bridge connects local stdio MCP clients to a running Talkies /v1/mcp endpoint. Set TALKIES_URL and, when authentication is enabled, TALKIES_AUTH_TOKEN.

Security in one minute

TALKIES_AUTH_TOKEN enables a shared bearer token for every HTTP and WebSocket route except /healthz. It is unset by default. Keep the port loopback-only or put Talkies behind TLS, authentication, and rate limiting. If untrusted callers can supply remote file_path URLs, set TALKIES_BLOCK_PRIVATE_DOWNLOADS=true. See Operations and security for the complete posture.

Development

make check                 # lint + unit tests in the dev image
make lint                  # flake8 + mypy only
make test-unit             # fast offline unit tests
make run                   # run the CPU image locally
make test-streaming        # real CPU native WebSocket ASR test
make test-streaming-custom # real CPU Sherpa/Vosk WebSocket + HTTP tests
make test-streaming-custom-cuda # real CUDA Sherpa WebSocket + HTTP test
make compile-heavy         # regenerate the hash-locked ML requirements
make build-all             # CPU and CUDA production images

make help lists every target.

Talkies is released under the WTFPL. Model weights are downloaded at runtime and have their own terms; image component notices are in THIRD_PARTY.md. Release notes are in CHANGELOG.md.

Reviews

No reviews yet

Be the first to review this server!