← Back to home@Gishi1

OpenAI-TTS-API-Relay-Server

No description

Stars
0
Language
Python
Created
Jan 29, 2026
Updated
Jan 29, 2026

Introduction

OpenAI TTS API Relay Server

A relay server that provides an OpenAI-compatible TTS API with support for multiple backend providers, both free and paid.

Features

  • OpenAI API Compatible - Drop-in replacement for OpenAI's /v1/audio/speech endpoint
  • Multiple Providers - Switch between different TTS backends seamlessly
  • Free Options - Edge TTS and gTTS work without any API keys
  • Paid Options - Support for OpenAI, ElevenLabs, Azure, and Google Cloud
  • Proxy Support - Route requests through HTTP/HTTPS proxies
  • Native Voice Names - Use provider's original voice names directly (default)
  • Optional Voice Mapping - Optionally map OpenAI voice names to provider-specific voices
  • API Key Profiles - Different configurations per API key for multi-tenant setups
  • Streaming Support - Stream audio for supported providers
  • Web UI - Next.js management interface for testing and configuration
  • Docker Ready - Easy deployment with Docker and Docker Compose

Supported Providers

ProviderFreeAPI Key RequiredStreamingQuality
Edge TTS (Microsoft)YesNoYesHigh
gTTS (Google Translate)YesNoNoBasic
OpenAINoYesYesHigh
ElevenLabsNoYesYesVery High
Azure Cognitive ServicesNoYesYesHigh
Google Cloud TTSNoYesNoHigh
Coqui TTS (Self-hosted)YesNoNoVaries

Quick Start

Using pip

# Clone the repository
git clone https://github.com/yourusername/openai-tts-relay.git
cd openai-tts-relay

# Install dependencies
pip install -r requirements.txt

# Run the server (uses Edge TTS by default - free, no API key needed)
python -m tts_relay.main

Using Docker

# Build and run
docker-compose up -d

# Or build manually
docker build -t openai-tts-relay .
docker run -p 8000:8000 openai-tts-relay

Usage

Basic API Usage (curl)

# Generate speech using native provider voice names (default behavior)
# Use the actual Edge TTS voice name directly
curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1",
    "input": "Hello, this is a test of the TTS relay server.",
    "voice": "en-US-AriaNeural"
  }' \
  --output speech.mp3

# Use a specific provider with its native voice names
curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-edge",
    "input": "Using Edge TTS provider.",
    "voice": "en-US-JennyNeural"
  }' \
  --output speech.mp3

# Specify provider in the request body
curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1",
    "input": "Using gTTS provider.",
    "voice": "en",
    "provider": "gtts"
  }' \
  --output speech.mp3

# List available voices for a provider
curl http://localhost:8000/v1/voices?provider=edge

Using OpenAI Python Library

from openai import OpenAI

# Point to your relay server
client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-needed"  # Unless you've configured authentication
)

# Generate speech using native provider voice names
response = client.audio.speech.create(
    model="tts-1",  # or "tts-edge", "tts-gtts", etc.
    voice="en-US-AriaNeural",  # Use native Edge TTS voice name
    input="Hello from the TTS relay server!"
)

# Save to file
response.stream_to_file("speech.mp3")

Using JavaScript/Node.js

import OpenAI from 'openai';
import fs from 'fs';

const openai = new OpenAI({
  baseURL: 'http://localhost:8000/v1',
  apiKey: 'not-needed',
});

async function generateSpeech() {
  const response = await openai.audio.speech.create({
    model: 'tts-1',
    voice: 'en-US-JennyNeural',  // Use native Edge TTS voice name
    input: 'Hello from JavaScript!',
  });

  const buffer = Buffer.from(await response.arrayBuffer());
  fs.writeFileSync('speech.mp3', buffer);
}

generateSpeech();

Configuration

Environment Variables

# Server
TTS_SERVER__HOST=0.0.0.0
TTS_SERVER__PORT=8000
TTS_SERVER__LOG_LEVEL=info
TTS_SERVER__API_KEY=your-secret-key  # Optional authentication

# Default provider
TTS_PROVIDERS__DEFAULT=edge

# Proxy (optional)
HTTP_PROXY=http://proxy:8080
HTTPS_PROXY=http://proxy:8080

# Provider API keys (optional - enables paid providers)
OPENAI_API_KEY=sk-...
ELEVENLABS_API_KEY=...
AZURE_SPEECH_KEY=...
AZURE_SPEECH_REGION=eastus
GOOGLE_APPLICATION_CREDENTIALS=/path/to/credentials.json

Configuration File

Create a config.yaml file (see config.example.yaml):

server:
  host: "0.0.0.0"
  port: 8000
  log_level: "info"

providers:
  default: "edge"

  edge:
    enabled: true
    default_voice: "en-US-AriaNeural"
    # By default, use native provider voice names directly
    use_voice_mapping: false
    # Optional: enable voice mapping to use OpenAI voice names
    # use_voice_mapping: true
    # voice_mapping:
    #   alloy: "en-US-AriaNeural"
    #   echo: "en-US-GuyNeural"

  elevenlabs:
    enabled: true
    api_key: "your-api-key"
    use_voice_mapping: false

API Endpoints

MethodEndpointDescription
POST/v1/audio/speechGenerate speech (OpenAI compatible)
GET/v1/modelsList available models
GET/v1/providersList available providers
GET/v1/voicesList voices for a provider
GET/v1/profileGet current API key profile info
GET/healthHealth check
GET/docsSwagger UI documentation

Request Format

{
  "model": "tts-1",
  "input": "Text to convert to speech",
  "voice": "alloy",
  "response_format": "mp3",
  "speed": 1.0,
  "provider": "edge"  // Optional: override default provider
}

Voice Names

Default Behavior: Native Provider Voice Names

By default (use_voice_mapping: false), the server passes voice names directly to providers. Use the provider's native voice names:

# Edge TTS - use Microsoft neural voice names
curl ... -d '{"voice": "en-US-AriaNeural", "provider": "edge"}'

# gTTS - use language codes
curl ... -d '{"voice": "en", "provider": "gtts"}'

# ElevenLabs - use voice IDs
curl ... -d '{"voice": "21m00Tcm4TlvDq8ikWAM", "provider": "elevenlabs"}'

To discover available voices, use the /v1/voices endpoint:

curl http://localhost:8000/v1/voices?provider=edge

Optional: OpenAI Voice Mapping

Enable use_voice_mapping: true in config to map OpenAI voice names to provider voices:

OpenAI VoiceEdge TTSgTTSElevenLabs
alloyen-US-AriaNeuralenRachel
echoen-US-GuyNeuralenDomi
fableen-GB-SoniaNeuralen-ukBella
onyxen-US-DavisNeuralenArnold
novaen-US-JennyNeuralenDorothy
shimmeren-US-AnaNeuralenAdam

You can customize mappings in the config file.

API Key Profiles

You can configure different settings for different API keys, allowing multi-tenant setups where each user/application has their own provider configuration.

server:
  # API key profiles - each key has its own configuration
  api_key_profiles:
    # User 1: Uses ElevenLabs with their own API key
    "sk-user1-abc123":
      name: "user1"
      description: "User 1 - Premium ElevenLabs"
      providers:
        default: "elevenlabs"
        elevenlabs:
          enabled: true
          api_key: "user1-elevenlabs-key"
          default_voice: "21m00Tcm4TlvDq8ikWAM"

    # User 2: Uses Edge TTS with Spanish voices
    "sk-user2-xyz789":
      name: "user2"
      description: "User 2 - Spanish Edge TTS"
      providers:
        default: "edge"
        edge:
          enabled: true
          default_voice: "es-ES-ElviraNeural"

    # User 3: Uses OpenAI TTS with a proxy
    "sk-user3-def456":
      name: "user3"
      providers:
        default: "openai"
        openai:
          enabled: true
          api_key: "sk-openai-key-for-user3"
      proxy:
        enabled: true
        http_url: "http://user3-proxy:8080"

Each profile can override:

  • providers - Different TTS providers and their settings
  • proxy - Different proxy settings

Check the current profile with the /v1/profile endpoint:

curl -H "Authorization: Bearer sk-user1-abc123" http://localhost:8000/v1/profile

Proxy Support

Configure proxy for all outbound requests:

# config.yaml
proxy:
  enabled: true
  http_url: "http://proxy.example.com:8080"
  https_url: "http://proxy.example.com:8080"
  no_proxy:
    - "localhost"
    - "127.0.0.1"

Or via environment variables:

HTTP_PROXY=http://proxy:8080
HTTPS_PROXY=http://proxy:8080

Adding Custom Providers

To add a new TTS provider:

  1. Create a new file in tts_relay/providers/
  2. Implement the TTSProvider interface
  3. Register the provider with register_provider()

Example:

from tts_relay.providers.base import TTSProvider, TTSResult
from tts_relay.providers.registry import register_provider

class MyCustomProvider(TTSProvider):
    name = "custom"
    display_name = "My Custom TTS"
    is_free = True
    supports_streaming = False
    supported_formats = [ResponseFormat.MP3]

    async def synthesize(self, text, voice, format, speed, **kwargs):
        # Your implementation here
        return TTSResult(audio_data=b"...", content_type="audio/mpeg", format=format)

    async def list_voices(self):
        return [VoiceInfo(voice_id="default", name="Default", provider=self.name)]

register_provider("custom", MyCustomProvider)

Web UI

A Next.js management interface is included in the web/ directory.

Features

  • Connect to any TTS relay server instance
  • Browse available providers and their status
  • Browse voices for each provider
  • Test TTS synthesis with any voice
  • Download generated audio files

Running the Web UI

cd web

# Install dependencies
npm install

# Run development server
npm run dev

# Build for production
npm run build
npm start

The UI will be available at http://localhost:3000

Screenshots

The web UI provides:

  • Dashboard - Server connection status and profile info
  • TTS Tester - Generate speech with any provider and voice
  • Providers - View all configured providers and their settings
  • Voices - Browse available voices for each provider

Development

# Install with dev dependencies
pip install -e ".[dev]"

# Run with auto-reload
python -m tts_relay.main --reload

# Run tests
pytest

License

MIT License - see LICENSE file.