OuteTTS

Pure language modeling approach to TTS without traditional audio codecs.

Open SourceSelf HostedOffline CapableGPU Required (4GB+ VRAM)
0.0 (0)

About

OuteTTS takes an unusual approach to speech synthesis, treating it purely as a language modeling problem: audio is produced through next-token prediction rather than a conventional pipeline of separate acoustic models and vocoders. The project from OuteAI ships models at 0.6B and 1B parameters across several releases up to version 1.0, supports voice cloning through custom speaker profiles created from reference audio in a few lines of code, and generates clips of roughly 42 seconds per run. It is distributed as both a Python package on PyPI and a JavaScript package on npm, and runs on many backends including llama.cpp, Hugging Face Transformers, ExLlamaV2, vLLM, and Transformers.js, with acceleration on CUDA, ROCm, Apple Metal, and Vulkan hardware. Because the models behave like ordinary language models, they also load in third-party runtimes such as KoboldCPP. The lightweight footprint and broad backend support appeal to developers who want local, open text-to-speech without heavyweight dependencies.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
Minimum VRAM
4 GB
Added
Apr 3, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate
0.0 (0)

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Browse all Text-to-Speech (TTS) tools