Models

EMA Lightning Outperforms ElevenLabs on Turkish TTS

Developer Canberk Aslan has released EMA Lightning, an open-source, 34 MB Turkish text-to-speech model that outperforms massive commercial alternatives at a fraction of the computing cost.

AlphaSignal22 hrs agoModels
Image: AlphaSignal

Developer Canberk Aslan has launched EMA Lightning, an 8.6 million-parameter text-to-speech (TTS) model designed specifically for the Turkish language. Released under the Apache 2.0 license, the entire model stack takes up just 34 MB of storage. Despite its tiny footprint, EMA Lightning runs locally on both CPUs and GPUs, eliminating the need for cloud APIs and offering developers a highly efficient, offline solution for speech generation.

In evaluations on the Freya-TR-Eval benchmark, which consists of 495 Turkish sentences, EMA Lightning achieved a 0.92% word error rate (WER). This performance surpasses ElevenLabs v4, Gemini 3.8 Flash, and Trendyol-TTS, a model that is roughly 277 times larger at 2.38 billion parameters and scored a 0.95% WER. While its UTMOS score of 3.30 matches ElevenLabs, it trails Gemini and Trendyol-TTS in predicted naturalness.

The model is built on a Diffusion Transformer (DiT) architecture with flow matching, distilled down to just four steps using DMD2. It also incorporates shared modulation and a windowed Gaussian aligner. This design allows EMA Lightning to generate its first audio in a mere 3.86 milliseconds. On an NVIDIA RTX 4090 GPU, it runs 440 times faster than real-time for single requests and 1,316 times faster when batched, translating to an estimated cost of just $0.0085 per million characters.

For AI practitioners, EMA Lightning represents a massive shift toward edge-based deployment for localized voice applications. By offering a single-voice, Turkish-only pipeline that can be installed simply via pip, it allows developers to bypass expensive cloud APIs. The model provides predictable latency and high intelligibility for resource-constrained environments, though developers should note the lack of emotional range and the need for broader independent evaluation.

This is our own summary of reporting by AlphaSignal

More in Models