Why this matters
Real-time, high-quality Turkish TTS usually requires large models or cloud APIs. EMA Lightning flips that trade-off: at just 8.6M parameters (~34 MB) it matches or beats larger systems on transcription-based accuracy while delivering millisecond-scale first audio and very high throughput on commodity GPUs — making private, low-cost, low-latency on-device TTS practical for production services.
What Sets It Apart
- Compact yet accurate: a 5.6M-parameter DiT acoustic model plus a 3M vocoder (total 8.6M) delivers 0.92% WER on Freya-TR-Eval, the best result reported on that benchmark against larger systems. That low WER means clearer, more faithful reading of normalized Turkish text.
- Extremely low latency and high throughput: first audio in ~3.86 ms from raw text on an RTX 4090, overall RTF ≈ 0.0023 (≈440× real time) and up to ≈1,316× with batching. Streams audio in 1 s chunks so playback can start immediately.
- Small, offline, and inexpensive: model artifacts are ≈34 MB and run offline on GPU or CPU; estimated cost on a rented RTX 4090 is about $0.0085 per million characters.
- Practical multi-caller scheduling: a built-in scheduler (Playhead) batches concurrent calls on the GPU so many simultaneous streams run efficiently without an external server queue.
Who It's For and Trade-offs
Great fit if you need private, low-latency Turkish TTS for apps like IVR, voice responses, on-device assistants, or batch voice rendering where minimizing model size, cost and latency matters. The model is especially useful when streaming start-time matters (live calls, robots, telephony) or when you prefer an offline Apache‑2.0 solution.
Look elsewhere if you require multiple voices, fine-grained emotional control, or highest-possible naturalness: EMA Lightning provides one neutral voice and UTMOS (~3.30) that is competitive but below some top-tier multi-voice systems. It is also optimized for Turkish; foreign words follow Turkish orthography.
Where It Fits
Compare to cloud voices and large TTS models: EMA Lightning trades model scale for operational advantages — tiny download, low GPU memory, offline privacy and dramatic inference speed — while matching or improving transcription-based accuracy on a Turkish benchmark. For multilingual or expressive multi-speaker needs, larger or specialized systems will be more suitable.
How it works (brief)
Text is normalized to spoken Turkish, encoded per letter, placed on a word–letter timeline by a duration predictor, aligned with a windowed Gaussian aligner, and turned into 64‑dim latents by a 4‑step distilled DiT generator. A HiFi‑GAN–style decoder upsamples latents to 48 kHz audio. The acoustic model and vocoder are shipped as two files (ema.pt 5.6M, decoder.pt 3.0M). Training used an internal ~1000‑hour single‑speaker Turkish corpus.
Practical notes
License: Apache‑2.0 (commercial use allowed). Responsible-use guidance recommends disclosing synthetic audio to listeners and vetting readouts for critical domains (banking, health, legal) because automatic normalization can misread unusual inputs.