Performs low-latency streaming speech-to-text, emitting one token per selectable 80/120/160 ms clock with configurable transcription delay and a 30s rolling KV cache for unlimited 24/7 transcription. Bilingual (zh/en) and includes semantic VAD.
Performs speaker diarization (who spoke when) for live and recorded audio using an open-weight, 100M-parameter streaming-capable model that supports up to eight anonymous speaker channels, overlapping speech, chunked processing, and configurable latency for ASR integration.