Streaming Speech-to-Speech & Voice Clone Simulator Zero-Shot d-Vector

Kotoba Tech simultaneous acoustic transfer & latency budget workbench

Presets:
Streaming Ingest & Timbre Parameters
350 ms
1.2 s
145 Hz
1.04x
Source: "We have been building speech to speech simultaneous translation for a while."
Target: "私たちはしばらくの間、音声から音声への同時通訳を構築してきました。"
Pipeline Latency Breakdown ● Simultaneous Real-Time
Chunk Buffer
350ms
Ingestion
Stream ASR
45ms
Conformer
Neural MT
20ms
Wait-k=2
d-Vector Enc
10ms
Zero-Shot
Vocoder
60ms
HiFi-GAN
Total Latency
485 ms
Timbre Cosine Sim
0.89
Embedding Stability
0.92
Est. BLEU Quality
38.4
Enjoy this tool? Build your own with Super