Google's Miipher-2 demo clips
English, noisy read speech
Quality predictor 2.83 to 4.58 (Miipher-2 4.52). Bandwidth 12 to 21 kHz (Miipher-2 renders are 24 kHz audio). Speaker similarity 0.982.
Swahili, FLEURS recording
Quality predictor 3.16 to 4.48 (Miipher-2 4.29). Bandwidth 8 to 23 kHz (Miipher-2 renders are 24 kHz audio). Speaker similarity 0.984.
Urdu, FLEURS recording
Quality predictor 2.69 to 4.53 (Miipher-2 4.42). Bandwidth 8 to 21 kHz (Miipher-2 renders are 24 kHz audio). Speaker similarity 0.981.
Catalan, FLEURS recording
Quality predictor 3.12 to 4.65 (Miipher-2 4.70). Bandwidth 8 to 23 kHz (Miipher-2 renders are 24 kHz audio). Speaker similarity 0.958.
English, the clip SAWT V4 lost
Quality predictor 3.57 to 4.44 (Miipher-2 4.60). Bandwidth 12 to 21 kHz (Miipher-2 renders are 24 kHz audio). Speaker similarity 0.988.
LibriTTS-R, against Google's Miipher renders
English audiobook (LibriTTS)
Quality predictor 3.75 to 4.65 (Miipher 4.49). Bandwidth 12 to 20 kHz (Miipher renders are 24 kHz audio). Speaker similarity 0.979.
English audiobook (LibriTTS)
Quality predictor 4.01 to 4.48 (Miipher 4.20). Bandwidth 12 to 22 kHz (Miipher renders are 24 kHz audio). Speaker similarity 0.995.
Quran recitation from the archive
Quran recitation, archive upload (surah 85)
Quality predictor 2.53 to 3.80 (SAWT V4 3.91). Bandwidth 8 to 16 kHz. Speaker similarity 0.936.
Quran recitation, archive upload (surah 38)
Quality predictor 2.88 to 4.05 (SAWT V4 3.79). Bandwidth 11 to 16 kHz. Speaker similarity 0.978.
The telephone clip from the SAWT V4 card
Danish over a telephone line
Quality predictor 3.06 to 4.68 (SAWT V4 4.60). Bandwidth 3.7 to 22 kHz (input: 300 to 3700 Hz telephone line). Speaker similarity 0.982.
Spectrograms
Inference pipeline