Turn speech into text—and text into natural-sounding audio—with Puget Voice
Turn speech into text—and text into natural-sounding audio—with Puget Voice. Record and transcribe live speech, or import audio and video files. Live Transcription supports two languages in the same session. Speaker ID distinguishes speakers, automatic language detection simplifies transcription, and custom profanity filters help keep transcripts clean. Supported ASR languages: English, Spanish, Chinese, German, French, Italian, Dutch, Portuguese, Polish, Arabic, Turkish, Russian, Japanese, Korean, Hindi, Persian, Swedish, Romanian, Danish, Czech, Finnish, Greek, Hungarian, Macedonian, Indonesian, Vietnamese, Cantonese, Thai, Malay, and Filipino. Supported audio formats: MP3, WAV, FLAC, AAC, M4A, OGG, OGA, OPUS, AIFF, AIF, CAF, AC3, EAC3, WMA, APE, and WV. Supported video formats: MP4, M4V, MOV, MKV, WebM, AVI, MPG, MPEG, TS, M2TS, MTS, VOB, FLV, F4V, OGV, WMV, 3GP, and 3G2. Puget Voice also provides multilingual Text-to-Speech with multiple voices and multi-speaker dialogue synthesis. Supported TTS languages: English, Chinese, Spanish, Italian, French, Japanese, Portuguese, and Hindi. Key features: - Live recording and transcription - Mixed-language transcription with up to two languages - Audio and video file transcription - Automatic language detection - Video subtitle generation - Speaker identification and separation - Custom profanity filtering - Multilingual Text-to-Speech - Multi-speaker dialogue synthesis