How to use
- Paste the Base64 audio, or a data URI such as data:audio/mpeg;base64,..., into the input box.
- Check the detected format and press play in the audio player.
- Download the file. It gets an extension that matches the detected format, such as .mp3 or .wav.
Decoding text-to-speech output
Text-to-speech services that answer with JSON cannot put binary audio in the response, so they encode it. Google Cloud Text-to-Speech, for example, returns the speech as a Base64 string in a field named audioContent, in the encoding you requested, such as MP3 or OGG Opus. Other voice and chatbot APIs use fields like audio, data or audio_base64.
Paste the value here to listen to the result while you compare voices, speaking rates or prompts. Copy only the string between the quotation marks; parsing the JSON first avoids picking up quotes or escaped characters such as \/.
Detected audio formats
The format is recognized from the first bytes of the decoded data, so a missing or wrong MIME type does not matter:
| Format | Signature | Base64 starts with |
|---|---|---|
| MP3 | ID3 tag or an MPEG frame header | SUQz or usually // |
| WAV | RIFF ... WAVE | UklGR |
| OGG (Vorbis, Opus) | OggS | T2dnUw |
| FLAC | fLaC | ZkxhQ |
| M4A (AAC) | ftyp box with brand M4A | varies (starts with a size field) |
Raw PCM: when the audio will not play
Some voice APIs, especially realtime and streaming ones, send raw PCM samples, often 16-bit mono at 16 or 24 kHz, with no file header. Telephony streams similarly carry headerless 8 kHz μ-law audio. Without a header, nothing in the bytes says what they are: the format cannot be detected, browsers cannot play the data, and this tool cannot guess the sample rate or channel count. You can still download the decoded bytes, but to hear them you need to add a WAV header using the parameters from the API documentation:
# FFmpeg: 16-bit little-endian, 24 kHz, mono
ffmpeg -f s16le -ar 24000 -ac 1 -i speech.pcm speech.wav
# Python
import base64, wave
with wave.open('speech.wav', 'wb') as w:
w.setnchannels(1)
w.setsampwidth(2) # 16-bit samples
w.setframerate(24000)
w.writeframes(base64.b64decode(pcm_base64))
If the result sounds too fast, too slow or like static, the sample rate, sample size or byte order does not match the data. Once it plays, Audio to Base64 can encode the WAV again if you need it as a string.
Frequently asked questions
Why can't the tool play my audio?
Either the data is raw PCM without a header (see above), or your browser does not support the codec. In the second case the downloaded file is still correct and opens in a media player such as VLC.
My API sends the audio in many Base64 chunks. How do I combine them?
Decode each chunk to bytes and append them in order. Joining the Base64 strings first only works if every chunk except the last encodes a multiple of 3 bytes, which you usually cannot rely on. This tool decodes one string at a time, so combine streamed chunks in your own code.
Is the decoded audio identical to the original?
Yes. Base64 is lossless, so the file contains exactly the bytes that were encoded, with the same quality and bitrate. The tool does not re-encode or convert the audio.
Can I paste a data URI?
data:audio/...;base64, string, or raw Base64 without a prefix. The format is detected from the decoded bytes either way.Is my audio sent to a server?
No. Decoding, format detection and playback all happen in your browser.