How to use
- Choose Encode to turn text or bytes into Base85, or Decode to turn Base85 back into data.
- Type or paste your input, or open a file. Switch the input to Hex if you are working with raw bytes.
- The result updates as you type. When decoding, choose Text or Hex output depending on what the data contains.
- Copy the result or download it as a file.
What is Base85?
Base85 is a family of encodings that turn every 4 bytes into 5 characters. Five base-85 digits can hold 85 × 85 × 85 × 85 × 85 = 4,437,053,125 values, just enough for any 32-bit number, so the output is only 25% larger than the input, compared with 33% for Base64.
Several Base85 alphabets exist and they are not interchangeable: Ascii85 (Adobe, PDF), Z85 (ZeroMQ) and the one on this page. It uses 0-9, A-Z, a-z and 23 symbols: !#$%&()*+-;<=>?@^_`{|}~. The quote characters, comma, period, slash, colon, backslash and square brackets are left out.
RFC 1924, Git and Python
The character set comes from RFC 1924, "A Compact Representation of IPv6 Addresses", published on 1 April 1996 as an April Fools' joke. It proposed writing a 128-bit IPv6 address as one base-85 number of exactly 20 characters. Nobody writes IPv6 addresses that way, but the alphabet was later reused for general binary data:
- Git stores binary files in patches (
git diff --binary,git format-patch) as zlib-compressed data encoded with this alphabet, with a length character at the start of each line. - Python 3.4 and later provides
base64.b85encode()andb85decode(), whose output matches this tool. - Mercurial uses the same encoding for binary files in Git-style diffs.
Note that RFC 1924 converts the whole 128-bit address at once, while Git, Python and this tool work in 4-byte groups, so a 16-byte value gives a different result from the RFC method.
Short final groups
When the input length is not a multiple of 4, the last group is filled with zero bytes, encoded, and only the first n + 1 characters are kept: 1 byte becomes 2 characters, 2 bytes become 3 and 3 bytes become 4. The decoder reverses this, so no padding characters are needed. This variant has no shortcut for zero bytes, unlike the z of Ascii85.
# Python 3.4+
import base64
base64.b85encode(b'hello') # b'Xk~0{Zv'
base64.b85decode(b'Xk~0{Zv') # b'hello'
Specification
| Alphabet | 0-9 A-Z a-z !#$%&()*+-;<=>?@^_`{|}~ |
|---|---|
| Output size | 5 characters per 4 bytes (125%) |
| Padding | None |
| Standard | RFC 1924 character set (Git, Python b85) |
| Case sensitive | Yes |
Examples
| Input (UTF-8) | Output |
|---|---|
Hello, World! | NM&qnZ!92JZ*pv8Ap |
Base64.is | LSb`dHZ(42a{ |
你好 | <h`KfrM& |
Frequently asked questions
Is Base85 the same as Ascii85?
z and may be wrapped in Adobe delimiters. Use the Ascii85 page for PDF and PostScript data.Was RFC 1924 a real standard?
No. It is an informational RFC published as an April Fools' joke and was never used for IPv6 addresses. Its character set became popular later through Git and Python.
How much larger is Base85 output?
Can I put Base85 in JSON or XML?
<, > and & must be escaped. Base85 is not URL-safe.