Base85 Encoder and Decoder

Convert text, hex bytes or files to Base85 and back. This page uses the RFC 1924 character set, the variant produced by Python b85encode and used in Git binary patches.

Runs entirely in your browser. Nothing is uploaded.

How to use

  1. Choose Encode to turn text or bytes into Base85, or Decode to turn Base85 back into data.
  2. Type or paste your input, or open a file. Switch the input to Hex if you are working with raw bytes.
  3. The result updates as you type. When decoding, choose Text or Hex output depending on what the data contains.
  4. Copy the result or download it as a file.

What is Base85?

Base85 is a family of encodings that turn every 4 bytes into 5 characters. Five base-85 digits can hold 85 × 85 × 85 × 85 × 85 = 4,437,053,125 values, just enough for any 32-bit number, so the output is only 25% larger than the input, compared with 33% for Base64.

Several Base85 alphabets exist and they are not interchangeable: Ascii85 (Adobe, PDF), Z85 (ZeroMQ) and the one on this page. It uses 0-9, A-Z, a-z and 23 symbols: !#$%&()*+-;<=>?@^_`{|}~. The quote characters, comma, period, slash, colon, backslash and square brackets are left out.

RFC 1924, Git and Python

The character set comes from RFC 1924, "A Compact Representation of IPv6 Addresses", published on 1 April 1996 as an April Fools' joke. It proposed writing a 128-bit IPv6 address as one base-85 number of exactly 20 characters. Nobody writes IPv6 addresses that way, but the alphabet was later reused for general binary data:

  • Git stores binary files in patches (git diff --binary, git format-patch) as zlib-compressed data encoded with this alphabet, with a length character at the start of each line.
  • Python 3.4 and later provides base64.b85encode() and b85decode(), whose output matches this tool.
  • Mercurial uses the same encoding for binary files in Git-style diffs.

Note that RFC 1924 converts the whole 128-bit address at once, while Git, Python and this tool work in 4-byte groups, so a 16-byte value gives a different result from the RFC method.

Short final groups

When the input length is not a multiple of 4, the last group is filled with zero bytes, encoded, and only the first n + 1 characters are kept: 1 byte becomes 2 characters, 2 bytes become 3 and 3 bytes become 4. The decoder reverses this, so no padding characters are needed. This variant has no shortcut for zero bytes, unlike the z of Ascii85.

# Python 3.4+
import base64
base64.b85encode(b'hello')      # b'Xk~0{Zv'
base64.b85decode(b'Xk~0{Zv')    # b'hello'

Specification

Alphabet0-9 A-Z a-z !#$%&()*+-;<=>?@^_`{|}~
Output size5 characters per 4 bytes (125%)
PaddingNone
StandardRFC 1924 character set (Git, Python b85)
Case sensitiveYes

Examples

Input (UTF-8)Output
Hello, World!NM&qnZ!92JZ*pv8Ap
Base64.isLSb`dHZ(42a{
你好<h`KfrM&

Frequently asked questions

Is Base85 the same as Ascii85?
They use the same arithmetic but different characters, so their output is different. Ascii85 also abbreviates four zero bytes as z and may be wrapped in Adobe delimiters. Use the Ascii85 page for PDF and PostScript data.
Was RFC 1924 a real standard?

No. It is an informational RFC published as an April Fools' joke and was never used for IPv6 addresses. Its character set became popular later through Git and Python.

How much larger is Base85 output?
Every 4 bytes become 5 characters, so the output is 1.25 times the input. Base64 is about 1.33 times. basE91 is slightly more compact still.
Can I put Base85 in JSON or XML?
In JSON, yes: the alphabet has no double quote or backslash, so nothing needs escaping. In XML and HTML the characters <, > and & must be escaped. Base85 is not URL-safe.