How text becomes binary: bits, bytes, ASCII and UTF-8

Computers store everything, from photos to this sentence, as numbers, and they store numbers as patterns of two states: on and off, written as 1 and 0. To store text, a computer needs an agreed list that gives every character a number. This guide explains how that works, step by step, with examples you can check in the binary translator.

Bits and bytes

A single 0 or 1 is a bit. Eight bits make a byte, which can hold 28 = 256 different patterns, from 00000000 to 11111111. That is enough for every English letter, digit and punctuation mark, with room to spare, which is why one byte per character became the standard for English text.

Reading a byte by hand

Binary works like ordinary numbers, except each position is worth twice the one to its right instead of ten times. From left to right, the eight positions in a byte are worth 128, 64, 32, 16, 8, 4, 2 and 1. Add up the positions that hold a 1:

BinaryCalculationNumberCharacter
0100100064 + 872H
0110100164 + 32 + 8 + 1105i
0010000132 + 133!
001000003232(space)

So 01001000 01101001 00100001 spells “Hi!”.

ASCII: the original character table

The table that links numbers to characters for English is ASCII, the American Standard Code for Information Interchange, first published in 1963. It defines 128 characters: the numbers 0 to 31 are invisible control codes (such as line feed, number 10), 32 is the space, 48 to 57 are the digits 0 to 9, 65 to 90 are capital letters, and 97 to 122 are lowercase letters.

The layout was chosen carefully. Capital A is 65 (01000001) and lowercase a is 97 (01100001): they differ only in the bit worth 32. Every letter works the same way, so early computers could change case by flipping a single bit. The digits were placed so that the last four bits of “0” to “9” are the numbers 0 to 9 themselves.

Beyond English: UTF-8

128 characters leave no room for é, ñ, Greek, Chinese or emoji. For decades different countries used different, incompatible tables for the upper 128 values of a byte, which is why text sometimes turned into gibberish when it moved between computers. The fix was Unicode, a single list that now gives a number to over 150,000 characters from almost every writing system, plus symbols and emoji.

Unicode numbers can be much larger than 255, so they need a way to be stored as bytes. The most common is UTF-8, designed in 1992. Its key trick is that the first 128 characters are stored exactly as in ASCII, in one byte each. Everything else uses two, three or four bytes, and the first bits of each byte say how many bytes belong together:

CharacterBytes in UTF-8Binary
A101000001
é211000011 10101001
😀411110000 10011111 10011000 10000000

A first byte starting with 0 is a complete character. One starting with 110 begins a two-byte character, 1110 a three-byte one and 11110 a four-byte one, and every following byte starts with 10. Because of this, a program can always find where a character starts, even in the middle of a file. Backward compatibility with ASCII made UTF-8 easy to adopt, and it is now used by the vast majority of websites.

Common questions

Why do binary translators sometimes give different results? Most use UTF-8, as Wordformat does, so plain English gives the same result everywhere. Translators that use a different encoding, such as UTF-16, give different results for accented letters and emoji.

Is binary code a secret code? No. Anyone with an ASCII or UTF-8 table can read it, so it’s fine for puzzles and fun but offers no privacy.

What about numbers in text? The character “7” is stored as the number 55 (00110111), not as the value 7. Programs convert between the two when you type a number into a calculator or spreadsheet.