A computer never stores the letter H. It stores a number that everyone has agreed stands for H, and it stores that number in binary - a row of 0s and 1s. Converting text to binary is therefore two steps, not one: first turn each character into its byte values using an agreed encoding, then write each byte as eight binary digits. Most confusion about text and binary, and most of the wrong answers online converters give, comes from skipping or fudging the first step.
Step one: characters to numbers
The agreement for the basic English characters is ASCII, which dates from the 1960s and assigns a number from 0 to 127 to every English letter, digit and common punctuation mark. Capital A is 65, B is 66 and so on to Z at 90; lowercase a is 97; the digit 0 is 48 and a space is 32. The gap of exactly 32 between a capital and its lowercase letter is deliberate - in binary it is a single bit, so changing case means flipping one digit.
Today the agreement is UTF-8, which is what nearly every web page, text file and API uses. For those first 128 characters UTF-8 is identical to ASCII, one byte each, which is why the two are so often confused. The difference only shows up once you leave plain English, and that is covered further down.
Step two: numbers to 8-bit binary
Each byte is a number from 0 to 255, and eight binary digits are exactly enough to hold one. The columns of a byte are worth, from left to right, 128, 64, 32, 16, 8, 4, 2 and 1. To write a number in binary, go through those columns from the left: if the number is at least the column's value, write 1 and subtract it; otherwise write 0. Keep the leading zeros so every byte is exactly eight digits - that fixed width is what lets a decoder split a long run of bits back into bytes.
Worked example: Hi!
- H is 72. 72 has no 128, one 64 (leaving 8), no 32, no 16, one 8 (leaving 0), then no 4, 2 or 1. That gives 01001000.
- i is 105. 105 = 64 + 32 + 8 + 1, so the 64, 32, 8 and 1 columns are set: 01101001.
- ! is 33. 33 = 32 + 1: 00100001.
So Hi! in binary is 01001000 01101001 00100001 - three characters, three bytes, 24 bits. Type the same text into the Text to Binary converter and it gives those three groups, with a character-by-character table underneath showing which byte belongs to which letter.
Converting binary back to text
Decoding is the same two steps in reverse. Split the bits into groups of eight, turn each group into a number by adding up the columns that hold a 1, and look each number up.
Worked example: 01000011 01100001 01110100
- 01000011: the 64, 2 and 1 columns are set, so 64 + 2 + 1 = 67, which is C.
- 01100001: 64 + 32 + 1 = 97, which is a.
- 01110100: 64 + 32 + 16 + 4 = 116, which is t.
The message is Cat. If the bits arrive with no spaces, count off eight at a time from the left; if the total is not a multiple of eight, a digit has been lost or added somewhere and no decoder can reliably recover which one.
One variation catches people out. Because every ASCII byte starts with a 0, some textbooks and older tools drop it and write 7-bit groups, so C appears as 1000011. Read in eights, that shifts every following byte and turns the message into nonsense. Zenoply's decoder spots a run that splits evenly into sevens but not eights and reads it as 7-bit ASCII, telling you it did so.
Why some characters take more than one byte
A single byte has only 256 possible values, and there are well over a hundred thousand characters in use. UTF-8 solves this by using one byte for ASCII and longer sequences for everything else: two bytes for accented Latin letters, Greek, Cyrillic, Hebrew and Arabic; three for most Indian scripts, Chinese, Japanese and Korean, and for symbols such as the rupee and euro signs; four for emoji.
The é in café, for example, is character number 233. In UTF-8 it becomes two bytes, 11000011 10101001 (C3 A9 in hex). The first byte's leading 110 says 'this character is two bytes long', and the second byte's leading 10 says 'I am a continuation'. The remaining eleven bits, strung together, spell out 233. The rupee sign ₹ is three bytes, 11100010 10000010 10111001, and the grinning face emoji is four, starting 11110000. That self-describing structure is why a decoder can tell where each character starts even in the middle of a stream.
The bug most converters ship
The quickest way to build a text-to-binary converter in JavaScript is to take each character's code number and print it in base 2. For English that gives the right answer, which is why the bug survives. For é it prints 11101001 - the number 233 in eight bits, which is the character's old Latin-1 value, not its UTF-8 bytes. Hand that single byte to anything that expects UTF-8 and it is an illegal sequence, so it shows up as a replacement mark rather than é. Emoji come out worse, as two 16-bit halves that are not bytes at all.
The converter here encodes through UTF-8 first, so a Hindi word, a price in rupees or an emoji all produce the bytes a file would actually contain, and decode back to exactly what went in. It also works the other way round: paste binary that some other tool produced and, if it is not valid UTF-8, it says so, gives the position of the first bad byte, and shows a second reading that treats each byte as one Latin-1 character - which is usually the text the other tool meant.
Doing it without the arithmetic
- Open the Text to Binary converter and type or paste your text on the Text to Binary tab.
- Choose what goes between bytes - a space, nothing, or a new line - and copy or download the result.
- Use the counts above the output to check the size: characters, bytes and bits are shown separately, because with UTF-8 they are not always in step.
- To decode, switch to Binary to Text. Anything already in the output is carried across, which is a quick way to confirm the round trip.
- The format menu - Write bytes as, or Input is when decoding - switches either direction to hexadecimal or decimal byte values, for escape sequences or byte arrays.
To convert a single number between binary, hex and decimal rather than a piece of text, the Base Converter is the better tool - it handles numbers of any length and fractions, where this one deals in bytes. Both run entirely in your browser, and nothing you type is uploaded.
Frequently asked questions
- What is Hello in binary?
- In ASCII and UTF-8, Hello is 01001000 01100101 01101100 01101100 01101111 - one 8-bit byte per letter, for H (72), e (101), l (108), l (108) and o (111). Both l's give the same byte because they are the same character; binary encoding is a lookup, not a cipher.
- How many bits is one character?
- Eight for any ASCII character - English letters, digits, spaces and common punctuation - because UTF-8 stores each of those in one byte. Other characters take more: 16 bits for an accented letter such as é, 24 for most Indian, Chinese and Japanese characters, and 32 for an emoji. So the number of bits in a message is eight times its byte count, not always eight times its length.
- Why does my binary decode to strange symbols?
- Usually one of three things. A digit is missing or extra, so every byte after it is shifted - check the total is a multiple of eight. The binary was written in 7-bit groups without the leading zero and is being read in eights. Or the bytes were produced as Latin-1 rather than UTF-8, so accented letters are single bytes that UTF-8 cannot read. A decoder that reports where the first bad byte is makes all three quick to find.