HTML Entity Encoder & Decoder
Free HTML entity encoder and decoder. Convert & < > and any character to named, decimal or hex references, decode them back, and search the full entity table — in your browser.
Runs in your browserNothing uploadedFree · no signup
Encodes only & < > " and ', which is everything a browser needs to read your text as text rather than as markup. Save the file as UTF-8 and an accent or an emoji can stay exactly as it is.
Entity reference
All 253 names HTML 4.01 and XHTML define, which is the set every browser and every XML parser has understood since 1999. Search by name, by character, or by number.
| Char | Name | Decimal | Hex | Description |
| & | & | & | & | ampersand |
| < | < | < | < | less-than |
| > | > | > | > | greater-than |
| " | " | " | " | double quote |
| ' | ' | ' | ' | apostrophe |
| · | |   |   | no-break space |
| ¡ | ¡ | ¡ | ¡ | inverted exclamation |
| ¢ | ¢ | ¢ | ¢ | cent |
| £ | £ | £ | £ | pound |
| ¤ | ¤ | ¤ | ¤ | currency |
| ¥ | ¥ | ¥ | ¥ | yen |
| ¦ | ¦ | ¦ | ¦ | broken bar |
| § | § | § | § | section |
| ¨ | ¨ | ¨ | ¨ | diaeresis |
| © | © | © | © | copyright |
| ª | ª | ª | ª | feminine ordinal |
| « | « | « | « | left guillemet |
| ¬ | ¬ | ¬ | ¬ | not sign |
| · | ­ | ­ | ­ | soft hyphen |
| ® | ® | ® | ® | registered |
| ¯ | ¯ | ¯ | ¯ | macron |
| ° | ° | ° | ° | degree |
| ± | ± | ± | ± | plus-minus |
| ² | ² | ² | ² | superscript two |
Frequently asked questions
- Which characters actually have to be encoded?
- Only five, and only in the places they would be read as markup: & < > and, inside a quoted attribute value, " and '. Everything else is optional. If your page is served as UTF-8 — which it should be — then an accent, a curly quote, an em dash and an emoji can all sit in the source as themselves, and the file will be smaller and easier to read for it. The ampersand is the one people forget, and it is the one that matters most: an unencoded & in front of a word can be swallowed as a character reference, which is exactly how ?a=1©=2 turns into ?a=1©=2.
- Should I use a named entity or a numeric one?
- A named entity is easier to read, and every parser has understood the 253 names in HTML 4.01 since 1999, so — is a safe and clear choice in a web page. A numeric reference is safer everywhere else: it needs no table at all, so it works in XML and in an RSS or Atom feed, where only & < > " and ' are defined and anything else is a parse error unless the document declares it. HTML5 adds around 2,000 more names, mostly MathML symbols and alternate spellings — those are fine in a browser but will break an XML parser, so numeric is the right default for a feed.
- Why does the decoder ask whether the text is element text or an attribute value?
- Because the answer genuinely differs, and this is the single most common way a URL mangles itself. A browser accepts 106 legacy names without their closing semicolon, so © is read as a copyright sign — but only in element text. Inside an attribute value the tokenizer refuses to consume a semicolon-less name when the next character is = or a letter, precisely so that a query string survives. So <a href="?a=1©=2"> keeps its literal ©=2, while the same characters written as page text render as ©=2. Pick the context your text actually came from and the tool follows the same rule the browser does.
- My text is full of &lt; — what happened?
- It was encoded twice. Something encoded < to <, then a second pass encoded that ampersand to &, leaving &lt;. One decode gives you < back, which still looks encoded, which is the confusing part. Tick "Keep decoding until nothing changes" and the tool unwinds it fully and tells you how many passes it took. The tool also flags the situation automatically, so you know it is double encoding rather than something stranger.
- What are the invisible characters it warns about?
- Non-breaking spaces, soft hyphens, zero-width spaces and joiners, directional marks and byte order marks. They look like nothing, or like an ordinary space, in every editor — but they decide where a line wraps, whether a find-and-replace matches, and whether two strings compare equal. They are usually a souvenir of a paste from Word, a CMS rich-text field or a PDF. The tool names each one it finds with its code point, and encoding above plain ASCII writes them as references so they become visible in the source.
- Why does — come out as an em dash and not a control character?
- Because that is what browsers do, and HTML5 now requires it. Code points 128 to 159 are C1 control characters in Unicode, but an enormous number of pages were authored in Windows-1252 and served as something else, so a numeric reference in that range is remapped to the Windows-1252 character instead: — is an em dash, ’ a right single quote, € a euro sign. The tool applies the same table and tells you when it has, because seeing that in your data usually means an encoding is wrong further upstream.
- Is my text uploaded anywhere?
- No. Every conversion runs in JavaScript inside your browser tab, so a page fragment, a customer name or an internal URL never leaves your device. The page keeps working with the network switched off.