Developer & encoding

How to Encode and Decode HTML Entities (and Why < Keeps Showing Up)

How HTML entities work: which characters must be encoded, when to use named, decimal or hex references, why &copy without a semicolon still renders, and how to fix double-encoded text.

6 min readUpdated Sep 9, 2026

An HTML entity - properly, a character reference - writes a character using nothing but plain ASCII. © is a copyright sign, — an em dash, < a less-than sign that will not be mistaken for the start of a tag. The whole system hangs off one character, the ampersand, and almost every entity problem is really a problem with that ampersand: one that should have been encoded and was not, or one that was encoded twice.

The five characters that actually need encoding

Modern advice is much shorter than the tables on older pages. If your file is served as UTF-8 - and it should be - accents, curly quotes, em dashes and emoji can all sit in the source as themselves. Only five characters have to be encoded, and only where a parser would read them as markup:

  • & becomes & - always, everywhere. An ampersand is the start of a character reference, so a bare one in front of a word can be swallowed.
  • < becomes &lt; - otherwise the parser starts looking for a tag name.
  • > becomes &gt; - not strictly required in text, but encoding it costs nothing and avoids the ]]> case in XML.
  • " becomes &quot; - only matters inside an attribute value delimited by double quotes.
  • ' becomes &#39; - only matters inside an attribute value delimited by single quotes.

The apostrophe is worth a note. &apos; is an XML entity that HTML did not define until HTML5, so it is undefined in HTML 4 and in XHTML served as text/html. The numeric &#39; is the one spelling every parser has always understood, which is why the HTML Entity Encoder writes that.

Named, decimal or hex

The same character has three spellings. A copyright sign can be written &copy;, &#169; or &#xA9;, and a browser treats all three identically. Which to choose depends on what will read the file, not on taste.

A named reference is readable, and the 253 names from HTML 4.01 and XHTML have been understood by every browser since 1999. HTML5 adds around 2,000 more, mostly MathML symbols and alternate spellings such as &excl; and &NewLine;, which are fine in a browser and nowhere else.

A numeric reference needs no table at all, which makes it the safe default outside HTML. XML predefines exactly five names - &amp; &lt; &gt; &quot; and &apos; - and treats any other as a parse error unless the document declares it in a DTD. So a feed containing &copy; or &mdash; fails to load in a strict reader, while &#169; and &#8212; are always fine. For a feed, a SOAP body or an SVG file, use numeric.

This is the failure that sends most people looking for an entity tool in the first place. You write a link with a query string:

<a href="/search?q=tea&copy=full">Full copy</a>

Inside the quoted attribute this survives: the href really is /search?q=tea&copy=full. Put the same characters in the page as text - a code sample, a support article, an error message - and the browser renders ?q=tea(c)=full, with a copyright sign where &copy was. No validator will complain; the text is simply wrong on screen.

The fix is one character: write &amp;copy=full and it renders as &copy=full in both places. Paste either version into the HTML Entity Encoder and switch between element text and attribute value to see the two readings side by side.

Why &copy without a semicolon still renders

Because 106 names are allowed to drop it: exactly the 96 Latin-1 names covering U+00A0 to U+00FF - &nbsp, &copy, &reg, &eacute, &pound - plus &amp, &lt, &gt, &quot and six all-caps aliases. Everything else needs its semicolon, so &hellip is literal text while &hellip; is an ellipsis.

There is a second half to the rule, and it is the reason query strings usually survive. Inside an attribute value, a semicolon-less name is not consumed at all if the next character is = or a letter or digit. That single exception is what keeps href="?a=1&copy=2" intact while the same characters as page text become ?a=1(c)=2. It is why a good decoder asks where your text came from before it answers.

Double encoding, and how to spot it

If your page is showing &lt;b&gt; on screen instead of bold text, or your database is full of &amp;lt;, something encoded the same string twice. The first pass turned < into &lt;. The second pass saw the ampersand in &lt; and encoded that too, giving &amp;lt;.

What makes this confusing is that one decode does not fix it: decoding &amp;lt; gives &lt;, which still looks encoded. The HTML Entity Encoder detects this, offers to keep decoding until nothing changes, and reports how many passes it took - a useful number, because three passes means the bug in your pipeline is running three times, not once.

The characters you cannot see

Non-breaking spaces, soft hyphens, zero-width spaces and joiners, directional marks and byte order marks all look like nothing, or like an ordinary space, in every editor there is. They arrive by way of a paste from Word, a CMS field or a PDF, and they decide where a line wraps and whether a find-and-replace matches.

Encoding them is the cheapest way to make them visible: a &#xA0; in the source is obvious in a way an invisible byte never is. It is the one case where encoding above plain ASCII earns its keep on a UTF-8 page.

Why &#151; decodes to an em dash

Code points 128 to 159 are C1 control characters in Unicode, so &#151; should strictly be an invisible control code. Every browser decodes it as an em dash instead, and HTML5 now requires that: a huge number of pages were authored in Windows-1252 and served as something else, so the specification defines a replacement table mapping that range to the Windows-1252 characters. &#146; is a right single quote, &#128; a euro sign, &#151; an em dash. Seeing these in your data usually means a Windows-1252 file was read as UTF-8 upstream, and that same mismatch will be corrupting characters with no entity form at all.

Encoding is not the same as sanitising

Encoding the five characters makes text safe in exactly two places: as element content, and inside a quoted attribute value. It does nothing for text inside a <script> or <style> block, an unquoted attribute, or an attribute that takes a URL - an href of javascript:alert(1) contains none of the five characters and is still an attack. Encode as the last step before output, and use a real sanitiser for markup you intend to keep.

For the neighbouring jobs there are separate tools: the String Escaper covers JSON, JavaScript, XML, CSV, SQL and shell quoting, and the URL Encoder handles percent-encoding, which is a different scheme entirely and the right one for a query string value.

It runs entirely in your browser

Page fragments carry customer names, internal hostnames and order numbers, so where a conversion happens is not a detail. Encoding and decoding both run in JavaScript inside your browser tab and nothing is uploaded, so you can paste a live template or a production error message without it leaving your device.

Frequently asked questions

Do I still need to encode accents and symbols on a UTF-8 page?
No. If the document declares UTF-8 and is served as UTF-8, then e-acute, an em dash, a Greek letter and an emoji can all appear as themselves, and the file will be smaller and much easier to read for it. Encoding above ASCII is worth doing in three situations: the output has to pass through a pipeline you do not trust with UTF-8, such as an old CMS field or a database column with the wrong collation; you are producing something that will be read by a strict XML parser; or you want invisible characters like a no-break space to become visible in the source. Otherwise the five markup characters are all you need.
Why does my text show &amp;lt; instead of a tag?
It was encoded twice. Something turned < into &lt;, and then a second pass encoded that ampersand as well, leaving &amp;lt;. Decoding once gives back &lt;, which still looks encoded, so it is easy to think the decoder failed - it just needs another pass. The usual cause is a value escaped on the way into storage and escaped again by a template engine on the way out, since Jinja, Twig, ERB and JSX all escape by default. Fix the pipeline so escaping happens exactly once, at output; to clean up the data you already have, decode repeatedly until the text stops changing.
Can I use HTML entity names like &copy; in an RSS feed or an SVG?
Not safely. Those are XML, and XML predefines only five entity names - &amp; &lt; &gt; &quot; and &apos; - so any other name is a parse error unless the document declares it in a DTD. A feed containing &copy; or &nbsp; will be rejected by a strict reader, and the failure often looks like an empty feed rather than an error message. Use numeric references there instead: &#169; and &#160; need no table and are valid in every XML document, as is the character itself if the file is UTF-8.