Text & writing

How to Remove HTML Tags and Get Clean Plain Text

How to strip HTML tags from text properly: why a regex leaves JavaScript and CSS behind, where line breaks belong, and the order entity decoding has to run in.

6 min readUpdated Sep 20, 2026

You have a chunk of HTML - a saved page, a CMS export, an email body - and you want the words without the markup. There is a one-line answer everybody reaches for, and it is wrong in three separate ways, each of which shows up on the first real page you feed it.

The one-liner, and what it leaves behind

Search for how to strip HTML and the answer is always some version of replace(/<[^>]*>/g, ""). Delete anything that looks like a tag. Take this fragment of a release-notes page:

<div class="post"> <h2>Release notes</h2> <style>.post{margin:0}</style> <p>Version&nbsp;2.1 ships <b>fast</b>er builds &amp; a new <a href="/docs">docs site</a>.</p> <ul><li>Cache warm-up</li><li>Smaller bundles</li></ul> <script>track("view", {a: 1 > 0});</script> </div>

Run the regex over it and you get this:

Release notes .post{margin:0} Version&nbsp;2.1 ships faster builds &amp; a new docs site. Cache warm-upSmaller bundles track("view", {a: 1 > 0});

Four things went wrong at once. The style sheet is in your text. The tracking call is in your text. The two list items have fused into Cache warm-upSmaller bundles. And the entities never decoded. Each is a category of mistake rather than a detail.

Script and style hold code, not markup

HTML has a handful of elements whose contents the parser does not treat as markup at all. Inside script, style, noscript and template, everything up to the matching closing tag is raw character data - so a regex that deletes things shaped like tags deletes a script's opening and closing tags and leaves the entire program between them.

On a hand-written fragment that is an eyesore. On a real page it is the whole output: a modern page carries far more bytes of minified JavaScript and CSS than of prose, so the stripped result comes out longer than the page did. That wall of gibberish an online stripper hands back is your own page's code.

The fix is a tokenizer, not a better regex: something that knows a script block ends at the first </script followed by whitespace, a slash or a greater-than sign, and discards what is inside. It also handles the other place the regex breaks down - a greater-than sign is legal inside an attribute value, so <a title="a > b">x</a> stops [^>]* early and leaks b">x into the output.

Deleting a tag deletes the boundary it stood for

This is the failure people notice last and care about most. Cache warm-upSmaller bundles is not a formatting problem; it is two facts that have merged into one. The li tags were the only thing telling you where the first item ended.

So put a space at every tag? That breaks it the other way. <b>c</b>at renders as the single word cat, and a stripper that inserts whitespace at every tag gives you c at. Bold, italics, links and spans sit inside a line of text, and whatever you put where they were is a word break that was never on the page.

The line is drawn by CSS, not by the tags. Whitespace between two inline boxes is rendered; whitespace at the edge of a block box is not. So a paragraph, heading, list item or table row is a real boundary and earns a line break, while b, i, span, a, em and strong earn nothing. An extractor carrying the browser's default display table gets that right for free, and can then put a blank line between paragraphs and a tab between table cells, so a row still pastes into a spreadsheet as a row.

The HTML Tag Remover follows exactly that model. On the fragment above it produces:

Release notes / (blank line) / Version 2.1 ships faster builds & a new docs site. / (blank line) / Cache warm-up / Smaller bundles

The heading is its own paragraph, faster is still spelled correctly, the two list items are two lines, and the CSS and the tracking call are gone with their tags.

Decode the entities - but only after the tags are gone

Character references have to be decoded or your text is full of &amp;, &nbsp; and &#8217;. That part is obvious. The order is not, and getting it backwards is a quiet, nasty bug.

If you decode first, an escaped &lt;script&gt; - text a page displays literally, and which any article or changelog about HTML is full of - becomes a real script tag. The stripping step then eats it, and on a page discussing HTML it can eat everything between two of them. Tags first, entities second.

Decoding also depends on where the text came from. HTML lets 106 legacy names drop their closing semicolon, which is why ?a=1&copy=2 reads as page text with a copyright sign in the middle - but inside an attribute value that is not a reference at all, so the same string survives intact in an href. Keep link URLs and each piece has to be decoded under its own rule; one pass over the finished document gets one of the two wrong.

And a no-break space decodes to U+00A0, not to a space. It looks like one and is not, and pages use runs of them to fake indentation, so folding them into ordinary spaces saves a confusing afternoon when a search over the pasted text refuses to match.

The honest default is what you would get by selecting the rendered page and pressing copy: no image placeholders, no URLs, just the words. Two exceptions are worth a checkbox each. Link URLs, because in a newsletter the links are half the content - in brackets after the link text they survive without the output becoming markdown. And image alt text, which is what a screen reader announces and what a browser shows when an image fails to load.

Stripping tags is not sanitising

This is a common reason people go looking for a tag stripper, and the one case where the answer is do not. Removing tags from user input as an XSS defence is a known-bad pattern: attacks get through on malformed markup your stripper and the browser disagree about, on attributes rather than elements, and on content re-encoded downstream. If you accept HTML from users, keep it as HTML and run it through a real sanitiser - DOMPurify is the standard answer - against an allowlist, on the server.

Doing it without writing any code

  1. Open the HTML Tag Remover and paste your HTML in. Everything runs in your browser and nothing is uploaded.
  2. Leave the layout on Readable. Switch to One single line for a word count, or to the third option to leave every byte outside a tag where it is.
  3. Tick Keep link URLs or Keep image alt text if you want them; off gives a clean copy-paste equivalent.
  4. Read the notes under the result: how many script and style blocks were dropped, how many no-break spaces were folded, how many images went.

If you want the markup tidied rather than removed, the HTML Beautifier reindents a document without changing a byte a browser renders. If the text you get back still has ragged spacing, the Whitespace Remover finishes the job.

Frequently asked questions

Why is the output of an online HTML stripper sometimes longer than the page?
Because it kept the code. The contents of a script or style element are character data rather than markup, so a tool built on a find-and-replace over anything between angle brackets deletes the opening and closing tags and leaves the entire JavaScript bundle and style sheet behind as what it calls plain text. A modern page carries far more bytes of code than of prose, so the result grows. A stripper that tokenizes the document the way a browser does removes those blocks together with their contents.
How do I keep paragraph breaks when I remove HTML tags?
By letting the element decide. Block-level boxes - paragraphs, headings, list items, table rows, divs - are real boundaries on screen, so they earn a line break; inline ones such as b, span and a are not, so they should leave nothing behind. That distinction is the whole job: insert breaks everywhere and a bold letter followed by text turns cat into c at, insert them nowhere and two list items fuse into one word. A blank line between paragraphs and a tab between table cells matches what you would get by selecting the rendered page and copying it.
Can I use a tag stripper to make user input safe?
No. Stripping tags is a text-extraction tool, not a security control, and removing tags as an XSS defence is a well-known anti-pattern - attacks get through on malformed markup, on attributes rather than elements, and on text that is re-encoded later in the pipeline. If you accept HTML from users, keep it as HTML and run it through a dedicated sanitiser such as DOMPurify against an allowlist, server-side. Use a stripper for reading a saved page, an email body or a CMS export, which is what it is built for.