Remove HTML Tags from Text
Strip HTML tags and get clean plain text online free. Removes script and style blocks with their code, decodes entities and keeps paragraph breaks — in your browser.
Runs in your browserNothing uploadedFree · no signup
Line breaks follow the page, not the source file: a paragraph, heading or list gets a blank line, a <br> or a list item gets one line, table cells are separated by a tab, and text inside <pre> keeps its own spacing. Tags that only ever sat between words — <b>, <span>, <a> — leave nothing behind, so “<b>c</b>at” stays “cat”.
Script and style blocks are removed with their contents, which is what a find-and-replace on <...> cannot do. Need the markup tidied rather than removed? Use the HTML Beautifier, or HTML Entity Encoder to go the other way.
Frequently asked questions
- Why does my text still contain JavaScript after using another HTML stripper?
- Because the contents of a <script> or <style> element are character data, not markup. A tool built on a find-and-replace over <...> deletes the opening and closing tags and leaves everything between them — the whole minified bundle, the entire style sheet — sitting in what it calls your plain text. It is the most common failure of a small stripper and the reason the output of one is usually longer than the page it came from. This tool parses the document the way a browser tokenises it, so script, style, noscript and template blocks are removed together with what is inside them.
- Will removing the tags run the words together?
- No, and that is the other half of the job. Deleting a tag also deletes the boundary it stood for: <p>one</p><p>two</p> is two paragraphs on screen but becomes 'onetwo' under a plain regex, and a table turns into one long word. The opposite mistake is just as bad — putting a space at every tag breaks '<b>c</b>at' into 'c at'. Line breaks here follow CSS, not the tags: a paragraph, heading or list gets a blank line, a <br> or a list item gets one newline, table cells are separated by a tab so a row still pastes into a spreadsheet, and inline tags such as <b>, <span> and <a> leave nothing behind at all.
- Does it decode &amp; and &nbsp; as well?
- Yes, and in the right order. Tags are removed first and the character references are decoded afterwards, because decoding first would turn an escaped <script> — text a page displays literally, which any article about HTML is full of — into a real tag that the stripping step would then eat. Each piece of text is decoded on its own and in its own context, so a URL kept from an href reads ?a=1©=2 exactly as a browser does rather than turning the © into a copyright sign. No-break spaces are folded into ordinary spaces and counted, since a run of them is how a page fakes indentation.
- Is this safe to use for sanitising user input?
- No. This is a text extractor, not a sanitiser. It is written to answer 'what does this page say', and it runs in your browser on text you paste in. Stripping tags server-side as a defence against XSS is a known-bad pattern — attacks get through on malformed markup, on attributes rather than elements, and on content that is re-encoded later. If you need to accept HTML from users, keep the HTML and run it through a real sanitiser such as DOMPurify against an allowlist, on the server.
- Can it keep the links and the image alt text?
- Both are optional and both are off or on separately. 'Keep link URLs' writes each href in brackets after the link text, which is what you want when you are pulling the text out of a newsletter or an email and the links are the point. 'Keep image alt text' puts the alt attribute where the image was, which is what a screen reader announces and what a browser shows when an image fails to load. Left off, images and link targets disappear the way they do when you select a page and copy it.
- Is my HTML uploaded anywhere?
- No. The parsing and the extraction run entirely in your browser and nothing is sent to a server, so a saved page, an internal email or an export from a CMS never leaves your device.