People convert HTML to Markdown for three reasons: to move a blog or help centre into a static-site generator, to turn a web page or a Google Doc into a README, and to feed a page to a language model without paying for thousands of tokens of markup. The conversion looks trivial - swap <h2> for ##, <b> for ** - and a naive converter really is twenty lines. The trouble is that Markdown gives meaning to ordinary characters depending on where they sit, so a converter that only swaps tags quietly rewrites your text.
What a converter has to decide
HTML can say more than Markdown can. Markdown has headings, paragraphs, emphasis, links, images, lists, quotes, code and - with GitHub's extensions - tables, strikethrough and task lists. It has no syntax for a superscript, underlining, a merged table cell or an embedded video. So every converter makes the same three decisions: how to write the things Markdown does have, how to protect text that merely looks like Markdown syntax, and what to do with everything else.
A worked example
Take this fragment of HTML:
- <h2>Setup</h2>
- <p>Run <code>npm install</code>, then copy <b>.env.example</b> to <b>.env</b>.</p>
- <ol>
- <li>Set <i>API_KEY</i></li>
- <li>2026. is the default year</li>
- </ol>
It converts to:
- ## Setup
- Run `npm install`, then copy **.env.example** to **.env**.
- 1. Set *API_KEY*
- 2. 2026\. is the default year
Most of that is the obvious swap. Two details are not. API_KEY keeps its underscore unescaped, because an underscore inside a word can never start emphasis. And the second list item gained a backslash - without it, a list item starting 2026. followed by a space is itself the start of a numbered list, so Markdown would render a list nested inside your list, numbered from 2026.
Why some characters get a backslash
In Markdown, a backslash before a punctuation mark means take this character literally. A converter has to add one wherever your text would otherwise be read as syntax. The characters that matter are the ones that start something: * and _ open emphasis, a backtick opens code, square brackets open a link, and at the start of a line #, >, - and a number followed by a full stop start a heading, a quote or a list.
Escaping all of them everywhere is safe and unreadable - file\_name\_here and 2 \* 3 in every sentence. Escaping none reformats your text. The HTML to Markdown converter escapes a character only where it could actually be misread, so snake_case, 2 * 3 and a year mid-sentence stay as written, while the same characters at the start of a line, or touching a word, are protected. Paste a page full of code identifiers and the difference is obvious.
When bold has to stay as HTML
Markdown only treats ** as bold when the characters around it pass a set of rules - the CommonMark flanking rules. Most of the time they pass. They fail when bold text ends in punctuation and runs straight into a word, which is common in marketing copy: <b>"Free"</b>forever. Written as **"Free"**forever, the closing asterisks are not allowed to close, and the page shows four literal asterisks instead of bold.
A converter that checks each delimiter against its real neighbours can catch this and write <strong>"Free"</strong>forever instead. Markdown allows inline HTML, every renderer shows it as bold, and the text is unchanged. The tool does this automatically and tells you how many spans needed it; everywhere else you get ordinary ** and *.
Tables
A simple HTML table becomes a GitHub Markdown table, with the columns padded so it reads as a grid in plain text, alignment carried over from align or text-align, and any pipe character inside a cell escaped so it does not split the cell in two. Markdown tables have two hard limits. They must have a header row - if yours has none, the first row is promoted and you are told - and every cell must be a single line with no merged cells.
A table with a colspan or rowspan, several header rows, or a list or code block inside a cell simply cannot be written as a Markdown table. The faithful option is to leave it as HTML, which Markdown permits and GitHub renders, and that is the default. If you need pure Markdown, the alternative is to flatten it: merged cells become empty cells and line breaks become spaces, which reads differently but at least is honest about it.
Tags Markdown has no syntax for
A superscript is the clearest case: drop the <sup> from x<sup>2</sup> and you have written x2, which means something else. So <sup>, <sub>, <u>, <mark>, <kbd>, embedded iframes and videos are kept as HTML by default. A <details> block is written the way GitHub expects, with the tags on their own lines and blank lines inside, so the Markdown between them still renders. Some things are dropped outright: scripts and style blocks, with their contents - code is not text - comments, hidden elements, and decorative icons marked aria-hidden.
Pasting from a web page or Google Docs
When you copy formatted text, your clipboard holds two versions: the plain text and the HTML behind it. An ordinary text box takes the plain text, which has already lost the headings, the links and the bold. The converter takes the HTML version, so you can select part of an article in your browser, or a whole Google Doc, and paste it straight in.
Google Docs needs one more trick. It does not use <b> and <i> at all - it marks bold and italic with inline styles on <span> tags, and wraps the whole paste in a <b> whose style says it is not bold. A converter that reads only tag names turns an entire Google Doc bold, or loses all of its emphasis. Reading the font-weight and font-style gets both right.
For a whole saved page, tick Main content only. If the page has a <main> element, or a single <article>, only that is converted, so the navigation, cookie banner and footer stay out of your Markdown.
Converting a page step by step
- Open the HTML to Markdown converter. It runs entirely in your browser, so nothing you paste is uploaded.
- Paste HTML source, or copy formatted text from a web page or Google Doc and paste that.
- Leave GitHub extras on for tables, strikethrough and task lists, or untick it for strict CommonMark.
- Decide whether tags Markdown cannot express should stay as HTML or become plain text.
- Read the notes under the result - promoted table headers, bold kept as HTML, scripts removed - then copy the Markdown or download it as a .md file.
To check a conversion, paste the Markdown into the Markdown to HTML converter and look at the preview. If you only want the words with no formatting at all, the HTML Tag Remover is the simpler tool.
Frequently asked questions
- How do I convert a Google Doc to Markdown?
- Select the text in the document, copy it, and paste it into an HTML to Markdown converter that reads the clipboard's HTML rather than its plain text. Google Docs puts both on the clipboard, and the plain text has already lost your headings, lists and links. The one thing to watch is emphasis: Google Docs marks bold and italic with inline font-weight and font-style styles on span tags, and wraps the whole paste in a b tag styled as normal weight, so a converter that only reads tag names either makes the whole document bold or drops every bit of emphasis. Headings, lists and links come across as real Markdown.
- Why does my converted Markdown have backslashes in it?
- They stop ordinary characters being read as formatting. In Markdown a line that starts with a number and a full stop is a list, a leading # is a heading, asterisks and underscores around a word are emphasis, and square brackets can start a link, so when those characters are part of your text they are escaped with a backslash. The backslash disappears when the Markdown is rendered. A careful converter escapes only where a character would genuinely be misread, so ordinary text like snake_case or 2 * 3 stays clean.
- Can every HTML page be converted to Markdown without losing anything?
- No, because HTML can express more than Markdown. Superscripts, underlining, merged table cells, colours, layout and embedded video have no Markdown syntax. Markdown does allow raw HTML, so the lossless option is to keep those parts as HTML inside the Markdown file, which GitHub and most static-site generators render correctly. If you need pure Markdown, they have to be reduced to plain text, and a good converter tells you how many elements it changed rather than dropping them silently.