Markdown is what you write; HTML is what a browser reads. Almost every complaint about the converter in between comes down to three surprises: the line breaks you typed are gone, a table came out as a row of pipe characters, and the HTML you pasted into the middle did something unexpected. None of those are bugs. They are the rules, and once you can see them the conversion stops being mysterious.
What the converter is actually doing
A Markdown parser works in two passes. It first splits the document into blocks by looking at the start of each line - a # makes a heading, a - or a digit a list, a fence a code block, anything left over a paragraph - then walks the text inside each block for inline markers. That shape explains most of Markdown's behaviour: a blank line ends a block, a code fence suspends the inline pass so asterisks inside it stay asterisks, and a paragraph reaches the inline pass as one run of text, which is where its newlines stop being structure.
The line break that disappears
This is the most common question, and the answer is that it is deliberate. Inside a paragraph, a newline is an ordinary space. Write this:
- Install the CLI, then run it.
- It writes a config file.
and you get one paragraph. The newline survives into the HTML, but a browser renders it as a space, so the two lines flow together on the page. The rule exists so you can hard-wrap a paragraph at 80 columns in your editor without the wrapping leaking into the output.
Markdown gives you two ways to force a real break: end the line with two spaces, or with a backslash. Both produce <br>. The two-space version is invisible in an editor and gets stripped by plenty of tools on the way, which is why people conclude Markdown has eaten their formatting.
If the text came from somewhere else - a changelog, a chat export, an address block - editing every line is not realistic. Tick Keep single line breaks in the Markdown to HTML converter and every newline inside a paragraph becomes a <br>. The tool also counts the breaks the default rule is about to swallow, so you hear about them before you spot them in the output.
A worked example
Take a short piece of documentation with a heading, a wrapped paragraph, a numbered list and a quote:
- ## Setup steps
- Install the CLI, then run it.
- It writes a config file.
- 1. Run `npm install`
- 2. Copy **.env.example**
- > Needs Node 20 or newer.
That converts to:
- <h2 id="setup-steps">Setup steps</h2>
- <p>Install the CLI, then run it.
- It writes a config file.</p>
- <ol>
- <li>Run <code>npm install</code></li>
- <li>Copy <strong>.env.example</strong></li>
- </ol>
- <blockquote>
- <p>Needs Node 20 or newer.</p>
- </blockquote>
Three things to notice. The heading picked up an id you can link to. The list items hold no <p>, because the list is tight - add a blank line between them and each gains a paragraph wrapper, which changes the spacing. And the paragraph kept its newline, which a browser shows as a space.
CommonMark and the GitHub extras
Markdown has a specification, CommonMark, which nails down the awkward cases the original 2004 description left open. It covers headings, emphasis, lists, links, images, blockquotes, code blocks and raw HTML. It does not cover four things people assume are Markdown:
- Pipe tables - the | Version | Date | syntax.
- Task lists - - [ ] and - [x], which become disabled checkboxes.
- Strikethrough with ~~two tildes~~.
- Autolinks: a bare URL turning into a link without any bracket syntax.
Those four are GitHub Flavored Markdown, an extension. A strict CommonMark converter leaves your table as a paragraph full of pipes - which is what it looks like when a README renders correctly on GitHub and wrongly somewhere else. The GitHub extras box is on by default; untick it only when you need plain CommonMark output.
Raw HTML inside the Markdown
Markdown allows HTML on purpose. A <details> block, a <kbd> key, an <img> with a width attribute - none of these have Markdown syntax, so the specification passes the tags through. That is the right default for a document you wrote and the wrong one for a document you did not.
The failure worth knowing about is not an exotic attack. Mention a <script> tag in prose without escaping it and the browser opens a script element there and treats the rest of your document as JavaScript source - everything after it vanishes. The converter flags raw HTML holding a script tag, an iframe or an inline event handler, so you find out before you publish.
When the Markdown came from a form or an API, switch Raw HTML in the Markdown to Show it as text or Remove it. Both also disable javascript: links written in ordinary link syntax - click) is valid Markdown and never passes through the raw-HTML branch, so a converter that only strips tags leaves it live. Treat this as tidying, not security: on a real site, sanitise the HTML server-side against an allow-list as well.
Heading anchors, and why they should match GitHub's
A heading id is what makes #installation work in a link. GitHub generates one from the heading text: lower-case it, delete everything that is not a letter, a number, a space, a hyphen or an underscore, then turn each space into a hyphen. Repeats get -1, -2 and so on.
- ## Setup steps -> id="setup-steps"
- ## Setup steps (again) -> id="setup-steps-1"
- ## What's next? -> id="whats-next"
Matching that rule matters because the anchors inside your own README already use it. Convert the file with a different slug rule and every internal cross-reference breaks the moment you publish the HTML elsewhere.
A fragment, or a whole page
A converter's natural output is a fragment: headings, paragraphs and lists with no <html>, <head> or <body> around them. That is what you want when the HTML is going into a template or a CMS field.
Save that fragment as a .html file and open it, though, and two things go wrong. Without a charset declaration the browser guesses the encoding, so an em dash or a curly quote turns into mojibake. Without a viewport tag a phone renders it at desktop width. Ticking Wrap it in a complete HTML page adds the doctype, the UTF-8 charset, the viewport tag and a title from your first heading, plus an optional stylesheet for tables, code and quotes in light and dark mode.
Doing it without writing any code
- Open the Markdown to HTML converter and paste your Markdown in. Everything runs in your browser and nothing is uploaded.
- Read the notes under the result: headings, links, images, code blocks, and any line breaks the paragraph rule is about to absorb.
- Leave GitHub extras on unless you need plain CommonMark, and tick Keep single line breaks if the source ignores Markdown's wrapping rule.
- If the Markdown is not yours, set Raw HTML in the Markdown to Remove it.
- Switch to Preview to see the rendered result - a sandboxed frame with scripts disabled - then back to HTML to copy or download it.
The output is indented rather than emitted as one unbroken line, which makes it readable in a diff. It uses the same rules as the HTML Beautifier: a line is only broken at a block boundary, where a browser does not render the whitespace, so a paragraph holding inline tags stays on one line. To go the other way and pull the words out of HTML you already have, use the HTML Tag Remover.
Frequently asked questions
- Why do my line breaks disappear when I convert Markdown to HTML?
- Because a single newline inside a paragraph is defined as an ordinary space, not a break. The rule lets you hard-wrap a paragraph in your editor without the wrapping showing up on the page, and it is the same in CommonMark, on GitHub and in every conforming converter. To force a break in the source, end the line with two spaces or with a backslash; both produce a <br>. To convert text that was never written with the rule in mind - a changelog, an address block, a pasted chat log - turn on Keep single line breaks and every newline inside a paragraph becomes a <br> instead.
- Is the HTML I get from a Markdown converter safe to publish?
- Not automatically, and no converter can make it so. Markdown deliberately allows raw HTML, so a <script> tag, an iframe or an onclick attribute written in the source passes straight through to the output - that is the specification working as intended, not a flaw in the tool. If the Markdown is your own, this is exactly what you want. If it came from a user, a form or an API, set raw HTML to Remove it, which also disables javascript: links written in ordinary Markdown link syntax, and then sanitise the resulting HTML server-side against an allow-list before it reaches a page. Stripping tags in the browser is a convenience, not a security boundary.
- How do I convert a whole README to an HTML page I can open?
- Paste the file's contents in, leave GitHub extras on so the tables and task lists survive, and tick Wrap it in a complete HTML page. That adds the doctype, a UTF-8 charset declaration so dashes and quotes do not turn into mojibake, a viewport tag so phones do not render it at desktop width, and a title taken from your first heading. Keep the small stylesheet on for something readable straight away, or switch it off if the page is going into your own design. Download the result as a .html file and it opens in any browser, offline, with the heading anchors from the README still working.