XML usually arrives unreadable: an API returns a SOAP envelope as one four-thousand-character line, a build tool writes a POM with no line breaks, a colleague pastes a fragment with three different indent widths in it. Making it readable is a ten-second job - and one of the easiest ways to quietly change a document, because whitespace in XML is not always decoration. This guide explains what an XML formatter may safely change, what it must leave alone, and how to use the XML Formatter to pretty-print, minify and check a document is well-formed.
What formatting is allowed to change
Indenting an XML document means adding newlines and spaces between its tags. To a parser those are not nothing: text between two elements is a text node like any other. When a document has a DTD or a schema declaring an element to hold only other elements, a validating parser knows that whitespace is ignorable and discards it. Without a schema - which is the normal case when you paste something into a formatter - nobody can be certain.
So a careful formatter works to a narrower rule that needs no schema: whitespace between two elements, in an element holding no real text, may be rearranged freely, because that is exactly what indentation is made of. Character data - the actual text inside your elements - is never touched. Everything below follows from those two rules.
A worked example
Here is a catalogue entry as an API might hand it to you, all on one line:
- <?xml version="1.0" encoding="UTF-8"?><catalog><book id="bk101"><author>Gambardella, Matthew</author><title>XML Developers Guide</title><price currency="USD">44.95</price><description>An <em>in-depth</em> look at creating applications with XML.</description></book></catalog>
Paste it into the XML Formatter and it comes back like this:
- <?xml version="1.0" encoding="UTF-8"?>
- <catalog>
- <book id="bk101">
- <author>Gambardella, Matthew</author>
- <title>XML Developers Guide</title>
- <price currency="USD">44.95</price>
- <description>An <em>in-depth</em> look at creating applications with XML.</description>
- </book>
- </catalog>
The declaration gets its own line, catalog and book each open a block, and the four leaves sit one level in. The tool also reports what it found - seven elements, four levels deep - a quick way to confirm you pasted the whole document and not half of it. Note the last line: description was not broken up even though it contains a tag. That is deliberate, and it is the part worth understanding.
The element a formatter must not break up
An element that contains both text and child elements has what the spec calls mixed content, and it is everywhere: a paragraph with a bold word in it, a description with a link, an XHTML fragment. Consider this one:
- <p>Hello <b>world</b>!</p>
A formatter that indents every child onto its own line turns it into a p element holding a newline, two spaces, the word Hello, a newline, the b element, a newline and an exclamation mark. The text was 'Hello ' and is now a line break plus indentation. Render that in a browser and the spacing is wrong; check it against a signature and it fails. It is the commonest way an XML beautifier corrupts a document, and the damage stays invisible until something downstream complains.
This tool takes the other route. Any element holding real text is written on one line, with its text copied through exactly as you wrote it, and only elements whose content is entirely other elements get the indent treatment. That is why description came back as a single line above. It means a heavily mixed document - an XHTML page, say - gets tidied less aggressively than you might expect, which is the correct outcome rather than a limitation.
Text values, and why trimming is a choice
One case sits on the boundary. Take a value that someone spread across three lines:
- <contact>
- <name>
- Priya Sharma
- </name>
- </contact>
The name element contains the string newline-spaces-'Priya Sharma'-newline-spaces. Rewriting it as a tidy one-liner is what almost everyone wants, and it is also, strictly, an edit to your data. So it is a switch rather than an assumption: leave 'Trim spaces and line breaks around text values' on, which is the default, and you get name holding exactly 'Priya Sharma'. Turn it off and the value survives byte for byte, which is what you want for a document that is signed, canonicalised, or compared against a checksum.
Either way the trim only removes whitespace at the two ends of a value; spaces inside it are never collapsed, so an element holding 'two spaces' keeps both. One thing overrides the switch entirely: any element carrying xml:space='preserve' is passed through untouched, along with everything nested inside it, because that attribute exists to say the whitespace here is data. Entities such as &#233; are passed through as written rather than decoded, and CDATA sections copied verbatim.
Well-formed is not the same as valid
Before it can format anything, the tool has to parse the document, and that is where it earns most of its keep. Because it will not guess, a document that will not parse comes back as a message naming the problem and its line and column rather than as mangled output. The checks are the ones a real parser makes:
- A closing tag that does not match the tag it closes - remember XML tag names are case sensitive, so </Title> will not close <title>.
- A tag that is never closed, or a closing tag with nothing open.
- A bare & in text or in an attribute value; it has to be written &amp;, and this is the single most common XML parse error there is.
- A < inside an attribute value, an attribute with no value at all, or one whose value is not quoted.
- The same attribute twice on one element.
- Text sitting outside the root element, an unterminated comment or CDATA section, or a comment containing --.
- An XML declaration that is not the very first thing in the file - not even a comment may come before it.
What this is not is validation against a DTD or XSD schema. Well-formed means the syntax holds together; valid means the document also matches a declared model - the right elements, in the right order, with the right attributes. A document can be perfectly well-formed and still be rejected by the service you send it to, and checking that needs the schema, which only you have.
When to minify instead
The Minify tab does the reverse job: it drops the whitespace between elements and puts the document on one line. That is worth doing for XML that travels - SOAP bodies, an RSS feed, a sitemap - where indentation is pure overhead, and the tool reports how many characters it saved. The two directions agree: minifying a document and beautifying it again gives the same result as beautifying the original, because both modes make the same promise about your content. Both also offer 'Remove comments', useful before publishing a config file that still carries a developer's notes.
It runs entirely in your browser
XML is the format of API payloads, invoices and export files, so it routinely holds hostnames, order numbers, customer names and API keys. The XML Formatter parses and rewrites everything inside your browser tab - nothing is uploaded, nothing is stored, and it keeps working with the network off. If your data is in a different format, the JSON Formatter does the same job for JSON, the HTML Minifier handles markup destined for the browser, and JSON to YAML and YAML to JSON cover the config-file pair.
Frequently asked questions
- Why did my paragraph stay on one line instead of being indented?
- Because it holds text as well as tags, and indenting it would change the text. An element with mixed content - a paragraph with a bold word inside, a description containing a link - has real character data between its child elements, so pushing those children onto separate lines inserts newlines and indentation into the string itself. The formatter writes any element holding real text on a single line and only indents elements whose content is entirely other elements, which is what keeps the document you get back identical to the one you pasted in.
- What is the difference between well-formed and valid XML?
- Well-formed means the syntax is correct: every tag is closed and properly nested, attribute values are quoted, and special characters like & are escaped. Valid means the document additionally matches a DTD or XSD schema - the expected elements, in the expected order, with the expected attributes and types. This tool checks well-formedness, which is what a parser needs before it can read the file at all, and reports the line and column of the first problem. It does not check validity, because that requires the schema your system defines.
- Is my XML uploaded anywhere?
- No. The document is parsed and formatted entirely in your browser and nothing is sent to a server, so an API payload, an invoice export or a config file holding hostnames and keys never leaves your device.