XML is what older systems speak: SOAP responses, RSS feeds, bank statements, government APIs, Android layouts, Maven builds. JSON is what everything written this decade expects. So converting between them is routine - and unlike JSON to YAML, there is no single correct answer, because XML can express several things JSON simply cannot. This guide explains how the XML to JSON converter maps one onto the other, which decisions you have to make yourself, and which parts of a document cannot survive the trip at all.
Why there is no official mapping
JSON has objects, arrays, strings, numbers, booleans and null. XML has all of that plus four things with no JSON equivalent:
- Attributes. An element can carry named values on its opening tag as well as content inside it, and both are data.
- Order between text and tags. In XML the position of a tag inside a run of text is meaningful; JSON object keys have no position.
- Comments, processing instructions and the XML declaration. JSON has no syntax for any of them.
- Namespace prefixes, which qualify a name rather than being part of it.
Every converter therefore invents conventions. The one used here is the most widely recognised: attributes get a marker prefix, an element's own text goes into a #text key, and the root element stays as the outermost key so its name is not thrown away. What matters is not which convention you pick but that the losses are visible - so the converter reports every place where a decision was made for you.
A worked example
Take a small order document with an attribute on the root, a nested object and a repeated element:
- <?xml version="1.0" encoding="UTF-8"?>
- <order id="ORD-0098">
- <customer>
- <name>Tom & Jerry Ltd</name>
- <zip>007</zip>
- </customer>
- <items>
- <item sku="A1">Widget</item>
- <item sku="A2">Gizmo</item>
- </items>
- </order>
Paste it into the XML to JSON converter and you get this:
- {
- "order": {
- "@id": "ORD-0098",
- "customer": {
- "name": "Tom & Jerry Ltd",
- "zip": "007"
- },
- "items": {
- "item": [
- { "@sku": "A1", "#text": "Widget" },
- { "@sku": "A2", "#text": "Gizmo" }
- ]
- }
- }
- }
Four things happened there. The id attribute became @id. The two item elements became an array. Each item had both an attribute and text, so its text went into #text. And Tom & Jerry Ltd came back as Tom & Jerry Ltd, because & is an XML entity and JSON has no entities - leaving it as written would have been a bug. Three of those four are worth understanding properly.
The one-item list problem
This is the single biggest trap in XML to JSON conversion, and it does not show up until production. XML has no way to mark an element as a list, so a converter can only count: two item tags become an array, one item tag becomes a plain object. Delete the second item from the example above and the same document converts to this instead:
- "items": { "item": { "@sku": "A1", "#text": "Widget" } }
The key is no longer an array. Code that did items.item.map(...) now throws, on the first order that happens to contain a single line. The shape of your JSON is being decided by your data rather than by your schema, and the failure only appears once real traffic produces a one-element list.
The fix is to stop counting. Switch on Always arrays and every child element becomes an array whether it repeats or not, so items.item is a one-element array in the case above and your loop works in both. It is more verbose to read, and that is the trade: use the default when you are eyeballing a document, and Always arrays whenever code will consume the output. The converter tells you which situation you are in - whenever repeat-counting actually decided a key's shape, it says so.
Attributes and why they need a prefix
An attribute and a child element are allowed to share a name. This is valid XML:
- <book id="1"><id>2</id></book>
Without a prefix both want the key id, and one silently overwrites the other. That is why the default is @id rather than id, and you can switch it to _, $ or nothing at all. If you do choose no prefix and a genuine collision occurs, the converter names the key rather than quietly dropping a value.
The same reasoning covers an element's own text when it also has attributes: it has to go somewhere, and #text is the conventional place. An element with no attributes and no children skips all of this and becomes a plain string, which is why name above is just a string rather than an object.
Keeping numbers, IDs and versions intact
XML is untyped: every value is text, and any typing is an assumption the converter makes. That assumption is where data gets destroyed. Turn a zip code of 007 into a number and it becomes 7. Turn a version of 1.50 into a number and it becomes 1.5. Turn a 20-digit order ID into a number and JavaScript rounds it to a double, replacing the last few digits with zeros.
So values stay strings by default - which is lossless, and why 007 survived the example above. When you do switch type conversion on, a number is only written if its digits come back identical, so 8080 becomes a number while 007, 1.50, +1 and 12345678901234567890 all stay text. true and false become booleans under the same switch.
Namespaces, SOAP and the soap: prefix
A SOAP response is mostly prefixes, and they make the JSON painful to navigate - every key is soap:Envelope or m:GetPrice, which means bracket syntax everywhere in your code. Turning on Remove namespace prefixes maps soap:Envelope to Envelope and drops the xmlns declarations along with it, since the prefixes they declare no longer appear. If two different names would collide once stripped - a:id and b:id both becoming id - you are told rather than losing one.
What cannot survive
Some things have no JSON form at all, and the honest answer is to report them. Comments, processing instructions and the XML declaration are dropped and counted. Mixed content - an element holding text and tags together - loses the order between them: <p>Hello <b>world</b>!</p> becomes a b key and a #text of "Hello !", which is enough for data but not for a document. And a named entity nobody declared, being the usual one, is left exactly as written, because only < > & ' and " exist in XML without a DTD and inventing HTML's table would be putting content into your data that was never there.
It runs entirely in your browser
API payloads carry customer names, order values, hostnames and tokens, so where a conversion happens is not a detail. The XML to JSON converter parses and converts inside your browser tab - nothing is uploaded, so you can paste a live response safely. If the XML will not parse, the XML Formatter reports the exact line and column of the problem, and the JSON Formatter will pretty-print or validate the result once you have it.
Frequently asked questions
- Why did my array turn into a single object?
- Because that element appeared only once. XML cannot mark an element as a list, so a converter counts occurrences: repeated names become arrays, a single name becomes a plain object. It means the shape of your JSON depends on the data, and code that loops over the list breaks the first time a real record has one entry. Switch on Always arrays to get an array for every child element regardless of how many times it appears - that is the setting to use whenever code, rather than a person, will read the output.
- Will a leading zero or a long ID be damaged?
- No. Every value stays a string unless you turn type conversion on, so 007 stays "007" and a 20-digit ID keeps all twenty digits. With conversion on, a number is only written when its digits come back exactly the same, so 8080 becomes a number while 007, 1.50 and anything past 2^53 stay text. Converters built on parseFloat get all three of those wrong.
- Is my XML uploaded anywhere?
- No. The XML is parsed and converted entirely in your browser and nothing is sent to a server, so a SOAP response, an API payload or a config file holding hostnames and keys never leaves your device.