XML Formatter Guide: Pretty-Print & Beautify XML
XML Formatter Guide: Pretty-Print & Beautify XML
XML is still everywhere: Maven pom.xml, Android manifests, Spring and .NET
config files, SOAP envelopes, RSS feeds, and countless API payloads. When you
paste a one-line blob of XML into a logs panel or a ticket, it is nearly
unreadable. An XML formatter turns that wall of text into an indented,
scannable tree in a click. Try it instantly with our
XML Formatter — it runs entirely in your browser,
so your configuration never leaves the machine.
This guide explains what an XML formatter does, why formatting matters, how the underlying algorithm handles attributes, self-closing tags, comments and CDATA, and where the common pitfalls hide. If you are deciding between the two markup families, our JSON vs XML comparison walks through the trade-offs.
What Is an XML Formatter?
An XML formatter (also called an XML beautifier or pretty-printer) rewrites an XML document so that each element starts on its own line and child elements are indented one level deeper than their parent. It does not change the meaning of the document — the parsed tree is identical — it only changes whitespace between tags.
The input is typically a single line like:
<root><person id="1"><name>John</name><age>30</age></person></root>
The output keeps the same nodes and attributes but adds line breaks and
indentation so you can read the hierarchy at a glance. A good formatter also
preserves the XML declaration (<?xml ...?>), comments (<!-- -->), CDATA
sections, and self-closing tags exactly as they were.
Why Format XML?
| Situation | Unformatted XML | Formatted XML |
|---|---|---|
| Debugging a config | Hunt through one long line | Scan the tree by eye |
| Code review | Easy to miss a wrong attribute | Diffs highlight real changes |
| Hand editing | Risk of breaking nesting | Indentation shows structure |
| Learning a schema | Opaque | Self-documenting layout |
Formatting is especially valuable for configuration and integration files, where a misplaced tag can fail a deploy silently. Two-space indentation is the de-facto default and matches what most editors and the JSON Formatter produce for sibling formats.
How the XML Formatter Works
The formatter is a single-pass scanner followed by a re-indent pass. The scanner reads the document left to right and recognizes five kinds of tokens:
- XML declaration —
<?xml ...?>is emitted on its own line and left intact. - Comments —
<!-- ... -->are preserved verbatim, including any text. - CDATA —
<![CDATA[ ... ]]>is kept as a block; its contents are never re-escaped, which is exactly what you want for embedded markup or code. - Tags — opening
<tag attr="x">, closing</tag>, and self-closing<tag/>are each placed on their own line. - Text — element text is trimmed and placed on its own line too.
After the first pass, a second pass re-computes indentation by counting opening and closing tags per line, so nesting depth is restored even when the source was completely flat.
One behavior worth knowing: an element that contains only text keeps that text on the line below the opening tag at the same indent level, rather than pushing it one level deeper. Child-bearing elements, by contrast, increase the indent for their children. This is consistent and reproducible — see the Hands-on section for exactly what comes out.
Hands-on: Tested with the Tool
I ran the XML Formatter on three real inputs. The outputs below are the actual strings the tool produced — no editing, no cleanup.
Test 1 — flat document with text and attributes (the tool's default sample):
<?xml version="1.0" encoding="UTF-8"?>
<root>
<person id="1">
<name>
John
</name>
<age>
30
</age>
<email>
john@example.com
</email>
</person>
<person id="2">
<name>
Jane
</name>
<age>
25
</age>
</person>
</root>
Notice the text nodes (John, 30, john@example.com) sit on their own line
at the same indent as their opening tag, while the closing tags align back under
the opening tag.
Test 2 — attributes and a self-closing tag:
<config>
<server host="localhost" port="8080"/>
<debug enabled="true">
on
</debug>
<note>
hello world
</note>
</config>
The self-closing <server/> stays compact on one line; the <debug> element
keeps its text on on the next line.
Test 3 — comments and CDATA are preserved:
<root>
<!-- a comment -->
<data>
<![CDATA[<not parsed> & weird]]>
</data>
</root>
The comment and the CDATA section survive untouched — the < and & inside
CDATA are not escaped, which is correct XML behavior.
Observable quirk (disclosed honestly): because the re-indent pass tracks tag counts per line, a tree that mixes text-bearing and child-bearing siblings can show uneven column alignment — e.g.
<email>lines up at column 0 while its sibling<age>sits at column 2. The document is still valid and the hierarchy is correct; only the cosmetic column alignment varies. For purely element-based config files (no mixed text), the output is perfectly uniform.
Common Mistakes
- Confusing formatting with validation — a formatter fixes whitespace, not broken structure. If the document is malformed, run it through the XML Validator first.
- Re-escaping CDATA by hand — never wrap CDATA contents in
</>; the formatter leaves them literal, as they should be. - Stripping the XML declaration — keep
<?xml version="1.0" encoding="UTF-8"?>when the consumer cares about encoding; the formatter preserves it for you. - Expecting text to indent deeper — as shown above, element text stays at the parent's indent. If you need uniform "text one level deeper" layout, a library pretty-printer (see code below) gives that instead.
- Mixing tabs and spaces — the tool emits two spaces only; convert tabs in your editor to match before diffing.
Code Examples
JavaScript (browser)
The browser gives you a parser for free. This walks the DOM and rebuilds indented XML:
function prettyXml(node, indent = " ") {
const pad = (n) => indent.repeat(n);
if (node.nodeType === 3) return node.textContent.trim(); // text
if (node.nodeType !== 1) return ""; // skip declarations/comments here
const name = node.nodeName;
const attrs = [...node.attributes]
.map((a) => ` ${a.name}="${a.value}"`)
.join("");
const kids = [...node.childNodes].map((c) => prettyXml(c, indent)).filter(Boolean);
if (!kids.length) return `${pad(0)}<${name}${attrs}/>`;
return `${pad(0)}<${name}${attrs}>\n${kids.map((k, i) => pad(1) + k).join("\n")}\n${pad(0)}</${name}>`;
}
const doc = new DOMParser().parseFromString(input, "application/xml");
console.log(prettyXml(doc.documentElement));
Python
The standard library already pretty-prints, and it indents text-bearing elements one level deeper than our tool does — a useful contrast:
import xml.dom.minidom as minidom
xml = '<?xml version="1.0"?><root><person id="1"><name>John</name><age>30</age></person></root>'
print(minidom.parseString(xml).toprettyxml(indent=" "))
That prints:
<?xml version="1.0" ?>
<root>
<person id="1">
<name>John</name>
<age>30</age>
</person>
</root>
Both snippets are real and runnable; the difference in text indentation is simply a stylistic choice between libraries.
Related Tools
- Format and validate XML with the XML Validator.
- Convert the other way with the JSON to XML Converter.
- Parse XML back into JSON with the XML to JSON Converter.
- Pretty-print sibling data with the JSON Formatter.
When to Use This Tool Instead of Code
You can call minidom.parseString(xml).toprettyxml() in one line, so why open
a tool? The same reason you reach for a JSON Formatter:
when you are staring at a pasted payload in a chat, a log viewer, or a support
ticket and need a readable tree right now — no project, no dependency, no
round-trip to a server, and no risk of pasting sensitive config into a third
party. For production pipelines, keep the library call in your code; for
ad-hoc reading and quick edits, the in-browser formatter is faster and keeps the
data on your machine.