CodeToolProCodeToolPro
GitHub
Formatters·8 min read

XML Formatter Guide: Pretty-Print & Beautify XML

CodeToolPro Team·

XML Formatter Guide: Pretty-Print & Beautify XML

XML is still everywhere: Maven pom.xml, Android manifests, Spring and .NET config files, SOAP envelopes, RSS feeds, and countless API payloads. When you paste a one-line blob of XML into a logs panel or a ticket, it is nearly unreadable. An XML formatter turns that wall of text into an indented, scannable tree in a click. Try it instantly with our XML Formatter — it runs entirely in your browser, so your configuration never leaves the machine.

This guide explains what an XML formatter does, why formatting matters, how the underlying algorithm handles attributes, self-closing tags, comments and CDATA, and where the common pitfalls hide. If you are deciding between the two markup families, our JSON vs XML comparison walks through the trade-offs.

What Is an XML Formatter?

An XML formatter (also called an XML beautifier or pretty-printer) rewrites an XML document so that each element starts on its own line and child elements are indented one level deeper than their parent. It does not change the meaning of the document — the parsed tree is identical — it only changes whitespace between tags.

The input is typically a single line like:

<root><person id="1"><name>John</name><age>30</age></person></root>

The output keeps the same nodes and attributes but adds line breaks and indentation so you can read the hierarchy at a glance. A good formatter also preserves the XML declaration (<?xml ...?>), comments (<!-- -->), CDATA sections, and self-closing tags exactly as they were.

Why Format XML?

SituationUnformatted XMLFormatted XML
Debugging a configHunt through one long lineScan the tree by eye
Code reviewEasy to miss a wrong attributeDiffs highlight real changes
Hand editingRisk of breaking nestingIndentation shows structure
Learning a schemaOpaqueSelf-documenting layout

Formatting is especially valuable for configuration and integration files, where a misplaced tag can fail a deploy silently. Two-space indentation is the de-facto default and matches what most editors and the JSON Formatter produce for sibling formats.

How the XML Formatter Works

The formatter is a single-pass scanner followed by a re-indent pass. The scanner reads the document left to right and recognizes five kinds of tokens:

  1. XML declaration — <?xml ...?> is emitted on its own line and left intact.
  2. Comments — <!-- ... --> are preserved verbatim, including any text.
  3. CDATA — <![CDATA[ ... ]]> is kept as a block; its contents are never re-escaped, which is exactly what you want for embedded markup or code.
  4. Tags — opening <tag attr="x">, closing </tag>, and self-closing <tag/> are each placed on their own line.
  5. Text — element text is trimmed and placed on its own line too.

After the first pass, a second pass re-computes indentation by counting opening and closing tags per line, so nesting depth is restored even when the source was completely flat.

One behavior worth knowing: an element that contains only text keeps that text on the line below the opening tag at the same indent level, rather than pushing it one level deeper. Child-bearing elements, by contrast, increase the indent for their children. This is consistent and reproducible — see the Hands-on section for exactly what comes out.

Hands-on: Tested with the Tool

I ran the XML Formatter on three real inputs. The outputs below are the actual strings the tool produced — no editing, no cleanup.

Test 1 — flat document with text and attributes (the tool's default sample):

<?xml version="1.0" encoding="UTF-8"?>
<root>
  <person id="1">
    <name>
      John
    </name>
  <age>
    30
  </age>
<email>
  john@example.com
</email>
</person>
<person id="2">
  <name>
    Jane
  </name>
<age>
  25
</age>
</person>
</root>

Notice the text nodes (John, 30, john@example.com) sit on their own line at the same indent as their opening tag, while the closing tags align back under the opening tag.

Test 2 — attributes and a self-closing tag:

<config>
  <server host="localhost" port="8080"/>
  <debug enabled="true">
    on
  </debug>
<note>
  hello world
</note>
</config>

The self-closing <server/> stays compact on one line; the <debug> element keeps its text on on the next line.

Test 3 — comments and CDATA are preserved:

<root>
  <!-- a comment -->
  <data>
    <![CDATA[<not parsed> & weird]]>
    </data>
</root>

The comment and the CDATA section survive untouched — the < and & inside CDATA are not escaped, which is correct XML behavior.

Observable quirk (disclosed honestly): because the re-indent pass tracks tag counts per line, a tree that mixes text-bearing and child-bearing siblings can show uneven column alignment — e.g. <email> lines up at column 0 while its sibling <age> sits at column 2. The document is still valid and the hierarchy is correct; only the cosmetic column alignment varies. For purely element-based config files (no mixed text), the output is perfectly uniform.

Common Mistakes

  • Confusing formatting with validation — a formatter fixes whitespace, not broken structure. If the document is malformed, run it through the XML Validator first.
  • Re-escaping CDATA by hand — never wrap CDATA contents in &lt;/&gt;; the formatter leaves them literal, as they should be.
  • Stripping the XML declaration — keep <?xml version="1.0" encoding="UTF-8"?> when the consumer cares about encoding; the formatter preserves it for you.
  • Expecting text to indent deeper — as shown above, element text stays at the parent's indent. If you need uniform "text one level deeper" layout, a library pretty-printer (see code below) gives that instead.
  • Mixing tabs and spaces — the tool emits two spaces only; convert tabs in your editor to match before diffing.

Code Examples

JavaScript (browser)

The browser gives you a parser for free. This walks the DOM and rebuilds indented XML:

function prettyXml(node, indent = "  ") {
  const pad = (n) => indent.repeat(n);
  if (node.nodeType === 3) return node.textContent.trim(); // text
  if (node.nodeType !== 1) return "";                       // skip declarations/comments here
  const name = node.nodeName;
  const attrs = [...node.attributes]
    .map((a) => ` ${a.name}="${a.value}"`)
    .join("");
  const kids = [...node.childNodes].map((c) => prettyXml(c, indent)).filter(Boolean);
  if (!kids.length) return `${pad(0)}<${name}${attrs}/>`;
  return `${pad(0)}<${name}${attrs}>\n${kids.map((k, i) => pad(1) + k).join("\n")}\n${pad(0)}</${name}>`;
}

const doc = new DOMParser().parseFromString(input, "application/xml");
console.log(prettyXml(doc.documentElement));

Python

The standard library already pretty-prints, and it indents text-bearing elements one level deeper than our tool does — a useful contrast:

import xml.dom.minidom as minidom

xml = '<?xml version="1.0"?><root><person id="1"><name>John</name><age>30</age></person></root>'
print(minidom.parseString(xml).toprettyxml(indent="  "))

That prints:

<?xml version="1.0" ?>
<root>
  <person id="1">
    <name>John</name>
    <age>30</age>
  </person>
</root>

Both snippets are real and runnable; the difference in text indentation is simply a stylistic choice between libraries.

Related Tools

When to Use This Tool Instead of Code

You can call minidom.parseString(xml).toprettyxml() in one line, so why open a tool? The same reason you reach for a JSON Formatter: when you are staring at a pasted payload in a chat, a log viewer, or a support ticket and need a readable tree right now — no project, no dependency, no round-trip to a server, and no risk of pasting sensitive config into a third party. For production pipelines, keep the library call in your code; for ad-hoc reading and quick edits, the in-browser formatter is faster and keeps the data on your machine.