XML has a reputation for being verbose, but most of its rules exist to remove ambiguity that plagues looser formats. Two terms worth knowing precisely — well-formed and valid — get used interchangeably in casual conversation but mean genuinely different things.
Well-Formed vs Valid
Well-formed means the XML follows the basic syntax rules every XML document must follow — every tag is properly closed and nested, attribute values are quoted, and there's exactly one root element wrapping everything else. Valid means the document also conforms to a specific schema (an XSD or DTD) that defines exactly which elements, attributes, and structures are allowed for that particular document type. A document can be perfectly well-formed — syntactically correct — while still being invalid against a given schema, for example by using an element name the schema doesn't recognize. General-purpose formatters typically only check well-formedness, since checking validity requires knowing which schema you're validating against.
Why Entity Escaping Matters
Characters like < and & have special structural meaning in XML — < starts a tag and & starts an entity reference — so if they appear literally inside text content, a parser can't tell whether you meant them as data or as markup. Escaping them as < and & tells the parser explicitly to treat them as literal characters rather than syntax. This is exactly the same underlying idea as escaping quotes inside a string in a programming language — the character itself is ambiguous without a signal for how to interpret it.
Attributes vs Child Elements
There's no strict technical rule for choosing between them, but a common convention is to use attributes for short, simple metadata about an element (like an id or a type) and child elements for the actual content — especially anything that might need its own nested structure, repeat multiple times, or contain extended text. Attributes also can't repeat on the same element and can't contain other elements inside them, which makes them a poor structural fit for anything more complex than a single simple value.
What CDATA Is For
A CDATA section (<![CDATA[ ... ]]>) tells the parser to treat everything inside it as raw literal text, skipping normal entity escaping entirely. It's most useful when embedding content that would otherwise require escaping many special characters — like a chunk of HTML or code — since it lets you paste that content in directly without converting every angle bracket and ampersand by hand.
Formatting Instantly
Paste your XML into our free XML Formatter & Validator to pretty-print it with proper indentation, using your browser's own built-in XML parser — with a real error message if something's genuinely malformed.
FAQ
What's the difference between XML being "well-formed" and "valid"? Well-formed means the XML follows the basic syntax rules every XML document must follow — every tag is properly closed and nested, attribute values are quoted, and there's exactly one root element. Valid means the document also conforms to a specific schema (an XSD or DTD) that defines exactly which elements, attributes, and structures are allowed for that particular document type. A document can be perfectly well-formed while still being invalid against a given schema.
Why do I need to escape characters like & and < in XML text content? Because those characters have special structural meaning in XML — < starts a tag and & starts an entity reference — so if they appear literally inside text content, a parser can't tell whether you meant them as data or as markup. Escaping them as < and & tells the parser explicitly to treat them as literal characters rather than syntax.
When should data be an XML attribute instead of a child element? There's no strict technical rule, but a common convention is to use attributes for short, simple metadata about an element (like an id or a type) and child elements for the actual content, especially anything that might need its own nested structure, repeat multiple times, or contain extended text. Attributes also can't repeat on the same element or contain other elements, which makes them a poor fit for anything more complex than a single simple value.
Are comments and CDATA sections preserved when formatting? Yes — comments and CDATA sections are kept in the formatted output, in their original position in the document structure.