XML Split
Split a document into separate files, one per repeating child element under a chosen parent.
What is XML Split?
The opposite of merging: a document holding many records becomes a file per record. This is what turns a nightly batch export into something a queue can consume one message at a time.
How it works
The repeating element is detected or specified by path. Each matched element becomes its own document: the original root — and any ancestors between it and the matched element — are rebuilt around it, so every part keeps its structure, its namespace declarations and any sibling fields that explain the batch. Names come from a key attribute when you name one, otherwise from a numbered prefix. A path that matches nothing is reported rather than producing empty files.
- Paste the document. Drop in XML, a schema, a feed or an XPath expression. Everything is parsed in the page: no upload, no server round trip, and the tools keep working with the network switched off.
- Set the options that match your document. Indent unit, whether to keep comments, which dialect, which output method. The defaults are the safe ones — nothing that changes meaning is enabled for you.
- Read the result, then copy or download it. Results appear as you type, copy straight to the clipboard, and download with a sensible filename (.xml, .xsd, .csv). Reset returns every field to its default.
Examples
A minified document from an API response
Paste one line of XML and the Formatter indents it, wraps long start tags one attribute per line and normalises empty elements — without touching text inside mixed content.
An XPath that works in Chrome and not in your servlet
The XPath Tester evaluates against a real tree and reports the axis or function it cannot support, instead of returning zero nodes and letting you guess why.
A document that will not parse
The Validator gives line and column for every mismatched tag, stray ampersand and duplicate attribute, with the reason in words: which open element was expected, which was found.
A schema that only half matches your data
The XSD Validator marks each violation with the path of the offending node, and the Schema Generator goes the other way — inferring a starting-point XSD from real documents.
Common mistakes
Assuming a pretty-printer cannot change meaning
Whitespace between elements is often insignificant, but inside mixed content it is not: `<p>Hello <b>world</b></p>` must stay on one line. Formatting here re-indents element-only content and leaves mixed content alone.
Treating a validation pass as proof of interoperability
The DTD and XSD validators implement honest subsets. A document reported valid here can still be rejected by a full validator that understands identity constraints, substitution groups or XSD 1.1 assertions — each page lists what it does not cover.
Reading the whole document because the parser did
A DOCTYPE is not inert. If your parser resolves external entities, pasting an untrusted document is enough to read a file or call an internal URL. The XXE Risk Checker shows what a document asks for; the fix is in the parser configuration.
Forgetting that attributes and elements are not interchangeable
`<id>1</id>` and `id="1"` mean different things to every schema and every XPath. When converting to CSV or YAML the distinction is preserved (`@id`), because flattening it away is the change that costs an afternoon later.
Frequently asked questions
Why do parts keep the original root element?
Because a bare `<shipment>` element with no root is not a well-formed document and most consumers will reject it. Rebuilding the ancestors means each part can be handed to the same code that processed the original — switch to bare mode if you really only want the element.
How are the file names chosen?
From the attribute you name — `SHP-1001` gives `shipment-1-SHP-1001.xml`, so a record can be found without opening anything. Without one, parts are numbered. Duplicate key values are reported, because two files with the same name is a silent data loss.
Is there a limit on how many parts?
Only the browser's memory, since every part is generated in the page. A few thousand parts is comfortable; beyond that, split with a command-line tool and keep the key-attribute naming convention, which is the part worth copying.