All guides/
Practical SEO workflow

Extract URLs from an XML sitemap without mixing in images or child maps

A sitemap URL list helps with a migration, page inventory or section comparison. First distinguish page URLs from child sitemap files and media references; otherwise a clean-looking export can describe the wrong things.

Identify the file you opened

Open the sitemap linked by robots.txt or supplied by the owner. Save its source URL and retrieval date. If the root element is urlset, the file contains URL entries. If it is sitemapindex, its loc values point to child sitemap files; they are not the site’s article or product URLs.
For an index, record child files and inspect each relevant child. State which sections you include. An export limited to the products sitemap is not a complete website inventory. Failed or inaccessible child files remain gaps, rather than empty sections.

Extract the page loc values precisely

Use an XML-aware reader or import workflow that respects the sitemap namespace. Within a urlset, select loc under each url entry. Avoid selecting every element named loc indiscriminately: image and other extensions can contain separate resource locations.
Preserve the original URL and decode XML entities correctly. An ampersand written as an XML entity belongs to the original query string. A line-based find-and-replace can break compact XML, multiline elements or encoded values. The structure matters more than how the file happens to be formatted.

Synthetic example: a shop with two child maps

Imagine an index referring to products.xml and advice.xml. The product child contains three page entries and an image extension under one entry. The advice child contains two page entries. A page-only inventory has five rows, not eight rows made from child-map and image locations.
The count is invented for illustration. Keep a source-file column alongside each URL. This lets you explain why a row is present and revisit the original entry when the owner questions the inventory. It also makes a later partial export distinguishable from the original complete selection.
ElementMeaningInclude in page list?
sitemapindex / sitemap / locChild sitemap fileNo; inspect the child
urlset / url / locPage URLYes
image:locImage resourceNo
lastmodDeclared update timeOptional separate column

Clean the list without destroying distinctions

Remove exact duplicate rows while retaining their source files in a separate note. Do not automatically lowercase paths or discard query strings: different addresses can represent different resources. Mark unexpected hosts, relative addresses and malformed entries for review.
When comparing inventories, retain both retrieval dates and define the comparison rule. “Missing from the new sitemap” means absent from that file set. It does not establish that the page has been deleted, blocked, deindexed or redirected. Each conclusion needs its own observation.

Separate extraction from a page audit

A successful XML parse does not prove that listed pages respond successfully or can be indexed. HTTP status, robots rules, canonical annotations and noindex directives require separate checks. A large crawl also needs explicit scope and limits rather than being started as an unnoticed side effect of extraction.
The public extractor in this section is planned. Today, use this manual inventory to scope technical work in your SEODatum project and review actual audit results when available. Keep the checklist below with the exported file so colleagues understand exactly what it establishes.
Source sitemap: Retrieved at: Included child files: Missing child files: Page URL count before/after deduplication: Normalization rule: Checks not performed:

Keep evidence and decisions in one project

Keep query groups and collected checks together in SEODatum, then connect the evidence to your next page decision.
Open SEODatum