What does the sitemap generator and validator do?

An XML sitemap lists the URLs you want search engines to discover, optionally with the date each page last changed. It does not guarantee indexing; it only helps discovery. This tool does two jobs: it turns a URL list into a protocol-compliant urlset file, and it checks an existing urlset or sitemapindex file against the rules.

It suits developers on custom CMSs, SEOs checking URLs before a migration and site owners reviewing plugin output.

How to use it

  • In Generate mode, write one URL per line. Optionally add lastmod, a language code and a group name: https://example.com/en/ 2026-09-28 en home.
  • If you do not know the date, leave the column out or write -; no lastmod is emitted for that row.
  • Rows that share a group name are language alternates of each other. For spreadsheet exports, pick a CSV delimiter.
  • Press Generate. Errors are listed with row numbers and copying stays disabled until they are fixed.
  • In Validate mode, paste sitemap XML or choose an .xml or .xml.gz file and press Validate.

Split output is downloaded file by file; no ZIP archive is built.

Creating XML from URL lists

Every URL is first validated as an absolute http or https address; relative paths, addresses with a user name or password and # fragments are rejected. The browser's URL parser then percent-encodes non-ASCII path characters, so the Turkish ç becomes %C3%A7, and converts internationalised host names to punycode. Sequences that are already %XX encoded are left alone, so nothing is double-encoded.

XML escaping is a separate second step: an & in the query becomes & in the file and turns back into & when the XML is read. Duplicates keep their first occurrence with a warning, and addresses outside the first URL's host are flagged.

URL and file-size limits

The sitemap protocol allows at most 50,000 URLs and 52,428,800 uncompressed bytes (50 MB) per file. The tool measures the UTF-8 byte length of every entry and starts a new file as soon as either limit would be exceeded. So 50,001 URLs are never written into one file; you get sitemap-1.xml, sitemap-2.xml and a sitemap-index.xml that lists them.

The index addresses come from the sitemap location field, or from the first URL's root when that field is empty. In Validate mode a file above 50,000 entries or 52,428,800 bytes is reported as an error. For gzip files the limit applies to the uncompressed size.

lastmod and hreflang

lastmod should describe when the page really changed, not when the sitemap was generated. Accepted W3C forms include 2026, 2026-09, 2026-09-28 and time-zoned values such as 2026-09-28T10:30+03:00. A time without a zone is an error, an impossible date such as 2026-02-30 is rejected and a future date gets a warning.

For hreflang, give codes such as tr and en within one group. Each page in the group receives every alternate, itself included, as an xhtml:link, which makes the set reciprocal by construction. Codes are checked against ISO 639-1 languages and ISO 3166-1 regions, and en-UK gets an en-GB suggestion. Validate mode checks self-references and return links within the pasted data; alternates outside it are counted as not verified. See Google's hreflang guide.

Worked example

The Load example list contains a Turkish product page with the path çiçek-saksı, its English counterpart, a dated blog post and a search page with the query ?q=seo&sayfa=2. The relevant part of the generated file looks like this:

<url>
  <loc>https://example.com/tr/urunler/%C3%A7i%C3%A7ek-saks%C4%B1</loc>
  <lastmod>2026-09-20</lastmod>
  <xhtml:link rel="alternate" hreflang="tr" href="https://example.com/tr/urunler/%C3%A7i%C3%A7ek-saks%C4%B1"/>
  <xhtml:link rel="alternate" hreflang="en" href="https://example.com/en/products/flower-pot"/>
</url>
<url>
  <loc>https://example.com/ara?q=seo&amp;sayfa=2</loc>
</url>

The Turkish letters are encoded, & became &amp; and both language versions point at each other. The Validate-mode example is broken on purpose: 2026-02-30 is an invalid date, /iletisim/ is a relative address and the English page has no return link to the Turkish one, and each problem appears as its own report row.

Limits and privacy

  • The tool does not check whether URLs are live, return 200, are canonical or indexable, or have been accepted by Google.
  • Child sitemaps listed in an index are not downloaded; validate each one separately.
  • priority and changefreq are checked for format only; no ranking effect is implied.
  • Image, video and news extensions are recognised, but their contents are not validated.
  • XML containing DOCTYPE or ENTITY declarations is rejected before parsing; no external schema or network resolution happens.
  • Inputs never leave your device; large files are processed in a separate Worker.

Frequently asked questions

How many URLs can a sitemap contain?

One file may hold up to 50,000 URLs and 52,428,800 uncompressed bytes. Beyond that the tool splits the list into several files and generates a sitemap index that lists them.

Should lastmod be today's date?

No. lastmod should reflect the page's real last change. Refreshing every date on each build makes it unreliable; leave it empty when unsure.

Are Turkish characters in URLs a problem?

No, but they must be encoded in the file. The tool turns paths such as çiçek into %C3%A7i%C3%A7ek and never encodes an already encoded address a second time.

Does the validator visit the URLs in my sitemap?

No. The checks run only on the text or file you provide; no URL is requested and no status code is checked.

Should hreflang go in the sitemap or in the HTML?

Both are valid; what matters is that the sets are consistent and reciprocal. On large multilingual sites the sitemap method avoids touching page templates.

Published: · Updated: