# XML sitemap checker and validator

URL: https://serpel.app/tools/sitemap-checker

Updated: 2026-10-10

A sitemap checker fetches your XML sitemap and tests it against the sitemap protocol. This free tool reads your robots.txt, validates the file and its child sitemaps, and requests up to 20 listed URLs to find 404 errors, redirects and noindex pages. It needs no account.

## What does a sitemap checker test?

An XML sitemap is a file that lists the URLs you want search engines to crawl, optionally with the date each page last changed. A sitemap checker fetches that file and tests it against the [sitemap protocol](https://www.sitemaps.org/protocol.html) and [Google’s sitemap guidelines](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap). This one works as an XML sitemap validator and as a live check of the URLs inside.

For every file it reads, the checker tests:

- **Delivery.** The HTTP status, the content type, and whether the file is compressed, either as a `.xml.gz` file or in transfer.
- **Structure.** Well-formed XML with the line number of the first syntax error, the `<urlset>` or `<sitemapindex>` root element, the sitemap namespace and UTF-8 encoding.
- **Limits.** No more than 50,000 URLs and 50 MB uncompressed per file.
- **URLs.** Every `<loc>` must be absolute, escaped, shorter than 2,048 characters and on the same host and protocol as the sitemap. Duplicates are flagged too.
- **Values.** Malformed or future `<lastmod>` dates, and `<changefreq>` or `<priority>` values outside the allowed range.

It then requests up to 20 of the listed URLs, spread evenly across the sitemap, and reports any that return 404 or another error, redirect or send a `noindex` header. A sitemap should list the URLs you want to appear in search, preferably the canonical version of each, so these are the entries that waste crawl effort.

## How do you check a sitemap with this validator?

1. **Enter a domain or a sitemap URL** Enter a domain such as `example.com` and the checker reads robots.txt for Sitemap lines, then tries `/sitemap.xml` and `/sitemap_index.xml`. Or enter the full URL of one sitemap, including a `.xml.gz` file.

2. **Run the check** The checker reads up to 3 sitemaps from robots.txt and up to 10 child sitemaps of an index. Large sitemaps can take up to half a minute.

3. **Read the verdict and the file table** Check the status, size and entry count of each file, then work through the findings, errors first.

4. **Fix the source and check again** Change the file in your CMS, framework or sitemap generator, then check again. Results are cached for 5 minutes, so wait a little before you recheck.

## What does a correct XML sitemap look like?

A sitemap is a UTF-8 XML file. The root element is `<urlset>` with the sitemap namespace, and every `<url>` needs a `<loc>`. Escape the characters `&`, `'`, `"`, `>` and `<` as entities, so `&` becomes `&amp;` inside a URL.

```xml
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://example.com/</loc>
    <lastmod>2026-10-01</lastmod>
  </url>
  <url>
    <loc>https://example.com/pricing</loc>
    <lastmod>2026-09-18</lastmod>
  </url>
  <url>
    <loc>https://example.com/blog?topic=seo&amp;page=2</loc>
    <lastmod>2026-09-30T08:15:00+00:00</lastmod>
  </url>
</urlset>
```

Large sites split their URLs into several files and list them in a sitemap index. The protocol allows up to 50,000 sitemaps in one index, and [Google](https://developers.google.com/search/docs/crawling-indexing/sitemaps/large-sitemaps) accepts up to 500 index files per Search Console property.

```xml
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://example.com/sitemap-pages.xml</loc>
    <lastmod>2026-10-01</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://example.com/sitemap-blog.xml.gz</loc>
    <lastmod>2026-09-30</lastmod>
  </sitemap>
</sitemapindex>
```

**Limits and valid values in an XML sitemap**
| Item | Rule |
| --- | --- |
| URLs per file | At most 50,000 |
| File size | At most 50 MB (52,428,800 bytes) uncompressed. Gzip is allowed, but the limit applies after decompression. |
| Sitemaps per index | At most 50,000 |
| `<loc>` | A fully qualified URL under 2,048 characters, on the same protocol and host as the sitemap |
| `<lastmod>` | A W3C date such as `2026-10-01`, or with a time such as `2026-09-30T08:15:00+00:00` |
| `<changefreq>` | always, hourly, daily, weekly, monthly, yearly or never |
| `<priority>` | 0.0 to 1.0, with a default of 0.5 |

## What are the most common sitemap errors?

- **An unescaped ampersand.** A URL such as `?a=1&b=2` makes the file invalid XML. Write `&amp;`. The checker shows the line of the first syntax error.
- **A wrong or missing namespace.** Start the file with `<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">`.
- **Relative URLs.** `/pricing` is not valid in `<loc>`. Use `https://example.com/pricing`.
- **http, https and www mismatches.** The [protocol](https://www.sitemaps.org/protocol.html) and [Google](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap) expect the URLs to use the same protocol and host as the sitemap. List the canonical version of each page.
- **URLs that redirect, return 404 or are noindex.** List the final URL, drop deleted pages and keep `noindex` pages out. A missing `/sitemap.xml` that answers 200 with a page is a soft 404, which our [soft 404 guide](https://serpel.app/blog/soft-404) explains.
- **A lastmod that is always today.** Google uses `<lastmod>` only if it is consistently and verifiably accurate. In Next.js, set real dates in `sitemap.ts`, as our [Next.js SEO guide](https://serpel.app/blog/nextjs-seo) shows.
- **A sitemap nobody can fetch.** A catch-all route that serves HTML, a login or a firewall rule that answers bots with 403 hides the file. Check the status and content type in the table.
- **Files over the limits.** Split the sitemap into several files and list them in a sitemap index.

## How do you submit a sitemap?

Add a Sitemap line with the full URL to your robots.txt, once per file, and submit the URL in the Sitemaps report in Google Search Console. The [robots.txt checker](https://serpel.app/tools/robots-txt-checker) lists the Sitemap lines it finds. [Google](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap) ignores `<priority>` and `<changefreq>`, so you can leave them out.

```text
User-agent: *
Allow: /

Sitemap: https://example.com/sitemap.xml
Sitemap: https://example.com/sitemap-blog.xml
```

## What does this sitemap checker not do?

- It reads XML sitemaps only, not text or RSS sitemaps, and it ignores image, video and news extensions.
- It checks at most 10 child sitemaps and 20 URLs per request, and it does not descend into nested index files.
- It requests each sampled URL once and does not render JavaScript. It reads the HTTP status, the redirect target and the `X-Robots-Tag` header, but not a `<meta name="robots">` tag in the HTML.
- It does not show whether Google has processed your sitemap. The Sitemaps report in Search Console does.
- Its requests come from Serpel’s servers with a SerpelBot user agent. A firewall may treat them differently from Googlebot.

For the whole site, Serpel’s [site audit](https://serpel.app/features/site-audit) includes indexing and crawling checks, among them robots.txt and sitemap problems.

## Frequently asked questions

### How do I check whether my sitemap is valid?

Enter your domain or the full sitemap URL in the checker. It validates the XML, the namespace, the limits and every `<loc>`, then requests a sample of the listed URLs. To see how Google processed the file, open the Sitemaps report in Search Console.

### How many URLs can a sitemap contain?

At most 50,000 URLs and 50 MB uncompressed per file, according to the [sitemap protocol](https://www.sitemaps.org/protocol.html). Gzip shrinks the transfer, but the limits still apply after decompression. Larger sites split the URLs into several files and list them in a sitemap index.

### Does Google use changefreq and priority?

No. [Google says](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap) it ignores `<priority>` and `<changefreq>`, and that it uses `<lastmod>` only if the value is consistently and verifiably accurate. Leave out the first two and set the last to the real modification date.

### Why does the checker report redirects, 404s or noindex in my sitemap?

A sitemap should list the URLs you want to appear in search, preferably canonical, as [Google’s guidelines](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap) put it. A URL that redirects, returns 404 or sends a noindex header works against that. Replace it with the final URL or remove it.

### Where should I put my sitemap?

Put it in the root of your host, for example `/sitemap.xml`. The protocol limits a sitemap to URLs at or below its own directory, unless you submit it through Search Console. Reference it from robots.txt with a Sitemap line as well.

### Can the checker read a .xml.gz sitemap?

Yes. Enter the URL of the `.xml.gz` file and the checker decompresses it, up to 50 MB uncompressed. It also reports whether a file is compressed, either as gzip or in transfer.