Sitemap Checker: Find Missing and Conflicting Website Pages

Open your sitemap (usually yoursite.com/sitemap.xml), copy it and paste it here, or upload the file. The checker validates the XML and every URL in it, then compares it with a list of pages you expect and with your robots.txt, so you can see what is missing, what should not be there and what conflicts. It does not visit your website: browsers block one site from reading another, so it checks what you paste. Everything runs in your browser. Nothing you type is sent to We.Inc or anyone else, and it is gone when you close the tab.

How it works

  1. Paste the contents of your sitemap.xml or upload the file (up to 10 MB). Sitemap index files are recognised too; the checker lists the child sitemaps for you to paste one by one.
  2. The XML is parsed in your browser. You see whether it is valid, whether it uses the standard sitemap namespace, and how many URLs it holds against the 50,000 URL and 50 MB limits.
  3. Every URL is checked: full https address, same host as the rest, no # fragment, no tracking parameters, no duplicates, and a valid lastmod date that is not in the future.
  4. Optionally paste the pages you expect to be listed, such as the links in your menu. The report shows expected pages missing from the sitemap and sitemap pages you did not expect.
  5. Optionally paste your robots.txt. The checker flags sitemap URLs that robots.txt blocks, and tells you if robots.txt does not point to a sitemap.
  6. Download the findings as a CSV. For the status code, canonical and noindex of a single URL, run the curl command shown on the page.

Missing pages and pages that should not be listed

Two kinds of mistakes matter. Missing pages are real pages you want found that the sitemap leaves out, often new service pages or blog posts added after the sitemap was made. Unwanted pages are addresses that should not be there at all: redirects, test pages, filtered versions of the same page, or pages marked noindex.

Comparing the sitemap with a list you trust, such as the links in your menu and footer, is the quickest way to find both. This checker keeps them in separate lists so you can fix one at a time.

What this check can and cannot see

Everything here is based on the text you paste. It can prove that the XML is valid and that the URLs are well formed, consistent and allowed by your robots.txt. It cannot see whether each page currently answers 200 or carries a noindex tag, because that needs a visit to each page. The report labels those as not checked rather than guessing.

Tips

Frequently asked questions

Does this sitemap checker crawl my site?

No. A web page cannot read another website's files because browsers block it, and we do not run a hidden crawler. You paste the sitemap, the page list and robots.txt, and the checks run on that text. The report says exactly how many URLs it checked.

How do I check status codes, canonicals and noindex?

Run curl -sIL followed by the page address in a terminal to see its status code and redirects. To see the canonical and robots meta tag, run curl -sL followed by the address and look for rel="canonical" and name="robots" in the output. Search Console's page indexing report shows the same evidence for your whole site.

What are the sitemap limits?

One sitemap file can hold up to 50,000 URLs and be up to 50 MB uncompressed. Larger sites split their URLs across several sitemaps listed in a sitemap index file.

Why are http or other-domain URLs flagged?

A sitemap should list the final addresses on the same site where the sitemap lives. http addresses on an https site redirect, and URLs on another host are ignored by search engines unless you have verified cross-site submission.

Does a valid sitemap mean my pages will be indexed?

No. A sitemap helps search engines find pages; it does not make them index or rank those pages. Quality, links and technical signals still decide that.

Is what I paste stored?

No. The check runs in your browser and nothing is sent anywhere.

Related: XML Sitemap Generator, Robots.txt Generator, Redirect Checker, Website Checker, Website builder.

Turn this result into your website or see the We.Inc website builder. The free plan gives you up to 3 template sites on a we.inc address, with no AI credits and no custom domain. Paid plans are Starter at $20 a month, Pro at $50 and Max at $99. See pricing.

More in Free Tools

We.Inc is an AI-powered website builder you can resell under your own brand. Launch a branded client dashboard, bill on Stripe Connect, and deliver AI-generated websites in minutes. White-label plans start at $99 a month for 25 client sites, with a 7-day free trial and no per-site fees.

Product

Who It's For

Features

Resources

Company

View Sitemap