All lintlab tools

The state of XML sitemaps on the CrUX top-1k origins (measured October 2026)

We fetched robots.txt and XML sitemaps for the 1,000 most popular origins in the Chrome UX Report (CrUX) from Google, its top-1k bucket, and checked them against the sitemaps.org protocol. Of the 916 origins that responded as websites, we found a sitemap on 400 (43.7%). 112 of those had no issue in the checks we ran on our sample.

Data from one crawl on 2026-10-01 (UTC), plus a same-day re-fetch of the origins it first counted as blocked (see Limitations). An origin is a scheme and host, such as https://www.example.com, and subdomains count as separate origins. Unless a row says otherwise, percentages are out of the 916 origins that responded as websites, not all 1,000. Why the denominator is 916.

Headline findings

  • We found a sitemap on fewer than half of the websites, and that's a floor. We found one on 400 of 916 (43.7%). 417 (45.5%) declared a sitemap in robots.txt, and 71 of the sitemaps we found came from /sitemap.xml or /sitemap_index.xml instead of a robots.txt line. Another 165 (18.0%) refused our requests (403, 429 or a challenge page), and on 77 (8.4%) robots.txt rules kept us from looking, so some of those may have a sitemap we couldn't see.
  • Most of the sitemaps we found start with an index file. 292 of 400 (73.0%) begin with a sitemap index, and 278 (69.5%) served at least one gzip-compressed file in our sample.
  • 112 of 400 (28.0%) had no issue in any check we ran on the files and URLs we sampled. 248 (62.0%) had at least one measured issue. We couldn't fully check 40 (10.0%).
  • The XML itself was rarely the problem. The URLs and dates in it were. The common issues were a sampled URL that answered with a redirect (92 sites, 23.0%), one lastmod value on every URL in a file (88, 22.0%), URLs on a different host from the listed origin (72, 18.0%) and the same URL listed more than once (69, 17.3%). A wrong or missing namespace turned up on 15 sites and XML that didn't parse on 4.
  • Most sampled URLs answered 200 directly. We status-checked 7,456 sampled sitemap URLs: 82.3% returned 200, 11.6% redirected, 2.4% returned a 4xx other than 403/429, and 0.6% returned a 5xx.

Bonus: how many of these origins publish an llms.txt?

On the same 1,000 origins (2026-10-02, one request each, no retries), 66 (6.6%) answered /llms.txt with HTTP 200 and a plain-text or Markdown content type. 627 (62.7%) returned 404, and 112 more returned status 200 with another content type (usually an HTML "not found" page), which we didn't count. Publishing the file doesn't show that any AI system reads it.

Method in brief

We took the 1,000 origins in the top-1k popularity bucket of CrUX's August 2026 global snapshot and measured each one exactly as listed, without switching between www and the bare domain. We read robots.txt first and followed it. We looked for sitemaps in robots.txt Sitemap: lines, then at /sitemap.xml and /sitemap_index.xml. We parsed what we found and checked it against the sitemaps.org protocol. Then we sent a status request to a sample of up to 20 listed URLs per origin. We sampled large sitemap indexes rather than fetching them in full. We bypassed no challenges and used no logins. Full method and how to reproduce it.

Outcomes: what we found per origin

Each origin gets exactly one outcome. 84 of the 1,000 origins gave us neither a robots.txt file nor a web page we could read. For 50 of them, robots.txt didn't answer within our 15-second limit. RFC 9309 treats an unreachable robots.txt as "disallow everything", so we didn't request their homepage either. For most of the rest, the robots.txt or homepage request failed with another timeout or with a connection, TLS, HTTP/2 or DNS error. We can't tell whether these are websites without a sitemap, so we left them out of the denominator.

OutcomeWhat it meansOriginsShare of 916
Sitemap, no issue foundEvery sampled file fetched and parsed, at least one URL sampled, no measured issue.11212.2%
Sitemap with issuesAt least one sitemap parsed and at least one issue from the table below.24827.1%
Sitemap, not fully checkedSitemap found, no issue detected, but a sampled child file failed to load or wasn't a sitemap, or no URL could be sampled.404.4%
No sitemap at the places we checkedThe website responded and we were allowed to look, but no robots.txt line or default path gave us a sitemap.25627.9%
Blockedrobots.txt or a sitemap request returned 403 or 429, a challenge header, or a page that matched a challenge signature. We didn't try to get past it.16518.0%
robots.txt rules prevented the fetchrobots.txt disallowed the sitemap paths for our user agent, or robots.txt returned 5xx or was unreachable, which RFC 9309 treats as "disallow everything".778.4%
Sitemap fetch errorA sitemap request timed out, failed in transport or returned 5xx.182.0%

Some other things we saw: robots.txt was reachable on 779 of 916 websites (85.0%). 50 origins answered at least one sitemap path with an ordinary HTML page and status 200 (75 responses in total). That's a "soft 404": the status says the request succeeded, but there's no sitemap at that address.

Under RFC 9309, a robots.txt that returns a 4xx means a crawler may fetch anything, and one that is unreachable or returns 5xx means it must assume everything is disallowed. We applied both rules. We treated 403 and 429 more cautiously: we counted them as blocked and didn't retry. Across all 1,000 origins, robots.txt returned an ordinary 4xx on 75, a 403 or 429 on 87 and a 5xx on 2, and couldn't be reached on 84. (RFC 9309, sections 2.3.1.3 and 2.3.1.4)

Issues among the 400 sitemaps we found

Each website counts once per issue type, and one website can have several. Rows show how many of the 400 websites with a found sitemap had the issue in the files and URLs we sampled.

IssueSitesShareHow we checkedWhy it matters for crawling
Sampled URL answered 3xx9223.0%The first response to a status request was a redirect. We counted it without following it.Google asks you to list the URLs you want shown in search results, and it generally shows canonical URLs. A URL that redirects isn't the final address.
One lastmod for every URL in a file8822.0%In a urlset with more than one URL, every entry has a lastmod and all of them are the same.Google uses lastmod only when it is consistently and verifiably accurate, and a single shared value may record when the file was generated, not when each page changed.
URL on another host7218.0%The hostname in <loc> differs from the listed origin's hostname. www and the bare domain count as different hosts.sitemaps.org says all URLs in a sitemap must be from a single host, and Search Console lists "URL not allowed" for URLs on a different domain from the sitemap file, unless cross-site submission covers them.
Same URL listed more than once6917.3%An identical <loc> appears twice or more across a site's sampled files.A repeated entry adds no new URL but still counts toward the 50,000-entry limit per file.
Sampled URL answered 4xx (not 403/429)4110.3%A status request returned a client error such as 404 or 410.The sitemap points crawlers to an address that had no page to show when we checked.
Wrong or missing namespace153.8%The root element of a sampled file isn't in http://www.sitemaps.org/schemas/sitemap/0.9.Search Console reports "Incorrect namespace" when the root element lacks the correct namespace or declares it incorrectly.
lastmod in the future143.5%A valid lastmod later than the moment we checked.A date that hasn't happened yet can't describe the last change, which works against the accuracy Google needs before it uses lastmod.
Scheme mismatch82.0%A <loc> uses a different scheme (http or https) from the listed origin.Google says it tries to crawl URLs exactly as listed, so a URL on the other scheme from your canonical homepage adds a redirect or points at a duplicate. (The protocol's own rule is narrower: URLs should use the sitemap's protocol.)
Malformed XML41.0%A sampled file isn't well-formed XML.Search Console reports a "Parsing error" when it can't parse the XML, and it names an unescaped character in a URL as a common cause.
Sampled URL answered 5xx41.0%A status request returned a server error.The page couldn't be served when we checked. One 5xx can be temporary, so recheck before acting on it.
Invalid lastmod30.8%Not a complete date (YYYY-MM-DD) or a full timestamp with seconds and a time zone, or not a real calendar date.sitemaps.org asks for W3C Datetime, and Search Console reports "Invalid date" otherwise.
URL with a #fragment20.5%A <loc> contains #.Google says Search generally doesn't support URL fragments, so a # variant doesn't name a separate page.
More than 50,000 entries in one file10.3%A sampled urlset or index lists more than 50,000 entries.sitemaps.org and Google both cap one file at 50,000 URLs, and Search Console asks you to split a larger one.
More than one <loc> in one entry00.0%A <url> or <sitemap> entry holds several <loc> elements. We counted only the first.The sitemaps.org schema allows one <loc> per entry, so a crawler may ignore the others.

About the "another host" row: sitemaps.org allows cross-submission through robots.txt, and Google allows several domains in one sitemap once you've verified them in Search Console. We didn't check either, so some of these URLs may be intentional.

We didn't score <priority> or <changefreq>. Google says it ignores both, and sitemaps.org calls them hints.

lastmod in numbers

45.6% of the 13,894,476 URL entries in sampled files had a lastmod (6,331,629). 99.3% of those parsed as valid dates. 53,804 valid dates, on 14 sites, were later than the time of the check.

Sources: sitemaps.org protocol · Google: Build and submit a sitemap · Google: Manage large sitemaps · Search Console Sitemaps report · Google: URL structure

Status sample: what listed URLs returned

We checked up to 20 sampled URLs per website, chosen by hash, with a HEAD request over HTTP/1.1. We fell back to GET only when HEAD returned 405 or 501, and we stopped that GET after 64 KiB. We counted redirects at the first hop. Before applying the 20-URL cap, we skipped URLs that robots.txt disallowed for our user agent and URLs on hosts whose robots.txt we couldn't read.

ResponseURLsShare of 7,456Sites with at least one (of 400)
2006,13582.3%350
Other 2xx10.0%1
3xx redirect86311.6%92
4xx (not 403/429)1822.4%41
5xx440.6%4
403 or 4291962.6%15
Request error (no complete response)350.5%4

We kept 403 and 429 separate from other 4xx because they can mean our checker was refused, not that the page is missing. The 35 request errors break down into 20 where the server closed the connection without replying (curl error 52), 13 GET fallbacks that ran past our 64 KiB cap after a 200 header had arrived (curl error 63), and 2 timeouts at 15 seconds. We count any transfer that fails as an error, even when headers arrived. Request errors aren't issues, so they don't change an outcome: 2 of the 112 sites with no issue found had at least one. While filling each site's sample, we skipped 5,470 URLs because of robots.txt: 5 were disallowed for our user agent and 5,465 sat on hosts whose robots.txt we couldn't read. 34 websites with a sitemap got fewer than 20 status checks, and 12 of those got none.

What to check on your own site

  1. robots.txt has a Sitemap: line with the sitemap's full, absolute URL, and robots.txt itself returns 200 or a plain 404, not a 5xx or a challenge page.
  2. The sitemap URL returns 200 with XML, or gzip-compressed XML, and not an HTML page with a 200 status.
  3. The file parses as XML, with characters such as & escaped, and its root element uses the namespace http://www.sitemaps.org/schemas/sitemap/0.9.
  4. No single file lists more than 50,000 URLs or exceeds 50 MB uncompressed. Split anything larger and list the parts in a sitemap index.
  5. Every <loc> is absolute, uses the same host (www or not) and scheme as your canonical URLs, and has no #fragment.
  6. Each URL appears once across all your sitemap files.
  7. lastmod uses W3C Datetime, is never in the future, and changes only when the page's content changes, not every time the file is rebuilt.
  8. A sample of listed URLs returns 200 directly. Replace redirecting URLs with their destinations, and remove URLs that return 404 or 410.

Guidance from Google: a site of about 500 pages or fewer with good internal links may not need a sitemap. A sitemap helps discovery but doesn't guarantee that every URL is crawled or indexed (Google: Learn about sitemaps).

Limits

  • We sampled. From each sitemap index we fetched at most 5 child sitemaps, chosen by hash, with at most 8 files and depth 3 per origin, and we used at most 20 Sitemap: lines per robots.txt. 248,513 index children went unfetched, and 1,296 Sitemap: lines on 20 origins went unused. File-level issue counts are floors. The redirect and 4xx/5xx counts depend on which URLs were sampled and could go either way.
  • Status checks are a sample too: at most 20 URLs per origin, a single request each, with redirects counted but not followed.
  • "No sitemap" means none where we looked: robots.txt lines, /sitemap.xml and /sitemap_index.xml. A site may submit a sitemap at another address directly to search engines.
  • Blocked is not the same as missing. We didn't try to get past 403, 429 or challenge pages. We counted a response as a challenge if it carried Cloudflare's cf-mitigated header or matched a challenge signature. On a 200 or 503 response the signature had to be strict: a challenge-page title such as "Just a moment" or "Access denied", or a known challenge-vendor element. A 404 or 410 never counted as a challenge, whatever its page said. On other error responses, words such as "captcha" or "access denied" were enough. Challenges on 200 responses decided 2 of the 165 blocked origins; the rest came from 403, 429 or other error responses. Detection can miss a challenge or mistake an ordinary page for one.
  • "No website" can hide a working site. When robots.txt didn't answer, we didn't request the homepage, so some of the 84 origins we counted as "no website" may serve pages to browsers.
  • Our lastmod check is stricter than W3C Datetime. We accept a full date, or a timestamp with seconds and a time zone. W3C Datetime also allows year-only, year-month and minute-precision values, which we would flag.
  • The "another host" check compares each <loc> hostname with the listed origin's hostname. It treats www and the bare domain as different hosts and doesn't look for cross-site submission.
  • The list is CrUX's view of popularity. It reflects Chrome users who opted in to sharing usage statistics. It counts each origin separately, so a site's www host and its subdomains are separate entries. Origins within the top-1k bucket have no order.
  • Live responses change. This is one crawl from one network location. Our first count treated some 404 pages that carried a Cloudflare script as blocked, so we fixed the rule and fetched the 191 origins it had called blocked again, about an hour after the crawl: 165 were still blocked, 23 had no sitemap at the places we checked, 2 gave no website and 1 timed out. Some of those moves may be the sites changing rather than the fix. The other 809 origins keep their first-crawl result. We ran a full repeat crawl against the same list less than two hours earlier, with the same settings but before a fix to how we count failed status checks. In it, 106 of the 1,000 origins landed in a different outcome, though no outcome's total moved by more than 8.

Full method and how to reproduce it

  • List: Chrome UX Report (CrUX) global snapshot for August 2026, file data/global/202608.csv.gz from the crux-top-lists repository (1,000,000 rows, SHA-256 4ae541994a7e28f2a9ddd33f711effde9346c4128c506782fa77c268d8f3905b). We kept the 1,000 origins in the top-1k bucket (rank field = 1000). The extracted list's SHA-256 is b8560b2372a7f72f412e1ab60dfa61482aa7d66299a90f69e4fb39e15d821f5a. Row order within the bucket isn't a popularity order.
  • Origins: each measured exactly as listed (scheme and host), with no switch between www and the bare domain. 1 listed origin uses plain HTTP. robots.txt came first on each host. We sent 1 request per second per host, ran up to 20 origins in parallel, used a 15-second timeout, and made at most two requests to any one URL per origin. No page rendering and no logins.
  • Discovery: robots.txt Sitemap: lines (up to 20 per origin), then /sitemap.xml and /sitemap_index.xml. We checked redirect targets against robots.txt too. Files were capped at 50 MiB uncompressed, and 0 origins reached the cap.
  • Parsing: We detected gzip from the file's first bytes (1f 8b), not from its headers. We parsed XML with saxes 6.0.0 and robots.txt with robots-parser 3.0.1. We parsed XML before checking for HTML or challenge pages, so a file with a valid sitemap root element counted whatever its content type. We counted a 200 HTML response at a sitemap address as a soft 404 unless it matched a strict challenge signature.
  • Status checks: HEAD over HTTP/1.1 with curl, and GET only after a 405 or 501, with the GET capped at 64 KiB. A curl failure counts as a request error even if a status line had already arrived.
  • Outcome order: We classify a found sitemap by issues first, then by whether it could be fully checked. Without a sitemap, the order is: no website, robots.txt disallow, blocked, fetch error, no sitemap.
  • Checks: 36 automated tests passed, including regression tests for requests that stall after the headers arrive. We also recorded the deciding response for 10 origins chosen by a fixed seed (spot-v7).

About the data

  • Source: "Chrome UX Report" (CrUX) by Google, licensed under CC BY 4.0. Changes: we took the August 2026 global snapshot, kept only the origins in its top-1k popularity bucket, and measured those origins ourselves. Every sitemap number on this page is our own measurement, not CrUX data. Google has not reviewed or endorsed this study.
  • Where the list came from: the dated file data/global/202608.csv.gz in the crux-top-lists repository on GitHub, which publishes snapshots of CrUX's public BigQuery data. It was the latest dated global snapshot there on 2026-10-01. The SHA-256 digests are under Full method.
  • What the list measures: CrUX groups origins into popularity buckets (top 1,000, top 10,000 and so on), based on Chrome users who opted in to syncing their browsing history and sharing usage statistics. It works at the origin level, and origins within a bucket have no order.
  • Run: 2026-10-01T22:15:40.212Z to 2026-10-01T22:47:14.556Z (UTC). The re-fetch of the 191 origins first counted as blocked finished by 2026-10-01T23:48Z.
  • User agent: lintlab-sitemap-study/1.0 (+https://lintlab.dev). This study's crawler identified itself with this user agent. Questions: [email protected]. To keep your site out of future runs, email the same address.
  • robots.txt respected on every host, following RFC 9309. No challenge was bypassed.
  • Volume: 1,677 sitemap files, 13,894,476 URL entries and 7,456 status checks.
  • We publish aggregates only and name no site.

Run these checks on your own sitemap

Several of these checks also run in lintlab's XML Sitemap Checker, Validator & URL Extractor on Apify. Give it a site or sitemap URL. It finds sitemaps through robots.txt, follows sitemap indexes and gzip files, and returns every URL as a JSON record with its lastmod and any issues it found: duplicate URLs, URLs on another host, invalid or future lastmod values, malformed XML, and files over the 50,000-URL or 50 MB limits. Turn on status checks and it also flags URLs that redirect or don't return 200. Four checks in this study aren't in the Actor: the exact namespace URI, scheme mismatch, #fragments and one lastmod on every URL. Pricing is pay per event: $0.001 per sitemap file parsed, $0.0003 per URL extracted and $0.0005 per URL status check, less on paid Apify plans, with Apify platform usage included. On the Free plan, extracting 10,000 URLs costs $3.00 plus $0.001 per file.

Check your sitemap on Apify