Free SEO tool
XML Sitemap Generator
Build a draft XML sitemap from links found in a public website's HTML.
Checked 0 URLs. Found 0 indexable HTML pages. Queue: 0.
About this tool
An XML sitemap lists URLs that you want crawlers to consider. It is especially useful when a site has many important pages or when internal linking does not make every page easy to find. A sitemap is a discovery aid, so its entries should point to canonical, indexable pages that you intend to keep public.
Use this generator when you need a small site's first sitemap, want a quick independent list of linked pages, or need to inspect the output of a lightweight crawl. Enter a public homepage, set a sensible page cap and watch the crawl progress. The tool follows links in returned HTML on the same host, checks the basic wildcard robots.txt disallow rules, and leaves out pages that declare noindex or a different canonical URL.
Review the XML before you publish it. Check that key pages are present, remove URLs that do not belong, and compare the result with the site's real content inventory. Links produced only by JavaScript and orphan pages may be absent. If your site already generates a maintained sitemap, use that as the source of truth and treat this download as a diagnostic draft.
How to use it
- Enter the site URL and choose a page limit.
- Start the crawl and watch the discovered page count.
- Review, copy or download the XML before submitting it.
What it checks
- Same-host HTML links
- Robots.txt Disallow rules for the wildcard user agent
- Noindex and canonical signals in returned HTML
Limits
The crawl stops at 500 URLs. JavaScript-rendered links are not discovered. Robots rules, redirects and canonicals can be more complex than this lightweight crawler can resolve. Review the output before publishing.
Frequently asked questions
Will it find every page?
No. It follows links in returned HTML only and cannot discover orphan pages or links added by JavaScript.
Does it respect robots.txt?
It applies wildcard Disallow rules found in robots.txt before fetching pages. Review complex rule sets separately.
Where does lastmod come from?
A valid Last-Modified response header is used where available. Otherwise the entry has no lastmod.
Can I stop a crawl?
Yes. Stop prevents the queue from starting more pages and preserves results already collected.
For a full crawl and indexability review, see my technical SEO audit.