Skip to main content
SEO·9 min read

Next.js 16 Sitemap, RSS, and Crawl Signals: A Production Setup

How to generate a Next.js 16 sitemap and RSS feed from the same post source, set honest lastmod dates, declare the feed in metadata, and avoid canonical leaks — without treating llms.txt as a ranking factor.

By Mussawar Hayat

Why Crawl Signals Still Matter in 2026

Sitemaps do not rank pages. They tell Google, Bing, and answer engines which URLs exist, when they changed, and which ones are canonical. An RSS feed does the same job for readers, feed readers, and some AI crawlers that prefer a small, dated list over a full HTML crawl. This guide is the production setup used on mussawarhayat.site: App Router sitemap, robots.txt, RSS at /feed.xml, and llms.txt as a citation map — without claiming any of those files are ranking factors.

What You Will Learn

  • What a sitemap is allowed to claim, and what lastmod must not invent
  • How to emit RSS from the same post source as the blog index
  • Where to declare the feed so browsers and crawlers can discover it
  • How llms.txt complements, and does not replace, HTML and schema
  • Canonical pitfalls when the root layout sets alternates.canonical

1. One Source of Truth for URLs

Do not maintain a hand-written sitemap and a separate blog index. On this site, getAllPosts() reads app/blog/content/*.ts. The blog page, the dynamic article route, app/sitemap.ts, and app/feed.xml/route.ts all call that helper. A new article file is enough to enter the index, the XML sitemap, and the RSS feed on the next build.

Portfolio case studies follow the same idea via getAllProjectSlugs(). Static routes (home, services, about, contact, privacy, terms) stay explicit because they are few and rarely change.

2. lastmod Must Be a Real Change Date

Google's sitemap documentation treats lastmod as the last significant update of the page. Setting every URL to new Date() on each build trains crawlers to ignore the field. Use the post date for articles. Use a stable date for legal pages. Reserve "now" for hubs that actually change when content is published, such as the blog index, and even then prefer the newest post date if you have it.

Do not put changefreq or priority expectations into your content strategy. Google has said it largely ignores both. They are harmless metadata, not levers.

3. RSS Without a Second CMS

A Route Handler at app/feed.xml/route.ts returns application/rss+xml. Each item needs a stable permalink guid, an absolute link, a title, a description (the excerpt, not the full HTML, so the feed stays small), and a UTC pubDate. Escape &, <, and quotes. Cache for an hour; the feed does not need to be dynamic per visitor.

Declare discovery in two places:

  • Root metadata alternates.types["application/rss+xml"] so the HTML head emits a feed link
  • robots.txt comment plus the sitemap line — robots is not a feed registry, but operators look there

Also link the feed from the blog index so humans can subscribe. Related reading: SEO for AI Overviews and generative search.

4. robots.txt, llms.txt, and Canonicals

robots.txt should allow public content, disallow /api/, and point at https://mussawarhayat.site/sitemap.xml. Naming AI user-agents and allowing them is a policy choice, not a ranking boost. Blocking GPTBot or Google-Extended is fine if you do not want training use; this site allows them so answer engines can cite the guides.

llms.txt is a plain-language map: who the entity is, canonical URLs, and which posts are worth citing. It is not a substitute for visible HTML, titles, or JSON-LD. If a claim is only in llms.txt, answer engines that only read the page will miss it.

A root alternates.canonical of / can leak onto pages that forget their own canonical, including privacy and terms. Set an explicit canonical on every indexable route. Article routes already use /blog/[slug].

Crawl Signals FAQ

Does an RSS feed improve Google rankings?

No. It helps discovery for subscribers and some crawlers. Rankings still depend on the page itself: intent, links, experience, and technical quality.

Should the feed include full article HTML?

Excerpts are enough for a portfolio blog. Full content increases payload and can diverge from the canonical page if the renderer and the feed escape differently.

Is llms.txt a Google standard?

No. Google documents sitemaps, robots.txt, canonicals, and structured data. llms.txt is an emerging convention for AI crawlers. Keep it accurate and secondary to the HTML.

What should lastmod be for a legal page?

The date the policy text last changed, not the deploy timestamp.

Summary

Generate the sitemap and RSS from the same post modules, put real dates in lastmod, set a canonical on every indexable route, and keep llms.txt as a citation index. That is the crawl setup, not a growth hack.

Need the same pipeline on a client Next.js app? See engineering services or hire Mussawar Hayat. Related: CSP and security headers and multi-site Next.js on Nginx.