Robots.txt and XML sitemaps solve different technical SEO problems.
robots.txt controls which URLs or paths crawlers may request.
An XML sitemap lists URLs you want search engines to discover and understand as important.
Neither file guarantees that a page will be indexed.
On Blogger, avoid enabling custom robots.txt or custom robots header tags merely because a generic SEO tutorial recommends them. Blogger exposes both controls under Settings → Crawlers and indexing, and Blogger's help documentation states that custom robots.txt text is used instead of the platform's default robots.txt content. Change these settings only for a defined crawl or indexing requirement, then verify the live output.
Use robots.txt to manage crawling, meta/X-Robots directives to manage indexing, and sitemaps to expose preferred canonical URLs. Do not mix these functions.
Last verified: September 24, 2026.
User-agent, Allow, Disallow, WordPress examples, URL parameters and broader crawl-control theory, use the general Robots.txt SEO Guide.
robots.txt Is a Crawl Control, Not a Reliable Removal Tool
If your goal is to keep a page out of Google Search, use the correct indexing control rather than assuming a robots.txt block will remove it.
1. robots.txt vs XML Sitemap
| File / Control | Primary Purpose | Does It Guarantee Indexing? |
|---|---|---|
| robots.txt | Controls crawler access to URLs/paths | No |
| XML sitemap | Lists preferred URLs for discovery | No |
| noindex | Requests exclusion from indexing | Index-control directive, not sitemap |
| rel=canonical | Signals preferred representative URL | No |
2. What robots.txt Actually Does
A robots.txt file tells compliant crawlers which URLs or paths they should not crawl.
A basic example:
User-agent: *
Disallow: /private-section/
This tells crawlers covered by the rule not to request URLs under that path.
It does not mean those URLs are automatically removed from search results.
3. Do Not Use robots.txt to Remove Pages From Google
Google's current guidance explicitly says robots.txt should not be used as the primary way to prevent a page from appearing in Search.
A blocked URL can still be known through:
- External links.
- Internal links.
- Historical crawling.
- Sitemaps.
If Google cannot crawl the page, it may also be unable to see page-level indexing directives.
4. noindex Must Usually Be Crawlable
If you want Google to process:
<meta name="robots" content="noindex">
Google generally needs to fetch the page.
Blocking the same URL in robots.txt can prevent Google from seeing the directive.
For the difference between noindex, canonicals and indexing diagnosis, use our Canonical URLs & Indexing Guide.
5. Blogger Has Built-In Crawler Controls
Blogger currently exposes crawler/indexing settings under:
Settings
→ Crawlers and indexing
Relevant settings include:
- Custom robots.txt.
- Custom robots header tags.
- Google Search Console access.
Do not enable custom crawler settings merely because an SEO tutorial says every site needs a custom robots.txt.
6. Default Blogger robots.txt vs Custom robots.txt
For many Blogger sites, the default crawler configuration is sufficient.
A custom robots.txt should be used only when you understand exactly what the new rule changes.
Common risks of unnecessary customization include:
- Blocking posts.
- Blocking pages needed for rendering.
- Breaking sitemap discovery.
- Creating conflicts with noindex behavior.
- Copying rules intended for WordPress or another CMS.
7. Do Not Copy Generic WordPress robots.txt Rules Into Blogger
Blogger does not have WordPress directories such as:
/wp-admin/
/wp-includes/
/wp-content/
Rules written for another platform can be irrelevant or harmful on Blogger.
8. XML Sitemaps: What They Do
A sitemap tells search engines where important URLs are located.
It can also include metadata such as last modification information where supported.
Google describes sitemaps as especially useful for helping search engines discover important parts of a site.
A sitemap is a discovery signal—not an indexing guarantee.
9. Blogger Sitemap Behavior
Blogger automatically provides sitemap functionality rather than requiring you to install an XML-sitemap plugin.
For a normal Blogger site, check the generated sitemap first before creating alternate feed-based submissions.
A Blogger site commonly exposes its generated sitemap from the site root. Depending on the site and platform output, you may also encounter a separate page sitemap.
https://www.example.com/sitemap.xml
https://www.example.com/sitemap-pages.xml
Do not assume both endpoints exist on every setup. Open the live files and submit only valid sitemap URLs that actually resolve for your blog.
10. Submit the Sitemap in Search Console
In Google Search Console:
Search Console
→ Sitemaps
→ Enter sitemap path
→ Submit
Submitting means telling Google where the sitemap is located. You are not uploading the sitemap file into Search Console.
11. What URLs Should Be in a Sitemap?
A sitemap should primarily list URLs you want indexed.
Prefer URLs that are:
- Canonical.
- Indexable.
- Returning HTTP 200.
- Useful and intentional.
Avoid intentionally listing:
- Redirecting URLs.
- 404 pages.
- noindex pages.
- Duplicate parameter variants.
- Old migration URLs.
12. Sitemap and Canonical Should Agree
If your sitemap contains URL A but the page canonical points to URL B, you are sending conflicting signals.
A clean setup looks like:
Sitemap URL
=
Canonical URL
=
Preferred internal-link URL
13. Sitemap Does Not Fix Weak Internal Linking
An XML sitemap can help discovery, but important pages should also have meaningful crawlable internal links.
If a page appears only in the sitemap and nowhere in site navigation or contextual links, treat that as an internal-architecture problem.
14. "Sitemap Could Not Be Read" or "Couldn't Fetch"
Check:
- Does the sitemap URL open publicly?
- Does it return HTTP 200?
- Is the property/domain correct?
- Is the sitemap blocked by robots.txt?
- Is the file valid XML or a supported sitemap format?
- Did you enter the correct path?
Do not fix this by submitting random feed URLs without first diagnosing the actual sitemap response.
15. Sitemap Contains URLs Blocked by robots.txt
Search Console can report when URLs listed in the sitemap are blocked from crawling.
This is a conflicting signal:
Sitemap:
"Please discover this URL."
robots.txt:
"Do not crawl this URL."
Review whether the URL should actually be crawlable and indexable.
16. Search Console Sitemap Status Does Not Equal Indexing Status
A successfully processed sitemap only means Google could process the sitemap.
It does not mean every URL was indexed.
For individual URLs, use URL Inspection and the Page Indexing report.
17. Do Not Block CSS and JavaScript Without a Reason
Google renders pages to understand their content and layout.
Blocking essential CSS or JavaScript can make that rendering less representative.
Do not add broad rules such as:
Disallow: /*.js
Disallow: /*.css
unless you have a very specific, tested reason.
18. Blogger Custom robots.txt Example: Use Only When Needed
If you have a real reason to replace Blogger's default file, keep the custom file minimal and test the public output immediately.
A permissive baseline looks like:
User-agent: *
Disallow:
Sitemap: https://www.example.com/sitemap.xml
An empty Disallow: means that rule is not blocking the crawler.
Do not treat this as a recommended template for every Blogger site. Enabling custom robots.txt replaces the default robots.txt content, so first inspect the live /robots.txt and make sure you understand which Blogger-generated rules you would be overriding.
19. Test the Live robots.txt File
Open:
https://www.example.com/robots.txt
Confirm the rules being served publicly match what you intended.
Do not assume the Blogger dashboard setting and the public output are the same until you verify them.
20. robots.txt Is Public
Do not put secrets in robots.txt.
Anyone can read the file.
A rule like:
Disallow: /secret-admin-backup/
does not secure that location.
Access control and authentication must protect sensitive resources.
21. Search Description, Robots and Structured Data Are Different
Blogger exposes several SEO-related settings, but they do different things.
- Search description: descriptive metadata.
- robots.txt: crawl control.
- robots header tags: indexing/snippet directives.
- JSON-LD: structured description of content/entities.
For structured-data implementation, use our Blogger JSON-LD Schema Guide.
22. Custom Domain Changes and Sitemaps
If a Blogger site moves from one hostname to another, make sure Search Console and sitemap references use the intended live hostname.
DNS and nameserver changes belong to a separate layer.
For custom-domain DNS troubleshooting, use our DNS Configuration Best Practices Guide.
23. Indexing Problems Are Not Always Sitemap Problems
If a URL is discovered but not indexed, possible causes include:
- Duplicate/canonical selection.
- noindex.
- Weak/thin content.
- Low internal-link importance.
- Redirects.
- Soft 404 behavior.
- Crawl/indexing quality decisions.
Do not keep resubmitting the sitemap as the only troubleshooting step.
24. Sitemap Resubmission After Every Edit Is Unnecessary
You generally do not need to manually resubmit the sitemap after every article update.
Keep the sitemap accessible and accurate.
Use URL Inspection for important individual URLs after a significant fix when you want to request another crawl. A request does not guarantee when Google will recrawl or reprocess the URL.
25. Large Sites and Sitemap Indexes
Large sites can use sitemap index files to reference multiple sitemap files.
Blogger manages much of this platform behavior automatically, so do not create a manual sitemap architecture unless the platform setup genuinely requires it.
26. robots.txt and Other Crawlers
Different search, AI, archival and service crawlers may identify themselves with different user-agent tokens and may not interpret every robots.txt rule exactly the same way.
Do not block or allow crawlers based on copied lists you do not understand. Check the crawler operator's current documentation before creating crawler-specific groups.
Before changing crawler access, decide:
- Which crawler is involved.
- What content you want it to access.
- What trade-off the change creates.
27. Blogger Search-Label and Archive Controls
Blogger can generate search and archive-style URL spaces that are different from normal post and page URLs.
Examples can include:
/search/label/Web%20Hosting
/search?q=example
Before changing crawl or indexing controls, decide what you actually want:
- Reduce crawling: use crawl controls only when there is a demonstrated need.
- Keep search/archive pages out of Google: use Blogger's custom robots header-tag controls where appropriate so Google can see the indexing directive.
- Preserve reader navigation: do not remove useful label pages merely because they are not intended as primary search landing pages.
Blogger exposes separate settings for Home page tags, Archive and search page tags, and Post and page tags. Use those deliberately instead of treating one robots.txt rule as the solution to every archive/indexing issue.
28. Blogger Mobile ?m=1 URLs
Blogger can expose mobile-style URL variants such as:
https://www.example.com/2026/08/example-post.html?m=1
Do not treat those variants as separate preferred URLs. Your canonical, sitemap and internal-link signals should continue to point to the normal canonical post URL unless there is a very specific platform reason not to.
If Search Console reports mobile variants, diagnose canonicalization rather than adding random robots.txt blocks that might create new crawl/indexing conflicts.
29. Blogger robots.txt Checklist
- Check the live
/robots.txt. - Keep custom robots.txt disabled unless there is a specific need.
- Do not use robots.txt as noindex.
- Do not block pages whose noindex directive Google must see.
- Do not block required CSS/JS broadly.
- Do not copy WordPress rules.
- Do not expose secrets and assume Disallow protects them.
30. Blogger Sitemap Checklist
- Open the live sitemap.
- Submit the correct sitemap path in Search Console.
- Check sitemap status.
- Keep preferred canonical URLs in the sitemap.
- Remove or avoid old/redirecting/non-indexable URLs where you control the sitemap.
- Investigate "couldn't fetch" from the actual HTTP/XML response.
- Use URL Inspection for specific indexing problems.
31. Common Blogger Robots and Sitemap Mistakes
- Blocking a URL and also trying to noindex it.
- Submitting non-canonical duplicates in a sitemap.
- Assuming a sitemap guarantees indexing.
- Replacing Blogger defaults with copied custom robots rules.
- Blocking rendering resources.
- Submitting random feed URLs to fix a sitemap error.
- Resubmitting the sitemap repeatedly instead of diagnosing the page.
- Using robots.txt to protect private data.
32. Final Blogger Crawl & Sitemap Decision Framework
Want crawler to fetch page?
↓
Allow crawl
Want page excluded from Search?
↓
Use index-control directive, not robots.txt alone
Want Google to discover preferred URLs?
↓
Use accurate sitemap + internal links
Sitemap error?
↓
Check HTTP response → XML → robots → property/path
Page discovered but not indexed?
↓
Move to canonical/indexing diagnosis
For Blogger, simplicity is usually safer: keep platform-managed crawl behavior unless you have a clearly defined problem that requires a custom rule.
If the issue is duplicate URLs, Google-selected canonicals, ?m=1 consolidation or indexability rather than crawler access, move the investigation to the Canonical URLs & Indexing Guide instead of adding more robots.txt rules.
robots.txt does not directly improve rankings. Its role is crawl control; ranking and indexing depend on broader content, technical and quality signals.
Frequently Asked Questions
Does robots.txt remove a page from Google?
No. robots.txt controls crawling. Google recommends using other indexing controls when you want a page excluded from Search.
Does a sitemap guarantee indexing?
No. A sitemap helps Google discover and understand preferred URLs, but indexing is not guaranteed.
Should I use custom robots.txt on Blogger?
Only when you have a specific Blogger-specific reason and understand the effect of each rule. Blogger provides built-in crawler/indexing settings, including custom robots.txt and custom robots header tags. For many blogs, leaving the platform defaults unchanged is safer than copying a generic template.
What sitemap should Blogger use?
Blogger provides generated sitemap behavior. Verify the live /sitemap.xml and any platform-provided page sitemap before creating alternate submissions.
Why does Search Console say a sitemap couldn't be fetched?
Check whether the sitemap URL opens publicly, returns a successful response, uses a supported format, belongs to the correct property and is not blocked.
Should noindex pages be in a sitemap?
Prefer listing URLs you actually want indexed. Intentionally including noindex URLs sends conflicting signals.
Can I block CSS and JavaScript in robots.txt?
You can create such rules, but broadly blocking resources Google needs to render the page can harm its understanding of the page. Do not block them without a specific reason.
Should I resubmit my sitemap after every new post?
Usually no. Keep the sitemap accessible and accurate. Use Search Console and URL Inspection for important troubleshooting or significant updates.
Can robots.txt protect private content?
No. robots.txt is public and is not an access-control mechanism. Protect private content with authentication or server-side access controls.
Abdul Shakoor
Founder of Digital Bhatti, an independent technical publication focused on web hosting and infrastructure, WordPress, technical SEO, web performance and automation.