How to Create a Robots.txt File: Complete Step-by-Step Guide
What Is a Robots.txt File?
A robots.txt file is a plain text file at the root of your website that tells search engine crawlers which pages to access and which to skip. It is the first file Googlebot reads when it visits your site. The file lives at yourdomain.com/robots.txt and follows a simple syntax that any website owner can understand and implement.
According to Google, robots.txt files help manage crawl budget. For large sites with thousands of pages, a proper robots.txt file ensures crawlers focus on your most important content.
Why Robots.txt Matters for SEO
Search engines have a crawl budget. If crawlers waste time on low-value pages (admin panels, login pages, internal search results, duplicate content), your important pages may not get crawled as often. A proper robots.txt file prevents crawl budget waste, protects sensitive areas, and guides crawlers toward your highest-priority pages.
Crawl budget optimization is especially important for:
- E-commerce sites with thousands of product pages
- News sites that publish frequently
- Large content sites with archives
- Sites with faceted navigation generating duplicate URLs
Robots.txt Syntax Explained
User-agent
Specifies which crawler the rules apply to. Use * for all crawlers, or specify a particular crawler like Googlebot, Bingbot, or Slurp.
Example: User-agent: * (applies to all crawlers)
Example: User-agent: Googlebot (applies only to Google)
Disallow
Tells crawlers which URLs they cannot access. The path is relative to the root domain.
Example: Disallow: /admin/ (blocks everything under /admin/)
Allow
Overrides a Disallow rule for a specific path. Useful when you want to block a directory but allow a specific page within it.
Sitemap
Points crawlers to your XML sitemap. This helps search engines discover your pages efficiently.
Example: Sitemap: https://example.com/sitemap.xml
Common Robots.txt Mistakes
Mistake 1: Blocking CSS and JavaScript
Google needs to render your pages to understand them. Blocking CSS and JavaScript files prevents Google from properly rendering your pages, which can hurt rankings. Always allow CSS, JS, and image directories.
Mistake 2: Using robots.txt to Remove Pages from Indexing
Robots.txt prevents crawling, not indexing. If a page is linked from other sites, Google may still index it even if robots.txt blocks crawling. Use noindex meta tags or X-Robots-Tag headers for pages you want to keep out of the index.
Mistake 3: Forgetting the Sitemap Directive
Always include Sitemap in your robots.txt. This helps search engines discover your sitemap and crawl your pages more efficiently.
Mistake 4: Blocking the Wrong Directories
Be careful with Disallow rules. Blocking /images/ prevents Google from understanding your images. Blocking /wp-admin/ is fine, but blocking /wp-includes/ can break your site.
Step-by-Step: Create Your Robots.txt File
Step 1: Plan Your Rules. Decide which areas should be crawled and which should not. List all directories and pages that should be blocked.
Step 2: Write the File. Start with User-agent, add Disallow and Allow directives, end with Sitemap. Use clear, logical paths.
Step 3: Test Your Rules. Use Google Search Console Robots.txt Tester to verify your rules work correctly. The tool shows which pages are blocked and which are allowed.
Step 4: Upload to Root Directory. Place at yourdomain.com/robots.txt. Ensure the file is accessible and not blocked by any other rules.
Step 5: Submit to Search Engines. Submit your robots.txt URL in Google Search Console. While Google automatically checks for robots.txt, submitting ensures it finds the file quickly.
Example Robots.txt Configurations
Basic Blog Configuration
For a standard WordPress blog, block /wp-admin/ (allow admin-ajax.php), allow all other paths, and include your sitemap URL. This ensures crawlers can access all public content while skipping admin areas.
E-commerce Site Configuration
For an online store, block /cart/, /checkout/, /account/, and search result pages. Allow /products/ and all public content. Include your sitemap. This focuses crawl budget on product pages that drive revenue.
Large Content Site Configuration
For a large content site, block tag pages, author archives, and internal search results. Allow /blog/ and all primary content paths. Include your sitemap. This concentrates crawling on your highest-value content.
Generate Your Robots.txt
Our free Robots.txt Generator creates proper robots.txt files with crawl directives, disallow rules, and sitemap reference. Choose your site type, customize the rules, and download a ready-to-use file.
Sources
- Google: Robots.txt Documentation — Official robots.txt specification and guidelines
- Moz: Robots.txt Guide — Comprehensive robots.txt implementation guide
- robotstxt.org: Robots.txt Standard — Official robots.txt protocol specification
Try Our Free Tools
Put what you learned into practice with these free SEO tools: