Free robots.txt Tester, Checker & Validator

A free robots.txt checker, tester and validator in one: fetch any site's robots.txt, validate every rule and sitemap line, and test whether a specific URL is allowed or blocked for Googlebot, Bingbot, or any other user-agent. It replaces the robots.txt Tester Google removed from Search Console.

Start your 7-day trial — no credit card, free plan after.

Test a robots.txt file

Enter a site to fetch its robots.txt. Optionally add a path and user-agent to test whether that path is allowed or blocked.

Free check. No signup. Results are not published or indexed.

How NorthDuty crawlability monitoring works

Enter a site to fetch its robots.txt. Optionally add a path and user-agent to test whether that path is allowed or blocked. Recurring checks are configured inside the NorthDuty app.

What this robots.txt checker reports

A single check parses the live robots.txt and tells you exactly how crawlers will treat your URLs.

Allow / disallow rules

Parses every user-agent group and counts the Allow and Disallow rules so you can see what's blocked at a glance.

Per-path testing

Enter a path and user-agent to get an allowed-or-disallowed verdict using longest-match with Allow-wins-on-tie, the behaviour major crawlers use.

Sitemap directives

Lists any Sitemap: lines so you can confirm search engines are being pointed at the right sitemaps.

Missing or empty files

Tells you when no robots.txt exists — in which case everything is crawlable by default.

Why test robots.txt

A single stray Disallow can deindex an entire section of a site, and it's easy to ship one by accident.

How the robots.txt tester works

No signup — enter a site and optionally a path to test.

1

Enter a site

Provide a domain or URL. NorthDuty fetches robots.txt from that site's root.

2

We parse the rules

The file is parsed into user-agent groups with their Allow and Disallow rules and any sitemap directives.

3

Optionally test a path

Add a path and user-agent to get an allowed-or-blocked verdict, with the matching rule shown.

4

Catch regressions automatically

NorthDuty's health checks cover SEO fundamentals including robots.txt, so an accidental block gets flagged on a schedule.

Google's robots.txt Tester is gone. What to use instead

Google retired the legacy robots.txt Tester from Search Console in late 2023. Its replacement, the robots.txt report, shows which robots.txt files Google found for your site, when they were last crawled, and any fetch or parse errors. What it no longer does is let you paste a URL and see whether a rule blocks it.

That per-URL check is what this tester does. Enter your site, a path, and a user-agent, and it applies the same matching Google documents: the most specific user-agent group wins, then the longest matching rule, and Allow wins a tie. Use it before you ship a robots.txt change, then use URL Inspection in Search Console to confirm how Google treats a specific live page.

robots.txt syntax cheat sheet

Every directive and pattern you are likely to need, and how Google treats it.

Directive or patternWhat it doesExample
User-agentStarts a group of rules for one crawler. * matches any crawler that has no group of its own.User-agent: Googlebot
DisallowBlocks crawling of URLs whose path starts with the value. An empty value blocks nothing.Disallow: /wp-admin/
AllowRe-opens a path inside a blocked section. The longer (more specific) matching rule wins.Allow: /wp-admin/admin-ajax.php
SitemapPoints crawlers at a sitemap. Must be a full URL and can appear anywhere in the file.Sitemap: https://example.com/sitemap.xml
* wildcardMatches any sequence of characters inside a path.Disallow: /*?add-to-cart=
$ end anchorMatches only when the URL ends exactly there.Disallow: /*.pdf$
Crawl-delayIgnored by Google. Some other crawlers, such as Bingbot, respect it.Crawl-delay: 10
NoindexNot supported in robots.txt. Google stopped honouring it in 2019; use a meta robots tag or X-Robots-Tag header.(use <meta name="robots" content="noindex">)

How rules match: worked examples

Run any of these through the tester above with your own paths to see the verdict and the rule that decided it.

RulesURL testedResult
Disallow: /wp-admin/ + Allow: /wp-admin/admin-ajax.php/wp-admin/admin-ajax.phpAllowed: the Allow rule is longer, so it wins
Disallow: /search/search-results/shoes/Blocked: rules are prefix matches, so /search also matches /search-results/
Disallow: /*.pdf$/files/guide.pdf?v=2Allowed: $ requires the URL to end in .pdf, and this one ends in ?v=2
Disallow: /Checkout//checkout/Allowed: paths are case-sensitive
User-agent: * with Disallow: / and User-agent: Googlebot with Allow: // as GooglebotAllowed: Googlebot obeys only its own group and ignores the * group
Disallow: (empty)Any URLAllowed: an empty Disallow blocks nothing

Common robots.txt mistakes that cost traffic

Most robots.txt damage is accidental and ships with a deploy or a migration.

Staging rules shipped to production

A staging site's User-agent: * / Disallow: / gets copied to the live site during a launch or migration, and search engines stop crawling everything. This is the first thing to check after any relaunch.

Using robots.txt to hide pages from Google

Blocking a URL stops crawling, not indexing. A blocked page can still appear in results if other pages link to it. To keep a page out of the index, allow crawling and add noindex.

Blocking CSS and JavaScript

If Google cannot fetch your theme's CSS and JS, it cannot render the page properly, which can hurt how the page is understood and ranked. Leave asset folders crawlable.

A robots.txt that errors

If robots.txt returns a 5xx server error, Google temporarily treats the whole site as disallowed. A 404 is treated as no rules at all. Monitor the file, not just the homepage.

Wrong location or host

robots.txt only works at the root of each host and protocol. https://www.example.com/robots.txt does not cover https://shop.example.com/. Each subdomain needs its own file.

Over-blocking WooCommerce URLs

Blocking cart and add-to-cart parameters is fine, but a broad rule like Disallow: /*? can also block filtered category pages and paginated product lists you want crawled.

WordPress and WooCommerce robots.txt

If there is no robots.txt file on the server, WordPress serves a virtual one: it disallows /wp-admin/, allows /wp-admin/admin-ajax.php, and since WordPress 5.5 adds a Sitemap line pointing at /wp-sitemap.xml. SEO plugins such as Yoast and Rank Math let you edit it from the dashboard, and a physical robots.txt file in the site root overrides the virtual one.

For WooCommerce stores, the pages worth keeping crawlable are products, product categories, and your main content. Cart, checkout, and account pages rarely need crawling, but they are better handled with noindex than with robots.txt, so Google can still see the tag. After any plugin or theme change, re-test a product URL and a category URL above.

Go Beyond One-Off Checks

Use the tool preview for a quick answer, then move into recurring monitoring for your most important pages and journeys.

Frequently Asked Questions

Answers about this diagnostic preview and when to move into recurring monitoring.

Is this a robots.txt tester, checker or validator?

All three, and the difference is mostly wording. It validates the file by fetching it and parsing every user-agent group, checks a live site rather than pasted text, and tests a specific path against those rules for the crawler you choose.

Is this robots.txt tester free?

Yes. It's free and requires no signup. Enter a site to fetch and parse its robots.txt, and optionally test a specific path.

How does the path test decide allowed vs disallowed?

It picks the most specific matching user-agent group, then applies the longest matching rule, with Allow winning ties — the same approach major search-engine crawlers use. Wildcards (*) and end-anchors ($) are supported.

What if a site has no robots.txt?

If no robots.txt is found, the tester reports that — and by the standard, everything on the site is crawlable by default.

Does this change my robots.txt?

No. The tester only reads the live file. To change crawling rules you edit robots.txt on your own server.

Does Google Search Console still have a robots.txt tester?

No. Google removed the legacy robots.txt Tester in late 2023 and replaced it with the robots.txt report, which shows fetch status and parse errors but does not test individual URLs. This tool fills that gap: enter a path and user-agent to get an allowed or blocked verdict with the matching rule.

Does blocking a page in robots.txt remove it from Google?

Not reliably. robots.txt controls crawling, not indexing. A blocked URL can still be indexed, usually without a description, if other pages link to it. To remove a page from results, allow crawling and use a noindex meta tag or X-Robots-Tag header.

Where does robots.txt need to be?

At the root of the host, for example https://example.com/robots.txt. It applies only to that exact protocol and host, so subdomains need their own file. Google reads up to 500 KiB of the file and ignores anything beyond that.

Are robots.txt rules case-sensitive?

The paths in Allow and Disallow rules are case-sensitive, so Disallow: /Private/ does not block /private/. Directive names like User-agent and Disallow are not case-sensitive.

Start monitoring your website with NorthDuty today.

A robots.txt mistake can quietly deindex pages for weeks. NorthDuty monitors SEO fundamentals continuously, so an accidental block gets caught fast.

7 days with Pro features and limits, no credit card — then keep one daily journey on the free plan.