These three get used interchangeably and they do completely different jobs. Two of them are about crawling, one is about indexing, and the fourth thing people lump in is about individual links.

Pick the wrong one and you either leave a page in search results you wanted hidden, or you block a page so thoroughly that the instruction to hide it never gets read.

What each one actually controls

Stops crawlingStops indexingNeeds the page crawlable
DisallowYesNoNo
NoindexNoYesYes
Meta nofollowNoNoYes
rel nofollow on a linkNoNoYes

Disallow controls crawling, not indexing

Disallow lives in robots.txt and tells bots not to request a URL.

User-agent: *
Disallow: /cart/

What it does not do is keep the page out of search results. If another page links to that URL, it can still be discovered and indexed, usually with no description because the crawler was never allowed to read it.

So robots.txt is the wrong tool if your goal is keeping something out of search. It is the right tool for saving crawl budget on pages that have no search value and nothing pointing at them, like faceted filter URLs or internal search results.

Noindex controls indexing

Noindex goes in the page itself, either as a meta tag or an HTTP header.

<meta name="robots" content="noindex">

For files that are not HTML, like PDFs, you send it as a header instead.

X-Robots-Tag: noindex

This is what you use when the goal is to keep a page out of search results. It works whether or not anything links to the page.

The mistake that breaks both

Here is the trap, and it catches experienced people.

For a noindex to work, the crawler has to read it. If you also disallow that URL in robots.txt, the crawler never fetches the page, never sees the tag, and the noindex is silently ignored.

The page can then sit in search results indefinitely, and nothing in your setup will tell you why, because on paper you did two things to hide it.

If you want a page out of the index, allow it to be crawled and let the noindex do its job. Add a disallow later, once the page has dropped out.

Nofollow is a different category

Nofollow gets grouped with these two, but it says nothing about the page it sits on.

A meta robots nofollow tells bots not to follow any link on that page. The rel attribute version applies to one specific link.

<a href="https://example.com" rel="nofollow">

There are two siblings worth knowing. Use rel sponsored for paid or affiliate links, and rel ugc for links in comments and forum posts.

On internal links, nofollow is almost always pointless. The old practice of sculpting where authority flows stopped working in 2009, and adding it today mostly just hides part of your own structure from crawlers. We covered when it makes sense in should you nofollow internal links.

Which one to use

Use disallow when

The URL has no search value, nothing links to it, and crawling it wastes budget. Filter combinations, internal search results, and cart or checkout paths on a large store.

Use noindex when

You want the page gone from results, full stop. Thin tag archives, thank-you pages, staging content that leaked, or duplicate paths that cannot be canonicalised.

Use nofollow when

You are labelling a specific link rather than managing a page. Paid placements get sponsored, comment links get ugc, and untrusted outbound links get nofollow.

The part none of this tells you

Every directive above is a request. It is not enforcement.

Well-behaved crawlers honour robots.txt. Others read it and ignore it. And a noindex only works if the bot fetched the page, parsed the head, and respected the tag.

So after you set any of this up, there is a question left over that the directive itself cannot answer. Did it work?

The only place that answer exists is your server logs, because they record what was actually requested rather than what you asked for.

Linkilo crawl log showing which search and AI bots requested each URL on a WordPress site
What was actually requested, rather than what you asked for.

Linkilo’s crawl log reads those logs inside WordPress and records every visit from Googlebot, Bingbot, GPTBot, and ClaudeBot, with reverse-DNS checks separating real bots from anything spoofing them.

That turns each of these directives into something you can verify. Disallow a directory, then check whether requests to it stopped. Add a noindex, then confirm the page is still being crawled, because if it is not, the tag is never being read.

It also matters more than it used to because of the AI crawlers. Blocking GPTBot or ClaudeBot in robots.txt is a request like any other, and the only way to know whether it held is to look at what still arrives.

The honest limit is that this is diagnosis, not enforcement. A crawl log tells you a bot ignored your robots.txt. Stopping it takes a rule at the server or firewall. Linkilo is also WordPress only.

The short version

Disallow stops crawling. Noindex stops indexing. Never use both on the same URL, because the first prevents the second from being read.

Nofollow is about links, not pages, and rarely belongs on internal ones.

Then check your logs, because all three are requests and only the logs tell you which were honoured.

Common questions

Will disallowing a page remove it from Google?+

No, and this is the most common misunderstanding. Disallow stops crawling, not indexing. If another page links to that URL it can still appear in results, usually with no description because the crawler was never allowed to read the page. Use noindex if the goal is removal.

Can I use noindex and disallow together?+

Not on the same URL. The crawler has to fetch the page to see the noindex tag, and a disallow prevents that fetch, so the tag is never read and the page can stay indexed. Allow crawling until the page drops out of results, then disallow it if you want to save crawl budget.

Does noindex work on PDFs and images?+

Yes, but not as a meta tag. Non-HTML files have no head section, so you send the directive as an X-Robots-Tag HTTP header instead. That works for PDFs, images, and any other file type your server delivers.

What is the difference between nofollow, sponsored, and ugc?+

They describe why a link exists. Sponsored marks paid or affiliate links, ugc marks links created by users in comments or forums, and nofollow is the general case for links you do not want to endorse. All three tell search engines not to pass endorsement through that link.

How do I know whether a crawler respected my robots.txt?+

Check your server logs. The directive is a request, not enforcement, so the only record of what actually happened is which URLs were requested and by which user agent. If a bot is still hitting disallowed paths, it is ignoring the file and you need a server-level block.