Technical SEO

How Do I Choose Between robots.txt, noindex and Canonical Tags on a Medical Website?

robots.txt vs noindex vs canonical on a clinic site: pick the right directive per page, avoid conflicts and check your own URLs with a free script.

You have a page you do not want in Google, and three tools that all sound like the answer. A developer says “just block it in robots.txt.” An SEO plugin offers a noindex checkbox. Someone else mentions a canonical tag. You pick one, and a month later the page is still showing up.

You are not missing a setting. These three tools solve three different problems, and using the wrong one for the job does nothing or makes things worse. Clinics hit this constantly because a medical site carries every awkward page type at once: booking confirmations, provider PDFs, location duplicates, directory filters and a portal login.

I use one robots.txt vs noindex vs canonical decision tree on every healthcare audit. In the next ten minutes you will learn which directive fits which clinic page, which combinations cancel each other out, and how to check your own site for conflicts with a script you can run today.

Directive decision tree for a medical website (diagram for: robots.txt vs noindex vs Canonical on a Medical Website)

robots.txt vs noindex vs Canonical: Which Directive Do I Use?

Use noindex to remove a page from Google, robots.txt Disallow to stop wasted crawling, and rel="canonical" to merge duplicates. Use authentication for anything private.

Google confirms the first rule in its documentation. For noindex to work, the page must not be blocked by robots.txt, as the Block Search indexing with noindex guide states. I explain the mechanism behind that in Crawling vs. Indexing vs. Ranking: Why a Blocked Clinic Page Still Shows in Google.

What Does Each Directive Actually Do?

Each directive acts on a different stage and carries a different strength. robots.txt and noindex are rules Google follows. rel="canonical" is a hint Google can override.

ToolStage it acts onStrengthWhat it cannot do
robots.txt DisallowCrawlingRuleRemove a URL from Google
noindex (meta tag or X-Robots-Tag)IndexingRuleWork on a page Googlebot cannot fetch
rel="canonical"Indexing (duplicate sorting)HintForce Google to pick your URL
AuthenticationAccessRuleNothing SEO-related, and that is the point
Two panels comparing directives Google follows with the canonical hint Google can override
Rules versus hints. Original diagram by Atiur Rahman.

Is a Canonical Tag a Command?

No. Google’s canonicalization guide calls a canonical preference “a hint, not a rule.” Google weighs your tag against other signals: HTTPS versus HTTP, redirects, whether the URL sits in your sitemap, and which page is the most complete and useful. When your tag disagrees with those signals, Google picks its own canonical, and Search Console reports the mismatch. I cover that diagnosis in Why Did Google Pick the Wrong Canonical URL for My Clinic Locations?.

How Long Does Google Take to See a robots.txt Change?

Up to 24 hours, usually. Google’s robots.txt documentation says it generally caches the file for up to 24 hours, and longer when it cannot refresh the copy. A rule you add at 9 a.m. does not act at 9:01.

Which Directive Fits Each Page on a Clinic Website?

Match the page type to the goal, then use the matching tool. This table covers the ten page types I see on medical sites.

Clinic pageGoalToolNotes
Appointment “thank you” pageKeep out of searchnoindexPage stays crawlable
Staging or test copy of the siteKeep out of search and privateAuthentication, plus noindexDisallow alone leaves URLs listable
Patient portal loginPrivateAuthenticationA public login page can carry noindex
Provider bio PDF or patient form PDFKeep out of searchX-Robots-Tag: noindexPDFs have no HTML head
Two near-identical location pagesMergerel="canonical"Better: write unique content per location
?utm_source= tracking URLsMergerel="canonical"Points to the clean URL
Directory filter combinationsStop crawl wasteDisallow plus a curated set of indexable pagesCovered in article #7
Internal site search results (/?s=)Stop crawl wasteDisallowLow value, endless variations
Appointment calendar ?date= URLsStop crawl wasteDisallowEndless date trail
Retired provider or service pagePage is gone301 to a close match, or 404/410None of the three tools applies

How Do I Write Each Directive?

The syntax for each tool is short. Copy the pattern that matches your goal.

noindex in the page head:

<meta name="robots" content="noindex">

noindex for a PDF, through the server (Apache):

<FilesMatch ".pdf$">
  Header set X-Robots-Tag "noindex"
</FilesMatch>

noindex for a PDF (Nginx):

location ~* .pdf$ {
    add_header X-Robots-Tag "noindex";
}

Canonical on a parameter or duplicate page:

<link rel="canonical" href="https://www.yourclinic.com/locations/downtown/">

Crawl control in robots.txt:

User-agent: *
Disallow: /?s=
Disallow: /book/*?date=

Google documents the * wildcard and the $ end-of-URL marker in the robots.txt specification. The same page explains that on conflicting rules Google applies the least restrictive one, so test every pattern before you ship it. Google’s robots meta tag specification covers X-Robots-Tag for non-HTML files.

Which Directive Combinations Cancel Each Other Out?

Four combinations cause most of the conflicts I find. Each sends Google two answers to one question.

CombinationWhat happensFix
Disallow + noindex on one URLGooglebot never fetches the page, never reads noindex, URL can stay listedRemove the Disallow, keep noindex
noindex + canonical to a different URLTwo opposite instructions: remove me, and treat that other URL as the main oneChoose one: noindex to remove, or canonical to merge
Canonical pointing at a noindex pageYou nominate a page that you then ask Google to dropPoint canonicals only at indexable URLs
noindex URL listed in the XML sitemapThe sitemap says “index this,” the tag says “do not”Remove it from the sitemap

I consider the first one the most expensive because it looks fixed. Someone adds noindex to the page, the team marks the ticket done, and nothing changes for months.

How Do I Find These Conflicts on My Own Site?

Run a script that checks each URL against robots.txt and reads its noindex signal. This version needs only Python 3 and no extra libraries.

import re, sys, urllib.request, urllib.robotparser
from urllib.parse import urlsplit

def check(url):
    parts = urlsplit(url)
    rp = urllib.robotparser.RobotFileParser()
    rp.set_url(f"{parts.scheme}://{parts.netloc}/robots.txt")
    rp.read()
    blocked = not rp.can_fetch("Googlebot", url)
    noindex = False
    try:
        with urllib.request.urlopen(urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0"})) as r:
            header = r.headers.get("X-Robots-Tag", "") or ""
            html = r.read(200000).decode("utf-8", "ignore")
        meta = re.search(r'<meta[^>]+name=["']robots["'][^>]*content=["']([^"']+)', html, re.I)
        noindex = "noindex" in header.lower() or bool(meta and "noindex" in meta.group(1).lower())
    except Exception as e:
        print("ERROR", url, e)
        return
    verdict = ("CONFLICT: blocked AND noindex (Google never reads the tag)" if blocked and noindex
               else "blocked from crawling" if blocked
               else "noindex (working)" if noindex else "indexable")
    print(f"{verdict:62} {url}")

for line in open(sys.argv[1], encoding="utf-8"):
    if line.strip():
        check(line.strip())

Save it as directive_conflicts.py, list your URLs one per line in urls.txt, then run:

python directive_conflicts.py urls.txt

Example output (format only):

CONFLICT: blocked AND noindex (Google never reads the tag)     https://www.yourclinic.com/staging/home/
noindex (working)                                              https://www.yourclinic.com/thank-you/
indexable                                                      https://www.yourclinic.com/providers/dr-lee/

Python’s built-in parser follows the basic robots.txt rules and ignores Google’s wildcard patterns. Treat a clean result on a wildcard rule as unconfirmed, and verify those URLs in Search Console’s URL Inspection tool.

What Pushback Do I Get When I Fix These, and How Do I Answer It?

Two objections come up in nearly every clinic engagement.

Developer: “Block everything in robots.txt to be safe.” The instinct is protective, and I respect it. The answer is that Disallow protects nothing from search results, and it hides noindex tags. Offer the alternative that does what they want: authentication for private pages, noindex for pages that are public but unwanted.

Marketing: “Noindex the duplicate location pages so we avoid duplicate content.” Here the answer is that noindex deletes the page from search, including the patients who would find that location. Write unique content for each location first. Use a canonical only when two URLs show one page.

What Do I Do This Week?

  • ☐ List every page you want out of Google, and write down its goal: remove, stop crawl waste, merge or private.
  • ☐ Pick the matching tool from the table.
  • ☐ Run the conflict checker against that list.
  • ☐ Remove every Disallow rule that covers a noindex page.
  • ☐ Remove noindex URLs from the XML sitemap.
  • ☐ Confirm each fix in URL Inspection after Google re-crawls.

Reusable asset: the decision tree in Figure 1 works as a printable one-pager. Pin it where your developers and marketers both see it.

Related reads in this hub:

Frequently Asked Questions

Can I Use noindex and a Canonical Tag on the Same Page?

Technically yes, and it gains you nothing. The two tags give opposite instructions: noindex says remove this page, and a canonical to another URL says treat that URL as the main one. Pick the tag that matches your goal and use only that one.

Does a robots.txt Disallow Pass Authority to Other Pages?

No. Googlebot cannot read a disallowed page, so it cannot follow the links on it. Use noindex instead when you want the page out of results while keeping its links readable.

How Do I Noindex a PDF Such as a Patient Intake Form?

Send an X-Robots-Tag: noindex HTTP header from the server. A PDF has no HTML head for a meta tag. The Apache and Nginx snippets above set the header for every PDF, and Google’s robots meta tag specification documents it.

How Do I Remove a Page From Google Fast?

Add noindex, keep the page crawlable and request a re-crawl in URL Inspection. For an urgent exposure, use the Removals tool for a temporary hide while the noindex takes effect. I cover the Removals tool in Crawling vs. Indexing vs. Ranking.

Is nofollow the Same as noindex?

No. noindex removes a page from search results. nofollow tells crawlers not to follow the links on a page or a single link. They solve different problems and do not substitute for each other.

Does a Canonical Tag Stop Google From Crawling Duplicate URLs?

No, and it reduces crawling only slowly. Google’s faceted navigation guidance says canonicals decrease crawl volume of non-canonical versions only over time. For fast crawl control, use robots.txt.

Facing unexplained indexation drops or broken booking funnels on your clinic website? Book a 30-minute technical consultation with Atiur.

Part of the Technical SEO for Healthcare Websites series. More guides are on the blog. Related case study: Multi-specialty dental group: +24% traffic growth through complex rebrand.

References

Share LinkedIn
Portrait of Atiur Rahman
Written by

Atiur Rahman

SEO and growth strategist with 13+ years in search. I build programs for B2B platforms, SaaS products, e-commerce stores, local service businesses and healthcare brands — from zero visibility to compounding demand.

Next project

Working on something similar?

Send me the site and what seems broken. I will take a look and give you an honest appraisal within one business day.