Skip to content

Technical SEO: what it is and how to check crawling, indexing, canonicals and redirects

Published 7 October 202611 min readBy the SmoothSeen editorial team

Technical SEO is the part of search optimisation that makes sure Google can crawl your pages, index them and know which version of each is the right one. It covers robots.txt and noindex, the sitemap, canonical URLs, status codes and redirects, hreflang and JavaScript. It does not make a page better: it stops a good page being left out.

Key points

  • Technical SEO makes sure Google can crawl your pages, index them and know which URL is the right one for each.
  • robots.txt controls crawling, not indexing. To keep a page out of Google you use noindex, and that page must not be blocked in robots.txt.
  • Google treats 301 and 308 as permanent redirects, follows up to 10 hops and recommends redirecting straight to the final destination.
  • An error page that returns 200 is a soft 404; you can show a helpful error page and return a 404 status code at the same time.
  • Google does not use the lang attribute to detect language, but hreflang helps it link to the right version, and every version must list all the others.

To check it on your own site: SEO audit

On this page

It is the foundation of SEO as a whole: content and links only count if Google reaches the page and stores it in its index. This guide goes through each piece with what Google documents and how to check it yourself.

What is the difference between crawling and indexing?

Crawling is Googlebot downloading a URL; indexing is Google analysing what it downloaded and storing it in its index. They are separate steps and they fail for different reasons. A page can be crawled and not indexed, and a URL blocked from crawling can still end up indexed without its content.

There is a step between the two, rendering: Google processes pages in three phases, crawling, rendering and indexing, and every page that returns a 200 goes through the rendering queue1.

The Page indexing report in Google Search Console is the place to start. It shows which URLs are out of the index and why: blocked by robots.txt, excluded by a noindex tag, a 404, a soft 404, a redirect or a duplicate without a user-selected canonical2.

robots.txt or noindex: which should you use?

robots.txt tells crawlers which URLs they may access; noindex tells Google not to show a page in its results. Mixing them up is the most expensive technical mistake, because the effect is the opposite of what you wanted.

Google puts it this way: robots.txt is mainly for avoiding overloading your server with requests, and it is not a mechanism for keeping a web page out of Google. A blocked URL can still appear in results, without a description, if other sites link to it3. For noindex to work, the page must not be blocked by robots.txt: if Googlebot cannot request it, it never sees the rule. And Google does not support noindex inside the robots.txt file itself4.

What you want
Stop Google spending crawling on internal search or the basket
What to use
Disallow in robots.txt
Why
It controls crawler access3
What you want
Keep a page out of Google
What to use
<meta name="robots" content="noindex"> or the X-Robots-Tag header
Why
It is the rule that removes a page from results4
What you want
Keep it from everyone
What to use
A password
Why
Google suggests it alongside noindex3

An annotated robots.txt for a service business with site search and a customer area:

# robots.txt for https://www.example.com/ (always at the root of the host)
# Anything after # is a comment and crawlers ignore it.

User-agent: *
# Internal search results create thousands of URLs with no value of their own.
Disallow: /search
# The basket and the customer area have nothing to rank.
Disallow: /basket/
Disallow: /my-account/
# Do NOT block a page here that you want removed from Google with noindex:
# Googlebot could not request it and would never read the tag.

# Sitemap location, always as an absolute URL.
Sitemap: https://www.example.com/sitemap.xml

Google supports # comments and the sitemap field, which must be a fully qualified URL including the protocol5.

Do you need an XML sitemap?

An XML sitemap is a file that lists the URLs you want Google to know about, with optional details such as when each was last modified. It helps discovery, but Google warns that it does not guarantee that everything in it will be crawled and indexed6.

Google says you might not need one if your site is small, about 500 pages or fewer, and comprehensively linked internally. Even so, it adds that in most cases a site benefits from having one6. If you have one, keep it consistent: listing a URL in a sitemap suggests it as canonical, although it is a weak signal7. So it should not contain URLs that redirect, return errors or carry noindex.

Canonical URLs: one version of each page

The canonical URL is the one Google picks to represent a group of duplicate or near-identical pages. The same page can respond on http and https, with and without www, with and without a trailing slash, or with campaign parameters. To Google, those are different URLs.

Google ranks the ways of stating your preference by how strongly they influence canonicalisation7:

  1. Redirects: a strong signal that the target should become canonical.
  2. rel="canonical": a strong signal. Google only accepts it in the <head> and recommends absolute URLs.
  3. Sitemap inclusion: a weak signal.

Google also asks you not to use robots.txt for this, because a disallowed URL can still be indexed without its content7. And these are preferences: Google may choose another canonical if the signals contradict each other. To avoid that, make redirects, the canonical tag, the sitemap and your menu links all point at the same URL.

This is how the technical tags look in the <head> of a page with two language versions:

<!-- Language goes on the html element. Google does not use it to detect
     the language, but screen readers and browsers do. -->
<html lang="en-GB">
<head>
  <!-- Title and description: the text usually shown in the search result -->
  <title>24-hour emergency plumber | Example</title>
  <meta name="description" content="A plumber at your door within the hour, every day of the year. A quote before any work starts and a guarantee on every job.">

  <!-- Canonical: absolute, with the protocol and host you have chosen -->
  <link rel="canonical" href="https://www.example.com/en/plumbing/">

  <!-- hreflang: each version lists itself and every other version -->
  <link rel="alternate" hreflang="en" href="https://www.example.com/en/plumbing/">
  <link rel="alternate" hreflang="es" href="https://www.example.com/es/fontaneria/">
  <link rel="alternate" hreflang="x-default" href="https://www.example.com/en/plumbing/">

  <!-- Only on pages you do NOT want in Google and that robots.txt does not block:
  <meta name="robots" content="noindex"> -->
</head>

How to write the title and description well is covered in the guide to the title tag and meta description.

Status codes and redirects

An HTTP status code is the number a server returns with every response: 200 if the page exists, 3xx if it has moved, 4xx if it is not there and 5xx if the server has failed. Google decides much of what it indexes from that number.

Code
200
What it means
The page exists
What Google does
Processes it and may index it
Code
301 and 308
What it means
Moved permanently
What Google does
Follows the redirect and treats it as a signal that the target should be canonical8
Code
302, 303 and 307
What it means
Moved temporarily
What Google does
Follows the redirect but does not use it as a canonical signal8
Code
404 and 410
What it means
Not there
What Google does
Does not index it and drops it if it was indexed; all 4xx codes except 429 are treated the same9
Code
429 and 5xx
What it means
Too many requests or server error
What Google does
Temporarily slows down crawling9

The difference between 301 and 308 lies in the protocol, not in Google: a 308 guarantees the browser repeats the request with the same method, while some older clients turn a request that hits a 301 into a GET10.

Three rules on redirects, all from Google:

  • Server-side first. Use JavaScript redirects only if nothing else is possible, because rendering can fail and Google might never see them8.
  • No chains. Googlebot follows up to 10 hops9, but Google advises redirecting to the final destination directly and, if a chain cannot be avoided, keeping it short: ideally no more than 3 hops and fewer than 511.
  • Keep them. When URLs change, Google recommends keeping redirects for as long as possible, generally at least a year11.

To see the chain a visitor goes through, curl is enough (replace example.com with your domain):

# Follow redirects and show only the status lines and destinations
curl -sIL http://example.com/services | grep -iE "^(HTTP|location)"

# Output from a site with a three-hop chain (example):
# HTTP/1.1 301 Moved Permanently
# Location: https://example.com/services         <- http to https
# HTTP/2 301
# location: https://www.example.com/services     <- bare domain to www
# HTTP/2 301
# location: https://www.example.com/services/    <- adds the trailing slash
# HTTP/2 200

The fix is one server rule that sends every variant, in a single step, to https://www.example.com/services/. Repeat the test with the other variants: they should all reach the 200 in one hop.

Custom 404 pages and soft 404s

A soft 404 is a URL that does not exist but responds with a code other than 404 or 410. Google describes the two usual cases: a friendly error page served with a 200, and a site that redirects every unknown URL to the home page. Both can harm Google's understanding and indexing of your site12.

The fix does not mean giving up a good error page. You can return a 404 status code while serving whatever content you like: a search box, links to the main sections and a sentence explaining what happened12. Search Console flags as "Soft 404" the pages that look like an error even though they return 2002.

In single-page apps, where the server answers 200 to everything, Google suggests a JavaScript redirect to a URL that does return a 404, or adding noindex to the error view1.

Having 404s is not a problem in itself. Google said so on its blog: some URLs returning 404 does not affect how your site's other pages perform in search results12. What is worth checking is your own links pointing at them, as explained in broken links: how to find and fix them.

hreflang and declaring the language

hreflang is an annotation that tells Google which versions of a page exist in other languages or regions, so it can show each person the right one. It can go in the <head>, in an HTTP header or in the sitemap. Google sets three conditions13:

  • Each version lists itself and all the other versions.
  • URLs are fully qualified, with https://.
  • Links go both ways. If page A points to B but B does not point back to A, Google may ignore the annotations.

The x-default value marks the page for anyone who matches none of the listed languages, such as a language selector13.

There is a common misunderstanding about the lang attribute. Google says it does not use the HTML lang attribute or hreflang to detect a page's language; it works it out from the visible content13. That is also why it recommends separate URLs for each language rather than switching the text with cookies or browser settings14. Declaring lang is still required for another reason: it is success criterion 3.1.1 of the WCAG accessibility guidelines, at level A, and it lets screen readers load the right pronunciation rules15.

Google can only crawl a link if it is an <a> element with an href attribute. A <span> with a click handler, an <a> without href or an href="javascript:…" may be missed16. If your menu or pagination works that way, there may be pages Google never discovers.

Google renders JavaScript, but through a queue and with limits: if the main content only appears after scripts run, it depends on that rendering succeeding1. AI assistants' crawlers are a different story, covered in the guide to JavaScript SEO and AI crawlers.

Technical SEO checklist

What
Indexing
How to check
Search Console Page indexing report
Warning sign
Important pages excluded by noindex, robots.txt or soft 404
What
robots.txt
How to check
Open /robots.txt
Warning sign
A Disallow over sections you want to rank
What
Sitemap
How to check
Search Console Sitemaps report
Warning sign
URLs that redirect, return errors or carry noindex
What
Canonical
How to check
Search Console URL Inspection
Warning sign
Google's chosen canonical is not yours
What
One version
How to check
curl -sIL on every variant
Warning sign
More than one hop, or variants that return 200
What
Status codes
How to check
curl -sI
Warning sign
Linked pages returning 404, 5xx or 302s used for permanent moves
What
404 page
How to check
Request a made-up URL
Warning sign
It returns 200 or redirects to the home page
What
hreflang and language
How to check
Source of each version
Warning sign
Versions that do not link to each other, relative URLs or <html> without lang

To prioritise what you find alongside the other areas, follow the method in the SEO audit checklist.

What does SmoothSeen check from this list?

Declaration of interest: this blog belongs to SmoothSeen, a web audit tool. On the URL you analyse, it checks a good part of this list: robots.txt and noindex, the sitemap, the canonical URL, HTTPS, redirects and redirect chains, the custom 404 page, the declared language and hreflang codes, plus the title, meta description and broken links. There is a summary of what the search visibility analysis covers.

It does not connect to Search Console, so it cannot see which URLs Google has indexed: the Page indexing report answers that. To see how the technical side fits with content and authority, go back to the guide to what SEO is.

Frequently asked questions

Can I use robots.txt to remove a page from Google?

No. robots.txt prevents crawling, not indexing: a blocked URL can still appear in results, without a description, if other sites link to it. To remove it, use noindex in a meta tag or in the X-Robots-Tag header, and leave the page accessible so Googlebot can read that rule. If it must be hidden from everyone, protect it with a password.

Is a 301 or a 308 redirect better?

For Google they are the same: both are permanent and both signal that the target should be the canonical URL. The difference is in the HTTP protocol, because a 308 requires the request method to be kept, which matters for forms and APIs rather than ordinary pages. Use whichever your server makes easier to configure.

Does every page need a canonical tag?

It is not mandatory, but it is simple and useful: a self-referencing canonical, absolute and in the <head>, makes clear which version is the right one when campaign parameters or trailing-slash variants appear. What matters most is that it agrees with your redirects, your sitemap and your internal links.

What to do next

Run curl -sIL on the four variants of your domain and request a made-up URL to see what status code your error page returns. Then open the Page indexing report in Search Console and note the important pages that are out of the index. If you would like the rest of the list checked on your site and compared with your competitors, run a free SEO audit with SmoothSeen.

Sources

  1. 1Understand the JavaScript SEO basics, Google Search Central, updated 4 March 2026.
  2. 2Page indexing report, Search Console Help, accessed 7 October 2026.
  3. 3Introduction to robots.txt, Google Search Central, updated 10 December 2025.
  4. 4Block Search indexing with noindex, Google Search Central, updated 10 December 2025.
  5. 5How Google interprets the robots.txt specification, Google Search Central, updated 31 August 2026.
  6. 6Learn about sitemaps, Google Search Central, updated 10 December 2025.
  7. 7How to specify a canonical URL with rel="canonical" and other methods, Google Search Central, updated 10 July 2026.
  8. 8Redirects and Google Search, Google Search Central, updated 14 April 2026.
  9. 9How HTTP status codes affect Google's crawlers, Google Crawling Infrastructure, updated 4 February 2026.
  10. 10308 Permanent Redirect, MDN Web Docs, updated 14 January 2026.
  11. 11How to move a site, Google Search Central, updated 20 August 2026.
  12. 12Do 404 errors hurt my site?, Susan Moskwa, Google Search Central Blog, published 2 May 2011.
  13. 13Tell Google about localized versions of your page, Google Search Central, updated 21 September 2026.
  14. 14Managing multi-regional and multilingual sites, Google Search Central, updated 10 December 2025.
  15. 15Understanding SC 3.1.1: Language of Page, W3C, accessed 7 October 2026.
  16. 16Link best practices for Google, Google Search Central, updated 10 December 2025.

How to cite this article

SmoothSeen. (2026, October 7). Technical SEO: what it is and how to check crawling, indexing, canonicals and redirects. https://smoothseen.com/en/blog/technical-seo/

Who writes this

SmoothSeen is a website audit tool that measures visibility in search engines and AI assistants and delivers reports under the agency's own brand.

This blog belongs to SmoothSeen: when an article discusses the product, it does so knowing the product is ours. Third-party figures link to their original source.

Change history

  • First version.

Keep reading

Technical SEO: crawling, indexing, canonicals, redirects