Technical SEO: what it is and how to check crawling, indexing, canonicals and redirects
Published 7 October 202611 min readBy the SmoothSeen editorial team
Technical SEO is the part of search optimisation that makes sure Google can crawl your pages, index them and know which version of each is the right one. It covers robots.txt and noindex, the sitemap, canonical URLs, status codes and redirects, hreflang and JavaScript. It does not make a page better: it stops a good page being left out.
Key points
- Technical SEO makes sure Google can crawl your pages, index them and know which URL is the right one for each.
- robots.txt controls crawling, not indexing. To keep a page out of Google you use noindex, and that page must not be blocked in robots.txt.
- Google treats 301 and 308 as permanent redirects, follows up to 10 hops and recommends redirecting straight to the final destination.
- An error page that returns 200 is a soft 404; you can show a helpful error page and return a 404 status code at the same time.
- Google does not use the lang attribute to detect language, but hreflang helps it link to the right version, and every version must list all the others.
To check it on your own site: SEO audit
On this page
- What is the difference between crawling and indexing?
- robots.txt or noindex: which should you use?
- Do you need an XML sitemap?
- Canonical URLs: one version of each page
- Status codes and redirects
- Custom 404 pages and soft 404s
- hreflang and declaring the language
- Crawlable links and JavaScript
- Technical SEO checklist
- What does SmoothSeen check from this list?
- Frequently asked questions
- What to do next
It is the foundation of SEO as a whole: content and links only count if Google reaches the page and stores it in its index. This guide goes through each piece with what Google documents and how to check it yourself.
What is the difference between crawling and indexing?
Crawling is Googlebot downloading a URL; indexing is Google analysing what it downloaded and storing it in its index. They are separate steps and they fail for different reasons. A page can be crawled and not indexed, and a URL blocked from crawling can still end up indexed without its content.
There is a step between the two, rendering: Google processes pages in three phases, crawling, rendering and indexing, and every page that returns a 200 goes through the rendering queue1.
The Page indexing report in Google Search Console is the place to start. It shows which URLs are out of the index and why: blocked by robots.txt, excluded by a noindex tag, a 404, a soft 404, a redirect or a duplicate without a user-selected canonical2.
robots.txt or noindex: which should you use?
robots.txt tells crawlers which URLs they may access; noindex tells Google not to show a page in its results. Mixing them up is the most expensive technical mistake, because the effect is the opposite of what you wanted.
Google puts it this way: robots.txt is mainly for avoiding overloading your server with requests, and it is not a mechanism for keeping a web page out of Google. A blocked URL can still appear in results, without a description, if other sites link to it3. For noindex to work, the page must not be blocked by robots.txt: if Googlebot cannot request it, it never sees the rule. And Google does not support noindex inside the robots.txt file itself4.
- What you want
- Stop Google spending crawling on internal search or the basket
- What to use
Disallowin robots.txt- Why
- It controls crawler access3
- What you want
- Keep a page out of Google
- What to use
<meta name="robots" content="noindex">or theX-Robots-Tagheader- Why
- It is the rule that removes a page from results4
- What you want
- Keep it from everyone
- What to use
- A password
- Why
- Google suggests it alongside
noindex3
An annotated robots.txt for a service business with site search and a customer area:
# robots.txt for https://www.example.com/ (always at the root of the host)
# Anything after # is a comment and crawlers ignore it.
User-agent: *
# Internal search results create thousands of URLs with no value of their own.
Disallow: /search
# The basket and the customer area have nothing to rank.
Disallow: /basket/
Disallow: /my-account/
# Do NOT block a page here that you want removed from Google with noindex:
# Googlebot could not request it and would never read the tag.
# Sitemap location, always as an absolute URL.
Sitemap: https://www.example.com/sitemap.xmlGoogle supports # comments and the sitemap field, which must be a fully qualified URL including the protocol5.
Do you need an XML sitemap?
An XML sitemap is a file that lists the URLs you want Google to know about, with optional details such as when each was last modified. It helps discovery, but Google warns that it does not guarantee that everything in it will be crawled and indexed6.
Google says you might not need one if your site is small, about 500 pages or fewer, and comprehensively linked internally. Even so, it adds that in most cases a site benefits from having one6. If you have one, keep it consistent: listing a URL in a sitemap suggests it as canonical, although it is a weak signal7. So it should not contain URLs that redirect, return errors or carry noindex.
Canonical URLs: one version of each page
The canonical URL is the one Google picks to represent a group of duplicate or near-identical pages. The same page can respond on http and https, with and without www, with and without a trailing slash, or with campaign parameters. To Google, those are different URLs.
Google ranks the ways of stating your preference by how strongly they influence canonicalisation7:
- Redirects: a strong signal that the target should become canonical.
rel="canonical": a strong signal. Google only accepts it in the<head>and recommends absolute URLs.- Sitemap inclusion: a weak signal.
Google also asks you not to use robots.txt for this, because a disallowed URL can still be indexed without its content7. And these are preferences: Google may choose another canonical if the signals contradict each other. To avoid that, make redirects, the canonical tag, the sitemap and your menu links all point at the same URL.
This is how the technical tags look in the <head> of a page with two language versions:
<!-- Language goes on the html element. Google does not use it to detect
the language, but screen readers and browsers do. -->
<html lang="en-GB">
<head>
<!-- Title and description: the text usually shown in the search result -->
<title>24-hour emergency plumber | Example</title>
<meta name="description" content="A plumber at your door within the hour, every day of the year. A quote before any work starts and a guarantee on every job.">
<!-- Canonical: absolute, with the protocol and host you have chosen -->
<link rel="canonical" href="https://www.example.com/en/plumbing/">
<!-- hreflang: each version lists itself and every other version -->
<link rel="alternate" hreflang="en" href="https://www.example.com/en/plumbing/">
<link rel="alternate" hreflang="es" href="https://www.example.com/es/fontaneria/">
<link rel="alternate" hreflang="x-default" href="https://www.example.com/en/plumbing/">
<!-- Only on pages you do NOT want in Google and that robots.txt does not block:
<meta name="robots" content="noindex"> -->
</head>How to write the title and description well is covered in the guide to the title tag and meta description.
Status codes and redirects
An HTTP status code is the number a server returns with every response: 200 if the page exists, 3xx if it has moved, 4xx if it is not there and 5xx if the server has failed. Google decides much of what it indexes from that number.
- Code
- 200
- What it means
- The page exists
- What Google does
- Processes it and may index it
- Code
- 301 and 308
- What it means
- Moved permanently
- What Google does
- Follows the redirect and treats it as a signal that the target should be canonical8
- Code
- 302, 303 and 307
- What it means
- Moved temporarily
- What Google does
- Follows the redirect but does not use it as a canonical signal8
- Code
- 404 and 410
- What it means
- Not there
- What Google does
- Does not index it and drops it if it was indexed; all 4xx codes except 429 are treated the same9
- Code
- 429 and 5xx
- What it means
- Too many requests or server error
- What Google does
- Temporarily slows down crawling9
The difference between 301 and 308 lies in the protocol, not in Google: a 308 guarantees the browser repeats the request with the same method, while some older clients turn a request that hits a 301 into a GET10.
Three rules on redirects, all from Google:
- Server-side first. Use JavaScript redirects only if nothing else is possible, because rendering can fail and Google might never see them8.
- No chains. Googlebot follows up to 10 hops9, but Google advises redirecting to the final destination directly and, if a chain cannot be avoided, keeping it short: ideally no more than 3 hops and fewer than 511.
- Keep them. When URLs change, Google recommends keeping redirects for as long as possible, generally at least a year11.
To see the chain a visitor goes through, curl is enough (replace example.com with your domain):
# Follow redirects and show only the status lines and destinations
curl -sIL http://example.com/services | grep -iE "^(HTTP|location)"
# Output from a site with a three-hop chain (example):
# HTTP/1.1 301 Moved Permanently
# Location: https://example.com/services <- http to https
# HTTP/2 301
# location: https://www.example.com/services <- bare domain to www
# HTTP/2 301
# location: https://www.example.com/services/ <- adds the trailing slash
# HTTP/2 200The fix is one server rule that sends every variant, in a single step, to https://www.example.com/services/. Repeat the test with the other variants: they should all reach the 200 in one hop.
Custom 404 pages and soft 404s
A soft 404 is a URL that does not exist but responds with a code other than 404 or 410. Google describes the two usual cases: a friendly error page served with a 200, and a site that redirects every unknown URL to the home page. Both can harm Google's understanding and indexing of your site12.
The fix does not mean giving up a good error page. You can return a 404 status code while serving whatever content you like: a search box, links to the main sections and a sentence explaining what happened12. Search Console flags as "Soft 404" the pages that look like an error even though they return 2002.
In single-page apps, where the server answers 200 to everything, Google suggests a JavaScript redirect to a URL that does return a 404, or adding noindex to the error view1.
Having 404s is not a problem in itself. Google said so on its blog: some URLs returning 404 does not affect how your site's other pages perform in search results12. What is worth checking is your own links pointing at them, as explained in broken links: how to find and fix them.
hreflang and declaring the language
hreflang is an annotation that tells Google which versions of a page exist in other languages or regions, so it can show each person the right one. It can go in the <head>, in an HTTP header or in the sitemap. Google sets three conditions13:
- Each version lists itself and all the other versions.
- URLs are fully qualified, with
https://. - Links go both ways. If page A points to B but B does not point back to A, Google may ignore the annotations.
The x-default value marks the page for anyone who matches none of the listed languages, such as a language selector13.
There is a common misunderstanding about the lang attribute. Google says it does not use the HTML lang attribute or hreflang to detect a page's language; it works it out from the visible content13. That is also why it recommends separate URLs for each language rather than switching the text with cookies or browser settings14. Declaring lang is still required for another reason: it is success criterion 3.1.1 of the WCAG accessibility guidelines, at level A, and it lets screen readers load the right pronunciation rules15.
Crawlable links and JavaScript
Google can only crawl a link if it is an <a> element with an href attribute. A <span> with a click handler, an <a> without href or an href="javascript:…" may be missed16. If your menu or pagination works that way, there may be pages Google never discovers.
Google renders JavaScript, but through a queue and with limits: if the main content only appears after scripts run, it depends on that rendering succeeding1. AI assistants' crawlers are a different story, covered in the guide to JavaScript SEO and AI crawlers.
Technical SEO checklist
- What
- Indexing
- How to check
- Search Console Page indexing report
- Warning sign
- Important pages excluded by
noindex, robots.txt or soft 404
- What
- robots.txt
- How to check
- Open
/robots.txt - Warning sign
- A
Disallowover sections you want to rank
- What
- Sitemap
- How to check
- Search Console Sitemaps report
- Warning sign
- URLs that redirect, return errors or carry
noindex
- What
- Canonical
- How to check
- Search Console URL Inspection
- Warning sign
- Google's chosen canonical is not yours
- What
- One version
- How to check
curl -sILon every variant- Warning sign
- More than one hop, or variants that return 200
- What
- Status codes
- How to check
curl -sI- Warning sign
- Linked pages returning 404, 5xx or 302s used for permanent moves
- What
- 404 page
- How to check
- Request a made-up URL
- Warning sign
- It returns 200 or redirects to the home page
- What
- hreflang and language
- How to check
- Source of each version
- Warning sign
- Versions that do not link to each other, relative URLs or
<html>withoutlang
To prioritise what you find alongside the other areas, follow the method in the SEO audit checklist.
What does SmoothSeen check from this list?
Declaration of interest: this blog belongs to SmoothSeen, a web audit tool. On the URL you analyse, it checks a good part of this list: robots.txt and noindex, the sitemap, the canonical URL, HTTPS, redirects and redirect chains, the custom 404 page, the declared language and hreflang codes, plus the title, meta description and broken links. There is a summary of what the search visibility analysis covers.
It does not connect to Search Console, so it cannot see which URLs Google has indexed: the Page indexing report answers that. To see how the technical side fits with content and authority, go back to the guide to what SEO is.
Frequently asked questions
Can I use robots.txt to remove a page from Google?
No. robots.txt prevents crawling, not indexing: a blocked URL can still appear in results, without a description, if other sites link to it. To remove it, use noindex in a meta tag or in the X-Robots-Tag header, and leave the page accessible so Googlebot can read that rule. If it must be hidden from everyone, protect it with a password.
Is a 301 or a 308 redirect better?
For Google they are the same: both are permanent and both signal that the target should be the canonical URL. The difference is in the HTTP protocol, because a 308 requires the request method to be kept, which matters for forms and APIs rather than ordinary pages. Use whichever your server makes easier to configure.
Does every page need a canonical tag?
It is not mandatory, but it is simple and useful: a self-referencing canonical, absolute and in the <head>, makes clear which version is the right one when campaign parameters or trailing-slash variants appear. What matters most is that it agrees with your redirects, your sitemap and your internal links.
What to do next
Run curl -sIL on the four variants of your domain and request a made-up URL to see what status code your error page returns. Then open the Page indexing report in Search Console and note the important pages that are out of the index. If you would like the rest of the list checked on your site and compared with your competitors, run a free SEO audit with SmoothSeen.
Sources
- 1Understand the JavaScript SEO basics, Google Search Central, updated 4 March 2026.
- 2Page indexing report, Search Console Help, accessed 7 October 2026.
- 3Introduction to robots.txt, Google Search Central, updated 10 December 2025.
- 4Block Search indexing with noindex, Google Search Central, updated 10 December 2025.
- 5How Google interprets the robots.txt specification, Google Search Central, updated 31 August 2026.
- 6Learn about sitemaps, Google Search Central, updated 10 December 2025.
- 7How to specify a canonical URL with rel="canonical" and other methods, Google Search Central, updated 10 July 2026.
- 8Redirects and Google Search, Google Search Central, updated 14 April 2026.
- 9How HTTP status codes affect Google's crawlers, Google Crawling Infrastructure, updated 4 February 2026.
- 10308 Permanent Redirect, MDN Web Docs, updated 14 January 2026.
- 11How to move a site, Google Search Central, updated 20 August 2026.
- 12Do 404 errors hurt my site?, Susan Moskwa, Google Search Central Blog, published 2 May 2011.
- 13Tell Google about localized versions of your page, Google Search Central, updated 21 September 2026.
- 14Managing multi-regional and multilingual sites, Google Search Central, updated 10 December 2025.
- 15Understanding SC 3.1.1: Language of Page, W3C, accessed 7 October 2026.
- 16Link best practices for Google, Google Search Central, updated 10 December 2025.
How to cite this article
SmoothSeen. (2026, October 7). Technical SEO: what it is and how to check crawling, indexing, canonicals and redirects. https://smoothseen.com/en/blog/technical-seo/
Keep reading
What is SEO? How search engine optimisation works and how to improve it in 2026
What SEO is, how Google decides which pages to show and a prioritised checklist to improve your rankings with free tools.
.htaccess force HTTPS: redirect to https, enable HSTS and add security headers in Apache
How to force HTTPS in .htaccess or an Apache VirtualHost, roll out HSTS safely and add security headers. Every snippet tested on Apache 2.4.69.
.htaccess gzip and Brotli: browser caching and blocking AI bots in Apache
How to enable gzip and Brotli, set browser caching and block AI training bots in Apache .htaccess without dropping out of ChatGPT search. Tested.