Broken links: what they are, how to find them and how to fix them
Published 7 October 20269 min readBy the SmoothSeen editorial team
A broken link is a link that leads to a page that no longer exists or responds with an error, almost always a 404. To find them, combine the Page indexing report in Search Console, a link crawler and checks with curl. To fix them, correct the link, add a 301 redirect if there is an equivalent page, or remove it.
Key points
- A broken link leads to a URL that does not exist or returns an error; it can sit on your site (internal), point elsewhere (external) or arrive from another site at a URL of yours that has gone.
- Google says URLs returning 404 do not affect how the rest of your site performs; the problem is the path that breaks for people and crawlers.
- Google treats 404 and 410 the same; a soft 404, an error page that returns 200, does waste crawling.
- Every broken link gets one of three fixes: correct the link, 301-redirect if there is an equivalent page, or remove it.
- According to Pew Research Center, 38% of web pages that existed in 2013 were no longer accessible in October 2023.
To check it on your own site: SEO audit
On this page
- What kinds of broken links are there?
- 404, 410 and soft 404s: what does a broken link return?
- Do broken links hurt SEO?
- How do you find broken links?
- How do you fix each broken link?
- Links to other sites: why do they decay?
- How often should you check?
- What does SmoothSeen check?
- Frequently asked questions
- What to do next
Broken links are a staple of SEO maintenance, and also one of its most exaggerated topics. This guide separates what Google actually says from what gets repeated without a source, and gives you a method for checking them without losing an afternoon.
What kinds of broken links are there?
A broken link is an <a href> whose destination returns an error or does not respond at all. Depending on where the link sits and where it points, there are three cases, and each is fixed in a different place:
- Type
- Internal
- Example
- Your menu links to
/services/renovations/, which you deleted - Who fixes it
- You, on your site
- What usually happens
- Visitors and Googlebot hit a dead end
- Type
- External
- Example
- One of your articles cites a study whose URL no longer exists
- Who fixes it
- You, by changing or removing the link
- What usually happens
- Readers cannot check what you claim
- Type
- Inbound
- Example
- Another site links to
/summer-offers-2024/, which has gone - Who fixes it
- You, with a redirect if there is an equivalent
- What usually happens
- You lose the visits that link used to send
Internal links matter most, because you control them completely and because they are how Google discovers your pages. Google uses links to find new pages to crawl and as a signal when working out how relevant they are1.
404, 410 and soft 404s: what does a broken link return?
A 404 says the server cannot find the page and does not say whether that is temporary or permanent; a 410 says the page has gone and probably will not come back. That is how the HTTP standard defines them, and it prefers 410 when the server knows the removal is permanent2.
For Google the difference does not matter. It treats all 4xx codes except 429 the same way: it does not index those URLs and drops the ones it had indexed3. Back in 2011 its blog already said it made no difference which of the two you returned4.
The case worth watching is the soft 404: a URL that does not exist but returns 200, or redirects to the home page. Google warns that such pages keep being crawled and waste crawl budget, and asks you to eliminate them5.
A word of caution before declaring a link broken: a 403 or 429 returned to an automated tool proves nothing. Many servers block or throttle crawlers that are not browsers, and the same URL opens fine in Chrome. Only 404 and 410 say the page does not exist; everything else needs a manual check.
Do broken links hurt SEO?
Not as a penalty, but they do break paths. Google explained on its Search Central blog that some URLs on your site returning 404 does not affect how your other pages perform in search results, and that 404s are a normal part of the web4. Search Console says the same: a 404 is not necessarily a problem if the page was removed without any replacement6.
What does have consequences is something else:
- A broken internal link to a page that now lives at another URL. The new page stops receiving that link, which is one of the ways Google finds and assesses it1.
- An inbound link that returns a 404. If another site links to you with a typo or to an old URL, Google suggests a 301 redirect to the correct URL to capture that traffic4.
- The visitor. Someone who clicks a menu link and lands on an error has to find another route or leave. We know of no study with a published methodology on how many leave, so we give no figure here.
- Crawling on large sites. Google recommends returning 404 or 410 for permanently removed pages, because a 404 is a strong signal not to crawl that URL again5. Its crawl budget guide is aimed at sites with a million pages, or ten thousand that change daily; on a small site this is not the main concern.
How do you find broken links?
Three sources, from the easiest to the most thorough. None of them sees everything on its own.
Google Search Console
In the Page indexing report, the "Not found (404)" reason lists the URLs on your domain that Google tried to crawl and got a 4046. To see where they are linked from, open each one in URL Inspection: under discovery it shows the "Referring page" Google may have used to find it7.
Its limit: it only sees URLs on your domain that Google knows about. Broken external links in your articles do not show up there.
A link crawler
A crawler goes through your site the way Googlebot would and checks every link, internal and external. The W3C Link Checker is free, runs in the browser and can check a whole site recursively, to the depth you choose8. Commercial desktop crawlers do the same with more filters and spreadsheet export.
curl, to check a list
If you already have the list of URLs (exported from a crawler or your CMS), this loop returns only the ones that fail or redirect:
# links.txt: one URL per line
while read -r url; do
# -L follows redirects; -o /dev/null discards the body
# %{http_code}: final status; %{num_redirects}: hops to get there
result=$(curl -s -o /dev/null -L --max-time 15 \
-w "%{http_code} %{num_redirects}" "$url")
echo "$result $url"
done < links.txt | awk '$1 !~ /^2/ || $2 > 0'
# How to read the output:
# 404 0 / 410 0 -> broken: fix or remove the link
# 200 2 -> works, but after two redirects: link to the final URL
# 403 0 / 429 0 -> inconclusive: open it in a browser
# 000 0 -> no response (domain down or timed out): try again laterThe second line matters more than it looks: an internal link that works after two redirects is not broken, but it makes every visit and every crawl take extra hops. The technical SEO guide covers redirect chains.
How do you fix each broken link?
There are three fixes: correct the link, redirect the URL or remove the link. Which one applies depends on whether the content still exists somewhere else.
- Situation
- Internal link with a typo or to an old URL
- What to do
- Correct the
hrefon the linking page - Why
- A redirect masks it, but a direct link avoids the hop; Google asks you to update internal links when URLs change9
- Situation
- The page has moved to another URL
- What to do
- 301 redirect from the old URL to the new one
- Why
- For moved pages, Search Console recommends a 301 to the new location6
- Situation
- The page was deleted and there is nothing equivalent
- What to do
- Let it return 404 or 410 and remove the links pointing at it
- Why
- This is what Google asks for content removed without a replacement4; the HTTP standard presents 410 as a notice that the owner wants links to it removed2
- Situation
- Temptation: send every 404 to the home page
- What to do
- Do not do it
- Why
- Google counts it as a soft 4044
- Situation
- External link to a page that has disappeared
- What to do
- Find the new URL of the same document, another source that says the same, or remove it
- Why
- Readers need to be able to check what you cite
- Situation
- Another site links to a URL of yours with a typo
- What to do
- 301 redirect to the correct URL
- Why
- You recover the visits that link sends4
One rule that saves trouble: redirect only if the destination answers what the person who clicked was looking for. A discontinued product page can go to the model that replaces it, not to the general category or the home page.
If the failure is your error page itself, because it returns 200, fix that first: until you do, no tool can tell your broken links from your working ones.
Links to other sites: why do they decay?
External links break on their own over time, even if you touch nothing. In October 2023, Pew Research Center checked a sample of almost a million pages collected by Common Crawl between 2013 and 2023. Of those that existed in 2013, 38% were no longer accessible10.
The same study checked the links on pages collected in 2023: 23% of news pages and 21% of government pages had at least one broken link, and 54% of the Wikipedia articles analysed had at least one dead link in their references10.
Three habits reduce the problem:
- Link to stable sources. Official documentation and reference pages last longer than a press release or a campaign page.
- Note when you consulted each source. If the link dies, you will know which version you cited and can look for an archived copy in the Internet Archive's Wayback Machine.
- Check the external links on your most visited pages every few months, not just the internal ones.
How often should you check?
A small site that changes little can get by with a pass every quarter and another after each redesign, migration or removal of sections, which is when dozens of links break at once. For a shop or a publisher posting daily, monthly is better. To fold the check into regular tracking of several sites, see how to organise SEO monitoring. If it is your first time, do it as part of a full SEO audit.
What does SmoothSeen check?
Declaration of interest: this blog belongs to SmoothSeen, a web audit tool. When it analyses a URL, it checks the links on that page and flags those that lead to a 404 or a 410. It does not count links that return 403 or 429 as broken, because that is usually a block on automated tools rather than a missing page. You will find it alongside the other checks in what the search visibility analysis covers.
It does not crawl your whole site or connect to Search Console: for the full list of 404s Google knows about, the Page indexing report is still the source. The rest of what counts in Google is in the guide to what SEO is.
Frequently asked questions
Does a 404 error get you penalised by Google?
No. Google has explained that having URLs that return 404 does not affect how the other pages on your site perform in search results, and that 404s are normal when content is removed. What is worth checking is whether an important page returns a 404 by mistake, and whether your own links, especially in the menu, lead to pages that no longer exist.
Is it better to return a 404 or a 410?
For Google there is no practical difference: it treats both the same and stops indexing the URL. A 410 is more precise under the HTTP standard, because it says the page has been removed permanently, so use it if your server makes that easy. If you are unsure whether the removal is permanent, a 404 is the correct response.
Can I redirect every 404 to the home page?
It is not a good idea. Google considers redirecting every non-existent URL to the home page a soft 404, which makes your site harder to understand and index, and the visitor lands on a page without what they were looking for. Use a 301 only when there is an equivalent page; otherwise return a 404 with a helpful error page.
Do broken external links affect my rankings?
Google does not document any penalty for linking to a page that has disappeared. The effect is on readers, who cannot check the source, and on how credible the content looks. That is why it is worth checking them on the articles that get the most traffic and replacing each dead link with the document's new URL or another source that supports the same point.
What to do next
Open the Page indexing report in Search Console, go to the "Not found (404)" reason and use URL Inspection to see which pages link to each URL. Fix the links in your menu and templates first, as they repeat across the whole site. If you would like to see a page's broken links alongside the rest of its technical analysis, run a free SEO audit with SmoothSeen.
Sources
- 1Link best practices for Google, Google Search Central, updated 10 December 2025.
- 2RFC 9110: HTTP Semantics, IETF, published June 2022, accessed 7 October 2026.
- 3How HTTP status codes affect Google's crawlers, Google Crawling Infrastructure, updated 4 February 2026.
- 4Do 404 errors hurt my site?, Susan Moskwa, Google Search Central Blog, published 2 May 2011.
- 5Optimize your crawl budget, Google Search Central, updated 22 July 2026.
- 6Page indexing report, Search Console Help, accessed 7 October 2026.
- 7URL Inspection tool, Search Console Help, accessed 7 October 2026.
- 8W3C Link Checker, W3C, accessed 7 October 2026.
- 9How to move a site, Google Search Central, updated 20 August 2026.
- 10When Online Content Disappears, Athena Chapekis, Samuel Bestvater, Emma Remy and Gonzalo Rivero, Pew Research Center, published 17 May 2024.
How to cite this article
SmoothSeen. (2026, October 7). Broken links: what they are, how to find them and how to fix them. https://smoothseen.com/en/blog/broken-links/
Keep reading
What is SEO? How search engine optimisation works and how to improve it in 2026
What SEO is, how Google decides which pages to show and a prioritised checklist to improve your rankings with free tools.
.htaccess force HTTPS: redirect to https, enable HSTS and add security headers in Apache
How to force HTTPS in .htaccess or an Apache VirtualHost, roll out HSTS safely and add security headers. Every snippet tested on Apache 2.4.69.
.htaccess gzip and Brotli: browser caching and blocking AI bots in Apache
How to enable gzip and Brotli, set browser caching and block AI training bots in Apache .htaccess without dropping out of ChatGPT search. Tested.