If a page is not showing up on Google, the cause is almost always one of eight things: the page returns noindex, robots.txt blocks crawling, the canonical tag points to another URL, the page sits in a redirect chain, Google decided the content is not worth indexing, no internal links point to the page, an empty page is treated as a soft 404, or the hreflang setup is broken.

You do not have to guess which one it is. A ten-minute checklist will tell you. Below, each cause is covered in order: how to confirm it and how to fix it.

Contents

Diagnose first: which stage is the page stuck in

Google processes a page in three stages: it discovers it, crawls it, then indexes it. The fix depends entirely on which stage the problem is in.

Animation showing how Google discovers, crawls and indexes a page, and which stage a page gets stuck in
A page goes through these three stages in order. The status message in URL Inspection tells you where it got stuck.

The fastest way to find the stage is the URL Inspection tool in Search Console. Paste the URL and note the status message:

Status message Stage Cause to check
URL is unknown to Google Not discovered 6
Discovered – currently not indexed Not crawled 2, 6
Crawled – currently not indexed Crawled 1, 5
Duplicate without user-selected canonical Crawled 3, 8
Page with redirect Crawled 4
Soft 404 Crawled 7

The tool has two tabs, and people mix them up all the time. The Google Index tab shows the page as it was on the last crawl, which may be old. The Live Test tab fetches the page right now. If you have made a fix, the live test will look clean while the index tab still shows the error. That is normal, and it means your fix worked.

1. The page returns noindex

The most common cause, and the easiest to miss. The page opens normally in the browser, the content is visible, no warning appears. But somewhere in the HTML there is a line telling Google “do not index me”.

How to check:

curl -s https://yoursite.com/page | grep 'name="robots"'

The expected output is index, follow. If you see noindex here, you do not need to check the other seven causes. This is the problem.

Where it comes from: a page setting in your SEO plugin, a condition in the theme template, a setting carried over from a staging site, or an X-Robots-Tag header sent by the server. Check the header too:

curl -sI https://yoursite.com/page | grep -i x-robots-tag

2. robots.txt blocks crawling

There is an important distinction here: robots.txt blocks crawling, not indexing. If a blocked page gets enough links from other sites, Google can index it by its URL alone, without ever seeing the content. The result is a search listing with no description.

The reverse is also true: if you want a page removed from the index, blocking it in robots.txt is the wrong method. If Google cannot crawl the page, it cannot see your noindex tag either. The right way is to allow crawling and add noindex.

How to check: first make sure the file exists. A robots.txt that returns 404 is more common than you would think.

curl -sI https://yoursite.com/robots.txt | head -1

Then use the robots.txt report in Search Console to see which rule blocks which URL.

3. The canonical points to another URL

The canonical tag tells Google where the original version of the content lives. Set it wrong and you push your own page out of the index.

How to check:

curl -s https://yoursite.com/page | grep canonical

The URL in the output should be exactly the URL you checked, including the trailing slash, www or not, and http versus https.

Three common mistakes:

  • No self-reference: every page that is not a duplicate should point to its own URL.
  • Blanket canonical: every page pointing to the homepage. Usually caused by a tag hard-coded into the template, and it pushes the entire site out of the index.
  • Canonical loop: page A points to B, and B redirects to A. Google cannot choose either.
Canonical loop animation: page A points to B with a canonical, B redirects back to A with a 307, Googlebot is caught between them, then the correct setup
When the canonical and the redirect point at each other, Google is stuck between two URLs. In the correct setup, the canonical points to the original URL that returns 200 directly.

A canonical is a strong hint, not a directive. If Google considers the pages genuinely different, it can ignore your tag. In that case the “Google-selected canonical” field in URL Inspection will differ from yours.

4. The page is in a redirect chain

A URL that redirects does not get indexed; its target does. The trouble starts when the target redirects somewhere else too.

How to check:

curl -IL https://yoursite.com/page

Count the HTTP/2 30x lines in the output. Zero or one is ideal. If the same URL appears twice, you are in a loop and that page will never be indexed.

Checking a redirect chain with curl -IL: a chain that returns 200 after 301 and 302 hops, and a clean URL that returns 200 directly
The first URL returns 200 after two redirects, the second one directly. Internal links and the sitemap should always point to the final URL.

The status code matters too:

Code Meaning When to use
301 Permanent The URL has permanently changed
302 Temporary Short-term redirect
307 Temporary, method preserved Server or framework level
308 Permanent, method preserved Modern equivalent of 301

Using 302 for a permanent move delays the transfer of signals. If the move is permanent, use 301 or 308.

5. Crawled but not indexed

This message is not a technical error. Google saw your page, read it and decided it was not worth adding to the index. Hard to hear, but the diagnosis is clear.

Typical reasons: the page is nearly identical to another page, the content is shallow compared with the other results for that search, or it is an automatically generated listing or tag page.

What to do: there is nothing to fix in the code. Add something to the page that a person searching for this cannot find anywhere else: your own data, your own measurements, your own screenshots. If you cannot, merge the page or remove it. One strong page beats five weak pages on the same topic.

6. Discovered but not crawled

Google knows the URL but has not got around to crawling it. There are usually two reasons: the crawl budget allocated to the site is low, or the page does not get enough links from within the site.

Crawl budget is overrated for small sites. On a site with a few hundred pages it is rarely the real problem. The real issue is usually internal linking: if nothing links to the page, Google treats it as unimportant.

What to do:

  • Add the page to your sitemap and submit the sitemap in Search Console.
  • Link to the page contextually from related content. A menu link in the footer is not enough; a link from within the text, with meaningful anchor text, is much stronger.
  • Request indexing once from URL Inspection. If there is no technical block, this usually works the same day.

7. Soft 404

The page returns 200, but Google reads the content as “there is nothing here”. It is most common on category pages with no posts yet, search result pages with no results and deleted product pages.

Modern JavaScript frameworks have a sneaky variant: the framework shows its 404 component, but the HTTP status stays 200. The user sees the right page; Google gets the wrong signal.

How to check: make up a URL that does not exist and read the status code.

curl -sI https://yoursite.com/this-page-does-not-exist | head -1

Expected: 404. If you get 200, your site is telling Google that every made-up URL exists.

What to do: if the page really does not exist, return 404. If it is only empty for now, either add content or set noindex until it is filled.

8. The hreflang setup is broken

On multilingual sites, hreflang tells Google which version to show to visitors in which language. If it is set up wrong, Google treats the language versions as copies of each other and leaves some of them out of the index.

Four rules:

  • Every target URL must return 200 directly. A single target that redirects invalidates the whole set on that page.
  • Every page must list itself too.
  • Links must be reciprocal: if page A points to B, B must point back to A.
  • Define an x-default; visitors who match no language are sent there.

How to check:

curl -s https://yoursite.com/ | grep -o 'hreflang="[^"]*" href="[^"]*"'

Verify each URL in the output one by one with curl -I.

In what order to check

Order matters, because a problem higher up makes the ones below it irrelevant. Internal linking will not help a page that returns noindex.

Step Check Expected
1 robots meta tag index, follow
2 robots.txt access 200, no block
3 HTTP status code 200
4 Redirect chain 0 or 1 hop
5 Canonical Points to itself
6 hreflang targets All return 200
7 Internal links At least one contextual link
8 Content depth Answers the search

The first six are technical and binary: they are either right or wrong. The last two require judgment and take time.

Conclusion

Most indexing problems are not about the content; they are about what the page tells Google. The good news is that every one of these signals can be measured on its own, and most of them are fixed with a one-line change.

What works in practice: first run through the order above and clear the technical blocks, then submit the page for indexing once, then give it a contextual link from within your site. Once the technical block is gone, Google usually moves fast. What takes months is the block itself, not what happens after it is removed.

Once the technical side is clean, authority comes next. If Google can see and index the page, the next factor that decides its ranking is the strength of the links pointing to it.

Frequently asked questions

Why is my website not showing up on Google?

There are eight common causes: a noindex tag, a robots.txt block, a wrong canonical, a redirect chain, content Google considers not worth indexing, missing internal links, a soft 404 and broken hreflang. The status message in Search Console URL Inspection tells you which one applies.

How long does it take for a new page to be indexed?

If there is no technical block, anywhere from a few hours to a few days. If you request indexing manually, it can happen the same day. On sites that Google crawls often, it is faster.

Does requesting indexing improve my rankings?

No. The request only puts the page in the crawl queue; it has no effect on ranking. Submitting the same page again and again does not move it up the queue either.

What is the difference between noindex and a robots.txt block?

robots.txt blocks crawling; noindex blocks indexing. If you want to remove a page from the index, do not block it in robots.txt: if Google cannot crawl the page, it cannot see the noindex tag.

Should every page have a canonical tag?

Yes, a self-referencing canonical is standard practice. It is especially important on sites where URLs with parameters can be created.

Is “Crawled – currently not indexed” an error?

It is not a technical error. It means Google saw the page and decided it was not worth indexing. The fix is to strengthen the content or merge the page with another one.

Does submitting a sitemap guarantee indexing?

No. A sitemap helps discovery; it does not guarantee indexing. After the page is crawled, it still has to be judged worth indexing.

What should I do if my page is treated as a duplicate?

First check the canonical tag and whether it points to the page itself. If the problem continues, the two pages may genuinely be too similar; merge them or make one of them clearly different.

Which language should be x-default on a multilingual site?

The version you want to show to visitors who match none of your languages. Usually this is the English version or a language selection page.

Next steps

If you have checked the technical foundation, the next step is building authority. You can find our link packages and prices on the Hacklink service page. If you get stuck anywhere, write to us.