Crawlable links, noindex and crawl budget: when can Google count a link?

Google can generally follow a link only if it is an a element with an href attribute, on a page it is allowed to crawl and index. A noindex tag left in place, a robots.txt block or a link built from script events can each stop a link from counting. Crawl budget matters mainly for very large sites.

In short

  • Google's link documentation says it can generally crawl a link only if it is an a element with an href attribute that resolves to a real web address.
  • Links inserted by JavaScript are crawlable when they use the same markup, but Google has to render the page first, and it may skip rendering a page that carries noindex.
  • Search Engine Roundtable reported in December 2017 that Google's John Mueller said a page left on noindex for a long time is removed completely and its links are no longer followed.
  • Google defines crawl budget as the set of URLs it can and wants to crawl, and addresses its guide to sites with more than a million pages or more than 10,000 pages that change daily.
  • Before paying for a link, check the raw and rendered HTML, the robots meta tag, the X-Robots-Tag header, the canonical and the site's robots.txt.

A crawlable link is one that a search engine’s crawler can read in the page’s code and follow to its destination. Noindex is an instruction, placed in a meta tag or an HTTP header, that tells search engines to keep a page out of their index. Crawl budget is Google’s term for the set of URLs on a site that it can and wants to crawl. The three belong together because a backlink only has a chance of counting when Google can fetch the page it sits on, read the link in that page, and keep the page in its index. A link can be visible to every human visitor and fail any one of those conditions, which is why checking them is part of accepting any link you earned or paid for.

Google’s link best practices page is direct about the markup. It says that generally Google can only crawl a link if it is an a element with an href attribute, and that most links in other formats will not be parsed and extracted by its crawlers. The page then lists what it does not recommend: an a element with a routerLink attribute and no href, a span carrying an href, and an a element that works only through an onclick event. It adds that the address in the href has to resolve into an actual web address, so an href that begins with “javascript:” is also on the not recommended list. Google says it may still attempt to parse these forms, which is not a promise that it will.

For a link builder the consequence is practical. Some publishers build their pages with front-end frameworks, link shorteners, tracking redirects or “click to visit” buttons, and the result can look and behave like a link without being one in the sense Google describes. A button that opens your site through a script is a way for readers to reach you and gives the crawler nothing to follow. A link that passes through the publisher’s own redirect page points, as far as the HTML is concerned, at the publisher, and whether anything reaches you depends on how that redirect is set up, a subject covered in redirects and canonicals. Being crawlable is also separate from being followed. A perfectly formed link can carry a nofollow, sponsored or ugc label, and those link attributes are a second question to settle once the first is answered.

Google’s documentation says links are also crawlable when JavaScript inserts them into a page, as long as they use the same a element and href markup. Its JavaScript SEO guide explains the cost of relying on that. Google processes such pages in three phases, crawling, rendering and indexing. In the first phase Googlebot fetches the HTML and parses it for URLs in the href attributes of links. Pages that return a normal response are then queued for rendering, and only after a headless browser has run the scripts does Google see whatever those scripts added. A link that is present in the raw HTML is found in the first pass. A link that exists only after rendering has to wait for the second, and it is never seen if the page is not rendered.

The same guide contains a trap that matters for the next section. It says Googlebot queues pages for rendering unless a robots meta tag or header tells Google not to index the page, and that when Google meets a noindex tag it may skip rendering and JavaScript execution. A page whose original HTML says noindex, and whose script later removes the tag, may therefore never be rendered at all. From the outside you cannot run Google’s own URL Inspection tool on a publisher’s page, because it works only for sites you have verified. The working method is to compare two things yourself: the page source as delivered by the server, and the page as it stands in the browser’s inspector after scripts have run. If your link appears only in the second, it depends on rendering, and you are entitled to ask the publisher whether the article body can be delivered as plain HTML.

Noindex and robots.txt on the linking page

Google’s page on blocking indexing says that when Googlebot crawls a page and finds a noindex rule, in a robots meta tag or in an X-Robots-Tag response header, it drops the page entirely from search results, whether or not other sites link to it. The two methods have the same effect, and the header version is invisible in the page source, which is how it gets missed. On the question of what happens to links on such a page, the fullest statement is one Search Engine Roundtable reported in December 2017. John Mueller said that if Google sees a noindex for a longer time, it concludes that the page does not want to be used in search and removes it completely, at which point the links on it are no longer followed. In effect a long-standing “noindex, follow” ends up as “noindex, nofollow”. He gave no timescale. Google’s statements on the point have varied over the years, so the cautious reading is the right one for anyone paying for a placement: a page set to noindex is a page whose links you should not expect to count.

Robots.txt works differently and fails differently. Google’s introduction to robots.txt says the file tells crawlers which URLs they can access and is not a mechanism for keeping a page out of Google. A blocked URL can still be indexed if other pages link to it, in which case it appears with no description, because Google has not read the content. For a link buyer the second half of that sentence is the important one. If Google does not crawl the page, it does not read the links on it, so your link is unseen even though the page’s address might turn up in a search. The two rules also interfere with each other. Google states that for a noindex rule to be effective the page must not be blocked by robots.txt, since a crawler that cannot fetch the page never sees the tag. Either rule alone is enough to stop a link from counting, and a canonical tag pointing to a different page is a third route to the same result.

These settings are sometimes mistakes and sometimes deliberate. A publisher that sells sponsored posts may keep them in a folder that is blocked or set to noindex to protect its own standing with Google, without telling the buyer. The lesson on link seller scams describes sellers who hide pages from Google on purpose, and hacked and cloaked links covers the case where Googlebot is served a different page from the one you see. A link bought as a niche edit deserves the same checks as a new article, because an old page can have been set to noindex years ago for reasons nobody remembers.

Crawl budget: what it is and who needs to care

Gary Illyes of Google set out the term in a January 2017 post on Google’s webmaster blog. Crawl budget, he wrote, is the number of URLs Googlebot can and wants to crawl, and it combines two things. The crawl rate limit is how much fetching a server can take without slowing down for its visitors. Crawl demand is how much Google wants the content, and the post says popular URLs tend to be crawled more often and that Google tries to stop URLs going stale in its index. The same post says crawl budget is not something most publishers have to worry about, and that a site with fewer than a few thousand URLs will usually be crawled efficiently. It also says that crawling is necessary for being in the results and is not a ranking signal.

Google’s current guide to crawl budget keeps the definition, the set of URLs that Google can and wants to crawl, and is addressed to large sites of more than a million unique pages whose content changes about weekly, and to sites of more than 10,000 pages whose content changes daily. For your own site, unless it is that size, crawl budget is very unlikely to be what stands between you and a ranking. Where it touches off-page work is on the publisher’s side, and here the reasoning is inference and not a published rule. The 2017 post lists the low-value URLs that drain crawling, including faceted navigation, duplicate content, soft error pages, hacked pages and low-quality or spam content. A publisher that produces large numbers of thin pages, with few internal links pointing to each new one, is giving Google little reason to fetch them quickly. That is one reason a bought link can sit undiscovered for weeks, and it is covered from the practical side in backlink indexing and indexers.

The checks take a few minutes per page, and they are best done before payment is released, when a problem is still the publisher’s to fix. Run them on the live URL of the page that carries your link.

  1. View the page source and find your link as an a element with an href that points straight to your URL.
  2. If the link is missing from the source, look for it in the browser’s inspector, and note that it depends on rendering.
  3. Search the source for a robots meta tag and read its content for noindex or nofollow.
  4. Read the response headers for an X-Robots-Tag line.
  5. Check that the canonical tag points to the page itself.
  6. Open the site’s robots.txt and confirm that the page’s folder is not disallowed for Googlebot.
  7. Confirm that the page is reachable through the site’s own navigation or category pages.
  8. Search Google for the exact URL a few weeks later to see whether it has been indexed.

The backlink status checker on this site does the first group of checks in one pass for up to ten pages: whether the link is there, whether it is followed and whether the page can be indexed. The publisher page checker is meant for the step before that, when you are still deciding whether to buy, and shows how many sites a page links out to and how those links are marked. Step seven is the one people skip. A page that no other page on the site links to depends on a sitemap or on luck to be found, and the reasons are set out in the lesson on internal links and orphan pages. A full acceptance routine, with what to do when a check fails, is in auditing vendor deliveries.

What the checks cannot tell you

Passing every check shows that Google is able to count the link. It does not show that Google does. An indexed page with a clean, followed link can still be discounted for reasons that have nothing to do with crawling, such as the site’s record of selling links, and the lesson on whether links from pages without traffic count goes into how far indexing can be trusted as a test. The reverse also holds. A link missing from the Links report in Search Console has not necessarily been ignored, because that report shows a sample. There is no public figure for what share of placed links sit on pages Google cannot index, so nobody can tell you how common the problem is, only how to look for it.

The other limit is time. Every item on the checklist describes the page on the day you looked. Publishers redesign their sites, move old articles into archives, change robots.txt rules and switch template settings, and any of those can turn a working link into one that no longer counts without the link itself being touched. A link that mattered enough to pay for is worth checking again on a schedule, and how to track backlinks sets out which checks to repeat and how often.

Common questions

What is a crawlable link?

It is a link Google's crawler can read and follow. Google's documentation says that generally means an a element with an href attribute pointing to a real URL. Links that work only through a script event may not be parsed.

Do links on a noindex page count?

Probably not for long. John Mueller of Google was reported in 2017 as saying that a page kept on noindex is eventually removed completely and its links stop being followed. Google's comments on this have varied, so treat a noindexed page as a link that will not count.

Does robots.txt stop a link from counting?

If robots.txt blocks the page your link sits on, Google does not crawl that page's content, so it cannot see the link. The blocked URL can still appear in search results without a description, which sometimes hides the problem.

What is crawl budget?

Google defines it as the set of URLs it can and wants to crawl on a site. It combines how much crawling the server can take with how much Google wants the content.

Does crawl budget matter for a small site?

Rarely. Google's Gary Illyes wrote in 2017 that a site with fewer than a few thousand URLs will usually be crawled efficiently, and Google's current guide is addressed to very large or very fast-changing sites.

Vendors to look at

  • BazoomEditor's pick

    Sponsored content and link marketplace, managed service

  • MotherlinkEditor's pick

    Backlink services, guest posts, niche edits, full SEO