12 July 2022

5 things you didn't know about Google Bot

Googlebot needs to crawl your website before users see it in search results. While this is an essential step, it doesn't get the same attention as many other topics. I think it's partly because Google doesn't share a lot of information about how exactly Googlebot crawls the web.

Google Bot Banners

Seeing that many of our customers have a hard time crawling and indexing their websites properly, we went through some Google documentation on crawling, rendering, and indexing to better understand the whole process.

Some of our results were extremely surprising, while others confirmed our previous theories.

Here are 5 things I learned that you might not know about how Googlebot works.

1. Googlebot skips some URLs

Googlebot won't visit every URL it finds on the web. The larger a website, the greater the risk that some of its URLs won't be crawled and indexed.

Why doesn't Googlebot just visit every URL it can find on the web? There are two reasons for this:

  1. Google has limited resources. There is a lot of spam on the web, so Google needs to develop mechanisms to avoid visiting low-quality pages. Google prioritizes crawling the most important pages.
  2. Googlebot is designed to be a good citizen of the web. Limit the scan to avoid server crash.

The mechanism for choosing which URLs to visit is described in Google's patent “ Method and apparatus for managing a backlog of pending URL crawls ”:

The pending URL scan is rejected from the backlog if the priority of the pending URL scan does not exceed the priority threshold

Various criteria are applied to requested URL scans, so that less important URL scans are rejected early from the backlog data structure.

These quotes suggest that Google is assigning a crawl priority to each URL and may reject crawling of some URLs that do not meet the priority criteria.

The priority assigned to URLs is determined by two factors:

  1. The popularity of a URL,
  2. The importance of crawling a given URL to keep the Google index fresh.

Priority may be higher based on the popularity of the content or IP address/domain name and the importance of maintaining the freshness of rapidly evolving content such as breaking news. Since crawl capacity is a scarce resource, crawl capacity is conserved with priority scores.

What exactly makes a URL popular? Google's patent, " Minimizing the visibility of obsolete content in web search, including reviewing web crawl intervals for documents, " defines URL popularity as a combination of two factors: view rate and PageRank.

PageRank is also mentioned in this context in other patents, such as Scheduler for search engine crawler.

But there is one more thing you should know. When your server responds slowly, the priority threshold that your URLs must meet increases.

The priority threshold is adjusted based on an updated probability estimate of satisfying requested URL crawls. This probability estimate is based on the estimated fraction of requested URL crawls that can be satisfied. The fraction of requested URL crawls that can be satisfied has as its numerator the average request interval, or the difference in arrival time between URL crawl requests.

To sum it up, Googlebot may skip crawling some of your URLs if they don't meet a priority threshold based on the URL's PageRank and the number of views it gets.

This has strong implications for any large website.

If a page is not crawled, it will not be indexed and will not appear in search results.

To do:

  1. Make sure your server and website are fast.
  2. Check your server logs. They provide you with valuable information on which pages of your website are crawled by Google.

 

2. Google divides pages into levels for re-crawling

Google wants search results to be as fresh and up to date as possible. This is only possible when a mechanism is in place to rescan already indexed content.

In the patent “ Minimizing the visibility of obsolete content in web search ” I found information on how this mechanism is structured.

Google is dividing pages into tiers based on how often the algorithm decides they should be re-ranked.

In one embodiment, documents are partitioned into multiple levels, each level including a plurality of documents that share similar web crawl ranges.

Therefore, if your pages aren't scanned as often as you want, they are most likely in a document layer with a longer scan interval.

However, don't despair! Your pages don't need to stay in that layer forever - they can be moved.

Each time a page is crawled it is an opportunity for you to prove that it is worth re-crawling more frequently in the future.

After each scan, the search engine reevaluates the web crawl range of a document and determines whether the document should be moved from the current level to another level .”

It is clear that if Google sees that a page changes frequently, it may be moved to a different level. But it's not enough to change some minor aesthetic elements: Google is analyzing both the quality and quantity of changes made to your pages.

To do:

  1. Use your server logs and Google Search Console to know if your pages are being crawled often enough.
  2. If you want to reduce the crawl interval of your pages, regularly improve the quality of your content.

 

3. Google doesn't re-index a page on every crawl

According to the patent Minimizing the visibility of outdated content in web search, including reviewing web document crawl intervals , Google does not reindex a page after each crawl.

If the document has been substantially changed since the last crawl, the scheduler sends an alert to a content indexer (not shown), which replaces the index entries for the previous version of the document with index entries for the current version of the document. The scheduler then calculates a new web crawl interval for the document based on its old interval and additional information, such as the document's importance (measured by a score, such as PageRank), update frequency, and/or click-through rate. If the document's content has not changed, or if the content changes are not critical, the document does not need to be reindexed.

I have seen it in nature several times.

Also, I did some experiments on existing pages on Onely.com. I noticed that if I was only changing a smart part of the content, Google wasn't re-indexing it.

 

To do:

If you have a news website and you frequently update your posts, check if Google re-indexes it quickly enough. If not, you can rest assured that there is untapped potential in Google News for you.

 

4. Click-through rate and internal link

In the previous quote, did you notice how click-through rate was mentioned?

The scheduler then calculates a new web crawl interval for the document based on its old interval and additional information, such as the document's importance (measured by a score, such as PageRank), update frequency, and/or click-through rate .”

This quote suggests that click-through rate affects a URL's crawl rate.

Let's imagine we have two URLs. One is visited by Google users 100 times a month, another is visited 10000 times a month. All other things being equal, Google should revisit the one with 10000 visits per month more frequently.

According to the patent, PageRank is also an important part of this. This is yet another reason to ensure you're properly using internal linking to connect different parts of your domain.

 

To do:

  • Can Google and users easily access the most important sections of your website?
  • Is it possible to reach all the important URLs? Having all your URLs available in the sitemap may not be enough.

 

5. Not all links are created equal

We just explained how, according to Google's patents, PageRank heavily affects crawl.

The first implementation of the PageRank algorithm was unsophisticated, at least judging by current standards. It was relatively simple: if you received a link from an * important * page, you would rank higher than other pages.

However, the first PageRank implementation was released over 20 years ago. Google has changed a lot since then.

I've found some interesting patents, such as the papers on Ranking Based on User Behavior and/or Feature Data , which demonstrate that Google is well aware that some links on a given page are more important than others. Furthermore, Google may treat these links differently.

“This reasonable navigation model reflects the fact that not all links associated with a document are equally likely to be followed. Examples of unlikely links may include links to "Terms of Service", banner ads and unrelated links to the document. "

So Google is analyzing the links based on their various characteristics. For example, they can examine the font size and position of the link.

" For example, the pattern generation unit may generate a rule that indicates that links with anchor text larger than a certain font size are more likely to be clicked than links with anchor text smaller than that particular font size. Additionally, or alternatively, the pattern generation unit may generate a rule that indicates that links positioned closer to the top of a document are more likely to be clicked than links positioned toward the bottom of the document."

It even appears that Google can create rules for evaluating links at the website level. For example, Google can see that links in "More Top News" are clicked more frequently so they can give them more weight.

“(…) the modeling unit may generate a rule indicating that a link placed under the “More Top Stories” heading on the cnn.com website has a high probability of being clicked. Additionally, or alternatively, the modeling unit may generate a rule indicating that a link associated with a destination URL containing the word “domainpark” has a low probability of being clicked. Additionally, or alternatively, the modeling unit may generate a rule indicating that a link associated with a source document containing a pop-up has a low probability of being clicked.”

As a side note, in a conversation with Barry Schwartz and Danny Sullivan in 2016 , Gary IIIyes confirmed that Google labels links, such as footer or penguin.

Basically, we have tons of link labels; for example, a footer link, basically, has much less value than an in-content link. So another label would be a Penguin real-time label .”

Summarizing the key points:

  • Google is prioritizing every crawled page
  • The faster the website, the faster Google crawls.
  • Google will not crawl and index all URLs. Only URLs with priority assigned above the threshold will be crawled.
  • Links are treated differently depending on their characteristics and positioning
  • Google doesn't re-index a page after every crawl. It depends on the severity of the changes made.

In conclusion

As you can see, crawling is far from a simple process of following every link Googlebot can find. It's truly complicated and has a direct impact on every website's search visibility. I hope this article has helped you understand crawling a little better and that you'll be able to use this knowledge to improve the way Googlebot crawls your website and consequently rank it higher. It also shows how important it is, in addition to having a properly structured website and good internal and external link building, to have fast, high-performance hosting and servers to best manage the Googlebot crawling process and thus maximize the profitability of your crawling budget.

Do you have doubts? Don't know where to start? Contact us!

We have all the answers to your questions to help you make the right choice.

Chat with us

Chat directly with our presales support.

0256569681

Contact us by phone during office hours 9:30 - 19:30

Contact us online

Open a request directly in the contact area.

DISCLAIMER, Legal Notes and Copyright. RedHat, Inc. holds the rights to Red Hat®, RHEL®, RedHat Linux®, and CentOS®; AlmaLinux™ is a trademark of the AlmaLinux OS Foundation; Rocky Linux® is a registered trademark of the Rocky Linux Foundation; SUSE® is a registered trademark of SUSE LLC; Canonical Ltd. holds the rights to Ubuntu®; Software in the Public Interest, Inc. holds the rights to Debian®; Linus Torvalds holds the rights to Linux®; FreeBSD® is a registered trademark of The FreeBSD Foundation; NetBSD® is a registered trademark of The NetBSD Foundation; OpenBSD® is a registered trademark of Theo de Raadt; Oracle Corporation holds the rights to Oracle®, MySQL®, MyRocks®, VirtualBox®, and ZFS®; Percona® is a registered trademark of Percona LLC; MariaDB® is a registered trademark of MariaDB Corporation Ab; PostgreSQL® is a registered trademark of PostgreSQL Global Development Group; SQLite® is a registered trademark of Hipp, Wyrick & Company, Inc.; KeyDB® is a registered trademark of EQ Alpha Technology Ltd.; Typesense® is a registered trademark of Typesense Inc.; REDIS® is a registered trademark of Redis Labs Ltd; F5 Networks, Inc. owns the rights to NGINX® and NGINX Plus®; Varnish® is a registered trademark of Varnish Software AB; HAProxy® is a registered trademark of HAProxy Technologies LLC; Traefik® is a registered trademark of Traefik Labs; Envoy® is a registered trademark of CNCF; Adobe Inc. owns the rights to Magento®; PrestaShop® is a registered trademark of PrestaShop SA; OpenCart® is a registered trademark of OpenCart Limited; Automattic Inc. holds the rights to WordPress®, WooCommerce®, and JetPack®; Open Source Matters, Inc. owns the rights to Joomla®; Dries Buytaert owns the rights to Drupal®; Shopify® is a registered trademark of Shopify Inc.; BigCommerce® is a registered trademark of BigCommerce Pty. Ltd.; TYPO3® is a registered trademark of the TYPO3 Association; Ghost® is a registered trademark of the Ghost Foundation; Amazon Web Services, Inc. owns the rights to AWS® and Amazon SES®; Google LLC owns the rights to Google Cloud™, Chrome™, and Google Kubernetes Engine™; Alibaba Cloud® is a registered trademark of Alibaba Group Holding Limited; DigitalOcean® is a registered trademark of DigitalOcean, LLC; Linode® is a registered trademark of Linode, LLC; Vultr® is a registered trademark of The Constant Company, LLC; Akamai® is a registered trademark of Akamai Technologies, Inc.; Fastly® is a registered trademark of Fastly, Inc.; Let's Encrypt® is a registered trademark of the Internet Security Research Group; Microsoft Corporation owns the rights to Microsoft®, Azure®, Windows®, Office®, and Internet Explorer®; Mozilla Foundation owns the rights to Firefox®; Apache® is a registered trademark of The Apache Software Foundation; Apache Tomcat® is a registered trademark of The Apache Software Foundation; PHP® is a registered trademark of the PHP Group; Docker® is a registered trademark of Docker, Inc.; Kubernetes® is a registered trademark of The Linux Foundation; OpenShift® is a registered trademark of Red Hat, Inc.; Podman® is a registered trademark of Red Hat, Inc.; Proxmox® is a registered trademark of Proxmox Server Solutions GmbH; VMware® is a registered trademark of Broadcom Inc.; CloudFlare® is a registered trademark of Cloudflare, Inc.; NETSCOUT® is a registered trademark of NETSCOUT Systems Inc.; ElasticSearch®, LogStash®, and Kibana® are registered trademarks of Elastic NV; Grafana® is a registered trademark of Grafana Labs; Prometheus® is a registered trademark of The Linux Foundation; Zabbix® is a registered trademark of Zabbix LLC; Datadog® is a registered trademark of Datadog, Inc.; Ceph® is a registered trademark of Red Hat, Inc.; MinIO® is a registered trademark of MinIO, Inc.; Mailgun® is a registered trademark of Mailgun Technologies, Inc.; SendGrid® is a registered trademark of Twilio Inc.; Postmark® is a registered trademark of ActiveCampaign, LLC; cPanel®, LLC owns the rights to cPanel®; Plesk® is a registered trademark of Plesk International GmbH; Hetzner® is a registered trademark of Hetzner Online GmbH; OVHcloud® is a registered trademark of OVH Groupe SAS; Terraform® is a registered trademark of HashiCorp, Inc.; Ansible® is a registered trademark of Red Hat, Inc.; cURL® is a registered trademark of Daniel Stenberg; Facebook®, Inc. owns the rights to Facebook®, Messenger® and Instagram®. This site is not affiliated with, sponsored by, or otherwise associated with any of the above-mentioned entities and does not represent any of these entities in any way. All rights to the brands and product names mentioned are the property of their respective copyright holders. All other trademarks mentioned are the property of their respective registrants. MANAGED SERVER® is a European registered trademark of MANAGED SERVER SRL, with registered office in Via Flavio Gioia, 6, 62012 Civitanova Marche (MC), Italy and operational headquarters in Via Enzo Ferrari, 9, 62012 Civitanova Marche (MC), Italy.

JUST A MOMENT !

Have you ever wondered if your hosting sucks?

Find out now if your hosting provider is hurting you with a slow website worthy of 1990! Instant results.

Close the CTA
Back to top