Crawler.

In a nutshell

A crawler is an automated program that retrieves web pages, follows links, and collects content for search engines or AI services—such as Googlebot for Google Search.

As of October 5, 2026 · Dustin Tatarowicz, marschfahrt

A crawler, also known as a bot or spider, is a program that automatically visits websites, downloads their content, and follows links to other pages. According to Google, the vast majority of pages in search results are found automatically in this way, without anyone having to submit them. In practical terms for your business, this means that anything the crawler can’t reach won’t appear on Google. In addition to search engines, AI providers are now also sending their own crawlers out across the web.

How a Crawler Works

Google's crawler is called Googlebot, and according to Google, it discovers new URLs primarily through links on pages it already knows. It retrieves the page and, according to Google, runs the JavaScript using a current version of Chrome. Because Google indexes the mobile version of most websites, the majority of requests come from the smartphone crawler, according to Google.

It’s important to understand the difference between crawling and indexing. Crawling is the process of retrieving content, while indexing is the process of understanding and storing it in the Google Index. According to Google, neither is guaranteed. You use the robots.txt file to tell crawlers which areas they are allowed to access. However, Google emphasizes that robots.txt is not a way to keep a page out of search results: if it’s linked to from elsewhere, the URL may still appear. The “noindex” directive or password protection are intended for this purpose.

What Your Business Should Keep in Mind

Link to every important page from your website. A service page that isn't linked to by either the menu or the text will be hard even for Googlebot to find. A sitemap is an additional help, but it doesn't replace a well-organized linking structure.

Check your robots.txt file before and after making any changes. In projects, we often see that a block set during the development phase is simply left in place after the site goes live. Search Console shows you whether Google can access and index your pages. To learn how to interpret this information, see “How to Use Google Search Console Correctly.”

Make a conscious decision about which AI crawlers you allow. According to its own documentation, OpenAI distinguishes between GPTBot, which collects content for training AI models, and OAI-SearchBot, which enables pages to appear in ChatGPT search results. Anthropic makes a similar distinction between ClaudeBot for training data and Claude-SearchBot for search results. Google offers its own tag, Google-Extended, which, according to Google, allows you to control usage for Gemini without affecting your visibility in Google Search. If you block training crawlers, you should generally leave search crawlers enabled.

Frequently Asked Questions About Crawlers

Can I pay Google to have my site crawled more often? No. Google explicitly states that it does not accept money to crawl a website more frequently or to rank it higher.

Does Googlebot put a strain on my server? Generally speaking, no. According to Google, Googlebot crawls most websites no more than once every few seconds, on average. If it becomes too much, you can reduce the crawling frequency.

Do all crawlers follow the robots.txt file? No. According to Google, reputable crawlers follow the rules, while others do not. Furthermore, according to Google, some programs falsely claim to be Googlebot. That’s why confidential information should be password-protected.

Our SEO check will show you whether your pages are accessible to crawlers and have clean titles and descriptions. If you don't want to tackle these issues yourself, our SEO support is here to help.

Can Google and AI bots even read your pages? The SEO check will show you—for free.

Go to the SEO Check →
← All terms in the glossary