Главная / Блог и руководства / Technical & Code / robots.txt and Crawler Access: What Webs...
Technical & Code

robots.txt and Crawler Access: What Website Owners Must Check for AdSense [ru]

Автор: Rankzio Editorial Team • Опубликовано: September 21, 2026 • ⏱️ 4 min read
robots.txt and Crawler Access for Website Owners

The Gatekeeper of Your Website

Every website accessible on the public internet is visited daily by hundreds of automated software agents known as web crawlers, spiders, or bots. Search engines like Google rely on crawlers (such as Googlebot) to index web pages, while advertising networks like Google AdSense utilize dedicated crawlers (primarily Mediapartners-Google) to inspect site content, determine contextual topics, and evaluate publisher policy compliance.

The gatekeeper that governs where these crawlers may and may not go is a simple text file located at your domain root: robots.txt. Although robots.txt is one of the oldest standards on the web, a single misplaced character or poorly conceived disallow directive can silently lock out review bots, causing instant AdSense rejections or preventing ads from rendering on your published pages.

Meet Mediapartners-Google: The AdSense Review Crawler

Many publishers mistakenly assume that because their articles appear in Google Search, their site must be completely open to Google AdSense. This assumption overlooks a critical architectural detail: Google Search and Google AdSense utilize distinct crawlers with separate user-agent signatures.

  • Googlebot crawls the web to build search engine index listings.
  • Mediapartners-Google crawls pages specifically to evaluate AdSense applications and match contextual ad creatives to page content.

If your robots.txt file permits Googlebot but blocks Mediapartners-Google, your articles will rank on Google Search, but your AdSense review will fail with an error stating that Google was unable to crawl your site. Understanding this distinction is crucial for every webmaster.

The Anatomy of a robots.txt File

A standard robots.txt file consists of directive blocks composed of three primary elements:

User-agent: [bot name or wildcard *]
Disallow: [path to restrict]
Allow: [path to explicitly permit]

1. The Wildcard Disallow Rule (The Danger Zone)

When web developers build a site on a staging server or test environment, they often include this rule to prevent premature indexing:

User-agent: *
Disallow: /

The forward slash (/) signifies your entire website. If you migrate your site to production and forget to remove this rule, all crawlers—including Googlebot and Mediapartners-Google—will obey the instruction and refuse to crawl any page on your site. This is one of the most common causes of instant "Site Not Ready" AdSense rejections.

2. The Recommended robots.txt Configuration for AdSense

For most informational blogs, content sites, and CMS installations (like WordPress), your robots.txt should be clean, lightweight, and welcoming to legitimate search bots. Here is an optimal configuration:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

# Explicitly welcome the AdSense crawler
User-agent: Mediapartners-Google
Disallow:

Sitemap: https://yourdomain.com/sitemap.xml

In this configuration, leaving the Disallow: field empty under User-agent: Mediapartners-Google explicitly signals to Google’s ad review bot that it has unrestricted access to crawl every public section of your site.

Common robots.txt Mistakes That Hurt AdSense

1. Blocking CSS, JavaScript, or Theme Assets

In the early days of the web, webmasters routinely blocked directories like /wp-content/ or /includes/ to save bandwidth. Today, Google crawlers render web pages like modern web browsers. If your robots.txt disallows access to CSS stylesheets or JavaScript files, Google cannot render your mobile layout or verify that your site has responsive design, leading to mobile usability warnings.

2. Blocking Staging Parameter URLs Without Canonical Tags

If your robots.txt attempts to block search query strings or filter parameters, ensure you do not inadvertently block core category pages or paginated article links.

3. File Encoding and Syntax Errors

Your robots.txt file must be saved in plain ASCII or UTF-8 text without a Byte Order Mark (BOM). Never use a rich-text editor (like Microsoft Word) to create robots.txt, as smart quotes or invisible formatting characters will invalidate the file.

How to Test Your robots.txt File

  1. Direct URL Verification: Open https://yourdomain.com/robots.txt in your web browser. Ensure the file loads cleanly as plain text without redirecting to a 404 error page.
  2. Google Search Console robots.txt Tester: Use the robots.txt report inside Google Search Console to check whether specific article URLs are permitted.
  3. Rankzio Automated Scan: Run your domain through Rankzio’s free AdSense Readiness Checker. Our engine inspects your live robots.txt, checks for crawler blocks, and alerts you if Mediapartners-Google is restricted.

For related technical indexing checks, explore our companion tutorial on Noindex Tags Explained Before You Apply for AdSense.

📋 Чек-лист внедрения

  • ✔ robots.txt exists at root domain (https://yourdomain.com/robots.txt)
  • ✔ No blanket Disallow: / directive blocking all crawlers
  • ✔ Mediapartners-Google explicitly permitted to crawl content
  • ✔ CSS and JavaScript asset files are not blocked from crawlers
  • ✔ XML Sitemap location declared at the bottom of the file
  • ✔ File saved in plain UTF-8 text without Byte Order Mark (BOM)
  • ✔ Verified crawler reachability with Rankzio Readiness Checker

Test Your Website Against 21 AdSense Signals

Use Rankzio’s free AdSense Readiness Checker to inspect your crawlability, legal disclosures, and content depth in seconds.

Check My Website Free →

Часто задаваемые вопросы

If a website has no robots.txt file, crawlers treat the absence as an open invitation to crawl all public pages. However, having a clean, properly configured robots.txt is considered a web development best practice and allows you to declare your XML sitemap.

Google generally caches robots.txt files for up to 24 hours. If you update the file, you can request a refresh via Google Search Console.

No. robots.txt is an advisory protocol respected by legitimate search engines. Malicious bots and scrapers routinely ignore robots.txt directives.

If our scanner detects Disallow: / under User-agent: *, it issues an urgent warning because this rule halts both AdSense and search engine indexation.

Related Publishing Guides

ads.txt Explained for New AdSense Publishers
How to Find and Check AdSense Code on Your Website
Noindex Tags Explained Before You Apply for AdSense
← Ко всем статьям