GI
GPT Insights
Technical GEO Check
Common Crawl Decoder

Technical GEO Check

How reachable is your website for AI crawlers?

Check historical Common Crawl signals, technical access restrictions and unusual changes, an important technical prerequisite for Generative Engine Optimization (GEO).

12 months are analyzed automatically. Protocol, www and subdirectories are removed from your input.

Up to 24 months of history Automatic drop detection Inspect URLs and status codes
What this check shows and what it does not

The Common Crawl Decoder checks whether and how regularly your domain has been captured in Common Crawl, whether unusual changes occur and whether current technical rules could restrict access for relevant crawlers.

What this check can and cannot tell you

The check shows

  • whether Common Crawl has captured your domain
  • how the number of captured pages develops over time
  • when unusual drops occur
  • which URLs and status codes were found in individual crawls
  • whether technical blocks should be examined as a possible cause
  • whether the observed crawl presence matches a deliberate or possibly unintended access policy

The check does not show

  • whether content was actually used to train a particular model
  • whether a particular AI system is crawling your website right now
  • whether your brand is visible in generative answers
  • how content is weighted or interpreted by a model
More on methodology and interpretation

Crawler blocks can be set deliberately, for example to protect content from unwanted automated use. They also arise unnoticed through hosting settings, CDN rules, firewalls or security plugins. For ordinary visitors the website often remains fully reachable.

Common Crawl is not the final word on AI training or AI visibility. The data is, however, a good proxy for a quick technical GEO check.

Whether an access restriction makes sense depends on a publisher's content, data and AI strategy. Some websites deliberately open their content to AI systems. Others deliberately limit training, search access or automated use. What matters is that the technical implementation matches the intended strategy.

Examples: what invisible crawler blocks look like

What invisible crawler blocks look like

A drop in Common Crawl data proves neither a technical fault nor an unintended block. The change can also go back to a deliberately chosen access restriction. It is, however, a strong signal to review robots.txt, CDN, hosting, firewall and security rules and to compare them with your own GEO strategy.

Abrupt stop

The domain was captured regularly over several months. After that the number of pages found drops almost completely. Such a pattern can point to a new robots.txt rule, a CDN configuration, a firewall or a change in hosting.

Possible technical block

Gradual decline

Crawl presence declines over several months. Possible causes are increasing error pages, redirect chains, changes to the site structure or gradually reduced reachability.

Review technical development
Background: causes and interpretation
Not every crawler block is wrong. What is wrong above all is not knowing that it exists, or what it can mean for your GEO strategy.

Technical foundation for GEO

Generative systems need content that is technically reachable and machine readable. Common Crawl does not represent every AI data source, but it gives a quick indication of whether a website arrives in an important open web dataset.

Deliberate and unnoticed blocks

Crawlers can be blocked deliberately in order to protect content and data. Blocks also arise unintentionally through robots.txt, CDN rules, firewalls, hosting settings or security plugins. For ordinary visitors the website often remains reachable.

Spot drops early

A sudden decline in captured pages can point to technical changes or new blocks. The timeline helps to narrow down the possible trigger.

Narrow down causes

Open individual months and inspect the captured URLs and their status codes. Frequent errors, redirects or blocked requests can indicate the cause.

More on this topic

In-depth reference material in the GPT Insights GEO Wiki.

Understand crawler accessibility

How real technical reachability differs from a declared crawl policy.

User-Agent and WAF-Based Access Restrictions

How user-agent rules, WAFs and challenges affect access for automated clients.