In plain English
A crawler is a program that fetches pages, extracts links and follows them. It identifies itself with a user agent string, which is how you can tell one from another in your logs.
The set of crawlers that matter has widened. Alongside search engine bots there are now AI training and retrieval crawlers, and deciding which to allow is a real editorial decision.
What to know
Why it matters
Knowing which crawlers visit and what they do lets you make deliberate decisions: which to welcome, which to slow down, and whether AI crawlers should be allowed to use your content. That is now a business question, not just a technical one.
Common mistakes
FAQs
Should I block AI crawlers?
It depends on whether you want to appear in AI answers. Blocking removes the content and the citation.
Do all crawlers obey robots.txt?
Reputable ones do. It is a request, not an enforcement mechanism.
I'm a marketing consultant, entrepreneur and content creator. I help businesses grow through practical marketing, websites, SEO, content and AI.
More About Tariq →