The short version
- A crawl is a simulation of what a bot could reach. Server logs are a record of what bots actually requested.
- Logs are the only way to see crawl frequency, wasted crawl on junk URLs, and pages bots never visit at all.
- For most small sites logs are overkill. For large sites they answer questions nothing else can.
- Search Console crawl stats give you a free, simplified version of the same picture.
Almost all technical SEO is done on crawl data, because crawling is easy and logs are awkward to get. That works until you hit a question about behaviour rather than structure — why a section never gets indexed, whether a crawler is burning its budget on faceted URLs, how often your important pages actually get revisited.
Those are questions about what happened, and only the server knows what happened.
The short answer
Fit for your workflow
Crawl data for structure. Logs for behaviour. Start with Search Console crawl stats.
Use a crawler to find broken links, duplicate titles, redirect chains and markup problems — the structural faults. Use logs when you need to know what crawlers are actually doing: which URLs they hit, how often, and what they are wasting requests on. Before setting up log analysis, check the crawl stats report in Search Console, which gives you a free simplified view that answers many of the same questions.
- Small site under a few thousand URLs
- Crawl data only. Logs will not tell you anything actionable.
- Large ecommerce with faceted navigation
- Logs. This is where crawl waste hides.
- Section not getting indexed
- Logs — find out whether bots ever request it.
- No log access
- Search Console crawl stats. Free and often sufficient.
What each one can tell you
The sixth row is the one that surprises people. A crawler starting from your homepage will find every page reachable by internal links, and will happily report them as fine. Logs can show that a search engine has never actually requested a third of them — which is a completely different problem with a completely different fix, and one no crawl will ever surface.
What to look for in logs
- Crawl distribution. What proportion of requests go to pages that matter versus parameters, filters and pagination?
- Status codes served to bots. A pattern of 404s or 5xx responses to crawlers is invisible in a clean crawl from your own machine.
- Crawl frequency on key pages. Important pages crawled rarely is a signal about perceived importance.
- Never-crawled URLs. Cross-reference your sitemap against log entries to find what bots have simply not visited.
- Verify the bots are real. Reverse DNS lookup on claimed Googlebot requests — spoofing is common and inflates your numbers.
One caveat about getting the data at all: logs are often harder to obtain than to analyse. Managed hosting may not expose them, a CDN in front of your origin will hold a different and more complete picture than the origin itself, and retention windows are frequently short. Establish where your logs actually live and how far back they go before planning any analysis around them.
Start with the free version
The crawl stats report in Search Console shows total crawl requests over time, response codes, file types and host status — no log access, no parsing, no tooling. For most sites it answers the crawl-health question adequately. Set up full log analysis when it tells you something is wrong and you need to know which URLs.