Skip to the content

Site Crawler — Complete Tool Guide

Why This Tool Exists

Every website eventually accumulates problems that no one notices by browsing it normally: a batch of pages quietly returning a 404, two blog posts with the exact same title tag, a product page that got orphaned when a menu was redesigned last year, or a whole city-page template that Google might treat as duplicate content. Finding these issues requires actually crawling the site page by page, the way a search engine would — and the tool that’s defined this category for over a decade, Screaming Frog, is a desktop application that has to be downloaded, licensed for anything beyond 500 URLs, and run locally on a machine that stays on for the length of the crawl.

That’s a real barrier for a huge number of people who need this exact kind of audit: a small business owner checking their own site, a freelancer auditing a client’s site from a laptop that can’t have software installed, an agency team member who wants to run a quick check without opening the desktop app, or anyone working from a Chromebook or a phone. The Site Crawler was built to close that gap — a full crawling and scoring engine that runs entirely in the browser, reachable from any device with an internet connection, with a free tier generous enough (500 pages) to cover the large majority of small and mid-sized business websites.


Google.com Site Crawl for 100 Pages

SEO Site Crawl - Full Guide

The Gap: What Existing Tools Don’t Do

Screaming Frog (the industry standard) It’s thorough and trusted, but it’s a desktop install, Java-based, and requires a paid license to unlock more than 500 URLs, custom extraction, and several other features. It has to run on a machine that’s turned on and connected for the whole crawl — not something you can kick off from a phone or a shared/locked-down work laptop.

Free online “SEO checker” tools Most single-page checkers only look at one URL at a time — they don’t spider a site, don’t build a duplicate-content picture across pages, and don’t track crawl depth or orphan pages, because that requires actually following links, not just scanning one page’s HTML.

Google Search Console Search Console shows you what Google has already indexed and its own crawl errors, but it’s reactive — it tells you what Google noticed, often weeks after the fact, and it doesn’t run the same real-time technical checks (title pixel width, near-duplicate clustering, schema property completeness) that a proper crawl does.

GEO / AI-visibility tools This is a newer need: as ChatGPT, Claude, Perplexity, and Gemini increasingly answer questions by pulling from live websites, whether those AI crawlers are even allowed to reach a site matters as much as traditional SEO. Almost no consumer-facing crawler checks this — most SEO tools were built before this became relevant and still only report on Googlebot.

Authenticated / staging sites Every server-side crawler — including Screaming Frog running through a proxy, and every hosted “enter a URL” tool — hits the same wall on a staging site behind a login, Cloudflare Access, or a VPN: the server has no browser session, so it just gets blocked. Nobody has a clean way to audit a site before it goes live without exporting cookies or fighting with VPN configs.

The common gap: nothing combines a full multi-page crawl, real SEO scoring across dozens of checks, near-duplicate detection, and now AI-crawler readiness, in a tool that needs no install, no license, and works from any browser — including, via a one-click bookmarklet, sites that require a login.


What We’re Providing: A Full Breakdown

1. Two Crawl Modes

Spider mode starts at one URL and follows every internal link it finds, exactly like a search engine would, up to a configurable depth and page limit. List mode crawls only a pasted list of specific URLs — useful for auditing a known set of pages (a migration list, a client’s priority pages) without touching the rest of the site.

2. 28 Real SEO Checks, Organized Into Six Categories

Every crawled page is scored against checks grouped into: Crawl & Links (status code, redirects, inlinks, crawl depth, response time), Snippet quality (title length in both characters and actual rendered pixel width, description length and pixel width, duplicate titles/descriptions across the whole crawl), Content & structure (single H1, heading hierarchy gaps, word count, Flesch readability score, boilerplate ratio), Indexing (status codes, canonical match, noindex directives from meta tags or HTTP headers, robots.txt blocking, soft-404 detection), Schema & social (JSON-LD presence and validity against Google’s actual required properties per type, Open Graph and Twitter Card coverage), and Technical hygiene (mixed content, missing alt text, viewport/charset/favicon, HTML lang attribute).

3. Pixel-Width Title and Description Checks — Not Just Character Counts

Most free tools only check title/description character counts (30-60, 70-160), but Google actually truncates based on rendered pixel width, and different letters take up different amounts of space. This tool calculates approximate pixel width for both title (≤580px) and description (≤920px), which is a meaningfully more accurate signal for whether a snippet will actually get cut off in search results.

4. Near-Duplicate Content Clustering (Not Just Exact-Match Duplicates)

Beyond flagging pages with byte-identical content, the crawler builds a 32-bit SimHash fingerprint of each page’s text (using word 4-shingles) and clusters pages whose fingerprints differ by only a few bits — the classic signature of templated or programmatic pages that only swap a city name, category, or product variant. This is genuinely rare in free tools and catches a category of duplicate-content risk that exact-match hashing completely misses.

5. Schema Validation Against Real Google Requirements

Rather than just checking whether JSON-LD exists, the tool validates each schema type against the actual properties Google requires to show rich results — for example, a Product needs image, offers.price, offers.priceCurrency, and offers.availability; a Review needs author.name and reviewRating.ratingValue. It covers over 25 schema types directly (LocalBusiness, Product, Review, Article, Event, Recipe, JobPosting, and more) plus dozens of common subtypes (Restaurant, Dentist, Electrician, HVACBusiness, and other LocalBusiness variants) that fall back to their base type’s requirements — built specifically with general small-business and local-service sites in mind.

6. GEO / AI Readiness Scoring

The tool checks whether a defined list of AI crawlers — GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot, Bytespider — are allowed or blocked in robots.txt, whether an llms.txt file exists, and whether a sitemap is declared. These feed into a weighted AI Visibility Score alongside schema completeness and answer-friendly content structure (headings phrased as questions). This directly addresses a blind spot in almost every other SEO tool: whether a site can even be found and cited by the AI assistants that are increasingly replacing traditional search for many users.

7. “Fix These First” — Plain-Language Priority List

Rather than handing over a wall of 28 checks and letting the reader figure out what matters, the tool surfaces the top five highest-impact issues in plain English — for example, “12 pages are hidden from Google (marked ‘noindex’) — people searching won’t find them at all,” or “Your site blocks 3 recognized AI tools (like ChatGPT or Claude) from reading it.” This is aimed squarely at the non-technical business owner who needs to know what to act on, not just a technical audit.

8. Orphan Page Detection

After the main crawl, an opt-in step reads the site’s XML sitemap(s) — including sitemap index files — and flags any URL that’s listed in the sitemap but was never reached by following links during the crawl. These are genuinely invisible pages: Google knows they exist because the sitemap says so, but nothing on the site actually links to them, which weakens their ranking potential.

9. Staging and Password-Protected Site Support

Since this crawler runs server-side, it hits the same wall as any outside visitor when a site is behind a login, Cloudflare Access, or a VPN. To solve this, there’s a one-time bookmarklet that runs the entire crawl engine directly inside the user’s own already-logged-in browser tab — inheriting their real session automatically, with no VPN configuration, no cookie exporting, and no console-pasting required after the first setup. A manual copy-paste console script is also offered as a fallback for anyone the bookmarklet doesn’t work for.

10. Custom Extraction and Source Code Search

Up to three custom fields can be defined with a label and a CSS selector (for example, pulling a price, an author name, or a rating from every page), and a separate regex-based source search can scan every crawled page’s raw HTML for a specific pattern — useful for confirming a tracking tag (like a GA4 or GTM ID) is present sitewide, or finding every page containing a specific phone number or phrase.

11. Crawl Comparison Over Time

A previously saved crawl (exported as JSON) can be loaded alongside a new one to produce a full diff: new pages, removed pages, and changed pages with a field-by-field breakdown of what changed (status, score, title, H1, canonical, noindex, word count), plus a before/after comparison of the AI Visibility Score. This turns one-off audits into an ongoing tracking process without needing a paid rank-tracking platform.

12. Site Visualization

A crawl-depth bar chart and a treemap of the site’s structure by URL segment (sized by page count, colored by average health score) give an at-a-glance picture of where a site’s problems concentrate — useful for spotting, for instance, that an entire /blog/ section is underperforming while /products/ is healthy.

13. Optional Real PageSpeed Data

For anyone willing to get a free Google API key, the tool can run Google’s actual PageSpeed Insights test (mobile) against crawled pages, rate-limited to stay within Google’s quotas, surfacing real Core Web Vitals data (LCP, CLS, TBT) per page rather than an estimate.

14. URL Rewriting, CDN Handling, and Include/Exclude Rules

Advanced crawl configuration lets a user strip tracking parameters, apply regex find-and-replace rules to discovered URLs before they’re crawled, treat specific CDN subdomains as internal, and restrict crawling to only URLs matching specific regex patterns — the kind of fine control usually reserved for paid desktop tools.

15. Full CSV and JSON Export

A completed crawl exports as a detailed CSV scorecard (summary stats, check pass rates, broken links, orphan pages, and full per-page detail across every field) or as a JSON file that can be reopened later or used as the “before” side of a comparison — with no data locked behind a paywall.

16. Free Up To 500 Pages, No License

The entire feature set above is available at no cost for crawls up to 500 pages — enough for the overwhelming majority of small and mid-sized business sites — with a waitlist for larger background crawls, saved crawl history, and white-label reporting for agencies working with bigger sites.


Use Cases in Practice

A business owner auditing their own site can run a full crawl from any browser, get an immediate plain-language “fix these first” list, and understand exactly what’s broken without needing to interpret 28 raw metrics themselves.

A freelancer or agency auditing a client’s site can use custom extraction to pull client-specific fields (prices, ratings), export a full CSV scorecard for a client report, and re-run the crawl monthly using the comparison feature to show measurable progress over time.

An in-house SEO or marketing manager can catch a staging site’s issues before launch using the bookmarklet, well before a redesign goes live and any problems become visible to real visitors and Google at the same time.

A team preparing for the AI-search era can check their AI Visibility Score to see whether GPTBot, ClaudeBot, and other AI crawlers are actually able to reach their site — a check that most SEO audits still don’t run, but which is quickly becoming as important as traditional search visibility.

A content team managing a large templated site (city-based service pages, programmatic SEO) can use near-duplicate clustering to catch templated pages that are similar enough to risk being treated as duplicate content by Google, before it affects rankings.


Getting Started

  1. Choose Spider mode (follow links from a start URL) or List mode (crawl only pasted URLs).
  2. Set a start URL, max pages, and max depth — defaults work well for most sites.
  3. Optionally customize which of the 28 checks to run, add custom extraction fields, or set a source-code search pattern.
  4. Click Start Crawl and watch the live progress bar; pause or stop at any time.
  5. Review the “Fix These First” summary at the top of the results for the highest-priority issues in plain language.
  6. Dig into the full Pages table by category tab (Titles, Schema, Content, and more) for line-by-line detail, or check the GEO/AI Readiness panel to see AI-crawler access.
  7. Run the optional Orphan Pages check against your sitemap, or PageSpeed Insights with your own API key, when ready.
  8. Export a full CSV scorecard, generate an XML sitemap directly from the crawl, or save the crawl as JSON to compare against next month’s results.

This content is generated from Claude based on the tool/platform performance.

TouchTargets Marketing Intelligence – Complete Tool Guide


Why This Tool Exists

Every marketer, business owner, and agency runs into the same problem every single week: the industry moves faster than any one person can track. Google pushes an algorithm update, a new ad platform changes its bidding logic, an email deliverability rule shifts, and a benchmark you quoted to a client three months ago is already stale. Staying current means checking Search Engine Journal, then Ahrefs’ blog, then Digiday, then HubSpot, then Sprout Social, then whatever is trending on LinkedIn that day — and doing it again tomorrow.

Most people solve this by either (a) subscribing to five different newsletters and skimming none of them, (b) relying on whatever their agency tells them, which may be biased toward selling more of that agency’s services, or (c) giving up and making decisions on outdated assumptions. None of those are good options, and none of them are free of friction.

TouchTargets Marketing Intelligence was built to replace all of that with one live page: a real-time aggregation of 20+ actual marketing publications, refreshed automatically every two minutes, organized by discipline, and layered with tools that a raw RSS feed simply doesn’t have — engagement ranking, live benchmarks, trend detection, and a personal reading list, all without requiring a login or a subscription fee.


The Gap: What Existing Tools Don’t Do

To understand why this tool is useful, it helps to look at what marketers and business owners are already using, and where each one falls short.

Google Alerts and generic RSS readers (Feedly, Feedspot, Inoreader) These tools are built for broad content monitoring, not marketing specifically. They don’t know the difference between an SEO story and an advertising story — everything lands in one undifferentiated stream. There’s no ranking by importance, no sense of what other people are actually reading, and the useful filtering features (categories, saved searches, alerts) are usually locked behind a paid tier. You end up doing the categorization work yourself, every day.

Single-publisher blogs (Ahrefs, SEMrush, HubSpot, Moz) Each of these is excellent at covering its own corner of the industry, but that’s the problem — it’s one company’s take, often written to promote that company’s own product. If you only read Ahrefs, you get an SEO-tool-shaped view of SEO. If you only read HubSpot, everything looks like an inbound marketing problem. None of them aggregate what the rest of the industry is saying in the same place, so you either read one source and get a narrow view, or you manually visit five sites and lose twenty minutes a day just navigating between them.

Social monitoring (Twitter/X lists, LinkedIn feeds) Social platforms surface a lot of noise — hot takes, self-promotion, and recycled screenshots — with no reliable way to separate a credible industry data point from someone’s opinion thread. There’s also no persistent structure: today’s insight scrolls away and is gone tomorrow, with no archive, no benchmark reference, and no easy way to search back through what was said last month.

Paid marketing-intelligence platforms Enterprise tools exist that do real aggregation and analysis, but they’re priced and built for agencies and large marketing departments, not for a solo marketer, a freelancer, or a small business owner trying to keep an eye on the industry without a five-figure annual contract. They also require account creation, onboarding, and usually a sales call before you can even see what’s inside.

The common gap across all of the above: nothing combines (1) real aggregation across many credible sources, (2) a signal for what’s actually resonating — not just what was published most recently, (3) live, referenceable industry benchmarks, and (4) zero cost and zero login. That combination is exactly what this tool provides.


What We’re Providing: A Full Breakdown

1. Live, Multi-Source Aggregation

The feed pulls from more than 20 real marketing publications — including Search Engine Journal, Ahrefs, Digiday, HubSpot, and Sprout Social — and refreshes automatically every two minutes. There’s a visible countdown timer and a “LIVE” badge so it’s clear the page is actively pulling fresh content, not a static snapshot from this morning.

2. Discipline-Based Categorization

Every article is tagged into one of five categories: SEO, Advertising, Analytics, Funnels & CRO, or Brand & Social. A business owner who only cares about paid ads doesn’t have to wade through SEO technical deep-dives to find what’s relevant — one click on the “Advertising” filter narrows the entire feed instantly, and the count next to each category shows how much content exists in that bucket right now.

3. Multiple Sort Modes, Not Just “Newest”

Most feeds only sort by publish date. This tool offers four sort modes: Latest first (chronological), Priority (an internal relevance score), Most reacted (based on reader reactions), and Most clicked (based on actual click-throughs). This matters because the newest article isn’t always the most important one — sorting by engagement surfaces what the broader audience actually found worth reading, which is a much stronger signal for “what should I know about this week” than raw recency.

4. Trending This Week

Rather than relying on a human editor to guess what’s hot, the tool automatically scans the past seven days of real headlines, strips out common stopwords, and surfaces the keywords that are recurring most often across multiple sources — with a visual bar showing relative frequency. Click any trending term and it searches the full feed for it instantly. This is genuinely useful for content planning: if “AI Overviews” or “zero-click search” is trending across the industry this week, that’s a strong signal for what your own content or client reporting should address next.

5. Live 2026 Benchmarks Panel

This is one of the most practical features for non-marketers. The sidebar shows current industry-average figures — Google CPC, average search CTR at position one, email open rates, average ROAS for ecommerce, the share of indexed content that’s AI-generated, and the zero-click search rate on mobile — each with a year-over-year change indicator. A business owner who’s being quoted a CPC or ROAS number by an agency or freelancer can check it against this panel in seconds, without digging through a dozen separate industry reports to find a comparable figure.

6. Google Trends Sidebar

A secondary trend signal pulled directly from Google Trends data for marketing-related search terms, giving a second, independent read on where search demand and public interest are heading — useful alongside the headline-based “Trending This Week” panel for cross-checking whether a topic is just being written about, or actually being searched for.

7. Weekly Poll

A single rotating question each week, with live results shown as a percentage breakdown once you’ve voted. It’s a lightweight pulse-check on where the community stands on a current debate (for example, a platform change or a strategy shift), and it resets every Monday so it always reflects the current conversation.

8. Save For Later (Personal Reading List)

Every article has a Save button. Saved articles are stored in the browser (no account, no server-side profile) and are available in a dedicated “Saved” view that persists across visits. If someone wants to read something later in the day, or build a small personal library of reference articles for a client presentation, this replaces the need to open ten browser tabs or email links to themselves.

9. Reactions and Most Clicked

Readers can react to any article with a thumbs up, fire, or neutral-face reaction, and those counts are visible to everyone — a lightweight crowd signal on top of the raw article. Separately, the “Most Clicked” sidebar tracks which articles actually got opened the most, which is a more honest popularity signal than reactions alone, since it reflects real reader behavior rather than a quick tap.

10. Share With Built-In Attribution

Every article can be shared directly (using the native share sheet on mobile, or a copy-to-clipboard link on desktop), and every outbound link — whether opened directly or shared — carries UTM parameters back to the source. This means anyone sharing an article for a client report or a team Slack channel is automatically passing along proper attribution.

11. Follow Topics + Browser Alerts

Readers can select which categories they care about and enable browser notifications so new articles in those categories trigger an alert while the page is open in any tab — useful for someone who wants to leave the tab open in the background and get pinged only when something relevant to their business lands, rather than checking back manually.

12. New Since Last Visit

A banner automatically appears showing how many new articles have been published since your last visit, with a one-click option to filter down to only the new items. This solves the common problem of not knowing what’s actually changed since yesterday.

13. Free Weekly Email Digest

For people who don’t want to check a page at all, a one-field email signup delivers the top 10 articles from all 20+ sources every Monday, no cost, no account required beyond the email address itself.

14. No Login, No Cost, No Paywall

Every single feature above — saving, reacting, voting, following topics, alerts — works entirely without creating an account. Preferences are stored locally in the browser. This removes the single biggest point of friction that keeps people from using competing tools regularly: nobody has to sign up, verify an email, or hit a usage limit to get full functionality.


Use Cases in Practice

A solo marketer or freelancer can do a five-minute morning scan across SEO, ads, and analytics in one place instead of checking five bookmarked blogs, and save anything worth revisiting for later without needing a note-taking app.

A small business owner without a marketing background can use the Benchmarks panel as a reality check — if an agency quotes a CPC or ROAS figure, it can be compared instantly against the current industry average shown in the sidebar, giving non-experts a fast, low-effort way to evaluate whether a number sounds reasonable.

An in-house marketing manager can pull directly from “Trending This Week” when planning the next month’s content calendar, using real, current cross-industry signal instead of guessing what topics are worth covering.

Agencies reporting to clients can lift the current benchmark figures directly into a client report as a citable industry comparison point, with a live, refreshable source behind it rather than a static number from an old report.

Content and social teams can treat “Most Clicked” as a free proxy for industry-wide attention — if a particular story is getting real click-throughs across the feed, that’s a reasonable signal that a related piece of original content might also land well with the same audience.


Getting Started

  1. Land on the page — the feed loads pre-sorted by latest, showing all five categories together.
  2. Narrow the view using the category pills or keyword chips at the top if you only care about one discipline.
  3. Switch the sort dropdown to Priority or Most Reacted if you want signal over raw recency.
  4. Tap Save on anything worth revisiting — it will be waiting in the Saved view, no account required.
  5. Open Follow Topics in the sidebar and enable notifications if you’d rather be alerted than check back manually.
  6. Subscribe to the weekly digest if you prefer a Monday email over checking the page at all.

This content is generated from Claude based on the tool/platform performance.