Guides
Tag audit

Tag audit

A tag audit crawls a site and lists the technologies on every page it visits: analytics and ad pixels, tag managers, chat, forms, consent banners, A/B testing. Agencies use it to check which pages carry which tags, for example that the conversion pixel is on every landing page and the old analytics snippet is gone.

Start one with POST /v1/crawl and a tech field:

curl -X POST https://api.superscraper.dev/v1/crawl \
  -H "x-api-key: ss_live_xxxxxxxxxxxxxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://www.acmeplumbing.com", "limit": 50, "tech": "deep" }'
techWhat it readsPrice
profileThe HTML each page returnsThe crawl price: 1 credit a page. Detection adds nothing
deepEach page loaded in a real browser: the requests it makes, the scripts it loads, the globals and cookies it leaves. Finds tags a tag manager injects after the page loads5 credits a page instead of the crawl price. A page that did not render but was fetched is 1 credit; a page that could not be fetched is free

A deep audit visits at most 200 pages (and no more than your plan's crawl limit), and stops after 5 hours of running; the pages it scanned by then are delivered and charged. Its credits for every page it may visit are held when it is queued (the answer carries max_credits) and settled to the pages actually scanned when it ends, so you pay only for what ran. A deep audit is refused with tier_not_available where deep scans are not switched on.

The crawl runs like any other: poll GET /v1/crawl/:id, then read GET /v1/crawl/:id/results. Each fetched page has a tech block:

{
  "url": "https://www.acmeplumbing.com/book",
  "ok": true,
  "tech": {
    "mode": "deep",
    "technologies": [
      { "name": "Google Tag Manager", "category": "Tag managers", "confidence": 0.95 },
      { "name": "Facebook Pixel", "category": "Analytics", "confidence": 0.9 }
    ],
    "rendered": true,
    "partial": false,
    "reasons": []
  }
}

The results also carry tech_summary: every technology found, how many pages it was on and which ones (up to 100 URLs each), most widespread first. A technology on 49 of 50 pages points you at the one page that is missing it.

{
  "tech_summary": {
    "mode": "deep",
    "pages_scanned": 50,
    "technologies": [
      { "name": "Google Tag Manager", "category": "Tag managers", "pages": 50, "urls": ["https://www.acmeplumbing.com/", "..."] },
      { "name": "Facebook Pixel", "category": "Analytics", "pages": 49, "urls": ["..."] }
    ]
  }
}

The completion webhook carries the same summary as techSummary. Both cover the pages you can read: if a very large crawl's stored results are cut at the storage limit, the summary and the charge cover the stored pages only. When results are held in object storage, GET /v1/crawl/:id/results redirects to the stored pages, each with its tech block; take the summary from the webhook, or count it from the pages.

To check one page, or a site's main pages and subdomains, without a crawl, use the tech lookup.