Glassrecord crawler

The Glassrecord crawler is a web browser run by [Company legal name]. It loads web pages the way a visitor's browser does and records which third parties the page loads and what data they receive under each consent choice. Glassrecord turns those records into findings for the owner of the site.

If you run a website and saw the crawler in your logs, this page explains why it came, what it did, and how to stop it.

Why it visits

The crawler loads a site for one of these reasons:

The crawler does not index your content for search, train models on it, or republish it.

How to identify it

The crawler runs Chromium and presents a normal Chrome user agent for the device it emulates, with our token appended:

<Chrome user agent for the emulated device> AtheveraScan/<version> (+https://glassrecord.com/crawler)

For example, the desktop profile sends:

Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/<major>.0.0.0 Safari/537.36 AtheveraScan/0.1.0 (+https://glassrecord.com/crawler)

It emulates three devices: a desktop Mac, an Android phone, and an iPhone. All three run Chromium. Requests made outside the browser, such as fetching robots.txt and sitemaps, send only the token:

AtheveraScan/<version> (+https://glassrecord.com/crawler)

The crawler loads pages from the United States and from Germany. Loads from the United States run on Cloudflare's network. Loads from Germany run on a server in Nuremberg. We do not yet publish a list of the crawler's IP addresses.

What it does on a page

For each page it loads, the crawler may:

To choose which pages to load, the crawler reads robots.txt, the sitemaps it names, and up to 10 pages of the site.

What it never does

How fast it goes

How to opt out

With robots.txt. The crawler's product token is atheverascan. It matches without regard to case. To keep the crawler off your whole site, add:

User-agent: atheverascan
Disallow: /

The crawler also obeys a Disallow: / in the group for all agents (User-agent: *).

What the crawler does with robots.txt depends on why it came:

By email. Write to [Crawler contact address] from an address at the domain, or from the site's listed contact, naming the domain. We add it to our opt-out list. A domain on the list gets no free scans. [Decide whether the opt-out list also stops customer scans of an unverified site; today it stops free scans only.]

If a customer scans your site without your permission. Write to [Crawler contact address]. Customers must be the site's owner or authorized by the owner to scan it. We will look into it and may stop the scans.

Free scan reports

A free scan's report is available to the person who asked for it, by a private link, for 30 days. The report is not listed on search engines. After 30 days the report, its findings, and its screenshots are deleted. The records of what each page loaded are kept as facts about the site.

If a free scan report about your site is wrong, open the finding and use "Report a wrong finding", or write to [Crawler contact address]. A report of a wrong finding goes to our staff for review.

Contact

[Company legal name], [Address]

Crawler questions and opt-outs: [Crawler contact address]

Everything else: support@glassrecord.com