Glassrecord crawler
The Glassrecord crawler is a web browser run by [Company legal name]. It loads web pages the way a visitor's browser does and records which third parties the page loads and what data they receive under each consent choice. Glassrecord turns those records into findings for the owner of the site.
If you run a website and saw the crawler in your logs, this page explains why it came, what it did, and how to stop it.
Why it visits
The crawler loads a site for one of these reasons:
- A Glassrecord customer monitors the site. The customer added the site to their account, and the crawler loads it on request or on a weekly schedule.
- Someone asked for a free scan of the site. Anyone can ask for a one-time scan of a public website at
glassrecord.com. A domain gets at most one free scan every 30 days. - We test our detection. We load a small set of public sites to check that the crawler reads consent banners and third parties correctly.
The crawler does not index your content for search, train models on it, or republish it.
How to identify it
The crawler runs Chromium and presents a normal Chrome user agent for the device it emulates, with our token appended:
<Chrome user agent for the emulated device> AtheveraScan/<version> (+https://glassrecord.com/crawler)
For example, the desktop profile sends:
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/<major>.0.0.0 Safari/537.36 AtheveraScan/0.1.0 (+https://glassrecord.com/crawler)
It emulates three devices: a desktop Mac, an Android phone, and an iPhone. All three run Chromium. Requests made outside the browser, such as fetching robots.txt and sitemaps, send only the token:
AtheveraScan/<version> (+https://glassrecord.com/crawler)
The crawler loads pages from the United States and from Germany. Loads from the United States run on Cloudflare's network. Loads from Germany run on a server in Nuremberg. We do not yet publish a list of the crawler's IP addresses.
What it does on a page
For each page it loads, the crawler may:
- Load the page up to four times from each region: once without answering the consent banner, once after rejecting, once with Global Privacy Control on, and once after accepting.
- Click the consent banner's accept or reject button.
- Record the network requests the page makes, the cookies and storage set in the crawler's own browser, and screenshots of what the page showed.
- Press the Tab key to move through the page, to check keyboard access.
- Type a random test value into up to three visible text fields, without pressing Enter or any button, and watch whether any third party receives the value. The value is generated for each load and contains no one's data.
- Open a chat widget by clicking its launcher and read its first message. It sends no message.
- Fetch the privacy policy, accessibility statement, and consent platform declaration the page links to.
- Look up the domain's mail records (SPF, DKIM, DMARC) over DNS and check the site's TLS certificate.
To choose which pages to load, the crawler reads robots.txt, the sitemaps it names, and up to 10 pages of the site.
What it never does
- It never submits a form on a site, except the sign-in form of a site whose owner gave Glassrecord a test account for that purpose.
- It never creates an account, makes a purchase, or sends a message through a site.
- It never signs in to a site unless the site's owner verified the site with Glassrecord and supplied a test account. Behind a sign-in, it types nothing else, submits nothing, and follows no sign-out or delete links.
- It never tries to solve a CAPTCHA or get past a bot challenge.
- It never requests private or internal network addresses.
- It never keeps cookie values or form contents, and it removes query strings from the addresses it stores, apart from a short list of parameters that show what a request sent to a third party.
How fast it goes
- One scan of a domain runs at a time.
- A scan loads at most 6 pages for a free scan, at most 12 pages for a site a customer has not verified, and at most 50 pages for a verified site. Each page may be loaded up to four times per region, as described above.
- A scan runs at most six page loads at a time.
- While choosing pages, the crawler waits at least one second between requests. It honors a
Crawl-delayinrobots.txtof up to 10 seconds. A site that asks for a longer delay gets fewer pages read, not more requests. - If
robots.txtanswers429 Too Many Requests, the crawler asks once more after theRetry-Aftertime, up to 10 seconds, and then stops asking.
How to opt out
With robots.txt. The crawler's product token is atheverascan. It matches without regard to case. To keep the crawler off your whole site, add:
User-agent: atheverascan
Disallow: /
The crawler also obeys a Disallow: / in the group for all agents (User-agent: *).
What the crawler does with robots.txt depends on why it came:
- Free scans read
robots.txtbefore loading anything. A free scan stops without loading any page when the group that applies to the crawler disallows the whole site, or when a comment in the file says automated access is not allowed. The crawler reads such comments in English, French, German, Spanish, and Italian. The person who asked for the scan is told why it stopped. - Sites a customer monitors follow the
AllowandDisallowrules when the crawler chooses pages. The crawler does not act on comments for these sites, since the customer has told us the site is theirs or that they are authorized to scan it.
By email. Write to [Crawler contact address] from an address at the domain, or from the site's listed contact, naming the domain. We add it to our opt-out list. A domain on the list gets no free scans. [Decide whether the opt-out list also stops customer scans of an unverified site; today it stops free scans only.]
If a customer scans your site without your permission. Write to [Crawler contact address]. Customers must be the site's owner or authorized by the owner to scan it. We will look into it and may stop the scans.
Free scan reports
A free scan's report is available to the person who asked for it, by a private link, for 30 days. The report is not listed on search engines. After 30 days the report, its findings, and its screenshots are deleted. The records of what each page loaded are kept as facts about the site.
If a free scan report about your site is wrong, open the finding and use "Report a wrong finding", or write to [Crawler contact address]. A report of a wrong finding goes to our staff for review.
Contact
[Company legal name], [Address]
Crawler questions and opt-outs: [Crawler contact address]
Everything else: support@glassrecord.com