RoktCrawl: Rokt's storefront crawler

If RoktCrawl appears in your server logs, this page explains what it is, what it does on your site, and how to reach us or ask it to stop.

Who we are

Rokt is an e-commerce technology company. We work with online retailers and other businesses to present relevant offers to their customers at the moment of a purchase. RoktCrawl is operated by Rokt and visits the public storefronts of businesses that work with us.

What the crawler does

RoktCrawl learns how a storefront's web addresses are structured: which URLs are search results, which are product pages and which are category pages, and how a search term or product identifier is written into the address. Rokt's systems use that knowledge to recognise the kind of page a shopper is on from its address alone, without inspecting the page itself.

A visit fetches the site's robots.txt, its sitemap and its homepage, then opens a real browser session that types a handful of ordinary search terms into the site's own search box and browses the results: sorting, filtering, paging, and a sample of product and category pages. The crawler records the addresses it saw, a copy of each page it loaded and a screenshot, so the patterns can be derived without visiting again.

How to recognise it

Every request the crawler makes carries this token at the end of an otherwise standard browser user-agent string:

RoktCrawl/0.1 (+https://www.rokt.com)

A full user-agent string from the crawler looks like this (the Chrome version follows the browser build in use):

Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36 RoktCrawl/0.1 (+https://www.rokt.com)

The token names the operator and links here. The rest of the string is the real browser the crawler drives; nothing about the platform or browser is disguised.

What it does not do

robots.txt

The crawler reads robots.txt before anything else and obeys it strictly. A URL that robots.txt disallows is not fetched: if the site's search form produces a disallowed address, the address is noted but the page is never loaded. The crawler matches rules addressed to the token RoktCrawl as well as the general * rules. To keep it off your site entirely, add:

User-agent: RoktCrawl
Disallow: /

The rule takes effect at the crawler's next visit; there is no cached copy of your robots file between visits.

Request rate

Page loads to one site are spaced at least 2.5 seconds apart, plus up to 1.5 seconds of random extra delay, and the crawler visits one page of a site at a time. Within a page load, the browser fetches the page's own images, styles and scripts as any browser would. A site's visit covers a small number of pages, and scheduled crawls run about once a week per region.

Egress addresses

The crawler leaves Rokt's network only through the fixed addresses below, so a request claiming to be RoktCrawl can be verified against them. The full list, with the policy that governs it, is on the crawler policy page.

RegionProductionStaging
US West (Oregon)16.145.117.88 32.186.91.66 184.33.228.452.38.96.62
US East (N. Virginia)3.212.239.12 34.195.230.14 100.56.59.133.93.202.206
Europe (Ireland)54.73.34.189 3.248.25.71 34.242.239.954.220.78.42
Asia Pacific (Sydney)13.54.60.54 32.236.51.124 13.211.111.19713.54.70.40

Contact and opting out

The quickest way to stop the crawler is the robots.txt rule above. To ask us to exclude your site, to report a problem with the crawler's behaviour on your site, or to verify that a request came from us, contact us at:

crawler@rokt.com

Please include your site's hostname and, if you have it, the time and source address of the requests you are asking about.

The crawler policy states these commitments formally.