RevisebergBot is the crawler behind Reviseberg's accessibility, quality and search checks. If you found this page through your server logs, it visited your site because somebody asked us to check it: a customer who added the site to their account, or a visitor who ran the free check on one of its pages. If the user agent in your logs says RevisebergBot-Index, it was the index crawl, described below.
How it identifies itself
Every request carries this user agent:
RevisebergBot/1.0 (+https://reviseberg.com/bot)
The token robots.txt is matched against is RevisebergBot.
What it does
- It loads pages in a real browser (Chromium), the way a visitor would, and runs automated accessibility checks on what it sees.
- On the pages of a journey a customer names, it presses Tab, Shift+Tab and Escape to test keyboard access. It does not fill in forms, log in (unless the site's owner gave it a login in their account), place orders or pay.
- It follows links within the site, up to the page limit of the account's plan. The free check looks at one page.
- It checks whether the links on a checked page still work, with one request per link.
How to limit or block it
RevisebergBot reads robots.txt before it crawls a site and obeys the group that names RevisebergBot, or the * group when none does, with the longest matching rule winning. It honours Crawl-delay up to 10 seconds between pages. If robots.txt itself answers 401 or 403, it treats the whole site as off limits.
To keep it out entirely:
User-agent: RevisebergBot
Disallow: /
The index crawl: RevisebergBot-Index
There is one crawl nobody on the site asked for: the measurement for the BFSG-Index (EAA Index), a planned benchmark of automatically detectable barriers on the websites of large consumer-facing organisations in Germany – online shops, banks, energy suppliers and travel portals. Nothing from it has been published yet. It uses its own name, so you can refuse it without refusing checks somebody asked for:
RevisebergBot-Index/1.0 (+https://reviseberg.com/bot/)
- It runs from a fixed list of large organisations, chosen by published criteria; micro-enterprises and public bodies are never on it.
- On each site it loads the homepage and a fixed sample of publicly linked pages – never more than 15 – and follows no links. On the homepage it presses Tab for at most 30 steps to check focus and skip links.
- It loads at most one page every 5 seconds per site, or less often when robots.txt asks for a longer
Crawl-delay, and it stops at the first HTTP 429 or 503. - It blocks every request other than GET, HEAD and OPTIONS – including the ones your own pages' scripts try to send – fills in no forms, logs in nowhere and does not get round bot protection. A consent banner is left as it appears.
- It obeys a robots.txt group naming
RevisebergBot-Index, one namingRevisebergBot, and otherwise*. A site whose robots.txt refuses it is listed as not tested, with no score.
To keep only the index crawl out:
User-agent: RevisebergBot-Index
Disallow: /
You can also write to hello@reviseberg.com to be left out of the crawl, or to be listed under a pseudonym instead of your name.
Questions
If the crawler caused a problem on your site, or you want to know who asked for a check, write to hello@reviseberg.com. Security reports go to security@reviseberg.com.