HostDesk
HostDesk website scanner

What our scanner reads, and how to stop it.

HostDesk reads a small sample of public pages on a business website to work out which customer questions the site answers clearly and which it doesn't. If you found one of our user agents in your logs, this page explains exactly what it did.

User agents

  • HostDeskProspectAudit

    Reads a small sample of public pages to work out which customer questions a business website answers clearly. Used for the free website scan and for research samples.

  • HostDeskProofAudit

    The same deterministic read, used when a visitor asks for a live preview of HostDesk on a specific website.

What the scanner does

  • Public pages only. It requests pages any visitor could open. It does not log in, submit forms, accept cookies, or attempt to reach anything behind authentication.
  • It respects robots.txt. A path disallowed for our user agent is not requested. A site-wide disallow stops the scan entirely.
  • It reads at most 10 pages. One homepage request plus a small number of prioritised follow-ups, fetched once. This is a sample, not a crawl of your site, and we never describe it as full-site coverage.
  • It runs rarely. A given domain is read when someone requests a scan of it or when it is part of a research sample — not continuously.

What we keep

  • Page URLs, page titles, a content hash and a word count.
  • Counts derived from the text: how many topics the pages evidence, how many customer questions we could not verify from them, and which categories those fall into.
  • The time we read the pages, and the version of the method we used to read them.

What we do not keep

  • Raw HTML from your site.
  • Personal data, contact details or anything identifying your visitors.
  • Credentials, cookies or session state.
  • An archive or mirror of your website's content.

Retention and versioning

A short-lived preview built for a specific outreach message is deleted after 30 days. The longer-lived record is the normalized observation described above — URLs, hashes and counts — which we keep so that a published statistic can still be checked later. Every observation carries the methodology and metric version that produced it, so an old number stays interpretable after the scanner changes, and we never rewrite an existing observation to match a newer one.

How to stop the scanner

Add this to your robots.txt:

User-agent: HostDeskProspectAudit
Disallow: /

User-agent: HostDeskProofAudit
Disallow: /

We honour that on the next request. You can also ask us directly to add your domain to our suppression list by writing to bot@hostdesk.ai from an address at that domain.

“Don't scan” and “don't contact” are separate requests. Asking us not to scan your site does not by itself remove you from outreach, and asking not to be contacted does not by itself stop the scanner — tell us if you want both and we will record both. We would rather ask than assume we know which one you meant.

Curious what the scanner found on your own site? Run it yourself.