The crawler
SurfingDogBot
SurfingDogBot is the crawler of the Surfing Dog network. It reads only what a business publishes for machines, to learn how an assistant can find the business and reach it: robots.txt, the homepage, up to three more of the site's own pages about what it offers (services, prices, opening hours, a menu, about; never a contact, legal, account or checkout page), and well-known discovery documents such as an agent inbox manifest, an MCP server card, an A2A agent card, an API description or llms.txt.
It never fills in a form, never logs in, never buys, books or orders anything, never calls a tool, and never solves a CAPTCHA. When a site asks it to prove it is human, it stops.
How to recognise it
Every request carries this User-Agent:
SurfingDogBot/1.0 (+https://network.surfingdog.ai/bot)
Its identity is checked by signature, not by address: every request is signed with Web Bot Auth (HTTP Message Signatures, RFC 9421, Ed25519), and the keys it signs with are published in its key directory at https://network.surfingdog.ai/.well-known/http-message-signatures-directory. No address ranges are published; a request that says it is SurfingDogBot and is not signed by one of those keys is not ours.
How it behaves
- It obeys robots.txt (RFC 9309), under the product token
SurfingDogBot, or the rules for every crawler (*) when there are none for it. It reads the robots.txt of every host it visits,wwwand other subdomains included, before asking that host anything else. If robots.txt cannot be read because of a server error, it reads nothing. - It honours Content-Signal:
search=nostops it; it recordsai-input=no(and Content-Usage that says no to AI), so that what it found on that site is not used as AI input; and it recordsai-train=no. If your site tells AI systems not to use or train on its content, we don't name you or list you publicly. - By default it sends at most one request every 10 seconds to a host, and slower when Crawl-delay asks. It backs off on 429 and 5xx answers and honours Retry-After.
- It reads a limited number of pages per site, each of limited size, and never follows a link off the site while it reads it. A door the business itself declared on another host (when it registered, in a registry or in its manifest) is checked there, read-only, as a door on the site would be: that host's robots.txt is read before anything else, and the same pace and limits apply.
- It records the names of the actions a site offers AI agents (the names of its tools and operations, and the short description a tool gives itself), to learn which kinds of actions sites offer. Never what an action returns: it calls none.
How it finds sites
Besides the sites it is asked to read, it finds sites on its own. When it reads a site, it notes which other sites that site's pages link to, and which other sites the site's documents for machines name (an AI catalog, an agent card, an MCP server card, llms.txt, a commerce profile). It fetches nothing extra to do this. It skips links marked nofollow, ugc or sponsored, and every link on a page that asks for that with <meta name="robots" content="nofollow"> (or none, or name="SurfingDogBot") or an X-Robots-Tag: nofollow header (for every crawler, or SurfingDogBot: nofollow).
- A site found this way is visited later, on its own, like any other: its robots.txt first, with the same pace and limits, and after that no more often than any other site.
- It adds at most 2000 new sites a day this way, and at most 10 from any one site.
- It never adds a site that opted out, and it leaves out social networks, search engines, link shorteners and other big platforms.
- Which site led it to yours is kept for the network's operators only, never published, and deleted if either site opts out below, or when the crawler next finds that either site's robots.txt keeps it out.
- To keep it away from your site, disallow it in robots.txt or opt out below. Either works whether or not it has visited you yet.
Listed on Surfing Dog Network?
The network lists a business it found on the business's own website only when an AI can already ask, book or buy there: the site publishes a machine door for agents (an agent inbox, an MCP server, an A2A agent card, an API, UCP or ACP checkout) that answers asks, bookings or orders. A site that offers only a phone, an email or a contact form is not listed.
- What is shown: the facts the business published about itself, on its own site or at its own doors, each with where it was read and when: its name, what it does, its town, region and country, its opening hours, its languages, the payments it shows, and its doors with what each takes. Never a phone number or an email address, no names of people other than the business's own name (a sole trader's business may carry theirs), and never the street or the exact spot: an entry the business has not claimed shows its town and country, and is found by distance only to within a few kilometres.
- Every such entry says so:
found on its own website · not a member · checked <date>, with a link to this page. - An entry found on a site whose robots.txt tells AI systems not to use or train on its content (
ai-input=noorai-train=noin Content-Signal, or the same in Content-Usage) is not listed while nobody has claimed it. - Some kinds of business are not listed this way at all, whatever their site publishes: among them health care, legal services, finance and insurance, places to stay, and trades that need a licence.
- To correct the entry, or claim it: from the business's own AI, call
update_businessorregister_businesson the network's MCP server, with proof that the domain is yours. - To remove it: opt out below. It needs no proof: the crawl stops, everything found on the site is deleted, and an entry nobody claimed goes with it. An entry its owner claimed is hidden while the site is not crawled; it goes for good when its owner takes it out, with its claim token or a proof as strong as its claim.
Claiming and correcting an entry
A business's own AI says "put my business on Surfing Dog" and calls register_business: it declares the business's doors and facts, and proves the domain is its own. A business with no entry yet registers the same way. The proof is one of these, and an entry shows its label, as in proof: domain:
- domain: a token the network gives, put on the domain: in the file
/.well-known/surfingdog-claim, in the inbox manifest's memberclaims, or in a DNS TXT record at_surfingdog-claim.<domain>; - key: the network's challenge, signed by a key the domain publishes (its manifest's receipt keys, or
/.well-known/jwks.json); - code: eight digits sent to an address at the domain itself (not at a subdomain, and never at a mail provider anyone can sign up to) that the owner types. The network writes to an address when it is typed in that call, and to no other: at most three codes a day, and none for 30 days once one is used. It keeps no copy of the address, just a keyed fingerprint, and never takes an address from a site. A code claim declares no door on another host, cannot turn off an entry found on the site, and cannot be taken over by a code to another address.
A stronger proof (domain, then key, then code) takes over an entry claimed with a weaker one, and the earlier claimant's declaration goes with it. A successful proof gives a claim token, which proves again for 90 days for that entry, to correct it with update_business, turn it off or on, or add and remove doors; a fresh proof ends the earlier tokens of the same proof or a weaker one. A declared door counts once the crawler has checked it, and a door on another host counts only when the site leads to it or the door's own document (its card, profile or manifest) names the business's domain. Taking a claimed entry out needs its claim token or a proof as strong as its claim; opting out below with no proof stops the crawl and hides it.
Checks and the agentic score
Anyone can ask for a site to be checked, at /check or with the check_business tool. A check is an ordinary visit of this crawler, the same pages and the same rules as above, made sooner: a share of each day's visits is kept for checks, and a site is read at most once a day. The result is a page at /b/<domain>: whether an AI agent can message the business, book, order, cancel or negotiate there, through which of its doors, and an agentic score by the published rules (/v1/score-rules), worked out from what the site and its doors publish and nothing else, each with its source and date. It is not a certification, it never changes the directory's search order, and nobody can pay for it.
- Checking a site is not claiming it, and changes nothing about its listing.
- Anyone with the link can see a result page. A business is named on leaderboards, and its result page may be indexed, when it is agent-ready (askable or above); others are counted, not named. The page of a business that is not named asks search engines not to index it.
- Its owner can hide the result, its ranks and leaderboard places included:
update_businesswithscore_pageset tohidden, with the listing's claim token or a proof that the domain is theirs.shownbrings it back. - A site that opts out below is never checked, and its score, rank and check are deleted with everything else.
- The badge (
/badge/<domain>.svg) is put on a site by its owner, if they want it; this network never places it anywhere.
How to stop it
Add this to your robots.txt:
User-agent: SurfingDogBot Disallow: /
Or opt out here. It needs no proof and takes effect at once: anything found on the site is deleted, its listing too unless its owner claimed it (then it is hidden until the owner decides), and a visit already under way stops within a minute.
What it keeps, and your rights
This is the notice under Article 14 of the GDPR, for information not obtained from you.
- Who is responsible
- The controller is Surfing Dog Lda, which runs the Surfing Dog network. Write to [email protected] about anything on this page.
- Where it comes from
- The business's own public website: what the site itself publishes, read as described above; the public registries where the business listed its own agent servers (such as the MCP registry), for the address of its site and its doors; and links to its site on other public websites the crawler read, for the address of its site only. Nothing is bought from anyone.
- What is collected
- Only facts a business has published about itself: its name, address, opening hours, categories, the payment and booking methods it shows, and its machine endpoints. Never personal contact details: no phone numbers, no email addresses, and no names of people other than a business's own name (a sole trader's business may carry theirs). At most it records whether a site offers only human contact.
- Why
- To help assistants find businesses and know how to reach them. The legal basis is legitimate interest.
- Who sees it
- The network's operators; processors that act for the network on its instructions, such as its hosting and the language-model provider described below; and, when the network lists the business, the assistants that ask it and the people using them. It is not sold, and not given to anyone else.
- Outside the European Economic Area
- A processor may handle some of it outside the European Economic Area. Where one does, the transfer rests on an adequacy decision of the European Commission or on the Commission's standard contractual clauses; write to [email protected] for a copy of the safeguards.
- Read by a language model
- When a site's own structured data leaves out its name, what it does or where it is, the text of a few of its own pages (never a contact page, with emails, phone numbers and messaging links taken out beforehand) may be read by a language model that a processor runs for the network, which copies the facts the site states about itself. A fact is kept when the model quotes it word for word from those pages, and dropped otherwise; it is shown marked as probably, and no filter reads it. The text is kept at most 7 days and deleted once read. A site whose robots.txt or Content-Signal says ai-input=no is never read this way.
- How long
- Until the facts are superseded by a later crawl, or 180 days after they were last seen; an entry found on a site and not checked again within 180 days is deleted with them. An opt-out deletes them at once.
- Your rights
- You may ask for access to what is kept about your site, have it corrected, have its use restricted, object to it, and have it erased. Opting out above is the quickest way to object and erase at once.
- Complaints
- You may complain to the data protection authority of the country where you live or work, or where you think the law was broken.
- Contact
- [email protected] for your rights; [email protected] for anything else about the crawler.
Listing terms (listing-1)
These are the terms a business agrees to when it registers or claims an entry (agree in register_business). A registration records the version it agreed to. Surfing Dog Lda runs the network.
- What the listing is
- Being listed is free. The order of the directory follows the published rules at /v1/ranking, which name its main parameters. Nobody can pay to be listed or to move in the order: no payment of any kind, to the network or to anyone else, has any effect on it.
- What the business declares
- Its facts and doors must be true and its own. The network checks each declared door before it counts, takes only what fits its vocabularies, and says in each answer what it did not take and why.
- When an entry is hidden or limited
- The network may hide an entry or a part of it, or refuse a registration, when its facts are false or misleading, when a door is not the business's, when the business is of a kind not listed (such as health care, legal services, finance and insurance, places to stay, and trades that need a licence), when it breaks the law or someone else's rights, or when it is used to abuse the network or other businesses. It gives the reason at the time it acts, in the answer to the business's next
register_businessorupdate_businesscall and as the entry's reason not to be shown. - When a listing ends
- The business may end its listing at any time:
listing offhides it, andopt_outtakes it out. The network gives 30 days' notice, with its reason, before it ends a listing for good, unless the law requires it to act sooner, or the business broke these terms repeatedly or in a way that harms others. The reason is given as above. - Changes to these terms
- A change is a new version, published on this page at least 15 days before it takes effect, and named in the listing's answers. A business that does not accept it may end its listing before then.
- Complaints
- Write to [email protected] with the domain: about a decision on an entry, its order, or these terms. Every complaint is read by a person and answered in plain words within 15 working days. Surfing Dog Lda is a small enterprise, so the duties of Regulation (EU) 2019/1150 to run a complaint system and to name mediators do not apply to it; it answers every complaint all the same, and will agree a mediator with a business when both wish to.
- What it keeps
- The declaration, the proof's label, a fingerprint of each claim token (never the token), when the business agreed and to which version, and the record of each change. An address typed for a code is kept only as a keyed fingerprint.