PageSourceSearch

About PageSourceSearch and its crawler

PageSourceSearch is a search engine over the source of web pages: the HTML of a site's pages and the JavaScript the site itself serves, searchable by exact bytes or by regular expression. The crawler that collects it identifies itself as PageSourceSearchBot and links here from its user agent.

What the crawler does

Its user agent is

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PageSourceSearchBot/1.0; +https://pagesourcesearch.com/bot) Chrome/128.0.0.0 Safari/537.36

It sends the request headers a current Chrome sends, so a site serves it the same content a visitor would get. For each site it fetches, in this order and never more:

Requests to one host are made one at a time, at least 1000 ms apart by default (or your Crawl-delay, when larger), and a 429 or 503 answer backs it off (Retry-After is honoured). It does not log in, submit forms, run your scripts or fetch images, fonts or stylesheets. A live site is revisited about once a week; a site that answers with a bot challenge is not tried again for a year.

robots.txt

The crawler reads /robots.txt before anything else and follows it (RFC 9309: the group for its token, then the * group; Allow, Disallow, * and $ patterns, Crawl-delay). Its token is

User-agent: PageSourceSearchBot

If robots.txt cannot be fetched because of a server error, the site is treated as fully disallowed and nothing else is fetched.

How to opt out

To keep the crawler off your site entirely, add this group to your robots.txt:

User-agent: PageSourceSearchBot
Disallow: /

To keep it out of part of a site, Disallow those paths instead. The change takes effect at the next visit, which starts with robots.txt: a disallowed root stops the visit before any page is fetched. To have already stored pages removed, or for any other question about the crawler, write to [email protected] with the domain name.

What is stored and shown

The crawler keeps the exact bytes of the pages and scripts it fetched, with the URL and the time of the fetch. A search result shows short snippets around each match, and the preview page shows the stored file as text. Nothing is executed, rendered or served as a live page, and the preview and match pages are marked noindex so search engines do not index your source through this site. Only a site's own code is searchable: library code (jQuery, React and the like) is recognised, stored for the preview and left out of the index.

Searching the index

How queries work, from exact bytes and several terms to regular expressions, filters and why a query can be rejected, is on the query help page, together with the FAQ and a note on the JSON API.