PageSourceSearch

Exact bytes or regex, over crawled HTML and JavaScript

Search the source code of the web

PageSourceSearch is a search engine over the source of web pages: the raw HTML and the JavaScript each site serves. Type an exact byte sequence or a regular expression and get every site that contains it, grouped by domain, with each match highlighted in the stored file.

Try a query

How to query

Query help: every rule, the filters and the API →

Which websites use a technology?

Ready-made lists: the highest-ranked sites whose source contains the code that reveals an analytics tag, a CMS, a framework or a widget, with the keyword that finds them.

What is a website built with?

Enter a domain for its profile: the JavaScript libraries and versions in its source, its analytics and tag manager IDs, the third-party hosts it loads, every stored page and script with its earlier versions, and a timeline of what changed between crawls.

Sites by tracking ID · Sites by third-party host

What makes PageSourceSearch different from a web search engine?

Exact bytes, not keywords

A literal query matches byte for byte: every bracket, quote and space counts, and nothing is tokenised, stemmed or corrected. What you type is what the page contains.

Regular expressions over source

Choose regex mode under Filters for patterns: an id that varies, a version range, a snippet with any spacing. Patterns run over the raw bytes of each file.

Grouped by site, highlighted in place

Results list the domains that hold every term, then the files, then each match with its byte offset and line. One click opens the stored file with the match highlighted.

What can you find with PageSourceSearch?

Anything a page or a site's own script contains, by its exact bytes or a pattern. The most common questions people answer with it:

How does PageSourceSearch work?

Three steps: a crawler stores the source, a classifier sets library code aside, and an index of 3-byte windows answers each query with a byte-for-byte check.

  1. Crawl. PageSourceSearchBot fetches a site's root page, a handful of further pages and the scripts they load, first-party hosts only, and stores the exact bytes.
  2. Classify. A classifier separates a site's own code from library code (jQuery, React and the like); libraries are stored and shown in previews but not indexed, so results are about the site, not its dependencies.
  3. Index and search. Every stored file is indexed by its 3-byte windows. A query looks those up, verifies the candidates byte for byte (or with your regex) and lists the sites that hold every term.

What the crawler fetches, its robots.txt token and how to opt out: About PageSourceSearchBot. What a query can be and why one is rejected: Query help.

Common questions

What is PageSourceSearch?
PageSourceSearch is a search engine over the source code of websites. Its crawler stores the raw HTML and first-party JavaScript of each site it visits, and a query (an exact byte sequence or a regular expression) lists every site that contains it, grouped by domain, with each match highlighted in the stored file.
How is it different from a web search engine?
A web search engine indexes the visible text of rendered pages and ranks them by relevance. PageSourceSearch indexes the source bytes, so tags, attributes, script code, tracking IDs, comments and markup patterns are searchable, and a match is exact: what you type is what the page contains.
Is the search case-sensitive?
Yes. A literal query is matched byte for byte, so Gtag and gtag are different queries. Choose regex mode under Filters and start the pattern with (?i) for a case-insensitive match.
What happens to a query of 1 or 2 characters?
The index is built from 3-byte windows of every file, so a shorter term has no window of its own. It is searched with a space added after it, and before it too for a single character: 'ab' finds 'ab ' and 'a' finds ' a '. A regex needs a fixed run of 3 characters.
Can I search across a whole site rather than one page?
Yes, and that is the default: with Same site selected next to the search box the words may be spread over different pages and scripts of one domain. Switch to Same page to require them all in one file.
Can I keep my site out of the index?
Yes. Add 'User-agent: PageSourceSearchBot' followed by 'Disallow: /' to your robots.txt, or write to [email protected] to have stored pages removed.

Full query help and FAQ · Regular expressions · How to opt out · JSON API