PageSourceSearch

Query help

A query is an exact byte sequence (the default) or a regular expression, run over the stored HTML and first-party JavaScript of every crawled site. This page explains what matches, how several terms combine, what the filters do and why a query can be rejected.

What does a literal query match?

In literal mode (the default) the query is matched byte for byte, case-sensitively (spaces separate terms, see below): gtag('config' matches gtag('config' and not gtag("config" or gtag( 'config'. There is no tokenisation, stemming or wildcard; every character, including (, ', ", $, < and >, is an ordinary byte. The text you type is sent as UTF-8, so a non-ASCII character matches the same UTF-8 bytes in a page.

How do several terms combine?

Words separated by spaces are separate terms, and a file matches only when it contains every one of them, anywhere, in any order: data=k124 r&=43-" finds the files holding both data=k124 and r&=43-". To keep the spaces of a phrase, put it in double quotes: "gtag('js', new Date" data=k124 is two terms, the first with its spaces (the quotes themselves are dropped). A quote anywhere else is an ordinary byte ("config": keeps both quotes, and a lone opening quote is literal), and a query without spaces is searched exactly as typed, quotes included ("config" finds the quoted word). A term of 1 or 2 characters is searched with a space added after it, and before it too for a single character (ab finds ab , a finds a : the index looks up 3-byte windows, and a word boundary is the likeliest neighbour of a short token); the results say so, and the search box then shows the term quoted with its space, data "at " for data at. A query has at most 8 terms. The matches of every term are listed together under the file, each linking to a preview that highlights that term. A regex is never split: a space in a pattern is part of the pattern.

By default (Same site, the toggle next to the search box) the terms may be spread over different pages and scripts of one domain: createRoot( wp-content/plugins/ then finds sites that mount React somewhere and run WordPress somewhere.

How do regular expressions work?

In regex mode (chosen under Filters) the query is a .NET regular expression run over the bytes of each file (each byte is one character, so \xFF matches the byte 0xFF). The pattern must contain a fixed run of at least 3 characters that every match has to contain, because that run is what the index looks up: gtag\(['"]config qualifies (config), [a-z]+\d does not. Alternations, character classes, optional or repeated parts and lookarounds never count towards the run. A match is limited to 200 ms per chunk; a chunk that times out is skipped and counted in a note. Options go at the front: (?i)gtag\('config' matches case-insensitively.

Why can a query be rejected?

What is a result?

Results are grouped by site (ranked sites first), then by file. Each match shows a snippet of up to 200 bytes around it, its byte offset in the file and its line (counting line feeds from the start of the file); the link opens the stored file with that match's term highlighted. 20 sites per page by default, up to 10,000 with per page. Only a site's own code is indexed: library files and library code inside bundles (jQuery, React and the like) are stored and shown in previews but never matched; the library filter (react, react@18) finds the sites that use a library instead. A literal that straddles two stored chunks of a file is not found, and more than 20,000 matching chunks (of the rarest term, when there are several) are cut off (truncated, with the site count a lower bound). Tick every domain (slower) to lift that cut-off and count and page through every site that holds the terms; a common term can then take tens of seconds.

The summary line says how the count was made: N domains is exact, at least N a lower bound after the cut-off, about N candidate domains an estimate from the index while the pages are filled forward and verified as you page. A page incomplete tag means the page's time budget ran out before it was full.

What do the filters do?

What is a site profile?

Every site the crawler has visited has a page at /site/{domain} (or through the lookup form) that gathers what its stored source reveals: the libraries and versions recognised in its files, the analytics, tag manager, pixel and site-verification IDs read out of its snippets, the third-party hosts its pages load scripts, frames and stylesheets from (and the hosts its own scripts call), every stored page and script with the crawls at which its content changed, and a timeline of the crawls with what changed between them. Each earlier version of a file opens in the same viewer as the current one: the crawler keeps every version's chunk map, and the chunks themselves are shared between versions, so nothing is copied when a file stays the same.

The identifiers and hosts link to reverse lookups: sites by tracking ID lists every site whose source carries the same id, sites by third-party host every site loading from the same host. Every domain name in the results and in a preview links to its profile.

Frequently asked questions

How do I see what a website is built with?
Open its profile at /site/{domain} or through the lookup form on /site: it lists the libraries and versions in its stored source, its tracking and verification IDs, the third-party hosts it loads, every stored page and script with its earlier versions, and a timeline of what changed between crawls. Every domain in the results links to its profile.
Can I find every website that shares one Google Analytics or Tag Manager ID?
Yes: /id/{kind}/{value}, for example /id/ga4/G-XXXXXXXXXX or /id/gtm/GTM-XXXXXXX, lists every site in the index whose source carries the id. /id lists every kind the crawler extracts; /hosts and /host/{host} do the same for third-party hosts.
Is the search case-sensitive?
Yes. A literal query is matched byte for byte, so 'Gtag' and 'gtag' are different queries. For a case-insensitive match choose regex mode under Filters and start the pattern with (?i), for example (?i)gtag\('config'.
What happens to a query of 1 or 2 characters?
The index is built from the 3-byte windows (trigrams) of every stored file, and a query is answered by looking those windows up, so a term shorter than 3 bytes has no window of its own. Such a term is searched with a space added after it, and before it too for a single character: 'ab' finds 'ab ' and 'a' finds ' a ', a word boundary being the likeliest neighbour of a short token. The results say when a term was padded. A regular expression must contain a fixed run of at least 3 characters.
Can I search for a phrase with spaces?
Yes. Put the phrase in double quotes: "gtag('js', new Date" is one term with its spaces. Without quotes, spaces separate terms, and a file matches only when it contains every term.
Does it search text or source code?
Source. PageSourceSearch indexes the raw bytes of the HTML and the first-party JavaScript a site serves, not the visible text of a rendered page. Tags, attributes, script code and comments are all searchable.
Why are library files such as jQuery or React not matched?
A classifier separates a site's own code from library code. Libraries are stored and shown in previews but left out of the index, so that a query about a site's code is not swamped by the dependencies every site shares. The websites-by-technology pages list the sites that use a given library instead.
Can I find every site that contains my query?
Yes. By default a search stops at the first 20,000 candidate chunks and marks the count 'at least N'. Tick 'every domain (slower)' under Filters to count and page through every site that holds the terms; a common term can then take tens of seconds.

The JSON API

The same search is available as a JSON API under /api/v1 for key holders: GET /api/v1/search?q=&mode=&kind=&domain=&since=&lib=&scope=&page= answers with the domains, files and matches the page shows, and /api/v1/preview and /api/v1/matches with a stored file and its matches. Keys are rate-limited per minute. Write to [email protected] to ask for one.