Exact bytes or regex, over crawled HTML and JavaScript
Search the source code of the web
PageSourceSearch is a search engine over the source of web pages: the raw HTML and the JavaScript each site serves.
Type an exact byte sequence or a regular expression and get every site that contains it, grouped by domain, with each match highlighted in the stored file.
How to query
jquery mapTwo words: sites whose source holds both, anywhere on the site (one may be in the HTML, the other in a script)
jquery map same page The same two words with Same page selected: both in one file
"jquery map"The exact phrase with its space, in quotes: the two words next to each other
gtag('config'Exact bytes, case-sensitive: every bracket, quote and space counts, nothing is stemmed or corrected
data atA word of 1 or 2 characters is searched with a space added, as "at ": the index needs 3 characters
hotjar\.com/c/hotjar-\d+ regex Regex mode (under Filters): a pattern over the source bytes; it needs a fixed run of 3 characters, hotjar here
(?i)gtag\('config' regex A case-insensitive match: regex mode with (?i) at the front; \( is a literal bracket, since a bare ( opens a group
Query help: every rule, the filters and the API →
Which websites use a technology?
Ready-made lists: the highest-ranked sites whose source contains the code that reveals an analytics tag, a CMS, a framework or a widget, with the keyword that finds them.
What is a website built with?
Enter a domain for its profile: the JavaScript libraries and versions in its source, its analytics and tag manager IDs, the third-party hosts it loads, every stored page and script with its earlier versions, and a timeline of what changed between crawls.
Sites by tracking ID · Sites by third-party host
What makes PageSourceSearch different from a web search engine?
Exact bytes, not keywords
A literal query matches byte for byte: every bracket, quote and space counts, and nothing is tokenised, stemmed or corrected. What you type is what the page contains.
Regular expressions over source
Choose regex mode under Filters for patterns: an id that varies, a version range, a snippet with any spacing. Patterns run over the raw bytes of each file.
Grouped by site, highlighted in place
Results list the domains that hold every term, then the files, then each match with its byte offset and line. One click opens the stored file with the match highlighted.
What can you find with PageSourceSearch?
Anything a page or a site's own script contains, by its exact bytes or a pattern. The most common questions people answer with it:
- Which sites use a tag or tracker an analytics ID, a tag-manager container, a pixel or a consent script, by the exact bytes it leaves in the page.
- Technology fingerprints framework markers, build-tool signatures, plugin paths and version strings in a site's own scripts.
- Structured data and SEO markup schema.org types, meta tags, canonical patterns and hreflang blocks as they are actually served.
- Security research leaked key formats, debug endpoints, vulnerable snippets and copied code across sites, with a regular expression.
- Competitive and market research which sites embed a vendor's widget, checkout, chat or affiliate code.
How does PageSourceSearch work?
Three steps: a crawler stores the source, a classifier sets library code aside, and an index of 3-byte windows answers each query with a byte-for-byte check.
- Crawl. PageSourceSearchBot fetches a site's root page, a handful of further pages and the scripts they load, first-party hosts only, and stores the exact bytes.
- Classify. A classifier separates a site's own code from library code (jQuery, React and the like); libraries are stored and shown in previews but not indexed, so results are about the site, not its dependencies.
- Index and search. Every stored file is indexed by its 3-byte windows. A query looks those up, verifies the candidates byte for byte (or with your regex) and lists the sites that hold every term.
What the crawler fetches, its robots.txt token and how to opt out: About PageSourceSearchBot.
What a query can be and why one is rejected: Query help.
Common questions
- What is PageSourceSearch?
- PageSourceSearch is a search engine over the source code of websites. Its crawler stores the raw HTML and first-party JavaScript of each site it visits, and a query (an exact byte sequence or a regular expression) lists every site that contains it, grouped by domain, with each match highlighted in the stored file.
- How is it different from a web search engine?
- A web search engine indexes the visible text of rendered pages and ranks them by relevance. PageSourceSearch indexes the source bytes, so tags, attributes, script code, tracking IDs, comments and markup patterns are searchable, and a match is exact: what you type is what the page contains.
- Is the search case-sensitive?
- Yes. A literal query is matched byte for byte, so Gtag and gtag are different queries. Choose regex mode under Filters and start the pattern with (?i) for a case-insensitive match.
- What happens to a query of 1 or 2 characters?
- The index is built from 3-byte windows of every file, so a shorter term has no window of its own. It is searched with a space added after it, and before it too for a single character: 'ab' finds 'ab ' and 'a' finds ' a '. A regex needs a fixed run of 3 characters.
- Can I search across a whole site rather than one page?
- Yes, and that is the default: with Same site selected next to the search box the words may be spread over different pages and scripts of one domain. Switch to Same page to require them all in one file.
- Can I keep my site out of the index?
- Yes. Add 'User-agent: PageSourceSearchBot' followed by 'Disallow: /' to your robots.txt, or write to [email protected] to have stored pages removed.
Full query help and FAQ · Regular expressions · How to opt out · JSON API