PageSourceSearch

Site profile

bigdataref.github.io

Built with bootstrap 3.3.1, jquery 1.11.1, google-analytics. 1 page and 6 stored files, last crawled 2026-10-03.

3 libraries 1 identifier 68 third-party hosts 1 crawl open the site

Technologies and libraries

Recognised by fingerprinting the stored files and bundles against known releases; a version is the release the bytes match.

LibraryVersionSeen
bootstrap 3.3.1 2026-10-03
jquery 1.11.1 2026-10-03
google-analytics 2026-10-03

Tracking, tag and verification IDs

Public identifiers read out of the source. Each links to every other site in the index that carries the same one.

WhatValueFound in
Universal Analytics property id (Google) UA-59841923-1 view source

Third-party hosts

Hosts outside bigdataref.github.io that its pages load scripts, frames or stylesheets from, and that its own scripts name in absolute URLs (connect). Each links to every site loading from the same host.

HostLoaded asFound in
accumulo.apache.org connect view source
adobe-research.github.io connect view source
ajax.googleapis.com script view source
ambari.apache.org connect view source
aws.amazon.com connect view source
cassandra.apache.org connect view source
cockroachdb.org connect view source
confluent.io connect view source
deeplearning4j.org connect view source
drill.apache.org connect view source
druid.io connect view source
ebay.com connect view source
falcon.apache.org connect view source
flink.apache.org connect view source
flume.apache.org connect view source
forge.gluster.org connect view source
freeman-lab.github.io connect view source
gethue.com connect view source
github.com connect view source
gopulsar.io connect view source
hadoop.apache.org connect view source
hbase.apache.org connect view source
hive.apache.org connect view source
hyperdex.org connect view source
ignite.incubator.apache.org connect view source
impala.io connect view source
influxdb.com connect view source
kafka.apache.org connect view source
mahout.apache.org connect view source
maxcdn.bootstrapcdn.com script stylesheet view source
mesos.apache.org connect view source
neo4j.com connect view source
nifi.incubator.apache.org connect view source
oozie.apache.org connect view source
parquet.incubator.apache.org connect view source
pig.apache.org connect view source
prestodb.io connect view source
redis.io connect view source
samza.apache.org connect view source
spark.apache.org connect view source
sqoop.apache.org connect view source
storm.apache.org connect view source
stratio.github.io connect view source
tachyon-project.org connect view source
tez.apache.org connect view source
thrift.apache.org connect view source
urbanairship.com connect view source
www.adobe.com connect view source
www.aerospike.com connect view source
www.cloudera.com connect view source
www.datastax.com connect view source
www.ebayinc.com connect view source
www.elasticsearch.com connect view source
www.elasticsearch.org connect view source
www.facebook.com connect view source
www.gluster.org connect view source
www.hortonworks.com connect view source
www.intel.com connect view source
www.kylin.io connect view source
www.linkedin.com connect view source
www.mongodb.com connect view source
www.mongodb.org connect view source
www.neo4j.com connect view source
www.ooyala.com connect view source
www.rabbitmq.com connect view source
www.redhat.com connect view source
www.stratio.com connect view source
zookeeper.apache.org connect view source

Stored pages and scripts

The newest stored version of each file, newest first@if (p.FilesTruncated) { (the 500 newest of 6) }. Versions are the crawls at which the file was new or its content changed; each opens the file as it was then, rebuilt from the same stored chunks.

FileKindSizeCollectedVersions
https://ajax.googleapis.com/ajax/libs/jquery/1.11.1/jquery.min.js js library not fetched 2026-10-03 2026-10-03
https://bigdataref.github.io/js/main.js js 2.1 KB 2026-10-03 2026-10-03
https://maxcdn.bootstrapcdn.com/bootstrap/3.3.1/js/bootstrap.min.js js library not fetched 2026-10-03 2026-10-03
https://bigdataref.github.io/js/i.js js 43.2 KB 2026-10-03 2026-10-03
https://bigdataref.github.io/ html 231.8 KB 2026-10-03 2026-10-03
https://bigdataref.github.io/js/lunr.min.js js 14.1 KB 2026-10-03 2026-10-03

Timeline

One entry per crawl, newest first, with what changed since the crawl before: libraries, identifiers, third-party hosts and files. Historical versions stay viewable because the stored chunks are shared between versions, never copied.

  1. first crawl

    1 page, 6 files, 291.2 KB; 3 libraries, 1 identifier, 68 third-party hosts

    3 libraries added
    • google-analytics
    • jquery 1.11.1
    • bootstrap 3.3.1
    1 identifier added
    69 third-party hosts added
    6 files added

Questions about bigdataref.github.io

How does PageSourceSearch know what bigdataref.github.io is built with?
From the site's own source code. The crawler stores the exact bytes of its pages and first-party scripts; the libraries are recognised by fingerprinting that code against known releases, the identifiers are read out of the tag snippets, and the third-party hosts are the script, iframe and stylesheet sources in the HTML and the absolute URLs inside the scripts. Nothing is inferred from headers or guessed.
How far back does the history of bigdataref.github.io go?
To the first crawl the timeline lists. Every crawl records what the site looked like; when a file's content changes, its earlier version stays viewable because the stored chunks are never deleted, only mapped. A crawl that finds a file unchanged adds no copy.
Can I see an older version of a script from bigdataref.github.io?
Yes. In the stored files list, every date under a file opens the version that was live at that crawl, with the same viewer as the current one. The raw bytes can be downloaded from there.
Which other websites use the same tracking IDs as bigdataref.github.io?
Each identifier links to its reverse lookup, the list of every site in the index whose source carries the same id: the sites one analytics property or tag manager container is shared across.

Look up another site, browse sites by tracking ID, third-party host or technology. Site owners: see the crawler page for how stored pages are removed.