PageSourceSearch

Site profile

datasetlist.com

Built with gtag, @vitest/utils 5.0.2, vue-axios-plugin 1.3.0, json-db-admin 0.0.1, udesly-ad-banner 0.0.4. 3 pages and 7 stored files, last crawled 2026-10-02.

rank 498,772 5 libraries 1 identifier 104 third-party hosts 1 crawl open the site

Technologies and libraries

Recognised by fingerprinting the stored files and bundles against known releases; a version is the release the bytes match.

LibraryVersionSeen
gtag 2026-10-02
@vitest/utils 5.0.2 2026-10-02
vue-axios-plugin 1.3.0 2026-10-02
json-db-admin 0.0.1 2026-09-24
udesly-ad-banner 0.0.4 2026-09-24

Tracking, tag and verification IDs

Public identifiers read out of the source. Each links to every other site in the index that carries the same one.

WhatValueFound in
Universal Analytics property id (Google) UA-55917238-5 view source

Third-party hosts

Hosts outside datasetlist.com that its pages load scripts, frames or stylesheets from, and that its own scripts name in absolute URLs (connect). Each links to every site loading from the same host.

HostLoaded asFound in
academictorrents.com connect view source
acdc.vision.ee.ethz.ch connect view source
agamenon.tsc.uah.es connect view source
ai.facebook.com connect view source
ai.google connect view source
ait.ethz.ch connect view source
allenai.github.io connect view source
allennlp.org connect view source
amazon-berkeley-objects.s3.amazonaws.com connect view source
arxiv.org connect view source
avdata.ford.com connect view source
bigearth.net connect view source
bimcv.cipf.es connect view source
bozcani.github.io connect view source
brixia.github.io connect view source
buildingnet.org connect view source
cadcd.uwaterloo.ca connect view source
components.one connect view source
cs.binghamton.edu connect view source
cv.iri.upc-csic.es connect view source
cvit.iiit.ac.in connect view source
dbarnes.github.io connect view source
deepfakedetectionchallenge.ai connect view source
diode-dataset.org connect view source
dramaqa.snu.ac.kr connect view source
drive.google.com connect view source
fashionpedia.github.io connect view source
generated.photos connect view source
gigaword.dk connect view source
github.com connect view source
hacs.csail.mit.edu connect view source
homes.cs.washington.edu connect view source
humaninevents.org connect view source
ieeexplore.ieee.org connect view source
imagemonkey.io connect view source
interaction-dataset.com connect view source
ixa.eus connect view source
jrdb.stanford.edu connect view source
kaldir.vc.in.tum.de connect view source
knowit-vqa.github.io connect view source
landcover.ai connect view source
leiainc.github.io connect view source
matterport.com connect view source
medmnist.com connect view source
movienet.site connect view source
nlp.cs.washington.edu connect view source
objectnet.dev connect view source
once-for-auto-driving.github.io connect view source
openaccess.thecvf.com connect view source
openreview.net connect view source
p-destre.di.ubi.pt connect view source
podcastsdataset.byspotify.com connect view source
products-10k.github.io connect view source
project.inria.fr connect view source
research.mapillary.com connect view source
rrc.cvc.uab.es connect view source
scale.com connect view source
sites.google.com connect view source
sites.research.google connect view source
storage.cloud.google.com connect view source
storage.googleapis.com connect view source
sviro.kl.dfki.de connect view source
tabfact.github.io connect view source
taodataset.org connect view source
textvqa.org connect view source
tvqa.cs.unc.edu connect view source
unsplash.com connect view source
use.fontawesome.com stylesheet view source
uwaterloo.ca connect view source
v-sense.scss.tcd.ie connect view source
vcl3d.github.io connect view source
vladlen.info connect view source
waymo.com connect view source
www.4seasons-dataset.com connect view source
www.a2d2.audi connect view source
www.aclweb.org connect view source
www.agriculture-vision.com connect view source
www.aiskyeye.com connect view source
www.astyx.com connect view source
www.atticusprojectai.org connect view source
www.biomotionlab.ca connect view source
www.cbsr.ia.ac.cn connect view source
www.crowd-counting.com connect view source
www.cs.albany.edu connect view source
www.cse.ust.hk connect view source
www.derczynski.com connect view source
www.eecs.yorku.ca connect view source
www.face-benchmark.org connect view source
www.googletagmanager.com script view source
www.ind-dataset.com connect view source
www.isprs-ann-photogramm-remote-sens-spatial-inf-sci.net connect view source
www.mapillary.com connect view source
www.medicalimageanalysis.com connect view source
www.mut1ny.com connect view source
www.objects365.org connect view source
www.panda-dataset.com connect view source
www.robots.ox.ac.uk connect view source
www.semantic-kitti.org connect view source
www.semanticscholar.org connect view source
www.si.edu connect view source
www.tau-nlp.org connect view source
www.uni-mannheim.de connect view source
www.w3.org connect view source
xview2.org connect view source

Stored pages and scripts

The newest stored version of each file, newest first@if (p.FilesTruncated) { (the 500 newest of 7) }. Versions are the crawls at which the file was new or its content changed; each opens the file as it was then, rebuilt from the same stored chunks.

FileKindSizeCollectedVersions
https://www.datasetlist.com/js/app.js js 14.5 KB 2026-10-02 current
https://www.googletagmanager.com/gtag/js?id=UA-55917238-5 js library not fetched 2026-10-02 current
https://www.datasetlist.com/js/datasets.js js 249.6 KB 2026-10-02 current
https://www.datasetlist.com/js/vue.min.js js library 83.9 KB 2026-10-02 current
https://www.datasetlist.com/privacy/ html 6.4 KB 2026-10-02 current
https://www.datasetlist.com/tools/ html 290.8 KB 2026-10-02 current
https://www.datasetlist.com/ html 118.1 KB 2026-10-02 current

Timeline

One entry per crawl, newest first, with what changed since the crawl before: libraries, identifiers, third-party hosts and files. Historical versions stay viewable because the stored chunks are shared between versions, never copied.

  1. first crawl

    3 pages, 7 files, 763.3 KB; 3 libraries, 1 identifier, 104 third-party hosts

    3 libraries added
    • @vitest/utils 5.0.2
    • gtag
    • vue-axios-plugin 1.3.0
    1 identifier added
    104 third-party hosts added
    7 files added

Questions about datasetlist.com

How does PageSourceSearch know what datasetlist.com is built with?
From the site's own source code. The crawler stores the exact bytes of its pages and first-party scripts; the libraries are recognised by fingerprinting that code against known releases, the identifiers are read out of the tag snippets, and the third-party hosts are the script, iframe and stylesheet sources in the HTML and the absolute URLs inside the scripts. Nothing is inferred from headers or guessed.
How far back does the history of datasetlist.com go?
To the first crawl the timeline lists. Every crawl records what the site looked like; when a file's content changes, its earlier version stays viewable because the stored chunks are never deleted, only mapped. A crawl that finds a file unchanged adds no copy.
Can I see an older version of a script from datasetlist.com?
Yes. In the stored files list, every date under a file opens the version that was live at that crawl, with the same viewer as the current one. The raw bytes can be downloaded from there.
Which other websites use the same tracking IDs as datasetlist.com?
Each identifier links to its reverse lookup, the list of every site in the index whose source carries the same id: the sites one analytics property or tag manager container is shared across.

Look up another site, browse sites by tracking ID, third-party host or technology. Site owners: see the crawler page for how stored pages are removed.