Precomputed tables of all revisions a content is in, and all origins a revision is in, indexed for reasonably efficient backward queries. Its current implementation is primarily a Parquet database, along with some external indexes for more efficient access. The swh-provenance Rust crate provides access to these indexes and a gRPC server to query the data remotely.
Referencing Software Heritage
If you use any of the datasets indexed on this website for research purposes, please acknowledge Software Heritage as recommended in the publications page, that is:
Add a footnote on the title page of your paper, formatted as: “This work was made possible by Software Heritage, the universal source code archive: https://www.softwareheritage.org”; and