Help From Our Friends · an open knowledge web experiment

Wikipedia is not alone.

Alongside every Wikipedia article there is a wider open world: libraries that lend, museums that publish their own collections, scientists who post their papers openly, mappers and naturalists who chart the planet for free. This experiment invites them in. Pick an article, and watch its friends arrive.

Press Enter. The article arrives in a second; its friends stream in behind it.

Current friends, what they bring, and how they’re licensed

Internet Archive

Lends the books. A footnote’s ISBN becomes a copy you can borrow tonight.

openness? Public-domain scans free to read; in-copyright books lent, not copied.

OpenLibrary

Knows every edition of every book — and which ones are actually open.

openness? Open bibliographic data, downloadable in bulk.

OpenAlex

Finds the free, legal copy of the paper behind the citation.

openness? Catalog CC0. Only papers with an open copy are shown — each card names its licence; closed ones are counted, not carded.

arXiv

Keeps whole sciences open by construction — the preprint is the publication.

openness? Metadata CC0; each paper names its own licence.

The Met

Shares its own record of its own objects, public domain wherever it can be.

openness? Public-domain works released CC0, images included.

Art Institute of Chicago

The painting’s home, telling you about the painting.

openness? Public-domain images CC0, served over open IIIF.

IIIF collections

One protocol, many doors: manuscripts and artworks served by whichever institution holds them.

openness? Terms set per object by its holding institution, stated in each manifest.

DPLA

America’s union catalog — tens of millions of items from local libraries, archives and museums.

openness? Metadata CC0; each item’s rights stated by its holder.

Europeana

Europe’s answer: three thousand museums, libraries and archives behind one door.

openness? Only openly licensed items are shown; each card names its licence.

iNaturalist

A community’s living field guide — photographs with the observer’s name on them.

openness? Each photo carries its observer’s chosen licence; only openly licensed ones are shown here.

GBIF

Draws where a species has been seen, from hundreds of millions of records.

openness? Records CC0 or CC BY, stated per dataset.

OpenStreetMap

Maps the world down to the building — by hand, by volunteers.

openness? Map data ODbL: share-alike, credit the contributors.

Free Law Project

Publishes the law itself. The opinion, not a paywall.

openness? Court opinions are public domain: nobody owns the law.

Wikimedia Commons

Brings the photographs, freely licensed, of very nearly everything.

openness? Every file free-licensed or public domain, credit shown on the file.

Wikidata & Wikipedia

The hosts: one writes the article that convenes everyone; the other makes the introductions — it knows every friend’s name for every thing.

openness? Article text CC BY-SA 4.0; Wikidata CC0.

How it works

The goal of this experiment is to demonstrate that open knowledge is not just hugely successful, but also increasingly hugely interlinked. So when you load an article, a small script on our server pulls existing linking information (from citations and Wikidata) and then grabs context from those sources to enrich the article — in one of two ways, and each card says which.

Identifier

The article states an ISBN, DOI, OCLC, LCCN, PMID or arXiv id, and a collection answers to exactly it. The strongest claim a card can make.

Statement

Wikidata states the connection outright — this painting is Met object 11417, this species is iNaturalist taxon 48662, this place is here. The card credits the property.

Each shelf says who asked, too: when one friend answers several of the article’s links, its cards split into one labelled shelf per link — and in the opening section, works by the subject never share a shelf with works merely cited there.

Challenges and future opportunities

This is a demo and not intended for production. Among other challenges:

Page layout

Arbitrary content means great layout is somewhere between difficult and impossible. Work with designers on this challenge would be necessary (though even rudimentary implementations would likely be very enjoyable for certain types of data nerds!)

Content curation

Sources can return thousands of responses. (Think the Smithsonian on the Apollo Program, for example.) A gallery with a thousand items is not very helpful to the reader, so some sort of curation (or at least ability to tune algorithmic prioritization) would be necessary before widespread deployment.

Source curation

Similarly, there are many collections of open content these days. Picking and prioritizing them would be an important challenge if we wanted to expand this.

Metadata gaps

The recent scan of thousands of theses from historical figures will be nice sources of context, but very little of it has metadata yet. Ideally the fix is to deploy Wikipedian energy to other repositories to improve the metadata, not have it curated only inside Wikipedia.

Bot volume and caching

Because of the volume of Wikipedia, to be deployable at any sort of scale, this would likely need extensive caching and likely formal agreements with the other data providers.