When The Metropolitan Museum of Art launched its Open Access initiative in 2017, releasing 375,000 public domain images under Creative Commons Zero (CC0), some questioned the return on investment. Why give away your most valuable digital assets for free?
The answer, four years later, was undeniable: page views of Met collection images on Wikipedia increased by 576% between April 2017 and April 2020.
The Discovery Problem
Museums have a discovery problem. Their websites are beautifully designed but depend on people knowing to visit them. Wikipedia, by contrast, is the world’s fifth most-visited website. When someone searches for “Henry VIII,” “Vincent van Gogh,” or even “pineapples,” they land on Wikipedia — and the images they see there increasingly come from museum open access programs.
The Met’s images now illustrate articles on topics ranging from Tudor England to the Japanese economy, from Down Syndrome to military history. These aren’t art history articles. They’re the articles people actually read, millions of times per day.
Wikimedia Commons: The Infrastructure Layer
The key insight is that Wikimedia Commons isn’t just a photo repository — it’s infrastructure. Over 100 million media files, all freely licensed, searchable, and connected to Wikipedia articles in 300+ languages. Upload once, appear everywhere.
When 375,000 Met images hit Commons, they became available to:
- Wikipedia editors in every language
- Educational publishers creating textbooks
- Documentary filmmakers seeking B-roll
- AI researchers training computer vision models
- Students writing papers
No museum website can replicate that reach.
Wikidata: The Structured Layer
Images get you visibility. Structured data gets you intelligence.
At The Met, we created Wikidata items for 14,000+ artworks, adding machine-readable metadata like “depicts,” “inception,” “material used,” and “creator.” These items connect to the broader Wikidata knowledge graph of 100 million+ concepts.
The result: when a voice assistant answers “What’s at The Met?”, the data comes from Wikidata. When Google Knowledge Panel shows facts about an artwork, Wikidata is the source. When a researcher queries “show me all oil paintings from 1600-1700 depicting cats,” the answer can be computed from structured data.
The Pattern
The Met wasn’t the first and won’t be the last. The Smithsonian, the Cleveland Museum of Art, the National Archives — the pattern repeats:
- Open access release: CC0 images, freely downloadable
- Bulk upload to Commons: Automated tools like Pattypan or GLAMpipe
- Wikidata enrichment: Create items, add statements, cross-link
- AI-assisted metadata: Machine learning suggests depicts tags, humans validate
- Impact: Increased visibility, new audiences, unexpected reuse
The Future
I’m now bringing this model to MIT Open Learning, connecting 25 years of OpenCourseWare materials with Wikimedia projects. The same pattern applies: freely-licensed educational content, uploaded to Commons, structured in Wikidata, discoverable by anyone.
Museums and universities aren’t giving away their assets. They’re investing in the discovery infrastructure of the 21st century. The returns — in visibility, impact, and mission fulfillment — speak for themselves.