← Research

Research note · August 10, 2026

Semantic search across institutions.

The Met's semantic search launch is a landmark moment for the field. Here is what becomes possible when that same capability crosses ten institutions, languages, and borders.

linkedculturesemanticsearchculturalheritagemultilingualmuseumtechdigitalhumanitiesopenaccess


On August 5, 2026, the Metropolitan Museum of Art published a clear and honest account of how they built semantic search into their online collection of more than 500,000 works. It is worth reading. The team notes that only about a third of metmuseum.org visitors search with a specific artwork in mind; most explore by theme, subject, material, and mood, and that is exactly where keyword matching runs out.

The article demonstrates the gap with queries like "mythical beasts," "people dancing in a circle," and "teapots shaped like produce." All of them work.

For anyone working in cultural heritage discovery, this is a significant moment. One of the world's great institutions has put the infrastructure and the argument in public.

LinkedCulture has been working on the same problem from a different angle.

The same technology, across ten institutions

LinkedCulture's semantic index now holds 525,298 open-access records across ten institutions: the Metropolitan Museum of Art, the Rijksmuseum, the Art Institute of Chicago, Paris Musées, the Cleveland Museum of Art, the Minneapolis Institute of Art, Harvard Art Museums, the Getty, a growing Joconde set from the French Ministry of Culture, and Museo Egizio.

When you search LinkedCulture, you are not searching one collection. You are searching all ten simultaneously, ranked by semantic similarity, with every result linking back to its home institution. The canonical record always stays with the museum that holds it.

525,298 objects across ten museums, ranked by size
525,298 objects across ten museums, ranked by size

The queries the Met demonstrates work in LinkedCulture too. But the results draw from Amsterdam, Paris, Chicago, Cleveland, and Turin in a single result set. A search for "mythical beasts" surfaces objects from cultures and centuries that no single institution holds together. "People dancing in a circle" crosses continents. "Cozy interiors" moves from Dutch Golden Age painting to Japanese woodblock prints to French decorative objects in one view.

Mythical beasts across ten institutions

People dancing in a circle across ten institutions

Cozy interiors across ten institutions

Stormy seas across ten institutions

Two layers the Met's account points toward

LinkedCulture's retrieval goes beyond text metadata in two ways that are relevant to where the field is heading.

The first is image description. LinkedCulture runs object images through a vision-language model that generates a full narrated description of each object, covering what it is, what you see, why it exists, and why it matters. These descriptions are embedded into a separate vector collection and fused into retrieval as a distinct scored leg alongside the structured metadata search. A query like "a figure in motion" can surface objects whose titles and catalog records contain no such language, because the visual content has been independently described and indexed. When the caption leg contributes to a match, users see an "AI-enhanced" badge on the result. The authoritative museum record remains what is displayed. The Met's article points toward this approach. LinkedCulture has implemented it across the cross-institutional index.

The second is multilingual retrieval. LinkedCulture holds records in English, French, and Dutch. Getting a query in one language to surface records catalogued in another is not a solved problem at cross-institutional scale. LinkedCulture Paper 3, Semantic Clusters and Language Partitioning in a Multi-Institutional Cultural Heritage Index, documents what actually happens when multilingual embedding models encounter natively multilingual metadata across institutions. The finding is that the document space partitions strongly by cataloging language even when the embedding model supports multiple languages, and that cross-language discovery requires orchestration rather than model substitution alone. That research is available at the link below.

The practical result is that LinkedCulture now translates queries across the cataloging languages present in the corpus, running retrieval in each language and merging the results. A search in English reaches French and Dutch records. A search in French reaches English and Dutch records.

What the cross-institutional layer reveals

Single-institution semantic search surfaces kinships within a collection. Cross-institutional semantic search surfaces kinships the field has never had infrastructure to see before.

A Persian textile at the Art Institute sits beside its counterparts at the Met and the Rijksmuseum, discoverable together for the first time. Works from Paris Musées appear alongside related objects from Cleveland and Chicago. The connections were always there in the objects themselves. The infrastructure to surface them was not.

This is what LinkedCulture is built to do. Not to replace any institution's catalogue or search system, both of which the Met has demonstrated can be excellent, but to build the layer that connects them.

Explore it: https://linkedculture.org/

The research behind this work is documented in the paper series:

Paper 1: Discovery Architecture for Cultural Heritage: Layered Retrieval, Institutional Authority, and the Limits of Keyword Search https://linkedculture.org/research/discovery-architecture-for-cultural-heritage

Paper 2: Embedding Cultural Heritage Metadata: Pipeline Design, Multilingual Retrieval, and Hybrid Search in LinkedCulture https://linkedculture.org/research/embedding-cultural-heritage-metadata

Paper 3: Semantic Clusters and Language Partitioning in a Multi-Institutional Cultural Heritage Index https://linkedculture.org/research/semantic-clusters-language-partitioning