← Research

Research note · June 24, 2026

Semantically derived collection clusters cross institutions effortlessly. Across langua...

That is the finding in Paper 3 of the LinkedCulture series, and it overturns something I assumed in Paper 2.

linkedcultureculturalheritagemuseumtechaccessibilitypublicdomaindigitalhumanitiesopenaccessmuseumcollections


Semantically derived collection clusters cross institutions effortlessly. Across languages, they do not cross at all.

Paper 3 is open access on Zenodo: https://linkedculture.org/research/semantic-clusters-language-partitioning

That is the finding in Paper 3 of the LinkedCulture series, and it overturns something I assumed in Paper 2.

When you build a multilingual museum search, you expect a multilingual model to create a shared space where a French object and an English object about the same thing sit near each other. They do not.

I analyzed a multi-institutional corpus of 302,798 cultural heritage records, with a French-language bloc big enough to see the structure clearly. The result was stark: a French record's nearest neighbors were 99.9% French, even though French records were only 17% of the collection. The embedding space had quietly split itself in two, by language, not by meaning. You can actually see it in the galaxy map I shared earlier. The two French-language museums sit off on their own, a separate continent, while everything else clusters together regardless of subject.

So I tested whether a better model would fix it, running a bake-off across several multilingual embedding models. They narrowed the distance between languages, but none dissolved the partition or produced reliable cross-language ranking. An English search for "blue bird" still could not reach the French "L'Oiseau Bleu."

What actually worked was not a smarter model. It was translating the query and searching again. Orchestration, not representation.

The paper also documents a quieter change in the retrieval layer: we replaced Reciprocal Rank Fusion with normalized-score fusion, after production use showed RRF could elevate weak candidates in sparse-metadata records. The takeaway: for multilingual cultural heritage discovery, multilingual retrieval and cross-lingual representation are not the same thing. Crossing the language line is, for now, an orchestration problem, not a solved representation problem.

Paper 3 is Open access on Zenodo: https://linkedculture.org/research/semantic-clusters-language-partitioning