Skip to content
Skip to content

Part one.


Anne Chen (AC): I think Wiki infrastructure creates the possibility of a next-generation data set. We often think of “freezing” data sets.

Dan Shick (WMDE): In the sense that it becomes a canonical set of data? Finalized?

AC: Yeah, exactly, but knowledge is always moving and evolving. Something I find exciting about what we’re building — and we’re shaping it from many different perspectives — is that ideally, the data set keeps pace with the research. So maybe an art historian would be able to see when epigraphers have had a breakthrough on the date of a certain inscription, and the reasons why.

Which is not to say that the technology replaces traditional scholarship — but knowledge graphs could help anchor the unwieldiness of the large language models (LLMs) we’re exposed to through generative-AI tools from big data companies. That might represent a path forward with the kind of AI that many humanists are concerned about or, frankly, that they’re seeing in the classroom, which really is destructive for the learning process. They’re concerned students will be incentivized to use AI to bypass the learning process.

But I think there’s a separate conversation to be had about accessibility, especially multilingual accessibility. Responsibly used LLMs with scholarly curated knowledge graphs would provide a whole different way forward.

WMDE: Absolutely. An LLM is a natural-language processor par excellence. Right now, what they’re being used for is being told what you want to hear… but what if you want to be told what you want to hear, but only from Wikidata?

AC: Exactly, that’s the vision that I see too. I understand why my colleagues are reluctant, but I worry that we’ll end up not teaching our students how to responsibly engage with the tools that they’ll inevitably be using for the rest of their lives — I don’t think this is going away — and I worry that in raising the alarm appropriately about environmental impact, about attention erosion, about —

WMDE: De-skilling?

AC: Yeah — that we might overlook ways it could be used responsibly. I sometimes think there’s so much moralizing around it that it becomes difficult to have that separate conversation.

But the work that we’re doing from this decolonial angle represents an important opportunity to explore together how we use technology to tackle real-world problems, like the fact that nobody is going back and translating a hundred years of scholarship, material sitting in dusty archives, for the benefit of communities who might like to learn from those materials. So can real human work steer the tools towards these persistent problems? It won’t be perfect, but maybe there’s a way that AI could help us share information, lower barriers even further — even in the Wikimedia community.

This was one of the things discussed at the AI Bridges symposium: how do we use these tools to do good work responsibly?

Dr. Anne Chen speaking at the AI Bridges conference in London, England (May 2026).
Dr. Anne Chen speaking at the AI Bridges conference in London, England (May 2026).

WMDE: Could you talk about how you use Wikidata and Wikimedia Commons together?

AC: Well, our project started before there was structured data for Commons, so we developed our workflows using Wikidata’s infrastructure. We had uploaded media into Commons, but we made the metadata statements directly in Wikidata. Now structured data has taken off — I’m so excited about it — but the structured data tab is tucked away. Even some experienced Wikimedians I’ve worked with were surprised to learn that there’s structured data on Commons.

We’re using Commons to store and share complex metadata about archival photographs, and we’re using it for geo-shapes. A particular building has a digital geographic footprint we can associate with it. That makes queries possible like, “Show me all the things found in this particular house and the footprint of that house.” So, in the example that I gave you before, with the pieces of the wall painting in two different museums, we also have objects at the Damascus Museum that were found in that same room, and we have photographs that document that particular context. We can pull all of that together, alongside the building’s location represented as a geoshape, with a query.

Our Syrian colleagues, who we trained in that early NEH pilot work I described, received training in oral-history methods, then held photo-elicited conversations (some in person, some via video call) with people in the community. They would show a photograph, say, from the 1930s and use it as a jumping-off point for a conversation about community traditions connected with the site, or alternative place names. Many people who live in those villages worked on the excavations for long periods of time and have a ton of valuable knowledge that they’re eager to share.

We’ll be able to weave together the scientific documentation and these local traditions. In the sub-discipline of community archaeology, there’s a framework for this called “braided knowledge”: local interpretations and scientifically grounded information seen as two sides of the same coin, mutually enriching. We’re using the Wiki ecosystem as a tool to achieve just that.

In trying to be sensitive to the digital divide, I think well-meaning people assume it’s not possible to do meaningful digital work. It requires flexibility and sensitivity to the circumstances, but our project shows it’s possible, and that we could be doing even more.
Dr. Anne Chen

We think of the knowledge graph, the open knowledge work, as important in and of itself, but I think sometimes looking at data makes laypeople’s eyes glaze over. We’ve been working with the Stories Collective; they’ve built a [more browsable] interface that sits in front of the portion of the knowledge graph our team is developing. It maps things based on the knowledge graph in the background, so you don’t have to click through all the individual links. That browser is linked from our web page — the Dura Stories interface.

WMDE: Why did you choose to do this project now? Was it because of the change in conditions in the conflict in Syria?

AC: That’s a great question. It’s really important to highlight that we’ve been doing this work since 2022. We all wanted to work together, and [because of the conflict] working digitally was one of the only ways it was possible. We’ve been able to continue this work digitally, even in really dire circumstances, when team members were forced into refugee circumstances.

In trying to be sensitive to the [all-too-real] digital divide, I think well-meaning people assume it’s not possible to do meaningful digital work. It does require flexibility and sensitivity to the circumstances, but our project shows it is possible, and that we could be doing even more. When we started, we were just trying to find ways to work together, across three continents, widely different time zones, and in two languages. We’ve been doing all the digital and oral-history training with the help of Asmaa Shehadeh, the team’s Arabic translation coordinator. I think it’s remarkable we’ve been able to build this skill set as a group, from the ground up.

In a way, it’s addressing digital colonialism in a different way. From the institutional side, we don’t know what we don’t know: I don’t know what vocabulary is missing from existing authority resources, or what my colleague might see in a given photograph that should be highlighted in the metadata. We need more conversations and skill building, at a slow enough pace that allows us to build relationships and, eventually, revise the metadata together in more inclusive ways that better acknowledge the multiple perspectives on a site like this.

Dr. Anne Chen presents "IDEA: Toward a FAIR-er Archive: Wikidata" and “Big Digs” at Wikidata Day 2023, New York City.
Dr. Anne Chen presents “IDEA: Toward a FAIR-er Archive: Wikidata” and “Big Digs” at Wikidata Day 2023, New York City.

WMDE: With some of this work, it sounds like you could only have done it digitally, especially when the conflict was at its peak. Is that a good assessment?

AC: Yeah, certainly at the height of the conflict. We designed things the way we did in part because, in 2022, there was no end to the conflict in sight. Things have shifted quite a bit since December 2024. There are new possibilities for supporting colleagues in documenting their collections, to ensure that their collections are interoperable with related international collections, and things like that.

WMDE: I saw that you took a WikiEdu course on Wikimedia. What did you learn in that course that enabled you to move forward with this project?

AC: First, a huge shout out to the folks at Wiki Education. They made my entry into the Wikimedia ecosystem a lot less scary. It was great knowing that I had somebody I could reach out to with questions, somebody who wasn’t going to judge me for not already knowing some particular thing. They paced things well and modeled a kind of low-lift [approach] — start with what you know — and getting to know the way the community works. I’ve modeled that now in the way I’m teaching other people to come into the ecosystem.

I’d been curious about how to implement Linked Open Data, and intellectually I understood some of the promise, but until I took that class, came into the community and began doing actual hands-on work, it hadn’t really clicked for me.

For Dura-Europos, we’ve built what I call a “learning laboratory” where we continue our work on the data associated with this important archaeological site, while also providing hands-on experience for colleagues and students. That hands-on piece gets people beyond just an intellectual understanding. Until you have your own data set to tinker with, it can be hard to get your head around.

The Dura data set is intersectional in many ways. It’s multilingual. It includes community perspectives and scholarly perspectives. Archaeology has brought material culture out of the ground for a lot of different humanities: the inscriptions that classicists work on; the wall paintings and sculptures that art historians work on; the languages that perhaps a Semitic language instructor teaches the specifics of.

When we bring all of those people together into a space where they can learn about the intersection of humanities and data and contribute from their own disciplinary perspective, it can show them what’s useful about this particular digital way of working, and the semantic web in general. I think that’s a valuable way to build digital literacy and a community of next-generation scholars who are thinking about how we can make our disciplines work better together.

All that was modeled through the Wiki Education course. The didactics they use in that course are open, so I’ve used them to bring other people along.

I’ve been thinking a lot about how the Wikimedia environment enables a kind of “cascading mentorship”. Will, my instructor at WikiEdu, gave me resources and a community, helped me to find my footing so that I felt comfortable teaching some graduate students. Now they’re teaching undergraduates and others. We’re starting to see that happening with our Syrian colleagues as well.

Will Kent, Scholars & Scientists Program Manager at Wiki Education.
Will Kent, Scholars & Scientists Program Manager at Wiki Education.

WMDE: That’s great. If there’s one thing that gives me hope, it’s that I keep hearing people say, “I got taught this thing, and I had the opportunity to pass it on, and I did it with great relish.” People are really interested in learning and then turning around and teaching.

A last question: in one of the papers you wrote, you discussed LOD blind spots. What are those blind spots? Maybe you have some advice for people interested in doing Linked Open Data work.

AC: There are definitely blind spots in the granular work that I was talking about before. Somebody who’s doing LOD work on a colonially entangled site will likely be connecting to the authority gazetteer for their particular disciplinary area. For me it’s the gazetteer Pleiades, the authority that disambiguates between, say, this Alexandria and that Alexandria. All these authorities have been shaped by the digitization of collections in the global north, so they’re going to have gaps and biases.

People don’t realize that those projects, and Wikidata too, are aware that they’re incomplete and can always be improved. There’s a broader community of scholars who could be making micro-contributions — like a transliteration that applies to a particular site name that hasn’t yet been captured — that could cause significant ripple effects into the broader LOD community.

That’s something that we should all be striving for. We could all be enriching the architecture that sits between the big institutional databases and standalone disciplinary projects. I see Wikidata as one way to do that.

We’re seeing some standardization in the LOD ecosystem, and we do need to build consistency for the sake of interoperability. But if we in the global north define a particular ontology as the only way for museums to structure their data, we risk uncritically reproducing top-down frameworks, and we’re missing out on some of the opportunity that linked data provides.

The CiDOC model, oriented on the museum and heritage domain. CiDOC CRM facilitates the structuring of data to represent real world phenomena.
The CiDOC model, oriented on the museum and heritage domain. CiDOC CRM facilitates the structuring of data to represent real world phenomena.

I’m talking here about CiDOC CRM, the de facto museum ontology. CiDOC is grounded in Western ways of thinking and organization. That’s not to say that it’s not valid — it is, and it’s a great standard, but it’s not available in Arabic. It’s very difficult to learn. It’s also a hierarchized and event-based ontology, which means that it doesn’t necessarily work in every context, especially when we’re trying to be more sensitive to different ways of knowing.

So by building knowledge graphs in flexible ecosystems, we’ve begun finding where there are points of commonality with CiDOC, and maybe other places where we need a separate property to capture a different way of knowing. In a place like Wikidata, we can define that together as a community, in a way that I don’t think exists in more rigid technology.

WMDE: Anything else you’d like to mention?

AC: Well, I’ve been thinking a lot about the relationship between Wikibase and Wikidata. First, let me give you a little bit of context around why our particular project looks the way that it does, and what I might do differently in a different life, knowing what I know now.

I come from these elite institutions where I’ve always had access to whatever I wanted. When I started the work in 2020, in the deep pandemic, I gained a visceral insight into what it was like not being able to access library resources. Also, there was no Wikibase Cloud at that time, and it was important to me to build a replicable pipeline, to be sensitive to what the replication barriers might be. I didn’t want to be showing other people, especially in Syria, how to do something they would ultimately have to host long-term.

Now that there’s Wikibase Cloud, certain kinds of content… you know, maybe not every loom weight for Dura-Europos belongs in Wikidata. I’m trying to signal that we should be careful about how much goes into Wikidata, what intersectional stuff should be enriching Wikidata as opposed to [stored in] Wikibases, now that that’s an option.

WMDE: Right. And some folks host their own Wikibase Suite instances as well.

AC: That might present too high a barrier for our colleagues in Syria, especially regarding long-term sustainability. Hosting with Wikimedia is a big value-add for data set sustainability. But certainly [Wikibase Suite] for collections without those particular parameters that have a little bit of IT support.

Working in Wikidata in particular has forced us to bump up against the different ways that institutions manage their data. I think a little bit of friction is good for moving the field forward, for building interoperability. If you can go into your Wikibase and do your own thing, and you’re not forced to think expansively about how other disciplines or institutions are doing things, it’s easy to end up just minting your own property.

WMDE: So you’re saying that interoperability should always be on your mind — even if you’re making your own Wikibase, make sure that it can be part of the LOD network.

AC: Right. Maybe there are certain things that you enrich in Wikidata, because it can benefit us all — people, places, vocabulary terms, things like that. Maybe your individual collection items live in a Wikibase. But I don’t think what should live in each place is perfectly clear to anybody yet.

WMDE: Agreed. Thank you so much for your time, and thanks for the work you’re doing!

AC: I appreciate it!


 

A headshot of Dr. Anne Hunnell Chen

Dr. Anne Hunnell Chen is Assistant Professor of Art History and Visual Culture at Bard College. Dr. Chen specializes in the art and archaeology of the globally-connected Roman world and is committed to exploring how low-barrier Linked Open Usable Data (LOUD) can provide more equitable access to archaeological data in the digital realm and empower stakeholder audiences as collaborative curators. She is the founder and co-director of the NEH-funded International Digital Dura-Europos Archive (IDEA), an archaeological data accessibility project whose documentation efforts are aimed at sharing out workflows that help to overcome disciplinary data silos and work to dislodge enduring impacts of colonialism.

Leave a Reply

Your email address will not be published. Required fields are marked *