Exploring the Richness of Culture and Technology

Mozilla AI at Internet Archive Europe: Owning Your AI Stack

On 25 June, Internet Archive Europe (IAE) hosted Davide Eynard and Thomas (toto) Bille from Mozilla AI at our Amsterdam space. The afternoon covered a question that sits close to the heart of what we do: when you depend on infrastructure you don’t control, what do you actually own?

Ada, Zangemann, and the case for tinkering

Davide opened with a children’s book: Ada & Zangemann. Ada is a girl who lives in a dumpster, salvages broken hardware, and builds things entirely her own. Zangemann builds beautiful, polished technology that no one else can modify or adapt. The clash between them drives the story, but Davide used it as a frame for something more immediate: most people’s relationship with AI today looks a lot more like Zangemann’s world than Ada’s. You use what you’re given, on the terms it’s offered, for as long as the provider decides to keep it available.

This has nothing to do with technology: it’s a power relationship. And that logic applies to memory institutions every bit as much as to individual developers.

The trade-offs are concrete

To make the point, Davide rebuilt a small tool he’d originally written by hand more than twenty years ago: a script to extract train timetables from a website too clunky to use directly. He produced the rebuilt version with an AI coding assistant, and it worked. But the result lived on someone else’s platform, not his own machine. And unlike the original, which taught him Perl and regular expressions he used for years afterward, this one taught him nothing. Convenience and ownership turned out not to be the same thing.

Search as activity, not action

One of the sharpest distinctions in Davide’s talk was between search as an action and search as an activity. When you type a question into a box and accept the answer, that’s an action. When you use an agent to follow a thread, evaluate what it finds, redirect it when it goes wrong, and build toward a conclusion over time, that’s an activity. The difference matters because the second approach keeps you in the loop. You catch mistakes. You steer.

The clearest example: Davide used an agent to track down the original source of a widely cited Bill Gates quote. The agent searched, hit dead links, found partial copies, and eventually installed a subtitle extraction tool autonomously, downloaded a YouTube video, and identified the exact moment Gates said the thing. Along the way, it returned to the Wayback Machine around a dozen times, working through broken URLs until it found a usable copy.

It was a quiet illustration of something IAE says often: preserved web history is not a nostalgia project. It is working infrastructure. AI systems now depend on it to check what is actually true.

A second demo used a locally run open source model to search the Rijksmuseum’s digital collection for images connected to alchemy and the pursuit of knowledge, producing usable results entirely on Davide’s own hardware. No cloud service, no rented compute, no data leaving the machine.

Otari: making ownership practical

Thomas (toto) Bille followed with a look at the infrastructure side of the problem. Mozilla AI has built Otari as an open-source LLM gateway: a single control plane for all your interactions with language models, whether you run a local model on your own server or route through a commercial API.

The features that generated the most discussion were practical ones. Budget controls granular enough to cap spending per user, per model, or per team. Guardrails that strip personal or sensitive data before it reaches any external provider. A federated router in development that will recommend which model to use based on real usage patterns across the community, with options to prioritise cost, quality, or energy use. Otari is fully self-hostable: if you don’t want your data passing through Mozilla AI’s servers, you don’t have to. The code is on GitHub.

The honest question

Someone in the room asked how Mozilla AI intends to stay financially viable if the tools are free and open source. Toto’s answer was direct: revenue from the hosted version, income from enterprise integration work, and a long-term bet on the community. It’s the same tension that runs through almost every public-interest digital project, including IAE’s own. There’s no clean resolution, but naming it honestly matters.

Why this conversation belongs here

The afternoon drew developers, researchers, and people from across the cultural heritage sector. The discussion after the talks ran for close to an hour.

What connected the room wasn’t a shared technical interest in language models. It was a shared unease about dependency. Memory institutions know what it looks like when access to knowledge sits on infrastructure you don’t own and can’t influence. IAE has spent years arguing that the rights archives have always held offline must be protected online too. The same holds for one layer up, to the AI systems now sitting on top of those archives.

Explore Mozilla AI’s open-source tools, including AnyAgent, AnyLLM, LlamaFile, and Otari, at mozilla.ai and github.com/mozilla-ai.

You can watch a replay of the presentation and conversation on Archive.org and see the slides online.

Mozilla AI at Internet Archive Europe: Owning Your AI Stack Read Post »

Rights on Paper Are Not Enough: Our Input to the EU Copyright Review

A copyright exception you cannot use is not really an exception.

That idea runs through the submission Internet Archive Europe (IAE) submitted to the European Commission this month. On 25 June, we responded to the Commission’s Call for Evidence on copyright, the review that will shape what users, libraries, archives and museums across Europe can do with digital materials for years to come.

The Commission asked for input on four areas. We answered each one, and we added a point the consultation left out: the basic rights memory institutions need to do their work online. Here is what we told them.

Protect the preservation work that the law already allows

European law already lets libraries, archives and research organisations mine text and data, including for training AI models, when the purpose is preservation or research. Lawmakers made a deliberate choice to protect that public-interest work without giving rightsholders a veto over it.

That protection is being quietly undone. Large rightsholders now apply blanket opt-outs, built for commercial AI, across the board. They draw no line between a company training a product and an archive preserving the record, and so they block the very preservation work the law set out to protect.

We asked the Commission to confirm what the Directive already implies: an opt-out designed for commercial use cannot override the preservation and research rights of cultural heritage institutions.

Tackle piracy without breaking the open web

Piracy of live events is a real problem. But the answer some governments have reached for does more harm than the problem it sets out to solve.

We pointed to two examples. Italy’s Piracy Shield blocks content so bluntly that the Commission itself wrote to Rome in 2025, finding the system out of step with fundamental rights. In Spain, courts have ordered VPN providers to block access in proceedings where those providers had no chance to speak, and ordinary websites get caught in the net.

These are enforcement tools built for commercial pirates, used without proper judicial oversight, and the damage falls on people and services that did nothing wrong. We asked the Commission to measure that damage before legislating further, and to rule out using live-event blocking against general, non-commercial websites.

One research exception, not twenty-seven

Right now, the EU’s research exception is optional. Member States have implemented it differently, leaving researchers facing 27 distinct legal environments and real barriers to working across borders. Publicly funded research often ends up locked behind the very paywalls the public already paid to overcome.

We backed a single, mandatory research exception across the EU, and a right for researchers to make publicly funded work openly available the moment it is published, with no embargo. Contracts and technical locks should not be allowed to override either.

The four rights every memory institution needs

The consultation’s four questions miss something larger. The Our Future Memory statement, which more than seventy organisations have now signed, including the International Federation of Library Associations and Institutions (IFLA) and the International Council on Archives (ICA), sets out four rights that libraries, archives, and museums need in the digital world: to collect, to preserve, to provide access, and to cooperate across borders. Today’s law falls short on all four.

Collect. No EU rule requires the deposit of born-digital and web-published material. Journalism, government records, and culture that exists only online are slipping out of the published record entirely. We asked for a common baseline for the legal deposit of digital materials.

Preserve. Current law allows institutions to copy works for preservation only if those works are already in their permanent collection. Most of the open web and most licensed content fall outside it. We called this the 21st-century black hole: material that exists today and will be gone tomorrow because no one holds the legal right to save it.

Provide access. Two old rules hold this back. One governs works that are no longer commercially available, where licensing can take two to four years, if it happens at all, leaving a century of out-of-print culture in limbo. The other still ties library access to physical terminals on the premises, with no route to secure remote access for readers who cannot travel. Both need fixing.

Cooperate. Libraries can lend to each other across borders on paper, but not in digital form. A researcher who cannot travel has no legal way to obtain a digital copy. We asked for a clear cross-border exception for the supply of digital documents between research and library institutions.

The gap we want closed

Across every one of these areas, the same pattern shows up. Exceptions exist in the law, but technical locks, one-sided contracts, and the fear of getting it wrong stop institutions from using them. Faced with legal risk, librarians and archivists hold back, and rights that look solid on paper quietly disappear.

The Commission has a real chance to close the gap between what the law says memory institutions can do and what they can actually do, day to day. We will continue to make that case as the review moves forward.

Read our full submission here and add your organisation’s voice to Our Future Memory at ourfuturememory.org.

Rights on Paper Are Not Enough: Our Input to the EU Copyright Review Read Post »

When Archives Speak Back: IAE Hosts Data CARE Festival Fellows in Amsterdam

More than 30 public AI researchers from around the world gathered at the Internet Archive Europe Amsterdam headquarters on June 9 to open the Data Care Festival organized by the Inclusive AI Lab, and to discuss the role of archives in supporting culture and society.

The discussion almost didn’t happen as planned. On 9 June, Internet Archive Europe (IAE) opened its Amsterdam space for the Fellows Soirée of the Data CARE Festival, a four-day gathering organised by the Inclusive AI Lab. The listed moderator, Kirthi Jayakumar, was not able to travel to the Netherlands because of a technical error in the passport. One of the listed panelists, Franklin Ozekhome, pop culture architect and founder of Pop Culture Varsity travelling from Nigeria, also encountered entry barriers and couldn’t make it that evening. Inclusive AI Lab founder Prof. Payal Arora, stepping in to moderate, named it plainly: a visceral reminder of who gets to speak, and what kind of passport shapes access to the very conversations about power and knowledge that the evening was there to have. 

The Data CARE Festival

The festival’s theme, “Reclaiming Techno-Optimism: Building for Context. Culture. Community,” reflects the core work of the Inclusive AI Lab, founded by Prof. Arora at Utrecht University. The lab incubates researchers, practitioners, and civic leaders from across the Global South and North, working on the concrete conditions under which AI can serve communities rather than extract from them. 

When Archives Speak Back

The evening panel, titled “When Archives Speak Back: Power, Data, and AI Storytelling,” brought together IAE Programme Manager Beatrice Murch; Dr. Jaswina Elahi, assistant professor at Utrecht University and principal investigator on heritage-building among postcolonial migrant communities in the Netherlands; Vincenzo Scagliarini, head of research at Logotel and editor of the collaborative economy project Weconomy (Italy); and Chux Daniels, who leads transformative innovation programmes across Africa. Rana Kuseyri, responsible AI researcher at the Inclusive AI Lab, had opened the evening by framing the festival’s core commitment: not optimism as a mood, but hope as a moral imperative.

The conversation that followed was grounded in lives, not abstractions. Elahi, whose research traces the heritage of postcolonial migrant communities in the Netherlands, pushed expanded the definition of what counts as culture. Growing up Surinamese-Hindustani in the Netherlands, she described a childhood shaped by Bollywood films, Surinamese Hindustani radio, and a recurring question: but where are you really from? Her doctoral work examined how digital platforms were allowing Hindustani communities in the Netherlands to construct and transmit cultural identity — research met, early on, with the assumption that young ethnic minority internet users must be at risk of radicalisation. The actual finding was that they were using the internet to feel connected to their communities, their home countries, and their culture.

That distinction matters for what archives do and don’t capture. Elahi was precise about it: heritage is not a building or a monument. It is a song that only exists in relation to another person. It is the way a grandmother cooks that her grandchildren can attempt to learn, but will always make it differently. The body carries heritage and passes it on. The recipes communities are writing down now, the YouTube searches for ingredients that are no longer available, the heritage books being compiled: these are not the heritage itself, but they are activating something that might otherwise be lost. Data alone is not heritage. But in the right hands, it can give communities a voice they were never offered elsewhere.

Scagliarini brought a different example: a team of five engineers from different countries, working for Cisco on a creative project that eventually made it to the Venice Biennale. The team included a designer who couldn’t code. From a conventional business perspective, that was a problem. What actually happened was that she and the engineers spent a month in conversation before the first GitHub push, and she came out of it having learned a new language, not to replace her own practice, but to build shared knowledge. Data, Scagliarini argued, is something living: it can always be broken apart, rearranged, and re-interrogated, even across centuries. The responsibility is to keep rewriting it, not to take any version of the record as fixed.

Daniels traced this across a different scale. The growing global presence of Afrobeats, with the nice touch of Nigerian music playing through Schiphol airport on his last arrival, is one signal of a cultural confidence that statistics about tech governance don’t yet reflect. Eighty-five percent of the world’s population lives in contexts where the biggest tech decisions get made without them. The UK Prime Minister meeting with Apple and Google to shape AI governance is not the same as the communities in Kenya, one of the world’s biggest social media user bases, having any say in how those systems work. The same asymmetry applies to knowledge-making more broadly: innovation in agriculture, finance, and mobility is being led in the Global South by people working without the infrastructure constraints that lock the Global North into old models. M-PESA exists because banks wouldn’t go to rural areas. The most interesting AI work may be happening in places the dominant platforms aren’t looking.

Beatrice connected this to the practical stakes of what IAE and the Internet Archive exist to do. Truth, she said, is fracturing. The archive’s job, namely establishing what was said and what happened at a given point in time, matters more when that fragmentation accelerates. Democracy’s Library, the Internet Archive’s project to gather government-funded public information and make it freely accessible, is one concrete response. Beyond that, the work comes down to choices made under constraint: you cannot archive everything. The guiding principle is not to let perfect be the enemy of done, and to be honest about the challenges, which are real: lawsuits, the rising cost of storage, legal frameworks that differ across EU member states, and journalistic organisations restricting Wayback Machine access. Thirty years in, the Internet Archive is still here, and still going.

Why IAE Supports the Inclusive AI Lab

IAE is a partner of the Data CARE Festival because the questions it poses are ones we share. Who controls the record? What gets preserved, and what gets lost? When AI trains on cultural heritage, whose heritage counts? These questions shape what libraries, archives, and memory institutions can do, and what communities can access and build on. The panel on 9 June was, among other things, a demonstration that these questions are not rhetorical. Two of the people who were supposed to be in the room didn’t make it, because of where they hold citizenship. That is the context in which memory institutions operate, and the context in which inclusive AI must be built. And this is what shapes what our understanding of the past will be in the future.

When Archives Speak Back: IAE Hosts Data CARE Festival Fellows in Amsterdam Read Post »

Most of the Renaissance Has Never Been Translated. Source Library Is Opening It.

Ninety percent of Renaissance Latin has never been translated into a modern language. More Latin was written after 1500 than survives from all of ancient Rome, and almost none of it has been read outside a specialist library. At the current pace of human scholarship, completing that translation work would take approximately 12,000 years.

On 4 June, Internet Archive Europe attended the Source Library BETA Launch at the Embassy of the Free Mind in Amsterdam. It was a milestone worth marking.

What Source Library Is

Source Library is the world’s largest freely available collection of translated historical primary sources from the Renaissance. At launch, it holds more than 15,000 books across 55 languages, including 6,000 first-ever English translations and roughly seven billion words of original text and translation, comparable in scale to the entire English Wikipedia. Works previously readable only by Latin scholars or locked behind expensive academic editions are now open to anyone.

The project is hosted at the Bibliotheca Philosophica Hermetica, the UNESCO Memory of the World-recognised collection at the Embassy of the Free Mind: more than 25,000 volumes on alchemy, Hermetica, Kabbalah, Rosicrucianism, and the roots of modern science. Many of these books were banned at various points in history. Now they are open.

The launch carries particular meaning in light of what followed. Joost Ritman, the Amsterdam businessman who founded the Bibliotheca Philosophica Hermetica and built it into one of the world’s great collections of philosophical, religious, and esoteric knowledge, died on 5 June 2026, the day after Source Library launched in the institution he founded . He had spent sixty years guided by a conviction he traced to the Florentine Medici: that those in a position of privilege carry an obligation to culture. In 2017, he donated the library, its research institute, and the House with the Heads Monument to a cultural non-profit foundation, publicly known as the Embassy of the Free mind. This was a gift to Amsterdam and the world: he made permanent what he had spent his life assembling, and he made it public. 

AI as Accessibility, Not Replacement

Source Library places AI-powered translations directly alongside images of the original source pages, so anyone can consult the original at any point. The goal, as project creator Dr. Derek Lomas of Delft University of Technology made clear at the launch, is not to replace scholarship but to make a vast body of untranslated material discoverable for the first time. Dr. Lomas is a cognitive scientist and human-computer interaction researcher, currently a professor of Human Centred Design at Delft University of Technology, who arrived at Renaissance philosophy through a long personal engagement with the Neoplatonic tradition. That combination gives Source Library a design sensibility that most digital archive projects lack: the design starts from how people actually discover and engage with material, not from how institutions prefer to organise it. 

Dr. Lomas and the team also maintain careful transparency about data sources throughout: content drawn from the library’s own catalogue is clearly distinguished from AI-generated material, and all translations record the model, date, and prompt used to produce them. That distinction matters. The AI output is treated as useful but revisable. The primary sources are  treated as the foundation it is, complimented by academic curatorial work.

The project carries an AGPL-3 licence, the same open source licence used by the Internet Archive. It draws on open digital image standards that allow libraries and archives to share their collections freely, and it acknowledges the institutions whose digitised holdings made the work possible.

Why This Matters

The question Source Library poses is one we encounter constantly: who gets to access knowledge, and on what terms?

For centuries, the thought documented during the Renaissance has been available only to those who read Latin, have access to specialist collections, or can afford expensive critical editions. Source Library removes those barriers. It doesn’t replace careful scholarship. It makes a vast body of human thought discoverable for the first time.

This is open access made concrete: 6,000 first translations, 55 languages, no paywalls. It also matters for AI. The training data available to language models shapes what they know and how they reason. A Renaissance that remains untranslated is a Renaissance that AI cannot draw on. Source Library is building the corpus that public-interest AI will need.

Explore Source Library at sourcelibrary.org.

Most of the Renaissance Has Never Been Translated. Source Library Is Opening It. Read Post »

Blue Sky Thinking, European Infrastructure: Internet Archive Europe Has Moved to Eurosky

We’ve moved our Bluesky presence. Our account looks the same. Our data now lives somewhere different, on European infrastructure, under European law. Here’s why we made that choice.

What changed and what didn’t

When you use Bluesky, your posts, followers, and interactions live on a server. By default, that server belongs to Bluesky. Moving to a Personal Data Server (PDS) changes the arrangement: you choose where your data is stored, by whom, and under which rules.

Eurosky is a European PDS provider. It operates within EU law, with stronger privacy protections and no commercial model built around exploiting what users share. We now host our Bluesky presence there via their EU-Haul migration service, which took less than an hour. Our content is still visible on Bluesky. Our data is no longer on Bluesky’s servers.

Part of something bigger

This isn’t a standalone technical decision. Internet Archive Europe and the Internet Archive have long been committed to building and supporting decentralised digital infrastructure: the kind that is resilient, publicly accountable, and not dependent on the choices of a small number of private companies.

Brewster Kahle, founder of the Internet Archive, has made this point clearly and consistently. At the Facebook Museum opening in Eindhoven this April, he described the stakes directly. If Europe doesn’t build its own public digital infrastructure, the choice will narrow to American or Chinese models. “That is not good enough,” he said, “and we have the technologies to do something about it.”

Moving our data to a European server is one small expression of that conviction. Choosing infrastructure that operates by European values, that we can migrate away from freely, and that doesn’t anchor our digital presence to a single company’s decisions is consistent with what we advocate for in policy and in practice.

DWeb Camp: Root Systems

That same conviction is why we’re co-presenting DWeb Camp this summer, as the gathering comes to Europe for the first time, at Alte Hölle in Germany.

DWeb Camp 2026: Root Systems brings together builders, researchers, artists, activists, and policymakers at Alte Hölle, an ancient forest one hour southwest of Berlin, from 8 to 12 July. Internet Archive Europe co-presents the event alongside the Internet Archive and the Department of Decentralization.

The theme captures something real. Like forest ecosystems, decentralised networks derive their strength from what lies beneath the surface: distributed connections that share resources without hierarchy and keep functioning even when individual nodes go down. DWeb Camp exists to build those root systems in practice, not just to talk about decentralisation, but to make it.

Brewster Kahle originated the DWeb project in 2016. In the decade since, it has grown into a global network of builders and dreamers united by shared principles: trust, human agency, mutual respect, and ecological awareness. This July, that community gathers in Europe for the first time. We’re proud to help bring it here.

If you build decentralised tools, research digital infrastructure, or simply believe that the web should be resilient and publicly accountable, this is the gathering to be part of. Tickets and details at dwebcamp.org.

You can do this too

If your organisation is on Bluesky, migrating your data to a PDS takes under an hour. Eurosky’s EU-Haul service handles it. Your content stays visible to everyone on the network. You gain more control over your own data. That is a reasonable trade.

Find us on Bluesky at @internetarchive.eu

Learn more and migrate your own account at eurosky.tech.

Blue Sky Thinking, European Infrastructure: Internet Archive Europe Has Moved to Eurosky Read Post »

The Wayback Machine holds 30 years of the web. News publishers are blocking it.

International Archives Day falls today, Tuesday 9 June. The theme for this year’s International Archives Week is #ArchivesForJustice: Rights, Memory, and Futures. The timing is pointed.

As reported by Nieman Lab and WIRED, a growing number of publishers have moved to block the Internet Archive’s web crawlers from preserving their content. For a full account of what that blocking involves and what it means, see the Internet Archive’s own FAQ on the issue

The stated reason is artificial intelligence. Publishers are worried that content preserved by the Wayback Machine can be accessed by AI companies looking for training data. That concern is understandable. The response is not.

Blocking the Archive is not the same as blocking AI

The Internet Archive is a nonprofit digital library. It is not building commercial AI systems. It is preserving a record of history. Its Wayback Machine holds more than one trillion archived web pages and is used every day by journalists, historians, researchers, and courts.

Archived pages are often the only reliable record of how a story appeared when it was first published. Articles get edited, changed, or removed, sometimes openly, sometimes not. The Wayback Machine often becomes the only source for seeing those changes. When publishers block it, they limit not just the Archive’s ability to preserve material, but anyone’s ability to access, verify, and study journalistic and historical records in the future. 

Over 250 journalists have signed the open letter

Fight for the Future has launched an open letter thanking the Internet Archive for its preservation work and calling on news organisations to reconsider. More than 250 journalists have signed.

As the letter puts it: “The freedom of journalists isn’t only the freedom to write, it’s also the freedom to have your work read and remembered for generations to come.”

The signatories include Rachel Maddow, who described the Archive as a national treasure she uses daily and cannot imagine working without. Glenn Kessler, the Washington Post’s fact-checker, described using it to examine the Trump administration’s false claims about USAID after the agency’s website was taken offline. Reuters journalist Bozorgmehr Sharafedin used it to uncover a covert CIA communication system, work that won the National Press Club’s Edwin M. Hood Award.

The Wayback Machine preserves permanent citations for nearly 5 million news articles referenced on Wikipedia. That is not a technical footnote. That is the record of public knowledge.

What #ArchivesForJustice means in practice

This year’s International Archives Week centres on accountability, memory, and the right to access the past. Archives for accountability. Archives for memory. Archives for future justice. The Wayback Machine is one of the clearest examples of those principles anywhere. It holds the web as it actually was, not as institutions later chose to present it. When publishers block it, they limit not just the Archive’s ability to preserve material, but anyone’s ability to access quality journalistic and historical records, as pointed out by the Electronic Frontier Foundation (EFF).

The Internet Archive has long worked collaboratively with publishers and respects their requests around access and preservation. What it asks is that publishers work with it, rather than against it, to ensure that the journalism being produced today remains accessible to historians, researchers, educators, and future generations.

Internet Archive Europe adds its voice to that call. The historical record belongs to everyone.Read the letter and add your voice at savethearchive.com/NewsLeaders.

The Wayback Machine holds 30 years of the web. News publishers are blocking it. Read Post »

Scroll to Top