July 2026

Mozilla AI at Internet Archive Europe: Owning Your AI Stack

On 25 June, Internet Archive Europe (IAE) hosted Davide Eynard and Thomas (toto) Bille from Mozilla AI at our Amsterdam space. The afternoon covered a question that sits close to the heart of what we do: when you depend on infrastructure you don’t control, what do you actually own?

Ada, Zangemann, and the case for tinkering

Davide opened with a children’s book: Ada & Zangemann. Ada is a girl who lives in a dumpster, salvages broken hardware, and builds things entirely her own. Zangemann builds beautiful, polished technology that no one else can modify or adapt. The clash between them drives the story, but Davide used it as a frame for something more immediate: most people’s relationship with AI today looks a lot more like Zangemann’s world than Ada’s. You use what you’re given, on the terms it’s offered, for as long as the provider decides to keep it available.

This has nothing to do with technology: it’s a power relationship. And that logic applies to memory institutions every bit as much as to individual developers.

The trade-offs are concrete

To make the point, Davide rebuilt a small tool he’d originally written by hand more than twenty years ago: a script to extract train timetables from a website too clunky to use directly. He produced the rebuilt version with an AI coding assistant, and it worked. But the result lived on someone else’s platform, not his own machine. And unlike the original, which taught him Perl and regular expressions he used for years afterward, this one taught him nothing. Convenience and ownership turned out not to be the same thing.

Search as activity, not action

One of the sharpest distinctions in Davide’s talk was between search as an action and search as an activity. When you type a question into a box and accept the answer, that’s an action. When you use an agent to follow a thread, evaluate what it finds, redirect it when it goes wrong, and build toward a conclusion over time, that’s an activity. The difference matters because the second approach keeps you in the loop. You catch mistakes. You steer.

The clearest example: Davide used an agent to track down the original source of a widely cited Bill Gates quote. The agent searched, hit dead links, found partial copies, and eventually installed a subtitle extraction tool autonomously, downloaded a YouTube video, and identified the exact moment Gates said the thing. Along the way, it returned to the Wayback Machine around a dozen times, working through broken URLs until it found a usable copy.

It was a quiet illustration of something IAE says often: preserved web history is not a nostalgia project. It is working infrastructure. AI systems now depend on it to check what is actually true.

A second demo used a locally run open source model to search the Rijksmuseum’s digital collection for images connected to alchemy and the pursuit of knowledge, producing usable results entirely on Davide’s own hardware. No cloud service, no rented compute, no data leaving the machine.

Otari: making ownership practical

Thomas (toto) Bille followed with a look at the infrastructure side of the problem. Mozilla AI has built Otari as an open-source LLM gateway: a single control plane for all your interactions with language models, whether you run a local model on your own server or route through a commercial API.

The features that generated the most discussion were practical ones. Budget controls granular enough to cap spending per user, per model, or per team. Guardrails that strip personal or sensitive data before it reaches any external provider. A federated router in development that will recommend which model to use based on real usage patterns across the community, with options to prioritise cost, quality, or energy use. Otari is fully self-hostable: if you don’t want your data passing through Mozilla AI’s servers, you don’t have to. The code is on GitHub.

The honest question

Someone in the room asked how Mozilla AI intends to stay financially viable if the tools are free and open source. Toto’s answer was direct: revenue from the hosted version, income from enterprise integration work, and a long-term bet on the community. It’s the same tension that runs through almost every public-interest digital project, including IAE’s own. There’s no clean resolution, but naming it honestly matters.

Why this conversation belongs here

The afternoon drew developers, researchers, and people from across the cultural heritage sector. The discussion after the talks ran for close to an hour.

What connected the room wasn’t a shared technical interest in language models. It was a shared unease about dependency. Memory institutions know what it looks like when access to knowledge sits on infrastructure you don’t own and can’t influence. IAE has spent years arguing that the rights archives have always held offline must be protected online too. The same holds for one layer up, to the AI systems now sitting on top of those archives.

Explore Mozilla AI’s open-source tools, including AnyAgent, AnyLLM, LlamaFile, and Otari, at mozilla.ai and github.com/mozilla-ai.

You can watch a replay of the presentation and conversation on Archive.org and see the slides online.

Mozilla AI at Internet Archive Europe: Owning Your AI Stack Read Post »

Rights on Paper Are Not Enough: Our Input to the EU Copyright Review

A copyright exception you cannot use is not really an exception.

That idea runs through the submission Internet Archive Europe (IAE) submitted to the European Commission this month. On 25 June, we responded to the Commission’s Call for Evidence on copyright, the review that will shape what users, libraries, archives and museums across Europe can do with digital materials for years to come.

The Commission asked for input on four areas. We answered each one, and we added a point the consultation left out: the basic rights memory institutions need to do their work online. Here is what we told them.

Protect the preservation work that the law already allows

European law already lets libraries, archives and research organisations mine text and data, including for training AI models, when the purpose is preservation or research. Lawmakers made a deliberate choice to protect that public-interest work without giving rightsholders a veto over it.

That protection is being quietly undone. Large rightsholders now apply blanket opt-outs, built for commercial AI, across the board. They draw no line between a company training a product and an archive preserving the record, and so they block the very preservation work the law set out to protect.

We asked the Commission to confirm what the Directive already implies: an opt-out designed for commercial use cannot override the preservation and research rights of cultural heritage institutions.

Tackle piracy without breaking the open web

Piracy of live events is a real problem. But the answer some governments have reached for does more harm than the problem it sets out to solve.

We pointed to two examples. Italy’s Piracy Shield blocks content so bluntly that the Commission itself wrote to Rome in 2025, finding the system out of step with fundamental rights. In Spain, courts have ordered VPN providers to block access in proceedings where those providers had no chance to speak, and ordinary websites get caught in the net.

These are enforcement tools built for commercial pirates, used without proper judicial oversight, and the damage falls on people and services that did nothing wrong. We asked the Commission to measure that damage before legislating further, and to rule out using live-event blocking against general, non-commercial websites.

One research exception, not twenty-seven

Right now, the EU’s research exception is optional. Member States have implemented it differently, leaving researchers facing 27 distinct legal environments and real barriers to working across borders. Publicly funded research often ends up locked behind the very paywalls the public already paid to overcome.

We backed a single, mandatory research exception across the EU, and a right for researchers to make publicly funded work openly available the moment it is published, with no embargo. Contracts and technical locks should not be allowed to override either.

The four rights every memory institution needs

The consultation’s four questions miss something larger. The Our Future Memory statement, which more than seventy organisations have now signed, including the International Federation of Library Associations and Institutions (IFLA) and the International Council on Archives (ICA), sets out four rights that libraries, archives, and museums need in the digital world: to collect, to preserve, to provide access, and to cooperate across borders. Today’s law falls short on all four.

Collect. No EU rule requires the deposit of born-digital and web-published material. Journalism, government records, and culture that exists only online are slipping out of the published record entirely. We asked for a common baseline for the legal deposit of digital materials.

Preserve. Current law allows institutions to copy works for preservation only if those works are already in their permanent collection. Most of the open web and most licensed content fall outside it. We called this the 21st-century black hole: material that exists today and will be gone tomorrow because no one holds the legal right to save it.

Provide access. Two old rules hold this back. One governs works that are no longer commercially available, where licensing can take two to four years, if it happens at all, leaving a century of out-of-print culture in limbo. The other still ties library access to physical terminals on the premises, with no route to secure remote access for readers who cannot travel. Both need fixing.

Cooperate. Libraries can lend to each other across borders on paper, but not in digital form. A researcher who cannot travel has no legal way to obtain a digital copy. We asked for a clear cross-border exception for the supply of digital documents between research and library institutions.

The gap we want closed

Across every one of these areas, the same pattern shows up. Exceptions exist in the law, but technical locks, one-sided contracts, and the fear of getting it wrong stop institutions from using them. Faced with legal risk, librarians and archivists hold back, and rights that look solid on paper quietly disappear.

The Commission has a real chance to close the gap between what the law says memory institutions can do and what they can actually do, day to day. We will continue to make that case as the review moves forward.

Read our full submission here and add your organisation’s voice to Our Future Memory at ourfuturememory.org.

Rights on Paper Are Not Enough: Our Input to the EU Copyright Review Read Post »

Scroll to Top