The term “open source” has a pretty well understood meaning in the software world. You can inspect the source code, modify it, redistribute it, and generally see what makes the software tick. Artificial intelligence, however, has made things considerably messier.
Mozilla is highlighting new research published in Communications of the ACM that attempts to bring some clarity to what “open” actually means when talking about foundation models. The basic problem is simple: two AI models can both be described as open while providing very different levels of access to the technology behind them.
The research, titled “Unpacking Open Source Artificial Intelligence: Toward a Framework for Openness in Foundation Models,” argues against treating AI as simply open or closed. Instead, openness can be examined across the individual components that go into creating and operating a foundation model.
Those components can include training data, model weights, code, documentation, evaluation information, and other pieces of the development process. A company might release model weights while keeping its training data private, for example. Another might publish code while revealing relatively little about how the model was trained or evaluated.
Both could potentially be described as “open,” but those differences matter.
Mozilla says the framework grew out of a 2024 gathering it organized with Columbia University’s Institute of Global Politics. More than 40 researchers, developers, and policy experts participated in discussions about openness and AI.
The resulting framework doesn’t attempt to create a single checklist that determines whether an AI model earns an “open source” badge. Instead, it provides a way to describe which portions of an AI system are open and which aren’t.
That distinction could become increasingly important as companies use open source terminology to promote their AI models. Downloadable weights can provide developers with considerable freedom, but having access to weights isn’t necessarily the same thing as having everything needed to understand how a model was created.
Anyone familiar with Linux and traditional open source software should recognize why that distinction matters. If the source code for an application is available, developers can inspect it, change it, build it, and study how it works. Reproducing a foundation model can require much more than obtaining a finished set of weights.
Training data is particularly complicated. Making every piece of training data publicly available can create its own problems involving copyright, privacy, security, and other concerns. That is one reason the researchers aren’t simply arguing that every component of every AI model should always be public.
The framework also complicates the common argument that more openness automatically makes AI either safer or more dangerous. Different components can introduce different benefits and risks when opened. How an AI system is deployed, moderated, and governed can matter alongside what developers choose to release.
In other words, “open AI” may be better understood as a spectrum of characteristics rather than a yes-or-no label.
That might sound like semantics, but the terminology has consequences. Developers deciding which models to build upon, researchers trying to examine their behavior, policymakers considering regulation, and ordinary users comparing AI products can all come away with very different impressions when something is simply called open source.
The research doesn’t settle the long-running debate over what should qualify as open source AI. It does, however, make it harder to pretend that the answer is simple.
As AI companies increasingly embrace the language of open source, perhaps the better question isn’t whether a model is open. It is exactly what has been opened, what remains behind closed doors, and why.
Support independent tech journalism
NERDS.xyz is independently owned and operated. If you enjoy my coverage of Linux, AI, hardware, cybersecurity, and tech culture, consider supporting the site on Ko-fi.
Support NERDS.xyz