Mozilla is spending $5 million to change how AI companies get their data

Artificial intelligence companies need enormous amounts of data, but where that data comes from and who gets paid for it remain uncomfortable questions. Mozilla thinks there is another way, and it is putting $5 million behind the idea.

Mozilla Data Collective has raised $5 million from Mozilla to expand its platform for sharing AI training data. The company wants to give developers access to more diverse material while allowing the people and organizations providing it to maintain control over how it gets used.

That is a notable contrast with the massive scraping operations that helped build much of today’s generative AI industry. Instead of treating anything accessible online as potential training material, the platform focuses on defined provenance, licensing, and consent.

It currently offers more than 1,700 datasets across more than 450 languages. The company says 350 organizations have been approved to contribute, covering multilingual, multicultural, and multimodal applications.

Money is part of the equation too. Mozilla Data Collective has introduced compensated datasets that allow owners to charge for access. Its stated business model charges downloaders a five percent platform fee, while uploaders keep the amount they choose to charge. Free and openly licensed material can continue to be offered as well.

That could become increasingly important as AI companies hunt for higher-quality material beyond the enormous quantities of information already collected from the web. Having more data is not necessarily the same thing as having better data, particularly when developers are trying to build systems that work across languages and cultures poorly represented online.

The company also claims commercial demand is growing quickly. According to Mozilla Data Collective, major AI labs, thousands of startups and scale-ups, and dozens of unicorns are already using material available through the platform. It says its annualized revenue run rate has reached nine times the milestone established for this stage of development, although it did not disclose the actual revenue figure.

The new funding will support expansion into multimodal cultural video and larger text collections covering European, African, and South Asian languages. Mozilla Data Collective also plans additional licensing options, subscriptions aimed at startups and scale-ups, and new security and data-management capabilities.

Mozilla Data Collective grew out of the Mozilla Foundation and became an independent UK-based entity in 2026. Its roots include Common Voice, Mozilla’s long-running project for building an openly available multilingual speech dataset.

There is obviously a business opportunity here. AI developers need quality training material, while organizations sitting on valuable archives may be more willing to participate when they can set the terms and get paid.

The challenge will be scale. Scraping the internet is cheap and easy. Building a marketplace around consent, provenance, and compensation is considerably harder.

Mozilla now has another $5 million to see if that approach can work.

Support independent tech journalism

NERDS.xyz is independently owned and operated. If you enjoy my coverage of Linux, AI, hardware, cybersecurity, and tech culture, consider supporting the site on Ko-fi.

Support NERDS.xyz
Written by

Brian Fagioli

Technology journalist and founder of NERDS.xyz

Brian Fagioli is a technology journalist and founder of NERDS.xyz. A former BetaNews writer, he has spent over a decade covering Linux, hardware, software, cybersecurity, and AI with a no nonsense approach for real nerds.

Leave a Comment