
In my last post, I discussed the tension between open data and AI training, emphasizing the need for a thriving open web to support transparent, open datasets. These datasets are crucial especially for smaller AI builders, nonprofits, and organizations working for public good. However, openness alone is insufficient —datasets must also be built responsibly, with an ethics-first approach.
If communities contribute their data; be it their voices or information about their lives; in good faith, it is the responsibility of an ethical AI builder to not only make sure they are not being exploited and their rights respected but go even further: make sure they are included in the direct benefits of AI.
The governance of LLM training datasets is now a critical issue, and while the field is complex, it’s essential to get it right. Concepts like Public Interest AI and Digital Public Infrastructure aim to honor the rights of individuals and communities—principles we should embed into dataset governance. Let’s unpack what that might look like in practice:
Ensuring meaningful consent is one of the biggest challenges in data governance. After all, if our data is being used, we should be able to say “ok” or “not ok” and define what we are agreeing to and what not. But at the moment traditional consent frameworks often fail for datasets derived from web scraping, or collaborative platforms. For example:
Addressing these challenges is key to respecting individuals’ and communities’ evolving rights in an incredibly complex data environment. Tools like Spawning are moving the needle in that direction.
In an ideal governance world, the power to influence what happens with the datasets, how they are curated and how they are used as a whole should be evenly distributed among those who contribute to, build the datasets and who finance it. Decentralized data governance, where decisions can be also made by members of communities that contribute to datasets or by empowered representatives of data subjects are another holy grail of ethical data governance. However, distributed decision-making is inherently complex and needs to clarify questions such as:
Examples like Mozilla’sCommon Voicegovernance documentand the BLOOM model governance paper demonstrate steps toward inclusive governance. Such frameworks ensure that data contributors retain agency and share in the benefits of AI’s development.
Data governance must also address the equitable distribution of AI’s benefits. Nations, particularly in the Global Majority, are increasingly concerned about data sovereignty, where their resources are used without fair returns.
Te Hiku Media provides a compelling model. By creating special licensing terms, they ensure that Māori language data aligns with their cultural values and goals. This approach underscores the importance of context-sensitive governance, particularly for datasets with cultural or heritage significance.
Reimagining dataset governance is not just about mitigating harm but about creating systems that actively shape a fairer, more inclusive world. When you think about the mounting lawsuits on behalf of creatives whose data was scraped into AI training by the big AI companies you can understand how timely and urgent this topic became.
I’d love to hear your thoughts about anything that might have sparked curiosity in you. How can we balance openness in AI research with protecting creativity, cultural rights and heritage? Are there models of distributed governance or data stewardship that inspire you?
And in my next post, I’ll explore something different but equally important: the anthropomorphization of AI and how it shapes our relationship with technology.