Reimagining dataset governance for LLM Training

Consent, Distributed Decision-Making, and Benefits for us all

kaodro 3 de diciembre de 2024
Reimagining dataset governance for LLM Training

In my last post, I discussed the tension between open data and AI training, emphasizing the need for a thriving open web to support transparent, open datasets. These datasets are crucial especially for smaller AI builders, nonprofits, and organizations working for public good. However, openness alone is insufficient —datasets must also be built responsibly, with an ethics-first approach.

If communities contribute their data; be it their voices or information about their lives; in good faith, it is the responsibility of an ethical AI builder to not only make sure they are not being exploited and their rights respected but go even further: make sure they are included in the direct benefits of AI.

The governance of LLM training datasets is now a critical issue, and while the field is complex, it’s essential to get it right. Concepts like Public Interest AI and Digital Public Infrastructure aim to honor the rights of individuals and communities—principles we should embed into dataset governance. Let’s unpack what that might look like in practice:

Ensuring meaningful consent is one of the biggest challenges in data governance. After all, if our data is being used, we should be able to say “ok” or “not ok” and define what we are agreeing to and what not. But at the moment traditional consent frameworks often fail for datasets derived from web scraping, or collaborative platforms. For example:

  • Data made available under licenses like Creative Commons may not have anticipated today’s AI use cases. We might have uploaded pictures of our faces decades ago but didn’t think they would be one day used for facial recognition.
  • Tools like robots.txt can be used to signal that AI scrapers should exclude certain content from crawling but they are insufficient for managing nuanced consent tied to specific purposes or users.
  • Web metadata is often incomplete and we don’t have yet standardized, scalable and reliable protocols to express our preferences in a nuanced way (“you can use my data for X but not for Y”) which makes it difficult both for open dataset builders to have legal certainty regarding consent and for those who want to express the consent regarding their data.
  • People should ideally be able to revoke their consent if harms arise or use cases change or they simply change their minds. While technically complex—especially for already-trained AI models—solutions are emerging to make consent reversible.

Addressing these challenges is key to respecting individuals’ and communities’ evolving rights in an incredibly complex data environment. Tools like Spawning are moving the needle in that direction.

2. Distributed Decision-Making

In an ideal governance world, the power to influence what happens with the datasets, how they are curated and how they are used as a whole should be evenly distributed among those who contribute to, build the datasets and who finance it. Decentralized data governance, where decisions can be also made by members of communities that contribute to datasets or by empowered representatives of data subjects are another holy grail of ethical data governance. However, distributed decision-making is inherently complex and needs to clarify questions such as:

  • Who decides which data gets included in training datasets?
  • How do we balance competing interests, such as openness versus cultural preservation?
  • What accountability mechanisms can ensure fairness across stakeholders?

Examples like Mozilla’sCommon Voicegovernance documentand the BLOOM model governance paper demonstrate steps toward inclusive governance. Such frameworks ensure that data contributors retain agency and share in the benefits of AI’s development.

3. Value and Sovereignty in Data

Data governance must also address the equitable distribution of AI’s benefits. Nations, particularly in the Global Majority, are increasingly concerned about data sovereignty, where their resources are used without fair returns.

Te Hiku Media provides a compelling model. By creating special licensing terms, they ensure that Māori language data aligns with their cultural values and goals. This approach underscores the importance of context-sensitive governance, particularly for datasets with cultural or heritage significance.

Reimagining dataset governance is not just about mitigating harm but about creating systems that actively shape a fairer, more inclusive world. When you think about the mounting lawsuits on behalf of creatives whose data was scraped into AI training by the big AI companies you can understand how timely and urgent this topic became.

I’d love to hear your thoughts about anything that might have sparked curiosity in you. How can we balance openness in AI research with protecting creativity, cultural rights and heritage? Are there models of distributed governance or data stewardship that inspire you?

And in my next post, I’ll explore something different but equally important: the anthropomorphization of AI and how it shapes our relationship with technology.