The Open Web-AI Data Paradox: Why we need to protect Data Commons

kaodro 19 de noviembre de 2024
The Open Web-AI Data Paradox: Why we need to protect Data Commons

đź‘‹ Dear NGI Community

At the core of almost any technology lies one fundamental ingredient: data.

AI systems are fueled by datasets that determine their capabilities, limitations, and societal impact. Yet assembling high-quality, equitable datasets isn’t an easy task.

The challenge is twofold:

✅ First, data is often unstructured, inaccurate, or incomplete—especially in humanitarian and developmental areas where we need it most—failing to represent the lived realities of diverse communities.
âś… Second, even when data is plentiful, it reflects the internet's inherent knowledge biases: predominantly English-speaking, disproportionately Western-centric, and deeply intertwined with existing social injustices baked into the systems that produce it.

The open web could play a pivotal role in creating and governing just AI datasets, but its value is often overlooked. Large Language Model developers rely on high-quality sources across the internet, yet predictions suggest we might exhaust such data in the next years.

🔥Ironically, generative AI contributes to a process of cannibalization of open data sources, becoming major information gatekeepers while scraping data without appropriate consent mechanisms.🔥

This practice has led to more websites blocking AI scrapers by default, effectively closing previously open data not only to major AI companies but also to open dataset creators acting in public interest and researchers alike.

This presents a complex challenge for open internet and open AI advocates: to keep AI open, we need an open internet, and to support the open internet, we need more ethical and sustainable open data practices for AI.

A flourishing open web holds intrinsic value for source diversity and verifiability—data sourced from it can be made transparent, with clear provenance. It's also key for keeping AI itself open, accountable, and accessible to technology builders worldwide. While the biggest AI companies have vast resources to collect, pay for and process data or exploit legal gray zones to acquire it, smaller corporate actors, nonprofits, and public institutions do not. To level the playing field in AI research and development, we need more open datasets—the concept of AI data commons comes to mind—created in ways that preserve, not undermine, the open web.

It's worth remembering that the open web remains incomplete—reflecting existing inequalities, with many languages, regions, and communities underrepresented. Investing in expanding the web's reach and inclusivity could amplify its value as a training ground for equitable AI systems.

Much of the open web operates on principles of decentralization, collaboration, and shared ownership. These principles could inspire how AI data commons are governed, with input from diverse stakeholders and as part of a Public Interest AI I mentioned in my last post.

Could collaboration between technical communities who build datasets and deploy artificial intelligence, and communities that build and depend on the open web for communication, knowledge sharing, and preservation be a way forward?

How might we design such collaborations to distribute power and value more equitably than in the past and prevent exploitation or existing power consolidation?

✨In my next post, I will explore the complexities of current AI data governance and why getting it right is both difficult and necessary.✨

👉 In the meantime, I would be curious to hear this community's opinions on the value of the open web going forward and whether such collaborations could be viable.