Module 06

Open source AI and digital commons

Why AI cannot be the exclusive property of corporations: the role of open source, open data and digital commons.

8 min3 resources

One of the most serious risks of contemporary AI is concentration: a few corporations control the most powerful models, the largest datasets and the infrastructure needed to train them. Google, Microsoft, Meta and OpenAI have a structural advantage that is self-reinforcing: more data produces better models, better models attract more users, more users generate more data. In this context, open source AI is not just a technical preference — it is a democratic necessity.

Open source AI means that the source code of models, trained weights and — ideally — training data are publicly available. Anyone can inspect how the model works, verify its biases, adapt it to their needs, improve it. Meta released Llama, Stability AI released Stable Diffusion, Mistral AI released models competitive with proprietary ones. These releases have democratised access to technology and accelerated innovation.

But "open source" does not automatically mean "for the common good". An open source model can be used to create deepfakes, produce disinformation or enhance surveillance systems. The question is not just code accessibility, but governance: who decides how it is used, who is responsible for harm, how to balance openness and safety.

The concept of "digital commons" offers a broader framework. Commons are neither private property nor state property: they are resources collectively managed by a community according to shared rules. Wikipedia is a digital common: anyone can contribute, but community rules (neutrality, verifiability, consensus) ensure quality. OpenStreetMap is another: a world map created and maintained by volunteers, an alternative to Google Maps.

Applying the commons model to AI means creating models, datasets and infrastructure that are collectively owned, democratically governed and oriented towards public benefit. BLOOM, a multilingual language model developed by BigScience — a consortium of over 1,000 researchers — is an example: it was created with a participatory decision-making process on ethical and technical choices. EleutherAI is a collective of researchers developing open language models and publicly documenting their choices.

Data itself can be a commons. "Data trusts" — legal structures that manage data on behalf of a community, as a trust manages assets — are a promising model. Instead of giving our data to corporations, we could contribute it to democratically governed trusts that negotiate its use in contributors' interests. It is still an experimental idea, but it could radically transform the relationship between citizens and data.

Key takeaways

  • AI concentration in a few corporations is a democratic risk that open source can mitigate
  • Open source is necessary but not sufficient: responsible governance is needed, not just open code
  • Digital commons (Wikipedia, OpenStreetMap, BLOOM) are models of collective governance applicable to AI
  • Data trusts could transform the relationship between citizens and data, restoring collective power

Reflection prompt

If your data (health, financial, behavioural) were managed by a democratic trust rather than corporations, what would change? Would you be willing to share more knowing governance is collective? What rules would you want for the trust?

Further reading