Stefaan Verhulst
Blog by Lea Gimpel: “The open web is undergoing a systemic structural transformation. Today, digital public goods (DPGs) and the broader knowledge commons face a critical challenge: large-scale, automated extraction by AI crawlers and scrapers operating without reciprocity. This leads to ballooning infrastructure costs for often volunteer-run, underresourced open projects, and the erosion of trust in open knowledge, among other challenges. DPG product owners have to make delicate decisions to protect their resources while keeping them as open as possible.
To help stewards of open resources navigate this complex landscape, the DPGA Secretariat is proud to share two complementary assets developed in consultation with our community:
What it is: A position paper and set of recommendations compiled directly by DPG product owners across major open initiatives like Wikipedia, Open Food Facts, Storyweaver, Govdirectory, The Turing Way, among others.
Core Focus: This note describes the problems faced by DPGs firsthand and outlines practical steps for engaging with commercial AI entities to ensure data users respect licenses, robot policies, and terms of service, and contribute back to the ecosystem, while also detailing technical mechanisms to protect against AI exploitation.
What it is: An overview of available mechanisms to protect against AI exploitation and how they map against the current open definition, and, by extension, the DPG Standard. The playbook’s aim is to deepen understanding and conceptual clarity around the tension between maintaining open access and ensuring the survival of an open resource.
Core Focus: The playbook helps understand the different layers of interventions available to DPG product owners, ultimately showing that code-level mechanisms are a first line of defence, but solving the challenge of AI exploitation requires ecosystem-level solutions at the policy, legal, and governance levels.
Layers of Defence: It introduces a 6-layer “defence and reciprocity pyramid” classifying technical, legal, regulatory and institutional measures..(More)”

Paper by Talia Caplan and Bilal A Mateen: “In 1600, Queen Elizabeth I granted the East India Company a royal charter and a monopoly over trade across much of the known world. For more than a century, the relationship was mutually advantageous: the Crown received revenue, access to rare commodities, and unparalleled geopolitical reach; the Company received protection, legitimacy, and the coercive backing of a state. By the mid-18th century, it became something very different. The Company maintained its own army, territories, and foreign policy, and increasingly operated in ways that were counter to the objectives of the Crown, resulting at times (due to its unconstrained, singular objective of maximising profits) in the deaths of millions.
Today, the world’s largest technology firms are at the precipice of a similar transition: towards becoming firms that function as corporate sovereigns whose power derives from the control of indispensable digital infrastructure and the algorithms that shape the attention economy, rather than from physical territory.2 The concentration of power is stark. A single firm (Nvidia) supplies between 80 and 90% of the chips on which advanced artificial intelligence (AI) models are trained; a handful of hyperscale providers (Microsoft Azure, Amazon Web Services, Google Cloud) operate the data centres in which those chips run; and a small number of laboratories (eg, OpenAI, Anthropic, Google DeepMind, DeepSeek), almost all American-owned, use those data centres to produce the frontier AI models on which others build.
These layers are less separate than they appear. A dense web of cross-investment now binds them, with chipmakers taking equity in the laboratories that buy their processors, and cloud providers and laboratories acquiring stakes in one another while committing to purchase each other’s services; analysts have estimated that interlocking arrangements of this type are worth more than US$800 billion. The effect is a tendency towards vertical integration achieved through finance rather than a formal merger, drawing nominally competing entities into a single, mutually dependent interest.
We expect that two structural features will characterise AI firms’ transition to pseudo-sovereign status. These firms will increasingly occupy every position in their own oversight, building the systems, funding the safety evidence (based on evidence from a preprint), staffing the advisory bodies, and drafting the standards meant to constrain them. They will also become jurisdictionally mobile in ways territorial regulators and sovereign governments are not, using digital infrastructure and complex corporate structures to obfuscate the residency of data and software in an effort to escape true oversight.
In that context, a health system that adopts a frontier AI solution for triage, imaging, or clinical decision support—such as the Horizon 1000 initiative seeks to roll out in Rwanda—does not necessarily acquire a substitutable tool but rather creates a dependency on a supply chain it neither controls nor can reproduce. The risks are especially acute in the global health context because upfront philanthropic support, especially when provided by the frontier AI companies that stand to benefit most, increases the likelihood of three specific risks…(More)”.
Paper by Polly Mackenzie et al: “There is now a broad consensus among practitioners, funders, and policymakers that Britain is paying the costs for a decades-long depletion of civic life. A great deal of excellent work has begun to reverse this decline, but a key question remains unanswered: how do we fund the work of civic renewal at the scale that is needed?
The Marshall Plan for Civic Life is a landmark programme of research and deliberation, convened by Kinship Works and Demos, to answer this question – identifying viable financing mechanisms, and developing an overall funding architecture. The initial work has been supported by This Day and the Joseph Rowntree Foundation but we are building a coalition – adding partners and modules to deepen our understanding of the problem and possible solutions.
In this opening Discussion Paper, we share initial reflections on a range of financial mechanisms that look promising. The paper is intended to stimulate discussion and we would welcome feedback. This will be followed in Autumn 2026 with a broader paper, describing a potential architecture and supporting philosophy for a Civic Marshall Plan, alongside an estimate of the scale of the funding gap that needs to be filled…(More)”.
Paper by Kevin A. Bryan & Joshua S. Gans: “AI predicts; humans use its predictions to make decisions. These predictions are combined with human verification and analysis, queries to other statistical models, and so on. The economic value of an AI, therefore, depends on how it interacts with the surrounding decision environment. We describe the value of AI as part of this “composite experiment” where AI makes a coarse prediction of the state of the world, show what this means for optimal model training via a geometric argument, explain why optimal training can be discontinuous in economic variables, and study how heterogeneous users or monopoly model trainers affect these results. In particular, maximizing the unconditional accuracy of AI predictions is generally suboptimal…(More)”.
Report by UNESCO: “Media freedom and journalism are under attack around the world. AI-fuelled disinformation is confusing and polarizing audiences, government censorship is sharply rising, physical and online violence towards journalists is increasing, and the journalism business model is failing. At the same time, funding for public media is declining, and drastic cuts to international aid and media development budgets have reduced the amount of reliable, independent journalism available to audiences globally.
This report explains why these cuts and attacks matter – and what is at stake when journalism declines and disappears. It synthesises the latest academic research on the value of journalism and its role in economics, national security and crises. The report is focused on public interest journalism, that is: reporting that is independent, accurate and ethical, and that seeks to inform the public about important issues affecting their lives, enable debate, and hold power to account.
The evidence collected here demonstrates that this journalism can have a profound and positive impact on societies and individuals globally. Its role supporting democracy is very well established: journalism shares the information citizens need to cast meaningful votes, it is a check and balance on power, and it acts as a conduit between citizens and elected officials.
But journalism’s impact goes far beyond that. As the evidence in the report indicates: Journalism is an economic enabler: it can reduce corruption, lower the cost of doing business, encourage the fair and transparent distribution of resources, and support economic growth and development. Journalism supports national security: it makes societies more resilient to disinformation and the interference of malicious actors, and reduces the risk of conflict and war.Journalism improves the response to crises: it provides information that saves lives, and improves preparation and response to disaster and crises.Journalism is also a surprisingly cost-effective way to achieve these positive outcomes, and it offers a high return on investment. As the evidence in this review shows, every $1 spent on journalism, can result in more than $100 in savings to the public through improved public services and reduced corruption. Journalism’s most significant impact, however, is that it plays a preventative and protective role – as part of a healthy information infrastructure…(More)”.
Blog by Stefaan Verhulst: “For two decades, the open data community has worked to articulate what it means to make data fit for use, converging on the now-canonical FAIR principles: Findable, Accessible, Interoperable, and Reusable. As we argue in Moving Toward the FAIR-R Principles, this framework, though foundational, is no longer sufficient in an era of artificial intelligence. It must be extended to include a fifth commitment: Ready for AI.
The conversation over data readiness for the AI age is not confined to the public-interest domain. Parallel discussions have unfolded in the corporate sector, where the imperatives of AI readiness have been confronted at scale, and it is worth asking what the open data movement might learn from them.
McKinsey’s recent article, AI Data Readiness: The Key to Scaling Impact, is instructive in this regard. On its surface, the article reads as a memo to chief data officers diagnosing why corporate AI pilots so often stall before reaching production. It also offers some valuable lessons and principles for data stewards operating in the public interest, and who may be considering how to apply FAIR-R principles in practice.
In what follows, we treat this private-sector-focused article as a source of transferable knowledge while still remaining attentive to its limits. The caveat about limits is consequential, and we return to it below: in general, while enterprises optimize primarily for reliable outputs and cost avoidance, public-interest stewardship must also account for equity, openness, and public accountability. Some lessons therefore transfer cleanly, while others may require translation (or simply not apply at all)…(More)”.
Paper by Neil D Lawrence and Jessica K Montgomery: “Public dialogues have produced clear demand signals for AI. These dialogues imagine innovations that improve our shared wellbeing and prosperity while developing under democratic control. AI development has largely advanced along another trajectory. We argue that this gap is a structural feature of an innovation system whose incentives are captured by the attention economy. Closing it requires a different driver of the innovation cycle. We propose an attention reinvestment cycle, in which efficiency gains accrue as freed time that can be invested in community innovation, with frontline professionals adapting and sharing tools that meet priority needs. Operating the cycle at scale requires institutional infrastructure—dialogic, absorptive, and distributive capacities—that the existing science-policy system has yet to develop…(More)”.
Release by University College London: “A major new resource that provides one of the most comprehensive pictures yet of what people are eating around the world has been introduced in a new study by a UCL and University of Oxford researcher.
From obesity and heart disease to climate change and food affordability, many of today’s biggest challenges are shaped by what we eat. But there is a surprisingly basic problem facing researchers and policymakers: We often don’t know with enough accuracy what people are actually consuming.
The new study, published in Nature Food and authored by Professor Marco Springmann (UCL Institute for Global Health as well as the Environmental Change Institute at the University of Oxford), introduces the Global Dietary Database for Impact Assessments (GDD-IA), which is freely available through an interactive online explorer, allowing users to investigate dietary patterns across countries and over time.
The GDD-IA combines information on food production, food waste, dietary surveys and human energy requirements to estimate what people eat from 1990 to 2020. It includes detail by age, sex and whether people live in urban or rural areas.
The resource has been designed to support research into some of the world’s most pressing questions: How do diets affect human health? What impact do they have on climate change and the environment? How affordable are healthy and sustainable diets for different populations?..(More)”.
Book by Andrew Guthrie Ferguson: “For consumers living in a digitally-connected world, smart technologies have built an inescapable trap of digital self-surveillance. Smart cars, smart homes, smart watches, and smart medical devices track our most private activities and intimate patterns. While these devices allow users to receive personal insights by monitoring their every move, that data can be accessed by police and prosecutors looking to find incriminating clues. Digital technology exposes everyone, everywhere, all at once, and we have few laws to regulate it.
In Your Data Will Be Used Against You, Andrew Guthrie Ferguson warns us of how the rise of sensor-driven technology, social media monitoring, and artificial intelligence can be weaponized against democratic values and personal freedoms. At the same time, that data will solve crimes, radically transforming how criminal cases are prosecuted. Ferguson explores how this proliferation of private data in combination with public surveillance networks promises new ways to solve previously unsolvable crimes, but also leaves us vulnerable to governmental overreach and abuse. He argues for legal interventions that address the threat of digital self-surveillance and provides concrete suggestions about how legislators, judges, and communities should respond…(More)”.
Article by Jack Hardinges, and Irina Bejan: “Attribution has always been a cornerstone of working with other people’s creativity and knowledge.
As Creative Commons reminds us, practicing attribution serves many important functions, from providing verifiable evidence of a claim, to paying respect to the work of others, and creating pathways for traffic and financial value to flow.
Despite the importance of attribution, many of today’s AI systems fail to acknowledge the sources of knowledge and creativity that make them possible.
A recent study led by the AI Disclosures Project found that more than 30% of responses by leading search-enabled Large Language Models (LLMs) provided no attribution whatsoever. Even where attribution is provided, it can be limited in ways that are hidden to users, and deep technical challenges persist in connecting outputs from AI to the sources they are derived from.
Despite this, we believe there are good reasons to be optimistic that AI systems can attribute the sources they use. Not perfectly, and not always, but enough and improving, such that we should reject the argument that a lack of attribution is an inevitable side effect of the way the technology works. The commons needn’t become “hidden substrate.”..(More)”.