Explore our articles
View All Results

Stefaan Verhulst

Paper by Kevin A. Bryan & Joshua S. Gans: “AI predicts; humans use its predictions to make decisions. These predictions are combined with human verification and analysis, queries to other statistical models, and so on. The economic value of an AI, therefore, depends on how it interacts with the surrounding decision environment. We describe the value of AI as part of this “composite experiment” where AI makes a coarse prediction of the state of the world, show what this means for optimal model training via a geometric argument, explain why optimal training can be discontinuous in economic variables, and study how heterogeneous users or monopoly model trainers affect these results. In particular, maximizing the unconditional accuracy of AI predictions is generally suboptimal…(More)”.

Training AI for When Humans Will Use It

Report by UNESCO: “Media freedom and journalism are under attack around the world. AI-fuelled disinformation is confusing and polarizing audiences, government censorship is sharply rising, physical and online violence towards journalists is increasing, and the journalism business model is failing. At the same time, funding for public media is declining, and drastic cuts to international aid and media development budgets have reduced the amount of reliable, independent journalism available to audiences globally.

This report explains why these cuts and attacks matter – and what is at stake when journalism declines and disappears. It synthesises the latest academic research on the value of journalism and its role in economics, national security and crises. The report is focused on public interest journalism, that is: reporting that is independent, accurate and ethical, and that seeks to inform the public about important issues affecting their lives, enable debate, and hold power to account.

The evidence collected here demonstrates that this journalism can have a profound and positive impact on societies and individuals globally. Its role supporting democracy is very well established: journalism shares the information citizens need to cast meaningful votes, it is a check and balance on power, and it acts as a conduit between citizens and elected officials.

But journalism’s impact goes far beyond that. As the evidence in the report indicates: Journalism is an economic enabler: it can reduce corruption, lower the cost of doing business, encourage the fair and transparent distribution of resources, and support economic growth and development. Journalism supports national security: it makes societies more resilient to disinformation and the interference of malicious actors, and reduces the risk of conflict and war.Journalism improves the response to crises: it provides information that saves lives, and improves preparation and response to disaster and crises.Journalism is also a surprisingly cost-effective way to achieve these positive outcomes, and it offers a high return on investment. As the evidence in this review shows, every $1 spent on journalism, can result in more than $100 in savings to the public through improved public services and reduced corruption. Journalism’s most significant impact, however, is that it plays a preventative and protective role – as part of a healthy information infrastructure…(More)”.

The value of journalism: global evidence on why media matters to economies, national security and crises

Blog by Stefaan Verhulst: “For two decades, the open data community has worked to articulate what it means to make data fit for use, converging on the now-canonical FAIR principles: Findable, Accessible, Interoperable, and Reusable. As we argue in Moving Toward the FAIR-R Principles, this framework, though foundational, is no longer sufficient in an era of artificial intelligence. It must be extended to include a fifth commitment: Ready for AI.

The conversation over data readiness for the AI age is not confined to the public-interest domain. Parallel discussions have unfolded in the corporate sector, where the imperatives of AI readiness have been confronted at scale, and it is worth asking what the open data movement might learn from them.

McKinsey’s recent article, AI Data Readiness: The Key to Scaling Impact, is instructive in this regard. On its surface, the article reads as a memo to chief data officers diagnosing why corporate AI pilots so often stall before reaching production. It also offers some valuable lessons and principles for data stewards operating in the public interest, and who may be considering how to apply FAIR-R principles in practice.

In what follows, we treat this private-sector-focused article as a source of transferable knowledge while still remaining attentive to its limits. The caveat about limits is consequential, and we return to it below: in general, while enterprises optimize primarily for reliable outputs and cost avoidance, public-interest stewardship must also account for equity, openness, and public accountability. Some lessons therefore transfer cleanly, while others may require translation (or simply not apply at all)…(More)”.

What Open Data Stewards Can Learn from Corporate Guidance on AI-Ready Data

Paper by Neil D Lawrence and Jessica K Montgomery: “Public dialogues have produced clear demand signals for AI. These dialogues imagine innovations that improve our shared wellbeing and prosperity while developing under democratic control. AI development has largely advanced along another trajectory. We argue that this gap is a structural feature of an innovation system whose incentives are captured by the attention economy. Closing it requires a different driver of the innovation cycle. We propose an attention reinvestment cycle, in which efficiency gains accrue as freed time that can be invested in community innovation, with frontline professionals adapting and sharing tools that meet priority needs. Operating the cycle at scale requires institutional infrastructure—dialogic, absorptive, and distributive capacities—that the existing science-policy system has yet to develop…(More)”.

Mind the gap: connecting AI innovation to widespread public value

Release by University College London: “A major new resource that provides one of the most comprehensive pictures yet of what people are eating around the world has been introduced in a new study by a UCL and University of Oxford researcher.

From obesity and heart disease to climate change and food affordability, many of today’s biggest challenges are shaped by what we eat. But there is a surprisingly basic problem facing researchers and policymakers: We often don’t know with enough accuracy what people are actually consuming.

The new study, published in Nature Food and authored by Professor Marco Springmann (UCL Institute for Global Health as well as the Environmental Change Institute at the University of Oxford), introduces the Global Dietary Database for Impact Assessments (GDD-IA), which is freely available through an interactive online explorer, allowing users to investigate dietary patterns across countries and over time.

The GDD-IA combines information on food production, food waste, dietary surveys and human energy requirements to estimate what people eat from 1990 to 2020. It includes detail by age, sex and whether people live in urban or rural areas.

The resource has been designed to support research into some of the world’s most pressing questions: How do diets affect human health? What impact do they have on climate change and the environment? How affordable are healthy and sustainable diets for different populations?..(More)”.

What do people really eat? New global database gives best answer yet

Book by Andrew Guthrie Ferguson: “For consumers living in a digitally-connected world, smart technologies have built an inescapable trap of digital self-surveillance. Smart cars, smart homes, smart watches, and smart medical devices track our most private activities and intimate patterns. While these devices allow users to receive personal insights by monitoring their every move, that data can be accessed by police and prosecutors looking to find incriminating clues. Digital technology exposes everyone, everywhere, all at once, and we have few laws to regulate it.

In Your Data Will Be Used Against You, Andrew Guthrie Ferguson warns us of how the rise of sensor-driven technology, social media monitoring, and artificial intelligence can be weaponized against democratic values and personal freedoms. At the same time, that data will solve crimes, radically transforming how criminal cases are prosecuted. Ferguson explores how this proliferation of private data in combination with public surveillance networks promises new ways to solve previously unsolvable crimes, but also leaves us vulnerable to governmental overreach and abuse. He argues for legal interventions that address the threat of digital self-surveillance and provides concrete suggestions about how legislators, judges, and communities should respond…(More)”.

Your Data Will Be Used Against You

Article by Jack Hardinges, and Irina Bejan: “Attribution has always been a cornerstone of working with other people’s creativity and knowledge.

As Creative Commons reminds us, practicing attribution serves many important functions, from providing verifiable evidence of a claim, to paying respect to the work of others, and creating pathways for traffic and financial value to flow.

Despite the importance of attribution, many of today’s AI systems fail to acknowledge the sources of knowledge and creativity that make them possible.

A recent study led by the AI Disclosures Project found that more than 30% of responses by leading search-enabled Large Language Models (LLMs) provided no attribution whatsoever. Even where attribution is provided, it can be limited in ways that are hidden to users, and deep technical challenges persist in connecting outputs from AI to the sources they are derived from.

Despite this, we believe there are good reasons to be optimistic that AI systems can attribute the sources they use. Not perfectly, and not always, but enough and improving, such that we should reject the argument that a lack of attribution is an inevitable side effect of the way the technology works. The commons needn’t become “hidden substrate.”..(More)”.

Attribution in the Age of AI

Article by Alexandra Bruell: “One of the richest sources of online information is re-evaluating its relationship with Google.

Reddit, the online message board that powers a swath of Google search results, has discussed shutting off the technology giant’s access to its content for AI use, according to people familiar with the matter.

It is part of a growing chorus of online media companies expressing frustration with the tech giant as AI changes the way people ask questions, siphons off search traffic and upends publishers’ revenue models. They say the search engine is no longer a reliable source of visitors, especially after Alphabet’s GOOGL  Google expanded its AI search features in recent months. USA Today, Politico, the Economist, People Inc. and Reuters are all evaluating how, or even if, they will continue to work with Google.

Reddit struck a $60 million-a-year deal in 2024 that allowed Google to use its material to train AI models. But with AI-generated answers to queries reducing clicks to outside websites, Reddit executives are assessing what the upside is of continuing to feed its content to Google, said the people familiar with the matter. The companies are in talks about potentially renewing their deal, which is ending soon.

“This is existential for some categories of publishers,” said David Buttle, CEO of media consulting firm DJB Strategies. “They are looking at more radical things.”..(More)”.

Google Was a Lifeline for Publishers. Now Some Are Thinking of Cutting It Off.

Paper by Zeynep Engin et al: “The digital substrate of states — data, algorithms, infrastructure, platforms, applications — is being governed without adequate conceptual foundations. The ability and legitimacy required to govern this substrate, and to govern with it, are simultaneously misaligned, contested, and structurally absent. We introduce digital statecraft as the organising concept for this emerging field, arguing that ‘digital’ reconstitutes the statecraft question rather than merely extending its domain. The concept operates on two dimensions – statecraft over digital systems, concerning the authority and capacity of the state in relation to the digital substrate itself, and statecraft with digital systems, concerning the deployment of algorithmic tools as instruments of governing authority. And it rests on two foundational requirements, technical coherence and legitimate authority, that are genuinely in tension. We derive ten principles of digital statecraft from these foundations, each naming a condition whose absence produces an identifiable and structural governance failure: public interest first, human-machine complementarity, governability by design, systemic coherence, hybrid institutions, adaptive governance, human centricity and civic agency, accountable and traceable authority, judgment across time, and the non-delegable core. This article takes the state as the starting point, the institutional form that developed historically in response to the problem of effective and legitimate public governance, and the only current candidate for which the full set of legitimacy conditions is institutionally available. But the digital statecraft programme holds open a deeper question than just whether states can reform themselves: governing well in the algorithmic age may require rethinking the boundaries, scale, and affiliative basis of statehood itself…(More)”.

Governing Well in the Algorithmic Age: The Foundations of Digital Statecraft

Article by Emanuel Maiberg: “As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing. “The world’s best AI training data is sitting on a shelf,” ISBNdb, a company that produces what it claims is “the world’s largest book database,” and that offers high-volume book acquisition services for AI companies, says on its site. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.” In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in “model collapse,” a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models. As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing…(More)”

AI Companies Are Buying Tons of Old Books Because They’re Free of AI Slop

Get the latest news right in your inbox

Subscribe to curated findings and actionable knowledge from The Living Library, delivered to your inbox every Friday