Stefaan Verhulst
Press Release: “The U.S. National Science Foundation announced today the launch of the NSF Open Knowledge Network (NSF OKN), a national open-data infrastructure that links 43 interconnected knowledge graphs and tens of billions of connected facts across health, the environment, justice, manufacturing, and national security. The network is live and open to the public at okn.us.
NSF OKN provides a structured, persistent, verifiable knowledge layer to complement contemporary artificial intelligence models, providing grounding of facts, attribution, and knowledge governance essential to critical AI applications in a number of fields. OKN’s value goes beyond machines querying structured knowledge: Its deeper value provides AI systems with a shared, explicit, interoperable, and auditable representation of the world on which they are expected to reason and act…(More)”.
Paper by Maojun Sun et al: “In recent years, data science agents powered by Large Language Models (LLMs), known as “data agents,” have shown significant potential to transform the traditional data analysis paradigm. This survey provides an overview of the evolution, capabilities, and applications of LLM-based data agents, highlighting their role in simplifying complex data tasks and lowering the entry barrier for users without related expertise. We explore current trends in the design of LLM-based frameworks, detailing essential features such as planning, reasoning, reflection, multi-agent collaboration, user interface, knowledge integration, and system design, which enable agents to address data-centric problems with minimal human intervention. Furthermore, we analyze several case studies to demonstrate the practical applications of various data agents in real-world scenarios. Finally, we identify key challenges and propose future research directions to advance the development of data agents into intelligent statistical analysis software…(More)”.
Paper by Katja Mayer: “Open Science has long been framed as a normative project grounded in values of transparency, accessibility and collaboration. Yet in the age of generative AI, many research data practitioners, who once championed openness, are increasingly questioning how these commitments can be upheld when openness also enables undesired large-scale data extraction, decontextualization and enclosure. This paper examines how the normativities of openness are being re-negotiated, and how researchers navigate the ethical and political tensions that arise when their work is repurposed for proprietary AI systems.
Drawing on in-depth qualitative interviews, I explore how situated ethical reasoning informs decisions about data sharing. Participants described new risks and dilemmas, from uncontrolled scraping and heightened re-identification threats to the reframing of FAIR (Findable, Accessible, Interoperable, Reusable) principles as pipelines for AI-readiness. Their accounts reveal openness as a relational and contested practice, embedded in specific institutional, infrastructural and community contexts, and continuously re-balanced between care, accountability and collaboration.
Rather than rejecting openness outright, interviewees described recalibrating it – through selective sharing, alternative infrastructures and advocacy for commons-based governance models…(More)”.
Report by the World Economic Forum: ‘Data sovereignty has become one of the defining governance questions of the digital age, yet the concept remains contested. Governments, individuals, indigenous groups, technology companies and regional blocs all define it differently, and localization alone cannot deliver it. Sovereignty in practice means aligning legal authority with technical control, across storage, compute, models and identity, while still enabling trusted, equitable collaboration across borders.
This briefing paper explores how governments, businesses and civil society can build and verify credible data sovereignty, balancing legitimate control with the trusted interdependence most actors need to thrive…(More)”.
Article by Quentin H. Riser et al: “Fragmentation across early care and education (ECE) systems in the United States obscures how many children access services and how these experiences shape school readiness. This study leverages Iowa’s Integrated Data System for Decision-Making (I2D2), which links health, social determinants of health (SDOH), and education records, to provide one of the first unduplicated counts of ECE participation statewide. Drawing on a matched cohort of 27,321 children eligible for kindergarten in 2017–2018, we examine three aims: (1) describe patterns of centre-based ECE participation; (2) assess associations between child/family characteristics and participation; and (3) evaluate links between ECE and kindergarten suspensions and attendance. Results indicate that 73\% of children participated in at least one ECE program prior to school entry, with substantial overlap across public and private preschool, and subsidised or unsubsidised child care. ECE participation varied as a function of poverty, race/ethnicity, maternal education, and cumulative early-life risks. ECE participation was associated with better kindergarten attendance but not suspension. Findings underscore both the promise of integrated data for advancing equity-driven research and the need for policies that target families with the greatest barriers to ECE access…(More)”.
Article by Katherine Loh, and Imad Aad: “On September 28, 2025, Swiss voters once again considered a referendum on a national digital identity (e-ID) law passed by Parliament. The government’s first attempt at legislating a national digital ID passed in 2019 but was roundly rejected by referendum in 2021, when 64% of voters opposed the law, citing concerns that it delegated operations to private companies. The popular backlash catalyzed years of national deliberation on the question of what kind of e-ID model could meet citizens’ demands for transparency, privacy, and security. Parliament passed a redesigned e-ID law in 2024, and this one stood the test of another referendum, but by the thinnest of margins. The new law was supported by just 50.4% of voters.
This referendum back-and-forth is not unfamiliar to the Swiss. By some accounts, Switzerland holds one-fifth of the world’s national referendums. Switzerland’s direct democracy enables any party that gathers more than 50,000 signatures to launch a nationwide referendum. Combining national, cantonal, and municipal-level referendum initiatives, Swiss citizens vote on around a dozen referendums a year.
What was unique about the e-ID debate was the way it provoked an unusual level of anxiety and mistrust. In a country where people have long relied on the digital world for banking, payments, communications, and travel planning, development of a national digital ID system brought into sharper focus a broader unease around the digitalization of society. For the Swiss, an e-ID represents more than an app; it is a bellwether for the future of long-cherished rights and values around identity and citizenship in the digital age…(More)”.
Paper by Stefaan Verhulst, Adam Zable, Begoña Gonzalez Otero and Jacqueline Lu: “Contemporary data governance has converged on participation. Citizens’ assemblies, deliberative polls, co-design workshops, and community advisory panels are now routinely suggested to build trust and secure a social license for the re-use of data. Yet this consensus rests on an unexamined assumption: that giving affected communities a voice will, on its own, change what data holders do. This paper argues that the field needs to complement voice with leverage-the capacity of individuals and communities to make their participation consequential by attaching credible legal, institutional, contractual, economic, or reputational costs to being ignored. Drawing on labor-relations bargaining theory, the sociology of disputes, and James C. Scott’s account of legibility, we develop leverage as an analytic category defined by three properties: it is counterfactual, enforceable, and durable. We further argue that leverage presupposes epistemic legibility-affected publics can rarely impose costs on a system they cannot see, document, or contest-and that most existing transparency instruments equip regulators and procurers rather than the governed. We then map six sources from which communities and data subjects can derive leverage (legal and regulatory; institutional decision; contractual; economic; data access; and political and reputational), examine the intermediary structures that aggregate and sustain it, and propose a six-dimension evaluative framework-obligation, enforceability, consequence, monitoring equipment, revocability, and institutional durability-for auditing whether any social license process actually redistributes power. The paper’s central claim is that a social license worth the name must share the properties of a legal license: conditions, a licensor, and enforceable remedies for breach…(More)”.
Report by UN Women: “…provides practical guidance for using data to design, implement and monitor gender-responsive care systems.
The publication is grounded in the 5Rs framework—Recognize, Reduce, Redistribute, Reward and Represent. It shows how governments and partners can connect diverse data sources, institutions and stakeholders to better understand care needs, address inequalities and translate evidence into more inclusive policies, services and investments across Asia and the Pacific…(More)”.
Article by Taylor Butler: “A sales receipt, a web search, a satellite image, a job posting that expired two years ago; the seemingly mundane multitudes of data being created by daily activities hold unique and valuable insights for researchers to parse.
What is alternative data, and what makes it so useful for research? We’re not talking about fringe science here, but the creative application of non-traditional sources of data to answering research questions. I spend a lot of time working with a catalog of datasets that captures a vast and varied array of human behavior and world phenomena, and I am constantly inspired by the innovative problem solving researchers employ to extract answers from new sources and diverse perspectives.
What makes data “alternative”?
Alternative data, with its many alternate names (organic [data], novel, found, digital trace, unstructured, observational, derived, etc.), is essentially any data that was generated for some purpose other than research – typically as a byproduct of some administrative, commercial, or technical process – and then utilized for research. Common sources of “alt data” include satellite imagery, mobile-phone records, usage data, administrative records, transaction records, job postings, social media and internet activity, sensor data, and much more.
This differs from traditional datasets that are designed, built, generated or measured, such as structured surveys, experimental data collection, or official government reporting. Alt data offers novel “proxy” measurements of real-time, real-life behavior and phenomena, giving researchers unique perspectives to new and old questions. While economics has paved the way in alt data usage, adopting large-scale administrative and private sector data in the study of market trends (Einav & Levin, 2014), many fields are also implementing new data streams into their research, from global health to climate change, public policy and more.
Take an age old question like: does spending time in nature improve mental health? There are many ways researchers go about answering this, the most traditional method typically being a survey or maybe even a random assignment experiment. But these researchers (Chandra Mouli et al., 2026) combined a traditional data source (CDC health measures) with real-life measures of behavior (cell phone mobility data) and neighborhood green space (NASA satellite imagery). Looking at nine US metro areas, they found that neighborhoods where residents actually visited nearby green space had lower rates of poor mental health; the presence of parks alone did not explain it. No single one of these three sources could have provided this nuanced insight; the combination of disparate data sources provides researchers with opportunities to explore old questions in creative new ways…(More)”.
Article by Nitasha Tiku: “At a beachfront resort in San Juan, Puerto Rico, SpaceX founder Elon Musk joined a three-hour discussion on humanity’s dire fate if artificial intelligence started to rapidly improve.
The afternoon panel was part of a private conference hosted in early 2015 by the Future of Life Institute, a new nonprofit whose founders — mostly outsiders to AI — believed the nascent technology would grow so powerful it could render humans extinct.
The first presenter was Oxford University philosopher Nick Bostrom, author of the recent bestseller “Superintelligence: Paths, Dangers, Strategies,” whose slides argued that AI could either help humanity spread across the cosmos or drive it to extinction.
The 80-person guest list was filled with bold-faced names, including the co-founders of DeepMind, acquired by Google the year before, and three future co-founders of OpenAI including Musk, who backed the nonprofit research lab’s launch later that year.
The conference aimed to legitimize concerns about AI’s risks within the industry and spur research into preventing them, physicist and Future of Life Institute co-founder Max Tegmark wrote in his book “Life 3.0: Being Human in the Age of Artificial Intelligence.” The elite gathering made it “harder to claim that people concerned about AI safety didn’t know what they were talking about,” he wrote.
Predictions involving human extinction, built on themes aired at the 2015 meeting, reached a wider audience in recent weeks after they were invoked by employees at leading AI firms who captured the world’s attention.
On Tuesday, the Future of Life Institute hosted a day-long event in Washington called the Pro-Human Assembly where Sen. Bernie Sanders (I-Vermont) and former Trump adviser Stephen K. Bannon in successive speeches called AI a threat to the species. “I had to metaphorically pinch my arm to make sure I wasn’t dreaming,” Tegmark told The Washington Post. “People are freaking out about it all across the political spectrum.”
Thought experiments once deployed to sway AI insiders are now shaping the public imagination, circulating through Washington and influencing how lawmakers plan to govern a multitrillion-dollar industry. But some AI and policy experts wonder if the parables could lead decision-makers astray…(More)”.