Explore our articles
View All Results

Stefaan Verhulst

Paper by Katinka den Nijs, Elisa Omodei & Vedran Sekara: “Large-scale human mobility datasets play increasingly critical roles in many algorithmic systems, business processes, and policy decisions. Unfortunately, there has been little focus on understanding bias and other fundamental shortcomings of these datasets and how they impact downstream analyses and prediction tasks. In this work, we study ‘data production’, quantifying not only whether individuals are represented in big digital datasets, but also how they are represented in terms of how much data they produce. We study one GPS mobility dataset which is collected from anonymized smartphones for ten major US cities and find that data points can be more unequally distributed between users than wealth. We build models to predict the number of data points we can expect to be produced by the composition of demographic groups living in census tracts, and find strong effects of wealth, ethnicity, and education on data production. While we find that bias is a ubiquitous phenomenon, occurring in all ten cities, we further find that each city suffers from its own manifestation of it, and that location-specific models are required to model bias for each city. This work raises serious questions about general approaches to debias human mobility data and urges further research…(More)”.

A tale of ten cities: data bias in human mobility is pervasive and highly location specific

Paper by Zhivar Sourati et al: “Language is far more than a communication tool; it encodes a wealth of information about a person’s identity, psychological state and social context, providing valuable insights for diverse fields including psychology, marketing and healthcare. Across three studies spanning seven datasets in different domains and over 880,000 texts, we show that the widespread adoption of large language models (LLMs) as writing assistants is linked to declines in linguistic diversity, interfering with the societal and psychological insights language provides. While core content is retained when LLMs polish and rewrite texts, LLMs also homogenize writing styles, reducing writing-complexity variance by a statistically significant 21–50% across datasets and models (P ≤ 0.05), and amplify patterns associated with dominant characteristics while suppressing others, emphasizing conformity over individuality. These trends hold across different LLMs, prompts and contexts, with potential implications for diagnostic processes, personalization efforts, hiring assessments and cultural preservation…(More)”.

The shrinking landscape of linguistic diversity in the age of large language models

Paper by Levin Brinkmann et al: “Intelligent machines have the potential to uncover problem-solving strategies beyond human discovery. Emerging evidence from competitive gameplay, such as Go and chess, demonstrates that AI systems are evolving from mere tools to sources of cultural innovation adopted by humans. However, the conditions under which intelligent machines transition from tools to drivers of persistent cultural change remain unclear. We identify three key dimensions that modulate machine influence on human problem-solving: the discovered strategies must be non-trivial, learnable, and offer a clear advantage. Using a cultural transmission experiment, we demonstrate that when these conditions are met, machine-discovered strategies can be transmitted, understood, and preserved by human populations, leading to enduring cultural shifts. Conversely, using agent-based simulations, we show how machine influence is constrained in the absence of these conditions. These findings provide a framework for understanding how machines can persistently expand human cognitive skills and underscore the need to consider their broader implications for human cognition and cultural evolution…(More)”.



Propagation and preservation of AI-discovered problem-solving strategies in human culture

Article by Tai Lung: “Can you estimate someone’s cancer risk from the air they breathe? Not exactly. Every person has different genetics, lifestyles, occupations, and environmental exposures that change over time. Air pollution varies from day to day, season to season, and even block to block. But scientists can make remarkably reliable estimates of a neighborhood’s cancer risk based on the type and amount of toxic air pollution.

After a year-long delay, this April, EPA released the latest air toxics data, which only included raw air data downloads. This year, for the first time in nearly 25 years, the air toxics data did not include cancer risk estimates

Without a public explanation or opportunity for input, one of the nation’s most important environmental health datasets has quietly gone dark. 

The disappearance of EPA’s cancer risk data continues a broader trend under this administration of environmental, public health, and other government datasets becoming less available, less complete, or more difficult to access.  

What’s Changed?

For the 2021 data (this most recent release), only raw air emissions and concentration data are available for download. Previously, the EPA also produced detailed explanations and a mapping tool that made the data easier to discover, understand, and use. Additionally, they included cancer risk estimates to help translate data and complex scientific models into information the public could understand. 

Now, if users want to understand cancer risk from air pollutants, they will need to download and analyze detailed air concentrations of nearly 200 toxic air pollutants for more than 8 million individual locations…(More)”.

Cancer Risk Is Still There, Even If the Data Isn’t

Paper by Theo Berger & Hannes Scheffter: “This study provides an innovative approach to identify the relevant competence profiles that data science education should promote for work in the social domain. We assess the supply side of European Data Science for Good (DSG) initiatives and analyse 335 projects. Based on project documentation, we identify 146 distinct data science methods and develop a taxonomy that distinguishes between statistical and machine learning approaches and classifies methods according to the scale of the target variables (nominal, ordinal, metric).

The results show a strong dominance of machine learning, particularly in classification-oriented problem settings, alongside a substantial role for exploratory and descriptive analysis in evaluation tasks. Drawing on these patterns, we provide a discussion on the implications for curriculum design of higher education programmes in social data science and artificial intelligence, arguing for an educational profile that aligns advanced methodological skills in machine learning with solid statistical literacy…(More)”.

What it takes to use AI for social good: evidence from European social good projects and implications for higher education

Report by Pew Research: “In November 2022, OpenAI released ChatGPT to the public for the first time. Less than four years later, around half of U.S. adults say they use chatbots powered by artificial intelligence, including 24% who say they use them daily. These tools’ ability to generate human-sounding text has raised a basic question about the modern web: How much content online is now written by AI rather than by other people?

To explore this question, we used the Common Crawl web archive to collect almost half a million English-language webpages from the past five years – starting a couple of years before the release of ChatGPT. We then ran the text of those pages through an AI detection tool called Open Pangram to see how many of them were likely written or substantially edited by AI…In the July 2026 snapshot, signs of AI authorship can be found in over one-third of pages published after ChatGPT was released. This is in line with other studies that have shown that large shares of recently published pages on the internet were likely written or substantially edited by AI…(More)”.

How Much of the Internet Is Written With AI?

Article by Caitlin Hayes: “Researchers have created an interactive map identifying 1.8 million individual trees in New York City, with a new approach that could optimize the placement of trees for cooling in cities around the world. 

In the study, published Aug. 14 in Scientific Data, researchers identify the types of trees using satellite imagery as well as on-the-ground datasets and 3D data collected with lidar, short for light detection and ranging. The approach was 82% accurate overall and higher for common types of New York City trees. 

It’s the first “wall-to-wall,” citywide map identifying trees, and it includes trees on private property, in natural areas and other unsurveyed areas, which together constitute an estimated 65% of the city’s canopy and have not been accounted for on previous maps. The research team will use the map to determine which trees are best for cooling, and the data could inform a host of other planting decisions where the type of tree makes a difference, from reducing allergens to managing invasive pests. 

“If we’re going to be managing these urban forests well, we need to know what’s there, and now we know that for another 1.4 million trees beyond the street trees that were already surveyed,” said Daniel Katz, a senior author and assistant professor in the School of Integrative Plant Science in the College of Agriculture and Life Sciences. “We can now use this information to make sure we get the most benefits from trees, from cooling to reducing air pollution and flooding. We also hope people will enjoy looking at their own neighborhood and seeing which trees are around them.” 

The research is part of the Cool Trees project, led by co-author Dr. Arnab Ghosh, M.S. ’19, ’23, associate professor of medicine at Weill Cornell Medicine. Optimizing trees’ cooling effects is increasingly important, as heat waves in New York City become more frequent and intense. An unrelated Cornell project, “The Generative Canopy,” aims to leverage street-tree data to determine where more trees are needed to provide shade to vulnerable New York City residents.  

The city recently set a goal of expanding tree cover from 22% to 30% by 2040, and the new map could directly inform those efforts, along with broader research from the Cool Trees team to determine the relationship between tree density and type, temperature and heat-related health risks…(More)”.

‘Wall-to-wall’ map of NYC trees could help cool cities worldwide

Paper by Sachit Mahajan & Dirk Helbing: “As AI systems increasingly influence everyday life, integrating diverse community values is both ethically essential and practically urgent. This paper presents Value-Sensitive Citizen Science (VSCS)—a systematic framework that combines Value-Sensitive Design (VSD) with citizen science to support meaningful public participation in AI development. VSCS addresses gaps in the existing approaches by integrating culturally grounded methods and cognitive scaffolding through the Participatory Value-Cognition Taxonomy (PVCT). Community members engage as co-researchers through iterative cycles guided by extended scenario reasoning (What-if, If-then, Then-what, What-now), translating local values into actionable technical requirements. The framework also embeds governance mechanisms to ensure adaptability, accountability, and ongoing oversight throughout the AI lifecycle. By bridging participatory design with algorithmic accountability, VSCS challenges monocultural and top–down approaches to AI. We discuss practical implications, including power asymmetries, scalability, and epistemic justice, and propose strategies for policymakers and practitioners seeking to advance inclusive, value-driven AI design across diverse sociotechnical contexts…(More)”.

Co-designing AI systems with value-sensitive citizen science

Paper by Aaron Witzki et al: “Access to sufficient and high-quality data has become a critical driver of innovation and competitiveness in today’s data-driven economy. However, many organizations, particularly smaller companies without an established user base, struggle to build or access adequately sized datasets. Yet, often a sufficiently large data base is not available. In this context, the concept of data ecosystems gained traction. Data ecosystems describe a distinct form of digital ecosystems in which data takes the role of the central product. Data ecosystems provide a potential solution to the need for data by facilitating the access. However, these ecosystems need to maintain the sovereignty of the data-generating, private consumer. It is of utmost importance to use or distribute data only with the consumers’ consent. Therefore, motivating users to share actively data with a data ecosystem is a relevant topic. For this reason, the aim of our work is to enable a more comprehensive understanding of the subject of incentive mechanisms fostering data sharing in data ecosystems. The study employs an extended systematic literature review to aggregate prior research in this domain, developing an extensive overview on the subject and highlighting future research avenues. We inductively identify five key dimensions of incentive mechanisms and five categories of influencing factors that may shape their effectiveness and users’ data sharing behavior. Building on these insights, we propose a conceptually grounded typology comprising eight ideal types of incentive mechanisms…(More)”.

What’s in it for me? a systematic literature review on incentive mechanisms fostering users’ data sharing behavior in data ecosystems

(Open access) book edited by Riku Neuvonen and Jukka Viljanen: “We live in a world where many things happen digitally. Ones and zeros determine everyday actions and the course of life. The current digital world is two-dimensional. We see, hear and influence it using tools that are clearly part of the physical world. In the current digital world, we see things on screens. During this millennium, screens have become smaller, and in addition to computer monitors, virtual worlds can now be viewed on the screens of tablets and smartphones. Hardware and devices delivering augmented reality experiences are already on the market, as are virtual reality glasses. Headsets of different types are already commonplace.

real virtual world is one that we experience like the physical world. In this book, we imagine the potential of such worlds and we examine these from the perspective of law and related disciplines. The current digital world has led us to questions about the applicability of rights. In theory, the regulation of the physical world also applies to the digital world. In practice, this is not always the case, as problems associated with the protection of privacy but also disinformation, online hate and others have emerged. In this context, we can talk about digital human rights, digital rights and digital constitutionalism.

The problems we face in the digital world are related to access, access to information, and privacy. The virtual world has the same problems, but virtuality as an all-encompassing experience also creates new challenges. In the age of platforms, algorithms have played a key role in displaying content and users are now shown what they or advertisers think they want to see. Various forms of artificial intelligence are also involved. These opportunities and problems of the platform era will also transfer to the virtual worlds of the future. This is especially the case when the advocates of virtual worlds, such as the metaverse, want them to be worlds where users spend a significant portion ofp. 2their time. These questions will be approached and explored from several different starting points.

This book is about considering, imagining and reflecting on virtual worlds. In this introduction, we will establish a picture of what virtual worlds can be, largely based on entertainment, films, books, television series, and games. This basis is logical since these forms of entertainment have been creating and experimenting with virtual worlds for decades. After establishing this picture, we will provide an overview of the research that deals with rights in virtual worlds before introducing the authors of the chapters and how each will approach this challenging theme from different perspectives…(More)”.

Real Rights in the Virtual World: Human Rights in the Age of Artificial Intelligence and Virtual Reality

Get the latest news right in your inbox

Subscribe to curated findings and actionable knowledge from The Living Library, delivered to your inbox every Friday