Explore our articles
View All Results

Stefaan Verhulst

Report by Pew Research: “In November 2022, OpenAI released ChatGPT to the public for the first time. Less than four years later, around half of U.S. adults say they use chatbots powered by artificial intelligence, including 24% who say they use them daily. These tools’ ability to generate human-sounding text has raised a basic question about the modern web: How much content online is now written by AI rather than by other people?

To explore this question, we used the Common Crawl web archive to collect almost half a million English-language webpages from the past five years – starting a couple of years before the release of ChatGPT. We then ran the text of those pages through an AI detection tool called Open Pangram to see how many of them were likely written or substantially edited by AI…In the July 2026 snapshot, signs of AI authorship can be found in over one-third of pages published after ChatGPT was released. This is in line with other studies that have shown that large shares of recently published pages on the internet were likely written or substantially edited by AI…(More)”.

How Much of the Internet Is Written With AI?

Article by Caitlin Hayes: “Researchers have created an interactive map identifying 1.8 million individual trees in New York City, with a new approach that could optimize the placement of trees for cooling in cities around the world. 

In the study, published Aug. 14 in Scientific Data, researchers identify the types of trees using satellite imagery as well as on-the-ground datasets and 3D data collected with lidar, short for light detection and ranging. The approach was 82% accurate overall and higher for common types of New York City trees. 

It’s the first “wall-to-wall,” citywide map identifying trees, and it includes trees on private property, in natural areas and other unsurveyed areas, which together constitute an estimated 65% of the city’s canopy and have not been accounted for on previous maps. The research team will use the map to determine which trees are best for cooling, and the data could inform a host of other planting decisions where the type of tree makes a difference, from reducing allergens to managing invasive pests. 

“If we’re going to be managing these urban forests well, we need to know what’s there, and now we know that for another 1.4 million trees beyond the street trees that were already surveyed,” said Daniel Katz, a senior author and assistant professor in the School of Integrative Plant Science in the College of Agriculture and Life Sciences. “We can now use this information to make sure we get the most benefits from trees, from cooling to reducing air pollution and flooding. We also hope people will enjoy looking at their own neighborhood and seeing which trees are around them.” 

The research is part of the Cool Trees project, led by co-author Dr. Arnab Ghosh, M.S. ’19, ’23, associate professor of medicine at Weill Cornell Medicine. Optimizing trees’ cooling effects is increasingly important, as heat waves in New York City become more frequent and intense. An unrelated Cornell project, “The Generative Canopy,” aims to leverage street-tree data to determine where more trees are needed to provide shade to vulnerable New York City residents.  

The city recently set a goal of expanding tree cover from 22% to 30% by 2040, and the new map could directly inform those efforts, along with broader research from the Cool Trees team to determine the relationship between tree density and type, temperature and heat-related health risks…(More)”.

‘Wall-to-wall’ map of NYC trees could help cool cities worldwide

Paper by Sachit Mahajan & Dirk Helbing: “As AI systems increasingly influence everyday life, integrating diverse community values is both ethically essential and practically urgent. This paper presents Value-Sensitive Citizen Science (VSCS)—a systematic framework that combines Value-Sensitive Design (VSD) with citizen science to support meaningful public participation in AI development. VSCS addresses gaps in the existing approaches by integrating culturally grounded methods and cognitive scaffolding through the Participatory Value-Cognition Taxonomy (PVCT). Community members engage as co-researchers through iterative cycles guided by extended scenario reasoning (What-if, If-then, Then-what, What-now), translating local values into actionable technical requirements. The framework also embeds governance mechanisms to ensure adaptability, accountability, and ongoing oversight throughout the AI lifecycle. By bridging participatory design with algorithmic accountability, VSCS challenges monocultural and top–down approaches to AI. We discuss practical implications, including power asymmetries, scalability, and epistemic justice, and propose strategies for policymakers and practitioners seeking to advance inclusive, value-driven AI design across diverse sociotechnical contexts…(More)”.

Co-designing AI systems with value-sensitive citizen science

Paper by Aaron Witzki et al: “Access to sufficient and high-quality data has become a critical driver of innovation and competitiveness in today’s data-driven economy. However, many organizations, particularly smaller companies without an established user base, struggle to build or access adequately sized datasets. Yet, often a sufficiently large data base is not available. In this context, the concept of data ecosystems gained traction. Data ecosystems describe a distinct form of digital ecosystems in which data takes the role of the central product. Data ecosystems provide a potential solution to the need for data by facilitating the access. However, these ecosystems need to maintain the sovereignty of the data-generating, private consumer. It is of utmost importance to use or distribute data only with the consumers’ consent. Therefore, motivating users to share actively data with a data ecosystem is a relevant topic. For this reason, the aim of our work is to enable a more comprehensive understanding of the subject of incentive mechanisms fostering data sharing in data ecosystems. The study employs an extended systematic literature review to aggregate prior research in this domain, developing an extensive overview on the subject and highlighting future research avenues. We inductively identify five key dimensions of incentive mechanisms and five categories of influencing factors that may shape their effectiveness and users’ data sharing behavior. Building on these insights, we propose a conceptually grounded typology comprising eight ideal types of incentive mechanisms…(More)”.

What’s in it for me? a systematic literature review on incentive mechanisms fostering users’ data sharing behavior in data ecosystems

Article by Tai Lung: “Can you estimate someone’s cancer risk from the air they breathe? Not exactly. Every person has different genetics, lifestyles, occupations, and environmental exposures that change over time. Air pollution varies from day to day, season to season, and even block to block. But scientists can make remarkably reliable estimates of a neighborhood’s cancer risk based on the type and amount of toxic air pollution.

After a year-long delay, this April, EPA released the latest air toxics data, which only included raw air data downloads. This year, for the first time in nearly 25 years, the air toxics data did not include cancer risk estimates

Without a public explanation or opportunity for input, one of the nation’s most important environmental health datasets has quietly gone dark. 

The disappearance of EPA’s cancer risk data continues a broader trend under this administration of environmental, public health, and other government datasets becoming less available, less complete, or more difficult to access.  

What’s Changed?

For the 2021 data (this most recent release), only raw air emissions and concentration data are available for download. Previously, the EPA also produced detailed explanations and a mapping tool that made the data easier to discover, understand, and use. Additionally, they included cancer risk estimates to help translate data and complex scientific models into information the public could understand. 

Now, if users want to understand cancer risk from air pollutants, they will need to download and analyze detailed air concentrations of nearly 200 toxic air pollutants for more than 8 million individual locations…(More)”.

Cancer Risk Is Still There, Even If the Data Isn’t

Paper by Theo Berger & Hannes Scheffter: “This study provides an innovative approach to identify the relevant competence profiles that data science education should promote for work in the social domain. We assess the supply side of European Data Science for Good (DSG) initiatives and analyse 335 projects. Based on project documentation, we identify 146 distinct data science methods and develop a taxonomy that distinguishes between statistical and machine learning approaches and classifies methods according to the scale of the target variables (nominal, ordinal, metric).

The results show a strong dominance of machine learning, particularly in classification-oriented problem settings, alongside a substantial role for exploratory and descriptive analysis in evaluation tasks. Drawing on these patterns, we provide a discussion on the implications for curriculum design of higher education programmes in social data science and artificial intelligence, arguing for an educational profile that aligns advanced methodological skills in machine learning with solid statistical literacy…(More)”.

What it takes to use AI for social good: evidence from European social good projects and implications for higher education

Article by Gillian Tett: “Worrying about the state of America’s statistics might not seem that urgent, given all the other explosive political rows swirling in Washington. But voters should pay attention. For at the heart of this story are three key questions: can anyone paint an accurate portrait of what America is today? If so, how should this shape voting processes? And can this data be trusted in an era of manipulation, cyber hacks and artificial intelligence?

This matters. America’s founding fathers realised 250 years ago that you cannot create fair voting systems without measuring the population. So Article I, Section 2 of the Constitution states that Congress should carry out a census in “such manner as they shall by Law direct” every 10 years. Since 1790 census workers have always done that, crossing the country to count people, whether by horse, foot or car (or, in Alaska, on snow sleds).

The Constitution also stipulates that these census results should be used for “apportionment”, to divide the 435 seats in the US House of Representatives among the 50 states, based on each state’s population count. (The allocation in the Senate is permanently fixed.)

This makes it highly political. But it also matters economically: the census shapes how $1.5tn of state aid is distributed each year and is critical to business activity too. Civil rights activists argue this enables a fair apportionment and representation for minorities…(More)”.

Who counts in Trump’s America?

Report by the U.S.-China Economic and Security Review Commission: “China watchers in 2021 would not have predicted the country would lead the world in commercializing data and treating it as an asset. After regulators in 2020 abruptly canceled the mega-IPO of Ant Group, then one of China’s largest and most prominent fintech firms, China’s homegrown tech giants spent the next three years in the crosshairs of the Chinese Communist Party (CCP). Curbing the private sector’s ability to collect and use data without state oversight was one focus of the tech crackdown, with the Cyberspace Administration of China (CAC) targeting technology firms and social media platforms. In those same years, China stepped up censorship of economic data, canceling official series and banning private estimates that painted an unflattering picture and limiting foreign access to Chinese data and its aggregators. 

Yet as it was reining in big tech’s data practices, China’s government was figuring out how to extract value from data at a national scale. Although Chinese officials designated data as a factor of production in 2020, China’s first significant action to reopen space for commercial innovation and data monetization policies came in 2022 with a sweeping development framework. Fast forward to 2026, and China’s data economy is beginning to flourish through government-led data exchanges, new data accounting rules, and pilot programs that treat data as a resource. As more actors refine and package data, they unlock productivity gains for themselves or outside buyers, much like land or capital goods are created, used, and transferred. 

Rather than leaving the development of a data economy up to market forces, China is building infrastructure where participants can exchange data—fostering third-party services like data valuation, analytics, and financing and encouraging wider participation in the data economy…(More)”.

The People’s Republic of Data: How China Is Turning Data into Capital

Article by JP Flores and Hannah Frank: “…The deeper problem is that a system built around journal prestige shapes what science gets done, how fast it moves, and whose work gets seen. 

We come to this argument from different directions. One of us, Hannah, conducts soil research on vineyards. Fundamental to the project are the relationships she has built with grape growers and winemakers. Yet when her work is finally published there is no guarantee those growers will be able to easily access the material, or that the wine industry beyond the academic paywall will be able to take the findings and apply them more broadly.

The other, JP, has committed years to experiments, failed analyses, and revised manuscripts that often seemed to be judged less by what they contributed to science and more by which journal ultimately agreed to publish them. Like many scientists, he was then asked to pay over $10,000 in article processing charges just to make that research publicly available. That money could have funded follow-up experiments, a student’s salary, or an entirely new project.

There are two central critiques of current publishing venues. First, systems that reward publication volume can encourage researchers to pursue smaller, lower-risk projects that generate a steady stream of papers rather than tackle more ambitious questions whose answers may take years to emerge…(More)”.

Science Doesn’t Need Better Journals. It Needs Better Incentives.

Climate Adaptation Knowledge Base by Urban Tech Hub: “Climate adaptation is a trillion-dollar challenge.Today the world spends $190 billion a year defending 1.2 billion people to developed-economy standards. Protecting everyone exposed to climate hazards would cost $540 billion…We’ve cataloged data on 2,667 adaptation actions, 831 solutions, and 12,868 outcomes in 1,309 cities globally…(More)”

Adaptbase

Get the latest news right in your inbox

Subscribe to curated findings and actionable knowledge from The Living Library, delivered to your inbox every Friday