This is the second in a three-part series sharing findings from our AI for Climate Action project, which is mapping community-centered AI tools for climate action across Latin America, Africa and Asia. In this post, we look at the data practices of organizations with AI projects for climate action. In the first post, we explored how organizations make decisions when considering AI for climate action and the final post will examine the role of AI in supporting social movements.
AI projects start with data. Data is needed to run models and to power prediction and analytics. While collecting large amounts of high quality data can improve environmental management, there are also detrimental environmental impacts of relying on massive models. Additionally, organizations must balance the need to collect large quantities of data and how to govern that data. Not every climate problem requires a large dataset.
Likewise, while smaller models (that use local and more contained datasets) offer the potential to offset environmental impacts and can be a better fit for specific climate issues, there are trade-offs in terms of their capabilities. Organizations must navigate these trade-offs when selecting and building their AI models.
In addition to the environmental implications, climate data can contain sensitive information (especially when tied to vulnerable communities, environmental defenders and protected areas), which can endanger people and the environment if exposed. When choosing to use AI for climate action, organizations must navigate knowing what data to collect and how to safeguard sensitive data (if opting to collect it). Further, organizations may need to grapple with governance and stewardship considerations, especially for projects concerning indigenous or marginalized communities.
Over the past few months, we have been interviewing researchers, CSOs and tech non-profits around community centered AI for climate action initiatives. As an organization that is heavily involved in responsible data, we were drawn to the specific data challenges and governance choices our interviewees made around privacy, epistemic justice, sovereignty and collection practices.
This blog highlights how interviewed organizations navigated some of the key data decisions featured in our case studies around where to find data, whose data and who governs data, particularly when available datasets did not account for local contexts and monitoring needs.
Accessing local data and adapting approaches to fit what data is available
Multiple interviewees spoke to the lack of contextually relevant data available and how they had to go about collecting and building their own datasets, meaning many relied on locally sourced datasets with which they developed purpose-limited AI systems. Organizations often collected and curated their own datasets based on what data was already available, what they had access to and what they had the ability to create for themselves.
When access to data was restricted, some organizations relied on building new relationships. Sauti Data Lab, a civic tech and data innovation hub in Uganda, encountered data sourcing as one of the key challenges in creating Flood Guard, an AI-powered flood mapping and waterborne disease prediction tool designed to address the prevalence of urban flooding and its implications on livelihoods. While some of the data could be easily sourced through free tools like Open Weather or found through satellite images from sources like Sentinel, the availability of historical flood, health and environmental data was more limited and kept in government offices.
Through stakeholder meetings, writing letters and a fortunate fellowship placement with the Uganda Bureau of Statistics (UBOS) funded by Global Partnership for Sustainable Development Data, the team was able to build the necessary relationships to access the restricted datasets. This network also led to introductions with the Ministry of Health and Kampala Capital City Authority, who were also instrumental in coordinating user interviews and feedback sessions with community members living in the most flood affected areas of the city.
While Sauti Data Lab dealt with access restrictions to data, Farmers for Forests and Roaya for Astronomy and Space Applications Foundation adapted their approaches in response to what data was locally available and applicable.
Farmers for Forests developed TreeLens, a drone-based system that creates aerial images of agricultural plots, identifies individual trees and estimates their carbon absorption. In creating the tool, they initially hoped they could use an existing model without having to collect and label large amounts of its own data. The team experimented with DeepForest, an open source tool designed to detect tree crowns, but it underperformed on the organization’s drone imagery. The team believed this was partly because the model had been trained using data from ecosystems that looked different from the Indian agricultural environments in which Farmers for Forests operated. They ultimately had to develop their own internal annotation tools and build their own datasets.
Likewise, Mozn, a youth-led weather-monitoring and early-warning initiative in Libya, developed by Roaya for Astronomy and Space Applications Foundation, faced locally relevant data constraints. Due to a lack of locally available historical weather data, the team adopted a post-processing approach, using machine learning to adjust global weather models to local variations and conditions. Creating an open source system that others could learn from was important to the team. The data is openly available to researchers, university students and others interested in weather and climate information.
“Almost 55 points across Libya are providing free, open-source data. It is not only our system that can use it. Researchers, university students, and anyone interested can use this data for good.”
Mozn also intends to make its code publicly available. Ayat carried the objective of replicability into the model’s design.
“I kept asking: with scarce resources, can this be rebuilt? If another country installed weather stations and collected its own data, could it put that data into the system? There would need to be adjustments for elevation, station information and thresholds, but the structure and idea could be implemented by another machine-learning engineer.”
In our research, organizations consistently considered the experiences of other organizations in their regions, noting how overcoming data scarcity and experimenting with approaches could provide templates and lessons learned for similarly placed initiatives.
Whose data and knowledge counts
Beyond what data is collected, our interviewees raised an important question of whose data and knowledge are accounted for in AI projects? Epistemic justice, fairness in who is allowed to create knowledge and be believed, is an important element of these projects. Researchers and technologists did not simply impose themselves on communities but relied on community expertise to shape models and tools around specific problems.
Multiple projects relied on co-design and community workshops to understand what needs and elements were important to community members before creating and designing tools. As we explored in the previous blog post, Diversa Studio employed active listening with Yaqui community members to understand the community’s data and what things they already knew to be true.
Likewise, for the Saving Voices Project in India, which develops lightweight voice models for Indigenous and under-represented languages (to preserve language, cultural heritage, and ecological traditions), community consultation was necessary to understand what types of knowledge should be represented in the archive and to build the speech model, as a repository of community memory. Instead of researchers providing sheets of unrelated sentences, community members would decide what stories, knowledge and vocabulary should be recorded and how the repository should be organized.
For PetaBencana.id, a Siti OSS disaster map, local knowledge is the bedrock of model design. They explain,
“One thing we really try to do in PetaBencana.id is emphasize that local knowledge needs to inform climate datasets. If we’re going to start with the global climate dataset or a generic one, we need to question whose knowledge is informing that in the first place and what assumptions are being carried through.”
Beyond epistemic justice, relying on local knowledge and expertise also shapes the purpose and impact of climate change mitigation and planning tools. Sometimes historic datasets do not account for local nuances or climate patterns or knowledge about a place or conditions that people know to be true and are essential to conservation. These projects prompt questions around reconciling different forms of knowledge and data, what data fits an AI model (and what to do when traditional knowledge does not fit a model or scientific process?) and how to create dialogue around these different ways of knowing.
Who is collecting: Involving citizens in data collection
For community-centered AI projects, participatory processes are part and parcel of design. In addition to co-design and community informed design, multiple projects used citizen science methodology to widen data collection efforts and ownership in the project. These initiatives were intentional around creating community and pedagogy around collecting data. In particular they were careful in navigating how to maintain the quality and robustness of data that would fit the AI model, while simultaneously widening access to who can collect data.
For Taxonomy Meets Tech, the Saving the Lebanese Biodiversity (SLB) Center Incubator at the Maroun Semaan Faculty of Engineering and Architecture (MSFEA) at American University Beirut, researchers ensured that non-specialists could participate in the taxonomy project through accessible smartphones and concise instructions. In taking this approach, they had to balance how they could best support them to produce data that was technically accurate and could be processed by an AI system, while ensuring local participants remained actors whose knowledge, relationships and motivations shaped how the monitoring system operates.
By taking a collaborative approach, the researchers not only expanded who collected the data, but also facilitated relationship building and community around conservation efforts, further advancing climate action (beyond simply gathering data).
Similarly, SKILIKET, an environmental monitoring project, conceived at the School of Architecture, Art and Design of Tecnologico de Monterrey in Mexico, took a citizen science approach to prompt more intentional observation from students, and encourage the people who are experiencing the environment to seek data and make sense of it. Student testimonials gathered after a semester with SKILIKET suggest the approach changes how they understand the spaces they move through every day.
“One of the activities we did was to start recording certain parameters, carbon, humidity, temperature, among others, which makes us sensitive to what we are living and feeling every day.” (Student testimonial)
“I take advantage of the fact that this knowledge can help me not only for the service, but also to take care of those I care about, like my family, and to apply it at home.” (Student testimonial)
These testimonials point to how widening access in who collects data also achieves SKILIKET’s goal of amplifying communities on the ground to notice and encounter different environmental conditions of the world around them.
Who governs the data: Data sovereignty, ownership and non-extraction
In addition to what data is collected (and who it is collected by), data governance is a crucial element to AI initiatives. These governance decisions impact where data is stored (e.g. opting to not use proprietary cloud services), who has access to the data, what data is not for public consumption and who stewards the data.
Many of those we spoke to work with Indigenous communities, where ecological and ancestral knowledge is under threat. Indigenous AI governance frameworks, which embed collectivism and data stewardship, encompassing elements of consent and mitigating harm, pave the way for more ethical frameworks for considering AI and the environment. Initiatives have tackled the implications of data colonialism and extractive nature of research projects, through prioritizing data privacy, retaining Indigenous knowledge (and keeping private what does not belong to everyone).
An AI of Our Own (AAOO) addresses questions around ownership and quality control in their work. They posit that AI’s potential to preserve languages and cultural practices was welcomed with the caveat that ownership of community knowledge(s) must always remain within the community. Another initiative, Saving Voices, follows the CARE Principles, an indigenous data governance framework, which cover collective benefit, authority to control, responsibility and ethics. It calls for primary data ownership to remain with communities, consent to be ongoing, and communities to retain the right to refuse collection, withdraw consent or require deletion.
Care, stewardship, shared ownership and being forthcoming around limitations and errors are some of the ways in which organizations are maintaining trust with communities and practicing data governance.
What this means for CSOs using AI in their climate action work: About our peer-learning cohort
For social justice organizations using AI in their work, the initiatives in this mapping suggest a set of starting questions to guide their reflections around data collection and governance:
- What data frameworks guide our governance practices and why have we chosen them?
- What does participatory data collection look like and how to maintain robust data in collective practices?
- How is data sovereignty practiced and maintained, especially amongst Indigenous communities? Who has ownership of data and how does this ownership take into account power dynamics and legacies of data colonialism?
These questions will be central to our upcoming peer-learning cohort, designed for practitioners at the intersections of AI, climate action and social justice who are developing, implementing, or exploring AI tools in their work. Guided by The Engine Room, participants will explore emerging trends, practical challenges, and community-centred approach to AI design, data practices, and governance. Together, these discussions will inform a Responsible AI for Climate framework, co-designed with participants to support more context-sensitive, care-centred, and responsible uses of AI for climate action.
Participation is now closed, but stay tuned for updates on our learnings and progress. If you have specific questions or would like to learn more about this initiative , get in touch with us at lesedi@theengineroom.org.
We would like to give special thanks to the twelve organizations who generously shared their stories for this study, as well as our regional researchers: Madhuri Karak, Radhika Jhalani, and Yosr Jouini, who conducted the desk research and interviews in Asia and North Africa.

