Building Africa's Human
Data Infrastructure
Tonative brought a community booth and a capacity-building workshop to DLI 2026, bringing together dataset creators, translators, validators, and researchers working to strengthen the people, skills, and systems Africa's AI future depends on.
About the Workshop
The rapid growth of generative AI has intensified demand for high-quality datasets, yet progress in African AI remains constrained by gaps in human data infrastructure - the people, skills, and coordinated pipelines required to create, translate, validate, and maintain datasets for African languages and contexts.
Hosted by Tonative Africa at the Deep Learning Indaba 2026 in Nigeria, this session brought together creators, translators, validators, and annotators for capacity building. Through a keynote from Dr. Lilian Wanzare, breakout discussions, and collaborative roadmap design, participants explored best practices for dataset creation, quality assurance, and long-term capacity development across African language communities.
Alongside the workshop, Tonative also ran a community booth throughout the conference, connecting with attendees, demoing the platform, and welcoming new contributors. Together, the two produced shared guidelines, surfaced priority challenges, and a roadmap for strengthening Africa's sovereign, sustainable, and locally owned AI ecosystems.
Location
Abuja Hall, Pan-Atlantic University, Lagos, Nigeria
Date
Thursday, 6 August 2026
Time
12:00 PM
Duration
1 hour, 30 minutes
Programme
Official Deep Learning Indaba 2026 programme
Format
In-person forum, dialogues, and community booth
Photo Gallery
Moments from the community booth, keynote, workshop, and breakout sessions at DLI 2026.
Tonative Group Picture
1 of 20 photos
Session Agenda
How the session ran, section by section.
Opening & Context
Welcome, session goals, and an introduction to the Tonative Data Academy and African AI data challenges. Facilitated by Cynthia Amol.
Guest Talk
A short invited talk highlighting challenges in African dataset creation, community capacity-building initiatives, and sustainable data pipelines.
Breakout Discussions
Participants split into groups around key pipeline stages, dataset creation & collection, translation & validation & annotation, and dataset usage & evaluation, to surface challenges, needs, and opportunities.
Collaborative Roadmap Building
Groups shared key insights and co-developed a shared roadmap for strengthening capacity, improving coordination across language communities, and designing scalable data pipelines.
Synthesis & Next Steps
Key takeaways, opportunities for collaboration, and post-Indaba follow-up plans, including a shared resource toolkit and cross-community collaboration network.
Guest Speaker
We were glad to welcome the following speaker.

Lilian D. A. Wanzare, PhD
KenCorpus · MCAAI, Maseno University
Guest Speaker
Co-founder of KenCorpus and Research Lead at the Maseno Centre for Applied Artificial Intelligence (MCAAI), with a decade of experience curating datasets across Kenyan languages including Dholuo, Kikuyu, Kalenjin, Maasai, Somali and Kenyan Sign Language (KSL). As Principal Investigator for KenCorpus, African Next Voices – Kenya and AI4KSL, she led the collection of some of the largest speech, text and sign language datasets for Kenyan languages, working directly with language communities throughout. Her research centres the empowerment of human data infrastructure as foundational to the AI development pipeline.
In her keynote, Dr. Wanzare shared insights on capacity building for African language data curators, the infrastructure needed for a scalable curation pipeline, and how to govern these datasets.
LinkedIn ProfileWorkshop Organisers

Alfred Kondoro
Head of Research, Tonative Africa
Lead Organiser
Alfred is a Tanzanian PhD researcher in Data Science at Hanyang University, Republic of Korea. He leads community-driven research initiatives at Tonative Africa aimed at strengthening African representation in AI, with work spanning NLP, HCI, and ICTD. His publications have appeared at EACL, AAAI, ACL, CHI, IMWUT, CIKM, CUI, and AfriCHI venues.
LinkedIn
Sharon Ibejih
Founder, Tonative Africa
Organiser
Sharon is a Senior Data Scientist at Ignite Energy Access and the Founder of Tonative Africa. Her work focuses on NLP, data curation, and AI pipeline design for low-resource African languages. She holds an MSc in Data Science and has presented at AfricaNLP, NeurIPS WiML, ICLR workshops, Deep Learning Indaba, and CVPR.
LinkedIn
Cynthia Amol
Co-Founder & Head of Data, Tonative Africa
Organiser
Cynthia is a PhD student in Computer Science and Google NLP Fellow at Maseno University, Kenya. A Deep Learning Indaba Alele-Williams Masters Award recipient, she leads the data validation pipeline at Tonative Africa and has co-organised workshops at NeurIPS, LREC-COLING, EACL, and CHI.
LinkedIn
Chinenye Anikwenze
Engineering Lead, Tonative Africa
Organiser
Chinenye is a Software Engineer and Automation Specialist focusing on defensive infrastructure and AI safety. As Engineering Lead at Tonative Africa, she manages technical infrastructure for 400+ contributors. Her research on Semantic Collapse and the security of tonal languages was recently presented at AFLC 2026 and Impact Fellowship Summit IREX 2026.
LinkedIn
Joy Olusanya
NLP Researcher & Training Manager, Tonative Africa
Organiser
Joy is a linguist and NLP researcher focusing on low-resource language technologies, multilingual NLP, and benchmark evaluation. She served as Workshop Chair for the CLRLC–LLMs Workshop at NeurIPS 2025 and is Founder and CEO of the Center for Low-Resource Languages and Cultures.
LinkedIn
Armand Bukama
Social Manager, Tonative Africa
Organiser
Armand is a Congolese computer scientist from the DRC with a degree from the Catholic University of Bukavu. He leads community engagement, outreach, and communications at Tonative Africa, and works at the intersection of AI, electronics, and sustainable energy for underserved communities.
LinkedIn
Faisal Muhammad Adam
Hausa Language Validation Lead, Tonative Africa
Organiser
Faisal is a lecturer and data science practitioner based in Kano, Nigeria, pursuing graduate studies in Applied Data Science at WorldQuant University. He serves as Hausa Language Validation Lead at Tonative Africa, coordinating contributors on multilingual dataset validation and quality assurance.
LinkedIn
Godspraise Okechukwu
Project Lead, Tonative Research Group
Organiser
Godspraise is a software engineer and NLP researcher based in Nigeria. As Project Lead in the Tonative Research Group, he contributes to dataset creation, validation, and multilingual resource development for African languages, building community-driven data pipelines for low-resource languages.
LinkedInWhat Came Out of It
Tangible, community-owned resources and connections from the session.
Community Guidelines
Shared guidelines for dataset creation and validation across African language communities.
Resource & Toolkit List
A curated list of tools, frameworks, and training resources for African language data pipelines.
Capacity-Building Roadmap
An actionable roadmap for scaling capacity-building initiatives and governance structures.
Collaboration Network
A cross-community network linking language teams, researchers, and practitioners.
Post-Indaba Summary Report
A public synthesis of insights and recommendations to inform future collaborations and publications.
Stay Involved
Missed us at DLI 2026 or want to keep building on what came out of the session? Reach out to the Tonative team - we'd love to hear from you.