How are education practitioners in India, Afghanistan, Malawi, Somalia etc. trying to use artificial intelligence (AI) in their assessment practices? What is working, what is falling short, and what do organisations need to move forward with confidence? These were the questions at the heart of our third and final AI in Education Roundtable, a discovery session hosted by the Global Schools Forum (GSF) to surface real experiences with AI in assessments and begin building a shared understanding of the opportunities and risks ahead.
The session was facilitated by Ekaterina Cooper and Habeeb Kolade and drew participants from non-state education organisations working across low- and middle-income countries (LMICs), as well as higher education contexts. Through plenary discussions and a structured reflection exercise, participants shared honest, grounded accounts of where AI fits in their assessment work. This article captures the full range of insights that emerged.
Setting the Scene
To anchor the conversation, Ekaterina offered a framing on assessments: evidence-gathering processes that inform teaching and learning progress. This definition mattered, because discussions about AI and assessments can quickly narrow to summative, high-stakes exams; in reality, most of the promising AI applications lie elsewhere, in the formative, ongoing, and diagnostic work that teachers do every day.
What Organisations Are Doing: A Snapshot of AI Use in Assessments
Across very different contexts; from Indian government schools to Afghan girls’ education programmes to competency-based learning in Malawi; organisations are actively exploring what AI can and cannot do in assessment.
- Scaling Literacy Assessment in India – One organisation working in Indian government schools has been exploring AI tools to assess foundational literacy; specifically, trying to go deeper than traditional paper-based tests to understand children’s comprehension across regional Hindi dialects. The team developed a prototype using letter and consonant-vowel combinations tailored to different Hindi-speaking states. However, progress stalled due to data privacy constraints. Under India’s Digital Personal Data Protection (DPDP) Act, questions about who owns data collected in government schools remain unresolved, and the team paused their pilot while navigating these governance challenges. Their experience reflects a tension many organisations would recognise: the technical possibility of an AI-powered assessment tool outpaces the institutional and legal frameworks needed to deploy it responsibly.
- Differentiated Assessments for Afghan Girls’ Education – One organisation works with girls and women excluded from secondary education in Afghanistan; a context defined by both acute need and acute constraint. The team has been using AI to add layers of complexity to teacher-created assessments, working up Bloom’s Taxonomy from recall and comprehension through to synthesis and evaluation.
- For multiple-choice questions, the process is semi-automated: AI generates options, but teachers preset the correct answers manually; a deliberate safeguard against unreliable AI marking. The team treats AI as a tool for expanding the range and complexity of questions, not as a replacement for teacher judgement on quality and accuracy.
- Open-ended and handwritten responses remain a significant challenge. Optical character recognition (OCR) for handwritten answers is costly and inconsistent, particularly in local scripts. When students type responses in English, accuracy improves; but typed, open-ended answers are more vulnerable to AI-generated student submissions, raising academic integrity concerns. The team also experimented with knowledge graphs to recommend questions based on prior student performance; a candid question was raised in the session about whether this qualifies as “true AI,” reflecting a broader need for clearer shared definitions in this space.
- Competency-Based Assessment – One organisation assesses students not through tests or grades, but through observed demonstration of specific competencies; such as resilience; during project work and real-life situations. AI plays a limited and carefully bounded role. In the early stages, the team uses AI to help break down and explain competencies so students know what they are working towards. Beyond that, teacher observation and professional judgement drive the process. A challenge familiar across the sector emerged here: some students use AI to write their reflective submissions, producing polished text that does not reflect their actual experience.
- Higher Education: Divergent Paths and Transferable Lessons – A participant from the higher education sector described how universities are responding to AI in two very different ways: some are pulling back, returning to traditional in-person examinations; others are leaning forward, experimenting with portfolios, oral assessments, and AI-assisted feedback. One university developed an internal AI tool that significantly reduced the time faculty spend on feedback while maintaining rubric alignment; faculty satisfaction with the tool has been high. The suggestion was made that higher education, with its relatively greater resources for piloting new approaches, offers lessons worth translating into K-12 and LMIC contexts.
- Other Experiments Across the Community – A participant from Somaliland shared an early-stage comparative study in which the same group of students completed an Early Grade Reading Assessment (EGRA) with both a human invigilator and an AI invigilator. The results have not yet been analysed, but the study raises important questions about comparability and what it means for assessment to feel human. From Tanzania, one participant noted that AI is being used to identify learning gaps and track progress; a key barrier flagged was that many educators lack the skills to use AI tools effectively, and that available tools do not reflect the rural Tanzanian context.
The Barriers: What Is Getting in the Way
Several barriers emerged consistently, even across very different organisational and geographic contexts.
-
Data privacy and governance. The experience in India offers the clearest illustration: a technically promising pilot was paused not because the tool did not work, but because legal frameworks for data ownership in government schools were unclear. As AI tools collect and process student data at scale, questions about ownership, storage, and accountability become unavoidable; and the cost of getting them wrong is high in contexts where trust between NGOs, governments, and communities is hard-won.
-
Academic integrity and AI-generated student work. Multiple participants described students submitting AI-generated work that meets the surface requirements of an assessment without reflecting any actual learning. Organisations are responding by designing assessments that require demonstrated performance; observed competency, oral explanation, or applied tasks that are harder to generate artificially.
-
Teacher trust and tool overload. Practitioners described a widely-shared frustration: teachers face a bewildering array of AI tools with little guidance on which are reliable, which save time, and which suit their context. For AI to earn teacher trust, it needs to demonstrate clearly that it saves time and produces outputs as good as; or better than; what the teacher would have done independently. Administrators, meanwhile, want reliable data that does not bias outcomes for certain students or response types.
-
Cost and accessibility. Premium AI tools tend to produce better outputs but are unaffordable for most organisations in this space. Free tools are more accessible but less reliable and more prone to producing content that requires extensive human revision. A practical suggestion raised in the session was a “smart buys” list; a curated, regularly updated matrix of AI tools organised by function, with honest assessments of cost, quality, and contextual fit.
- Internal champions. Successful AI adoption rarely happens through top-down mandates; it is driven by individuals willing to experiment and build the internal case for change. Identifying and supporting these champions, and creating conditions for their experiments to be shared and scaled, is as important as any external resource.
What Would Help: What Practitioners Are Calling For
-
Curated guidance on tools. Practitioners do not want to navigate the AI landscape alone. A practical, regularly updated guide; honest about limitations and calibrated to LMIC contexts; would be far more useful than abstract frameworks.
-
Shared examples of AI funding policies. Some donors are hesitant to fund AI work due to concerns about student harm or uncertainty about evaluation. Organisations would benefit from seeing how other donors and institutions have developed AI funding policies; both to build internal confidence and to support conversations with their own funders.
-
A community for sharing risks. When practitioners share honest accounts of what is not working; not just what is; the whole community learns faster. A sustained, safe space for sharing experiences, failures, and questions is as valuable as any formal training.
- A living resource microsite. GSF announced plans to develop a microsite consolidating resources, recommendations, funding opportunities, and session summaries from all discovery roundtables; a practical navigation aid for a landscape that changes quickly.
Reflections: What This Session Tells Us
The most honest and useful AI use cases in assessments are currently upstream; in question design, rubric development, and feedback generation; rather than in the high-stakes, summative work that dominates public debate. Organisations that have found the most traction are those that have been precise about which assessment tasks AI can help with, rather than attempting wholesale automation.
The barriers are real and structural. Data governance, tool costs, teacher capacity, and localisation are not problems that better prompting will solve; they require institutional, policy, and community-level responses.
The question of academic integrity deserves more attention than it often receives. AI-generated student submissions are not just a technical problem; they are a signal that assessment design needs to evolve. Assessments that privilege authentic, observed, and contextualised performance are not just more resistant to gaming; they are also often better aligned with what we actually want students to learn.