AI in education: what generative AI already does today
Keynote, 40–60 minutes. Delivered at ICAILY 2026, Cape Town.
The urgent conversation is no longer what artificial intelligence might one day do for education. It is what these systems are already doing in real classrooms, with real teachers and real institutional consequences: where the gains are genuine, where they remain unproven and what responsible adoption actually demands.
Scaling AI in schools without buying licences
Keynote or workshop, 45–90 minutes. Delivered as the Day 1 closing keynote at the Summit de Inteligência Artificial do Brasil 2026.
Schools do not scale artificial intelligence by distributing licences. They scale it when teachers have confidence, repertoire, criteria and autonomy. Built on training 50 teachers across institutions, and on what changed when the institutional stance changed.
Trustworthy AI agents in high-stakes operations
Talk or panel, 30–45 minutes. Prepared for MiningTech South America 2026.
Tool permissions, data grounding, exception handling, audit trails and human supervision. How to tell an agent that works from one that only appears to, and what evidence to demand before an agent touches a production process.
Evaluating AI when the benchmark is lying to you
Technical talk, 30–45 minutes. Based on published research.
Hidden information channels inflate apparent performance. Drawing on a forensic study of answer leakage in agentic code repair, this talk covers restricted feedback, same-pool comparison and immutable provenance as practical controls for anyone buying or building AI systems.
Assessment after generative AI
Talk or staff training, 45–90 minutes. The argument behind the Folha de S.Paulo piece.
A detector can flag text; it cannot establish understanding, judgement, authorship or the ability to revise. What each discipline should count as evidence of learning, and how version histories, oral defence and commentary on errors make thinking visible again.
Who waits, and for how long
Policy talk, 30–45 minutes. The argument behind the GovInsider and EUobserver pieces.
When an AI service says a human reviews ambiguous cases, someone is queuing. Public institutions should measure who waits, how long and what happens meanwhile. Also covers multilingual evaluation: a system that works in English and fails in another language can still look compliant in aggregate.