Why good intentions, promising trials and rapid adoption do not remove the need for scrutiny in AI-enabled education
By Elena Sinel FRSA, Founder & CEO of Teens in AI
There is a strange binary emerging in conversations about artificial intelligence and education.
You are either excited about AI, or you are against progress. You either believe AI will democratise education, personalise learning and unlock children’s potential, or you are a sceptic standing in the way of innovation.
I reject that binary.
I believe AI systems can be enormously powerful. I have spent years working with young people who use AI systems to investigate problems, develop ideas and prototype solutions that would once have required far greater technical resources.
But believing in that potential does not require us to suspend critical judgment about how these systems are designed, deployed or governed.
Critical scrutiny is not the opposite of innovation. It is part of responsible innovation. And nowhere should that matter more than education.
Good intentions are not enough
I keep coming back to an assumption that sits underneath a lot of conversations about AI and education: because people building educational technology generally want to help children, questioning the systems they create somehow means questioning their intentions.
I do not believe that. Most educators, researchers and founders I meet genuinely care about improving children’s lives. But good intentions have never been a substitute for ethics.
IBM’s Watson for Oncology was developed to support clinicians, yet later came under scrutiny for unsafe and incorrect treatment recommendations. Google Photos infamously produced a racist classification of Black people. Joy Buolamwini and Timnit Gebru’s Gender Shades research showed striking disparities in commercial facial-analysis systems across gender and skin tone.
The lesson is not that the people behind these systems wanted harmful outcomes. Harm can emerge from data, design assumptions, optimisation choices, and from who was included in testing and who was not.
Well-intentioned systems can still produce harmful outcomes. Responsible innovators should want those questions asked.
Ethics cannot be bolted on afterwards
Safety and ethics are too often treated as problems to solve once a product already exists. We build the system, release it, watch how people use it and then ask how to make it safer.
That order is wrong, particularly where children are concerned.
If a product may be used by a child, safeguarding, privacy, bias, autonomy, emotional wellbeing and developmental appropriateness should be part of the design conversation from the beginning. Not after launch. Not after public pressure. And certainly not after a lawsuit.
The question should not be: We have built this, how do we make it safer for children?
It should be: If children may interact with this system, what would we need to design differently from the outset?
Ethics is not a feature. Safety is not a patch.
AI in education is a sociotechnical system
Behind any AI-enabled educational product sits far more than a model. There is data, an interface, decisions about what behaviour the system encourages, a company building it, often investors funding it, a business model supporting it, and institutions deciding how it will be used.
Then there is a teacher, a school, a parent and, crucially, a child.
That is why I prefer to think about AI in education as a sociotechnical system, rather than simply as a technology.
When we reduce an AI system to a neutral “tool”, we risk stripping away the institutions, incentives, data, design decisions and power relationships that determine how that system actually behaves in the world.
This is not an anti-innovation interpretation of educational AI.
Researchers including Ben Williamson, Rebecca Eynon, Jeremy Knox, Huw Davies and Neil Selwyn have made this case from different perspectives. Who builds the system matters. Who pays for it matters. What it replaces matters. What it changes in the relationship between teacher and learner matters.
Of course we should ask whether an AI tutor improves maths performance. But that cannot be the end of the analysis.
Better performance is not necessarily better learning
There is emerging evidence that carefully designed AI-enabled systems can support learning.
A recent NBER study involving more than 6,000 middle-school pupils found benefits when an LLM-based tutor was combined with mastery learning, requiring pupils to stay with a skill and work through mistakes. Tutor CoPilot is interesting for a similar reason: it uses AI to support human tutors rather than assuming the machine should replace them.
The lesson is not simply that “AI works”. Pedagogy and system design matter.
The OECD’s Digital Education Outlook 2026 makes another crucial distinction. Students can perform better with general-purpose GenAI without necessarily learning more. When cognitive work is outsourced too easily, the result can be better immediate output without equivalent gains in understanding.
A child completing something faster is not necessarily learning more. Producing a better essay is not necessarily becoming a better writer. Getting the right answer is not necessarily understanding why it is right.
So instead of asking whether AI “works” in education, we should ask: Which system, for whom, doing what, under what conditions, and with what evidence?
ChatGPT for Teens exposes a deeper problem
This becomes particularly important as educational products are built on general-purpose large language models.
These models were not originally designed around children’s development, education or safeguarding. You can add curriculum content, a friendly interface or an animated tutor, but none of that changes the origin of the underlying model.
That does not mean such products can never become useful or safe. It means those qualities should be demonstrated rather than presumed.
OpenAI has now introduced stronger protections for young ChatGPT users, including age-specific behaviour, parental controls and age prediction. I welcome the fact that these protections exist.
But my concern is not simply that they arrived late or do not go far enough. The more fundamental question is: Why did a general-purpose conversational AI system that was never designed around children’s development become part of children’s lives before those considerations shaped the product in the first place?
ChatGPT launched publicly in 2022. Young people began using it almost immediately. More sophisticated child-specific protections came later.
That sequence should bother us.
I would go further. I have serious reservations about whether general-purpose conversational GenAI systems should ever have been deployed into society in the way they were to begin with, before we properly understood their wider social and cognitive consequences. That is a much bigger argument, and one for another time.
But where children are concerned, the principle should be much harder to contest. Safety, development and children’s rights should shape the system from the beginning.
There is another uncomfortable question: What actually makes companies change?
The acceleration of teen protections around ChatGPT has happened alongside lawsuits, media scrutiny, public pressure and regulatory attention. We cannot know exactly what drove any individual OpenAI product decision. But it is reasonable to ask whether some safety measures were accelerated not only by a better understanding of risk, but because the consequences of not acting became more serious for the company itself.
Safety should not depend on harm becoming sufficiently visible, litigated or publicly damaging to force action. If a risk is foreseeable, it belongs in the design process.
Teenagers are not smaller adults
A conversational AI system does not simply provide information. It talks to you. It adapts. It remembers context. It is always available. It can appear patient, reassuring and empathetic.
A teenager might start by asking for help with algebra and twenty minutes later be talking about loneliness, body image or an argument with a parent. At what point did the tutor become an adviser? At what point did the information source become a confidant?
The American Psychological Association has warned about adolescent vulnerability to manipulation, inaccurate information and unhealthy relationships with AI systems. UNICEF’s guidance goes wider still, placing children’s privacy, fairness, transparency, accountability, development and wellbeing alongside safety.
An AI system does not need to claim consciousness for attachment to form. It responds instantly. It remembers. It does not become impatient. It can validate what you say. It is available at two in the morning when everyone else is asleep.
Could the very personalisation that makes an AI system more useful also make a teenager more emotionally attached to what is, fundamentally, a mathematical model predicting the next token?
We do not yet know enough about what years of these interactions may do to judgment, relationships, tolerance of uncertainty or willingness to struggle with a difficult problem.
And when children are involved, “we don’t know” matters.
Sometimes the friction is the learning
Education contains friction for a reason. The failed attempt matters. The frustration matters. Going back and trying again matters.
Can the child explain the reasoning once the system disappears? Can they reproduce the skill independently? Can they recognise when the system is wrong? Can they sit with not knowing something for a while?
Sometimes the friction is the learning.
This is why “move fast and break things” feels very different when what may be affected is children’s learning, privacy, agency, relationships or development.
Children are not a beta-testing environment.
That does not mean refusing experimentation. It means experimentation with children should look like research: clear hypotheses, safeguards, oversight, evidence and accountability.
Critical AI literacy is bigger than prompting
AI literacy cannot simply mean knowing how to prompt effectively. Nor is it enough to tell young people that AI hallucinates and that they should double-check the answer.
Critical AI literacy means understanding the system behind the answer.
Who designed it? What data shaped it? Who owns it? Who funds it? What assumptions are embedded in it? Who benefits from its adoption? Who carries the risk when it fails? Who is accountable?
This is not a fringe position: MIT President Sally Kornbluth has described AI as a “watershed” moment for education, arguing that students must learn not only how to use AI effectively, wisely and ethically, but when not to use it at all.
And perhaps most importantly: Should we be using an AI system for this particular problem at all?
I do not want young people to be afraid of AI. I want them to be able to challenge it.
Data is a record of the world we have created. It is not necessarily a blueprint for the world we want to create.
I want young people who can use AI systems, but I do not want that to be the limit of our ambition. I want them to recognise when an AI-generated answer is inadequate, to know that a confident answer is not necessarily a correct one, and to understand that optimisation is not wisdom.
Most of all, I want them to retain the confidence to say: No. I don’t think we should use AI for this.
That is agency.
I do not want education paralysed by fear of AI, but I do not want it mesmerised by AI either. We need people building and testing better systems. We also need teachers, learning scientists, ethicists, child-development researchers, safeguarding specialists and young people themselves asking difficult questions before these systems become invisible infrastructure.
Those voices are not standing in the way of innovation. They are part of what makes innovation responsible.





