The Johns Hopkins School of Education and Whiting School of Engineering co-hosted an AI Synergy Summit earlier this week, and it was a fascinating dynamic. It was some mix of youthful energy and seasoned insight given an even footing that made the event feel different. It struck a measured tone of reflective, cautious optimism. It wasn’t the hype-train of ASU+GSV, but it also wasn’t the Audrey Watters/Gary Marcus/Benjamin Riley expressway of doom (although all three were mentioned and you would do well to read their work). While I am more sympathetic to the views of the expressway over the views of the train, I came away feeling even more committed and excited by the work in front of us all.
I plan to write another post on an interesting thread that emerged in some of the discussions at the summit on Small Language Models, local deployment, distributed computing, and self-hosting. For now, I’d like to share a short talk on some of the constraints facing AI’s potential in closing the research-to-practice gap in Education.
What I am going to say here is not new. Rather, I owe a great debt to those many thinkers who came before me: Daniel Willingham, Carl Hendrick, John Dunlosky, Paul Kirschner, Robert Bjork, Patricia Alexander, William James, Jerome Bruner, Daniel Kahneman, Richard Mayer, the organization Dean’s for Impact, and many others.
The application of their ideas, with slight modification, is enough to address in some meaningful way many of the problems and challenges that come up at the intersection of AI, Teaching, and Learning.
That said, before we can begin discussing how AI may interact with the research-to-practice pipeline in Education, we must sketch the boundaries of the discussion.
For starters, these models are epistemically hollow. That is, they achieve a remarkable facsimile of knowledge—text outputs that are fluent, coherent, and even seemingly reasoned, but remain a facade. A Large Language Model (LLM) does not know anything. It is a collection of impressive statistical algorithms that process the patterns embedded in language, trained to guess the next word based on the trillions of words and word combinations it has seen. It reflects language, but it does not understand it.*
The philosopher Ludwig Wittgenstein cautioned that philosophical problems arise when language “goes on holiday.” In the discourse around AI, it seems that language has indeed packed its bags. We’re borrowing words from one language game* to play another, and this semantic problem gives rise to philosophical ones which will haunt generations far beyond our own.
Take the term “hallucination.” It evokes a mental lapse as if the AI were an individual who momentarily took leave of its senses. An AI that generates a false statement isn’t deviating, it’s doing exactly what it was built to do: predict the statistically probable sequence of words.
A hallucination, then, is a statistically plausible response that also happens to be factually incorrect.
The error here lies in our expectations, not in the system’s “mind” (it has none). By calling these errors hallucinations, we smuggle in a metaphor of human perception that obscures more than it reveals. We start asking the wrong questions.
This line of thinking begins to reveal the three great threats of AI in Education: shallow learning, misinformation, and cognitive overload.
Education is no stranger to ideas that sound right yet ultimately prove false. In fact, the field is littered with convenient myths and intuitive misunderstandings. We shouldn't be too hard on ourselves about this either. Learning is complex, and to all the parties involved, invisible. We have many surrogates for learning that help us measure and think about it aspects of it, but the change in the memory of an individual is, by its very nature, a deeply internal process, and our instruments for measuring it—assessments, classroom observation checklists, and even brain imaging—are very blunt indeed.
What we call "Education" then arises from a panoply of effects. The rigid, dogmatic application of a single method, framework, or technique is misguided and almost always results in diminishing returns. If you try to use inquiry-based learning every day, you run risks. Students may not have the necessary prior knowledge to engage in the inquiry and can thus form misconceptions and misunderstandings that are difficult to get out of. If you try to use game-based learning all of the time, you are substituting a kind of intrinsic motivation for the extrinsic engagement of the game.
All this to say, every method and technique has its affordances and boundary conditions depending on contextual factors. Even still, teachers can, and do, wield a massive influence on student learning.
To provide an example of an intuitive misunderstanding, or stubborn myth in Education, let’s talk briefly about Learning Styles. For decades, teachers have been told that students learn best when taught in their preferred modality (visual, auditory, kinesthetic, etc.), leading to well-intentioned efforts to “personalize” lessons by style. It sounds wonderfully empathetic and intuitive. Yet rigorous reviews have found no evidence that matching instructional style to a learner’s supposed preference boosts learning outcomes.
Students have preferences, sure, but humans are not so neurologically pigeonholed that a “visual learner” can only learn through images. The concept of Learning Styles never needed AI to entrench itself, it spread on the power of a tidy, hopeful narrative, but one can easily imagine, and in fact test for themselves, an AI system offering “custom learning style lessons” if fed that narrative, and references to learning styles are certainlyabundant in the training data of all the foundation models. Without a grounding in research, technology ends up amplifying our fads and dampening our blind spots. And again, in Education, there are many of those.
Let's take one more. Bloom’s Taxonomy, originally a simple heuristic for thinking about the complexity of cognition and knowledge, morphed in practice to a rigid hierarchy. A Pyramid (or ladder or steps or arrow or ice berg, depending on the visualization) to scale from mere facts at the bottom to the exalted realm of creativity at the top.* Teachers were admonished to always push students toward analysis, evaluation, and synthesis, often at the expense of foundational knowledge. The result is an unintended denigration of factual knowledge, as though one could analyze or create in a vacuum. Yet, analysis and critical thinking, for that matter, are not generic skills standing alone in a vacuum. They’re an emergent property of deep domain knowledge.
Higher-order thinking requires a bedrock of factual knowledge. In the pragmatic words of a popular adage: “You can’t connect the dots if you don’t have any dots to connect.” Facts and concepts stored in long-term memory are the very dots that insight connects.
Carl Hendrick, in a recent Substack post, put this in stark language worth repeating here: "If students [and I would add teachers] rely on AI to generate ideas, structure arguments, or retrieve facts, they run the risk of cosplaying domain knowledge while bypassing the effortful processes that lead to genuine learning. The result is a kind of intellectual outsourcing that short-circuits schema construction and leaves students with shallow fluency but fragile understanding."
The question is, how do we avoid a future in which students and teachers become only prompters and collectors of AI outputs without the requisite disciplinary thinking to make sense of the AI outputs? In short, and I stress this is the nothing new part, AI tools should support retrieval practice, elaboration, and spaced repetition, to make actionable the conception of a mental latticework in which recall, explanation, application, and creation intertwine and overlap.
AI, as we've established, doesn’t think, but it does pattern-match with extraordinary scale and speed. This affordance makes it well-suited to generate varied practice problems, draft examples, surface analogies, give step-by-step instructions, and provide feedback on well-defined tasks. I again agree with and would like to echo Carl Hendrick when he writes, "AI's genuine educational potential lies not in mimicking human cognition but in amplifying distinctly human learning processes."
Heuristics and rules of thumb, like interleave concepts or space practice and knowledge retrieval opportunities, can be effective tools, but they’re also difficult in their vagueness. Interleave in what instructional context? Spaced practice for what content domain? Retrieval with what form of knowledge (declarative, procedural, conceptual)? When should I interleave X concept? What about in Y circumstances?
The effectiveness of these methods hinge on timing, sequencing, and individual adaptation. All tasks that are cognitively demanding and logistically difficult to manage at scale in real classrooms. An AI, however, may hold potential in its ability to dynamically orchestrate instruction. Not in terms of total automation and the removal of the teacher, but in very targeted terms, guided by the teacher. Small-scale interventions like adjusting spacing intervals, question sequencing, or concept mixing in response to each learner’s progress or the teacher’s need. In algebra, it may be the creation of varied problem sets employed in dynamic spaced intervals to help consolidate procedural fluency. In history, it might be elaborative interrogation questions to deepen causal reasoning.
The challenge ahead is one of scaffolding understanding. That work will not be led by engineers alone, nor by educators alone, but by a new kind of partnership that respects the affordances and limits of both human and machine.
To return to my use of the term language game at the beginning of this talk, if we think of Education itself as a kind of language game—shifting between instruction, exploration, feedback, and performance—then the key design challenge is ensuring that AI participates appropriately. It must follow the rules, even if it doesn’t know the game.
*Possible critique. This invites a deeper question: What counts as understanding in pedagogical contexts?
*Wikipedia put it well: “The central component of language games is that they are uses of language, and language is used in multifarious ways. For example, in one language-game, a word might be used to stand for (or refer to) an object, but in another the same word might be used for giving orders, or for asking questions, and so on.”
*Possible critique. Not all practitioners misuse Bloom’s. In fact, Anderson and Krathwohl’s revision aimed to resolve exactly this.
