Revisiting a Taxonomy of LLMs for Education Applications

A deep dive into the 2024 LLMs-in-education taxonomy by Wang et al., this essay challenges the reductive framing of teaching and learning as automatable tasks. It asks: What pictures of education do we adopt when we let performance become the measure of understanding?
Revisiting a Taxonomy of LLMs for Education Applications

Wang et al., 2024

picture held us captive. And we could not get outside it, for it lay in our language and language seemed to repeat it to us inexorably (PI, §115).

The 2024 paper “Large Language Models for Education: A Survey and Outlook” proposes a comprehensive taxonomy to capture the breadth of LLM use in education. Like all taxonomies of such and such domain of education, it forms a particular picture of what teaching and learning are. This picture centers on performance as the benchmark for quality. Teaching and learning in this scheme consist of discrete, automatable units, including “question solving,” “error correction,” “material creation,” and so on. This captures neither the full grammar of educational activity nor the lived context in which those activities gain meaning. I do not say this to critique the work of the researchers involved, they state clearly that they are developing a technology-centric taxonomy. Rather, I mean to surface a deeper problem: that in articulating this particular taxonomy, we see a field emerging that structures our perception of educational work in ways that are shallow and epistemically flattening. The danger, in other words, is not that the taxonomy is inaccurate within its own frame, but that its frame is becoming a more dominant lens through which technology in education is interpreted.

Let’s take each domain and subdomain of the taxonomy in kind before launching into a critique.


Domain 1: Study Assisting

Subdomain 1.1: Question Solving (QS)

In this subdomain, LLMs are positioned as advanced problem solvers across various fields, including math, law, medicine, and programming. The authors also survey the techniques used to enhance LLMs’ abilities in question-solving, including Chain-of-Thought prompting, in-context learning and few-shot prompting, and multi-agent collaboration. It is in this question-solving affordance that LLMs can provide “high-quality answers to [students] blocking questions in a timely manner” (p. 2). What is the picture of study implied here? The learner seems to be reduced to a consumer of “high-quality” solutions. Does it follow that learning occurs via exposure to correct answers?

I should clarify that if we have any aspirations toward AI tutors, we should, of course, test whether they can solve the questions that the student would be working on. However, it does not follow that the necessary end state is that the AI simply provides answers. More vital, and harder to gauge, is the ability to guide a student to correct methods of interrogating and solving the problem at hand. These methods used by students for interrogating and solving problems seem to come from deliberate practice and experience gained through spaced repetition and the interleaving of new and related concepts. An interesting area of study would be how an AI system can adapt and respond to changes in a learner’s cognitive states, allowing us to optimize the spacing of practice intervals and the types of questions we interleave.

Subdomain 1.2: Error Correction (EC)

This section of the paper discusses how Large Language Models (LLMs) are being applied to correct student errors in grammar and programming. The authors state that such error correction is particularly important during the early stages of learning when timely correction is most impactful. Models like Codex and GrammarGPT are shown to outperform prior systems in error detection. Effectiveness in error correction is defined by fix-rate, patch size, and improved outputs. So, what is the implicit assumption of what learning is by LLMs as error correctors? Learning is cast as the elimination of error, and the model is the corrector. This is of vital importance as I would hazard a guess that in the provision of feedback and correction, a teacher would not view themselves primarily as a corrector, but perhaps as a coach or guide. These two roles carry with them a motivational domain not accounted for in the paper. Just how much of the role of the teacher is to perform a motivational role? That question is open to debate, and for a more thorough handling of the question, I refer you to this recent piece featured on ACX.

In short, this subdomain is primarily couched in terms of performance metrics—the number of errors fixed, patch size, and alignment with human corrections. However, from an instructional perspective, the value of error correction lies not just in correctness, but also in productive struggle (the student taking steps to correct errors), and in how we direct attention (the actionable steps for improvement). Without this pedagogical framing, there's a risk that LLM-based correction becomes mere repair, and not instruction.

Subdomain 1.3: Confusion Helper (CH)

Unlike QS and EC, this subdomain is explicitly pedagogical in nature. CH avoids solving the problem for the student and instead generates hints, guiding questions, or explanations to facilitate learning. Tellingly, this section reports mixed results:

  • LLM-generated explanations are often less effective than those provided by human tutors.

  • Synthetic hints may be too general or use vocabulary that is inaccessible to students.

  • Adaptive explanations to different student subgroups show some promise.

This subdomain is explicitly less performance-centered and recognizes that learning happens in the process of resolving confusion, not just in getting the right answer. More fittingly, the evaluative emphasis shifts toward instructional value, student comprehension, and even pedagogical fit.

Domain 2: Teach Assisting

This section surveys how LLMs support educators by automating routine tasks such as generating questions, grading assignments, and producing instructional materials. Across these domains, LLMs are framed as productivity tools that unburden teachers from repetitive work and streamline the preparation process. The underlying claim is that LLMs enable teachers to focus on the “human” aspects of instruction by outsourcing rote or clerical tasks.

Subdomain 2.1: Question Generation

The primary preoccupation of this subdomain is with LLM-generated questions. This makes sense, as questions are among the most powerful tools in a teacher’s toolbox, and having a diverse array of them at hand is generally useful. The authors survey fine-tuning and prompt-engineering methods to calibrate outputs to match textbook content, curricular goals, or instructional taxonomies. For example, Doughty et al. (2024) examine GPT-4' ’s ability to generate high-quality multiple-choice questions aligned to Python programming objectives, while Lee et al. (2024) map LLM-generated prompting questions in reading comprehension tasks.

The LLM is praised for its ability to produce balanced questions of “clear language” and “high-quality distractors,” but nowhere is there inquiry into the sequence, function, or cognitive demand of those questions within a learner’s trajectory through a course. Here again, the underlying theory of instruction appears underdeveloped. The questions discussed exist as discrete objects, detached from the lessons or learners they are meant to serve. They exist at the level of fluency, question structure, and alignment to curricular goals and objectives.

Questions for elaborative interrogation? Those are more difficult to pin down.

Subdomain 2.2: Automatic Grading

I don’t have much to say here. This subdomain addresses the use of LLMs to score student work, particularly open-ended writing and short-answer responses. Previous grading models relied heavily on similarity-based comparisons to “golden answers,” lacking sensitivity to student reasoning or alternative expressions of correctness. With the advent of LLMs, however, the models are incorporating rubrics, prompting for reasoning, and producing both scores and narrative feedback of “satisfactory” quality. Moreover, just as earlier domains treated learning as the accumulation of correct answers, this domain risks treating teaching as the accumulation of judgments. It abstracts away the relational, motivational, and developmental aspects of assessment, reducing the act to something transactional and post hoc.

Subdomain 2.3: Material Creation

In this final subdomain of Teach Assisting, LLMs are deployed to generate instructional content: asynchronous course modules, EFL materials, worked examples, and other learning resources. Several studies emphasize the importance of a “human-in-the-loop” process, combining zero-shot prompting with human review and revision. Among the subdomains, this one is the least reductive in spirit. The authors acknowledge the need for contextual relevance, adaptation to learner needs, and integration with the instructor’s existing materials and teaching style. Still, the primary value of LLM use is framed in terms of efficiency: streamlining content creation, reducing friction, and minimizing time investment by teachers.

Again, the constant appeals to efficiency remind me of the error in the study of Open Educational Resources (OER), which consistently made appeals to access and cost savings at the expense of understanding the instructional implications and outcomes of adopting OER materials. What is the instructional opportunity cost of using LLMs for material creation? Do teachers have reduced fluency and command over the materials they are teaching? You could imagine a study in which we have two groups of novice teachers: those using an AI system like SchoolAI or Magic School, and those planning with traditional materials. Then, after a planning session, we ask them to explain the lesson material to us. Who do you think would perform better at the explanation?

The hypothetical study raises a question worth considering. Is the process of materials creation an expression of teachers’ content and pedagogical knowledge and a form of practice? My hunch is that they are, and that despite the professional development offered in school, we often overlook the fact that the most effective forms of development for teachers are peer collaboration, mentorship, and the consistent practice of simply doing the job with timely feedback and actionable steps for improvement. Perhaps the administrative burdens in school are the problem, and that while what a teacher mostly does is teach, how much of the other side of the job is filled with what the anthropologist David Graber would describe as beurecratic bullshit? I leave that question for you to reflect on.

Domain 3: Adaptive Learning

Unlike previous domains, which focused primarily on task automation or material generation, Adaptive Learning enters more ambitious conceptual territory: diagnosing what students know and adapting instruction accordingly. Even here, however, the center of gravity remains squarely in the realm of performance optimization. Learning is treated as a kind of data-driven navigation. The goal is to measure, map, and match the “mastery” of “concepts” to content via a dynamic delivery system. If the picture of teaching in Domain 2 was that of a clerical worker made more efficient, the picture of learning in Domain 3 is that of a GPS unit routing a user from confusion to correctness.

Subdomain 3.1: Knowledge Tracing

Knowledge Tracing attempts to model a student’s mastery of content over time, typically by analyzing patterns of correct or incorrect responses to determine what they "know." Of the studies mentioned, Lee et al. (2024) warrants a further note in that they use LLMs to predict question difficulty based on question stems and associated concepts, allowing for better estimations of challenge levels even for unseen items. This is a clever enhancement, especially when considered alongside Sonkar and Baraniuk (2024) exploring whether LLMs can simulate incorrect student reasoning if given a learner’s “knowledge profile.” Despite the inversion and enhancement, we cannot absolve LLMs of the above critiques. Even here, we’re still tracking observable performance against a latent model of “mastery” inferred through proxy indicators.

Subdomain 3.2: Content Personalization

We finally arrive at the perennial unfulfilled promise of edtech—personalization. I have a few ideas about why personalization doesn’t seem to be a generalizable instructional intervention, but that’s for a different essay, as this one is already too long. For more on the failure and complications of “personalization,” I again refer you to the review of school on ACX for some more flavor on this conversation.

Also, as I type the phrase “more flavor,” Grammarly is telling me to change the phrasing to “additional context,” and now I’m lamenting how AI-generated text makes us more boring writers. This topic is dry enough as it is. If I want to add some flavor and taste, I should be able to without a pop-up bugging me to try the “beta” version of my sentence.

With that aside aside, let’s not dismiss the authors’ treatment of personalization out of hand. This subdomain presents the most aspirational vision in the paper: an intelligent system that not only diagnoses a learner's current level but also guides and motivates them using personalized content. But how far does this personalization go? Currently, most efforts appear to be limited to surface-level adaptations, such as swapping in keywords or redirecting learners based on concept mastery (or, to unmoor ourselves from jargon, how many correct answers the student got). The idea of personalization is reduced to tailoring content, not reshaping the path, process, or product of the learning exercise. An algebra problem mentioning LeBron James to a student who likes basketball is still an algebra problem. True personalization involves operationalizing heuristics such as spacing practice and interleaving concepts, tailored to the specific classroom context.

Critiques and Conclusions

The taxonomy also provides an overview of commercial tools, including chatbots, content creation tools, teaching aids, quiz generators, and collaboration tools. However, given the length of this post, I’ll address these in a separate post. Let’s center ourselves on what has been presented so far and consider it rationally.

The taxonomy treats LLMs as tools. Tools for solving problems, generating questions, and creating materials. However, tools gain meaning only within specific human activities. The primary difference between a griddle spatula and a paint spreader lies in their intended use. In descriptive terms, they are both flat pieces of metal affixed to a handle of some sort. The difference, again, is in their use in a particular form of life—cooking and painting. Similarly, error correction or hint generation are not simply functions; they are roles embedded in the form of life we call education. This is not to say the taxonomy is wrong on its own terms. It is technically accurate, even comprehensive, but it is grammatically confused. It takes the surface of tasks (inputs and outputs) and treats them as the essence of teaching and learning. It overlooks the deeper grammar of educational activity: why we ask questions, how we correct, when we guide, and what it means for something to be understood. Again, this is not the fault of the authors, who do a commendable job with their stated task. It is an issue of the enterprise of LLMs in Education. As such, the taxonomy risks perpetuating a picture of education as performance optimization—learning as the accumulation of correct answers, teaching as the efficient delivery of prompts and feedback. It is a picture that is hard to see beyond—a picture of correct answers without meaningful learning.

If LLMs are to become part of the educational form of life, the real challenge is not to automate what teachers do, but to preserve, and perhaps even deepen, the intentionality and human responsiveness that define teaching as a practice.


Thanks for reading. Talk to you soon!

Where Good Ideas Become Great Lessons

Start with an objective, a topic, or a source. Leave with classroom-ready materials you can teach from.