AI in Lesson Planning: Helpful, But Can't Replace Teachers

Higher Education Press

A narrative review of 53 studies finds that while LLMs can generate well-structured lesson plans and save teachers time, their output often lacks the contextual awareness, differentiation, and pedagogical depth needed for direct classroom use.

As generative AI tools such as ChatGPT, Gemini, and Claude become increasingly prevalent in educational settings, teachers are experimenting with using them to draft lesson plans. But how good are these AI-generated plans, really? A new narrative review published in Frontiers of Digital Education offers a systematic synthesis of the current evidence, concluding that large language models (LLMs) can serve as useful support tools for lesson planning—but are not yet capable of replacing the teacher's pedagogical role.

The study , conducted by Vassilis A. Failadis, Sotiris K. Tasoulis, Spiros V. Georgakopoulos, and Vassilis P. Plagianakos from the University of Thessaly, Greece, was published on September 11, 2026.

A Focused Look at an Underexplored Area

While much existing literature on AI in education takes a broad approach, this review narrows its scope specifically to lesson planning—an essential part of a teacher's daily work and a core element of professional preparation. Following the search–appraisal–synthesis–analysis (SALSA) framework, the researchers analyzed 53 recent publications, organizing their findings around four main themes: quality and pedagogical value, challenges and limitations, teachers' attitudes and perceptions, and improvements and future directions.

Well-Structured but Lacking Pedagogical Depth

Across the reviewed studies, LLM-generated lesson plans are frequently described as well-organized, with clear instructional formats. Most studies acknowledge that platforms such as ChatGPT can produce structured plans following common instructional models—for example, organizing activities into brainstorming, group discussion, guided exploration, presentation, and evaluation stages.

However, empirical evidence suggests that while these plans are often clearly structured, they typically provide only a general framework and require significant adaptation before they can be effectively implemented in real classrooms. Key limitations identified across the literature include:

  • Limited pedagogical depth: Teaching practices suggested by LLMs tend to promote procedural knowledge rather than conceptual understanding, offering little support for complex learning processes such as reasoning, reflection, and justification.
  • Weak differentiation: Plans typically do not include strategies for individualization or alternatives for students with different learning profiles—a limitation that is particularly pronounced in specialized contexts such as special education.
  • Unrealistic time planning: One study found that a plan designed for one hour would require at least three hours in practice, with LLMs assigning very short durations to essential tasks while giving more time to reminder stages.
  • Inconsistent output quality: High variance has been reported even within the same prompt groups, with outputs generated from identical prompts differing by at least five rubric elements.

"These findings suggest that producing LLM-generated lesson plans is not a one-step process but involves iterative refinement," the authors note. "The quality of the generated material depends partly on the clarity and specificity of the prompts."

Teachers Remain Central to the Process

The review found that while most teachers acknowledge the usefulness of AI tools in lesson planning, they simultaneously express reservations about their pedagogical adequacy. Teachers use AI to generate content, but many still prefer to work collaboratively with colleagues and emphasize the need for clear policies to regulate its use.

Some educators see LLMs as a helpful aid but do not believe they can replace teachers' pedagogical judgment. Many agree that AI can assist with the initial organization of a lesson, but stress that the final responsibility for adapting and applying the plan lies with teachers.

"There appears to be a difference in how teachers respond to AI, depending on their level of experience: Younger educators tend to be more open to using AI, while more experienced ones remain cautious," the authors observe. "This difference illustrates how teachers' perspectives are influenced by professional experience and familiarity with classroom practice."

Four Directions for Improvement

The review identifies four main directions that appear frequently in the literature for improving the reliability and pedagogical value of LLM-generated lesson plans:

  1. Teacher training and AI literacy—Educators need appropriate preparation to understand what AI can do and to develop the ability to adapt and assess the content it produces.
  2. Prompt engineering and reflective use—The pedagogical value of generated content depends directly on how clear, precise, and well-targeted the prompts are.
  3. Policy development and ethical guidelines—Clear policies are needed to ensure responsible and educationally appropriate use of AI tools.
  4. Model improvement and pedagogical alignment—AI tools need enhanced features and greater flexibility in instructional design, including guided interfaces and options for teachers to provide contextual information.

Research Gaps and Future Directions

Despite the growing number of publications on LLMs in lesson planning, the review identifies four important gaps in the literature. First, most studies remain theoretical or rely on teachers' opinions without examining what happens when these plans are actually applied in the classroom. Second, most studies are based on qualitative observations, interviews, or questionnaires, with limited direct comparisons between teacher-created and LLM-generated lesson plans. Third, the literature is geographically concentrated, with the United States and Türkiye appearing most frequently. Fourth, limited attention has been given to the processes through which teachers interact with AI systems during lesson planning.

Recent empirical evidence also points to a difference between perceived and actual quality. One study found that although teacher-created lesson plans were rated higher in quality overall, teachers were unable to reliably distinguish between LLM-generated and teacher-created plans. Another found that refined LLM-generated mathematics lesson plans received high evaluation scores and, in several cases, outperformed teacher-created lesson plans—although teacher-generated lesson procedures were often more closely consistent with classroom practice.

A Tool for Support, Not Substitution

The review concludes that LLMs can serve as useful tools for lesson planning, but they are not yet capable of replacing the teacher's pedagogical role. While LLM-generated lesson plans are usually well-structured, organized, and goal-oriented, they tend to fall short in adaptation, differentiation, and pedagogy.

"LLMs should currently be viewed as support for lesson planning rather than as ready-made planning tools and should not yet be considered a substitute for pedagogical expertise," the authors write. "Their use makes sense when they function as support tools and not as substitutes for pedagogical judgment."

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.