Experts call for stronger procurement of AI and better measurement of how it impacts student learning

Most artificial intelligence (AI) tools in K-12 classrooms are still judged by adoption, engagement and how much time they save educators, rather than the impact they have on teaching and learning, according to a Sept. 16 panel discussion hosted by Brisk Teaching, a platform that provides educators with AI-driven tools.

Jean-Claude Brizard, president and CEO of the nonprofit Digital Promise, highlighted the recent shift in public attitude toward AI from that of acceptance to skepticism, as many large local educational agencies in recent weeks have halted adoption of AI tools.

“The natural sort of reaction to a lack of evidence, to lack of understanding, is to ban, to slow down adoption,” Brizard said. “We still have this magical thinking of AI that needs to really go away, and we go back to the fundamentals of teaching and learning. We know how kids learn, we know what we need to do. So, the question for us basically is how technology will support that.”

Panelists discussed findings of a white paper released in August by Brisk Teaching, which provided a framework for how LEAs could utilize factors beyond time saved by teachers to measure whether AI is meaningfully improving teaching and learning in classrooms.

Drawing on insights from educators, district leaders and researchers, the document offers a vision for an evaluation framework that captures how many high-quality instructional moments AI tools enable, how often and with what relative weight — noting that “when a teacher connects an AI action to specific academic standards, that interaction earns additional weight in the scoring framework, creating an incentive for intentional, aligned use rather than AI for AI’s sake.”

A strong evaluation framework would also show if students received low-quality content or feedback from AI programs or constructive feedback that they otherwise wouldn’t have received; and measure if an AI tool is providing content that is accurate, pedagogically sound, rooted in district standards and curricula, and provides data to educators that is easy to access and understand.

Specifically, the framework provided by Brisk “maps each type of AI-enabled interaction to an evidence-based category of instructional practice and then weights it according to its relative impact on student learning,” the white paper states. For instance, direct feedback carries the highest weight, while administrative tasks that save time but don’t constitute instructional moments carry the lowest. And when a teacher “connects an AI action to specific academic standards, that interaction earns an additional standards alignment bonus, creating an incentive for intentional, curriculum-aligned use,” according to the white paper.

Arman Jaffer, founder and CEO of Brisk Teaching, noted during the webinar that while saving teachers time can be a benefit to AI education tools, that alone won’t tell the whole story.

“I think what often gets lost in the headlines is the complexity and that nuance of specifically how AI can actually impact learning,” Jaffer said. “So much of the impact of AI currently is couched in time saved, and I think efficiency is a really core component and a real huge value add, but that also doesn’t really tell the story of how students can benefit from a teacher who’s leveraging AI in the classroom. So, our hope was to add an additional measure beyond efficiency to really talk about efficacy on student outcomes.”

Panelists emphasized that LEA leaders must articulate what high-quality instructional moments look like in their context. The white paper provides a place to start.

Questions to ask of staff include:

  • What instructional practices do we most want to strengthen, and how will we know if they’re improving?
  • How do we define high-quality instructional moments in our context, and are our AI tools aligned to that definition?

Questions for technology vendors include:

  • Do you have a framework for measuring instructional impact that goes beyond engagement and efficiency metrics?
  • What does your measurement dashboard show, and how will we be able to access and act on that data?
  • Can we change how the framework measures and weights different types of instructional moments to ensure it meets our district priorities? How customizable is it?