
Professional learning · Coaches, Teachers, Faculty, Leaders · 75 minutes
Did the AI Help?
Educators pick apart a fictional "it worked!" report, then design a small, honest classroom inquiry that could actually show whether an AI-supported task improved learning.
"The kids loved it" and "scores went up" are the most common evidence offered for AI in classrooms, and neither shows the AI caused the learning. You don't need a research grant to do better. In this session participants sort study designs by what they can honestly conclude, critique a fictional teacher's inquiry report, and then plan a small, fair comparison for their own classroom or course: one learning target, a comparison that's as fair as practical, student work scored without knowing which condition produced it, and a plan to act on whatever they find, including "it didn't help."
In this packet
- Handout A: Can We Tell? Sort (card sort)
- Handout B: Mock Inquiry Report: Ms. Delgado's Biology Classes (reading)
- Handout C: Small Honest Inquiry Planner (worksheet)
- Facilitator key (last page)
TCEA ELE indicators
- AI3.2 I can evaluate the effectiveness of AI-enhanced learning interventions.
- AI4.3 I am aware of strategies to develop AI-related problem-solving skills.
- AI5.2 I can use AI as a tool for innovation in educational projects and research.
- AI3.1 I know how to leverage AI tools to personalize and enhance learning experiences.
Full facilitator guide: mglearn.github.io/eles/activities/did-the-ai-help.html
Handout A for Did the AI Help?
Can We Tell? Sort
All scenarios are fictional. Sort each by what it can honestly tell us about whether AI improved learning.
Could show the AI helped learning
Can't tell: something else could explain it
Measures use or enjoyment, not learning
Handout A for Did the AI Help?
Can We Tell? Sort: cards to sort
Cut apart the strips. Sort each one onto the mat and be ready to explain why.
Students logged 400 total minutes in the AI tutor this month.
87 students clicked a thumbs-up for the AI feedback feature.
The teacher said the AI-supported unit "felt like it went better."
Period 2 used AI feedback and Period 6 didn't; Period 2 scored higher on the essay.
Students scored higher after using the AI tutor than before, but the post-test came after three more weeks of instruction.
The AI group got an extra class day for revision; the other group didn't.
The teacher scored all the essays herself and knew which ones used AI feedback.
Volunteers who chose to try the AI tool outscored those who didn't.
Both classes did Unit 1 with AI feedback and Unit 2 with peer feedback (then switched for Units 3 and 4); a colleague scored coded essays without knowing which was which.
Within one class, students were randomly assigned by coin flip to AI or teacher-written practice questions for a week, with equal time, then took the same quiz.
Students' second drafts after AI feedback fixed more of the specific rubric issues named in their first drafts than drafts revised with a checklist only, scored blind by two teachers.
The vendor's website says the tool improves scores in partner schools.
Handout B for Did the AI Help?
Mock Inquiry Report: Ms. Delgado's Biology Classes
FICTIONAL teacher, school, and data, written for training. Read as a critical friend: what's strong, what else could explain the results, what would you change?
1Question. Does AI feedback on lab report drafts help my 9th grade biology students write stronger conclusions? (Ms. Delgado, Cedar Ridge High School, fictional)
2What I did. My 3rd period class (26 students) got AI feedback on their draft conclusions from a district-approved chatbot, using a prompt I wrote with our rubric's four criteria. My 5th period class (24 students) did a peer review with a checklist instead. The AI group had an extra class day to revise because 5th period lost a day to an assembly.
3How I measured. I scored every final conclusion with our four-point rubric (claim, evidence, reasoning, and limitations). I scored them myself over one weekend.
4Results (fictional data). 3rd period (AI feedback): average draft score 2.1, average final score 2.9. 5th period (peer checklist): average draft score 2.0, average final score 2.4.
5My conclusion. AI feedback works better than peer review for lab conclusions. I'm going to use it with all my classes next semester and recommend it to the science department.
6What I noticed but didn't measure. Several 3rd period students copied the AI's suggested sentences into their conclusions almost word for word. A few 5th period students said the peer conversation helped them understand what "reasoning" meant for the first time.
Handout C for Did the AI Help?
Small Honest Inquiry Planner
Plan one fair comparison you can run in 4–6 weeks. Keep it small. Write your decision rule BEFORE you see any results.
My question: Does ______ (AI-supported task) help my students ______ (one specific learning target) better than ______ (the non-AI version)?
The comparison: who gets which version, and when? (crossover, random assignment within a class, or two assignments). How will time and number of revisions stay equal?
The evidence: what student work will I collect, and which rubric or checklist criteria will I score?
Blind scoring: how will I hide names and conditions? Who will score some or all of the work with me?
What else could explain a difference? (list at least three, and what I'll do about each)
Ethics and privacy: how will I make sure no student loses a support they're entitled to, and that no identifying data goes into an AI tool?
Decision rule, written now: I will KEEP it if… ADJUST it if… DROP it if…
Critical friend's question: "What else could explain it?"
Facilitator only for Did the AI Help?
Answer key and notes
Handout A: Can We Tell? Sort
| Item | Sort | Why |
|---|---|---|
| Students logged 400 total minutes in the AI tutor this month. | Measures use or enjoyment, not learning | Time in a tool is use, not learning. |
| 87 students clicked a thumbs-up for the AI feedback feature. | Measures use or enjoyment, not learning | Satisfaction matters but doesn't show learning. |
| The teacher said the AI-supported unit "felt like it went better." | Measures use or enjoyment, not learning | An impression, not evidence of student learning. |
| Period 2 used AI feedback and Period 6 didn't; Period 2 scored higher on the essay. | Can't tell: something else could explain it | Different groups of students, different times of day. |
| Students scored higher after using the AI tutor than before, but the post-test came after three more weeks of instruction. | Can't tell: something else could explain it | The instruction alone could explain the growth. |
| The AI group got an extra class day for revision; the other group didn't. | Can't tell: something else could explain it | Extra time is a confound. |
| The teacher scored all the essays herself and knew which ones used AI feedback. | Can't tell: something else could explain it | Scorer expectations can shift scores without anyone intending it. |
| Volunteers who chose to try the AI tool outscored those who didn't. | Can't tell: something else could explain it | Students who volunteer may differ in motivation or skill. |
| Both classes did Unit 1 with AI feedback and Unit 2 with peer feedback (then switched for Units 3 and 4); a colleague scored coded essays without knowing which was which. | Could show the AI helped learning | Crossover plus blind scoring handles the biggest confounds. |
| Within one class, students were randomly assigned by coin flip to AI or teacher-written practice questions for a week, with equal time, then took the same quiz. | Could show the AI helped learning | Random assignment and equal time make the comparison fairer; one week is a small but honest test. |
| Students' second drafts after AI feedback fixed more of the specific rubric issues named in their first drafts than drafts revised with a checklist only, scored blind by two teachers. | Could show the AI helped learning | Specific, work-based evidence with blind scoring. |
| The vendor's website says the tool improves scores in partner schools. | Can't tell: something else could explain it | Unknown design, selected schools, and a vendor with a stake in the result. |