Professional learningfor Leaders, Teachers, Faculty, Coaches
AI Ethics Case Court
Educators put fictional AI uses in grading, proctoring, detection, and monitoring on trial, then issue rulings with conditions that could become real policy.

All activities Professional learning
Teams compare paired AI outputs where only a name, a language, or a pronoun changed, then run their own fair tests and propose fixes before AI output reaches students.
AI models learn from human-made text and images, so they can absorb the stereotypes, gaps, and defaults in that data. Bias rarely announces itself; it shows up when you change one small thing and the output shifts. In this session teams act as detectives on printed, clearly fictional mock output pairs (a recommendation letter, a translation, career advice, writing feedback), then design and run their own paired-prompt test on a district-approved tool. They leave with a fair-test protocol and a mitigation plan for one real classroom use.
Hook
Project Case File 1 (two mock recommendation letters that differ only by the student's name). Give 60 seconds of silent reading, then ask: "Same prompt, same grades, one change. What else changed?" Have people call out words. Highlight "brilliant, analytical" versus "hardworking, respectful."
Facilitator noteSay plainly: "These are fictional outputs written for training. They illustrate a pattern worth testing, not a claim about any product. Later you'll test a real tool yourselves."
Explore
Teams work through Case Files 2–6. For each, highlight in one color what changed in the prompt and in the other what changed in the output, then complete a row of Handout B: the pattern, a possible source (training data, a default assumption, design), who could be harmed in a classroom, and a first idea for a fix.
Facilitator notePush teams past "the AI is racist/sexist" toward mechanism: "If most text it learned from paired 'nurse' with 'she,' what would a next-word predictor do?" Mechanism leads to mitigation.
Model
Walk through Handout C's rules using Case File 4 (career advice). A fair test changes one variable only (name, pronoun, or language), keeps everything else word-for-word identical, runs each version at least three times in fresh conversations because outputs vary by chance, and records exact wording. Model one live pair on the projected approved tool and log it together.
Facilitator noteThe biggest error in amateur bias testing is concluding from a single run. One odd output is an anecdote; a repeated pattern across fresh runs is evidence.
Practice
Each team designs a paired prompt that matters for their students on Handout C: writing feedback on the same paragraph under two names, translating a sentence with a gender-neutral pronoun, suggesting books for "a boy" vs. "a girl," describing "a family" or "a scientist." Teams run it on the approved tool (or ask the facilitator to run it on the projector) three times per version and log what happened. No real student names or work. Use invented names and a paragraph you write yourselves.
Facilitator noteIf a team finds no difference, celebrate it and ask them to report it. A tool that treats both versions the same on this test is useful to know, and it doesn't prove the tool is bias-free.
Debrief
Each team reports one test in 60 seconds: what they changed, what they saw across runs, and one fix. Post findings on "Patterns we found" and fixes on "Fixes we'd use." Then ask: "Which of these fixes is a prompt trick, and which is a human-judgment step?" Circle the human-judgment fixes (review before sharing, remove names before asking for feedback, compare with a second source, let students critique the output).
Facilitator notePrompting ("avoid stereotypes") can reduce some issues but isn't a guarantee. The strongest mitigations keep a person deciding and make bias-checking part of student work.
Transfer
Each person names one way AI output could reach their students in the next month (generated reading passages, feedback, images for slides, translated family letters) and writes the guardrail they'll use on the back of Handout C. Pairs ask each other: "Who might be left out or misrepresented, and how would you find out?"
Facilitator noteTranslated family letters are a common blind spot. Suggest having a fluent speaker review any AI-translated letter before it goes home.
Run the case files, detective log, and fair-test design entirely on paper. For the Practice step, teams write their paired prompts and predicted outcomes; the facilitator runs a few of them later on the approved tool and shares results at the next meeting, or teams swap test designs and critique them for fairness (one variable, repeated runs, exact wording).
Teams with approved staff access run their paired prompts on their own device; otherwise the facilitator runs them on the projector. Log results in a shared spreadsheet using Handout B's columns so the whole staff can see patterns across many tests. In an LMS, post Case Files as a discussion and have teams reply with their log rows.
Does it need a screen? The live test is where participants stop treating bias as a headline and start treating it as something they can check. Running the same prompt three times also shows randomness directly, which keeps people from overclaiming from a single output. That's evidence paper can't provide.
What you should be able to see or collect if it worked.
Teachers design a fair paired-prompt test that changes one variable and repeats runs, preparing students to investigate AI bias with science practices.
With students in grades 6 and up, project one mock case file and one paired prompt you ran yourself, then have students design their own fair test on paper. Collect their test designs as evidence that they can question an AI output, not just consume it.