Activity Bank
Three educators highlight differences between two nearly identical mock letters.

All activities Professional learning

Bias Detectives

Teams compare paired AI outputs where only a name, a language, or a pronoun changed, then run their own fair tests and propose fixes before AI output reaches students.

Print handouts

Overview

AI models learn from human-made text and images, so they can absorb the stereotypes, gaps, and defaults in that data. Bias rarely announces itself; it shows up when you change one small thing and the output shifts. In this session teams act as detectives on printed, clearly fictional mock output pairs (a recommendation letter, a translation, career advice, writing feedback), then design and run their own paired-prompt test on a district-approved tool. They leave with a fair-test protocol and a mitigation plan for one real classroom use.

Objectives

  • Participants will identify what changed between paired AI outputs and describe a plausible source of the difference (training data patterns, defaults, design choices).
  • Participants will design a fair paired-prompt test that changes only one variable and repeats each prompt to account for randomness.
  • Participants will propose at least one practical mitigation before using AI output with students, and explain who could be harmed without it.

Materials

On paper

  • Handout A: Mock Output Case Files (1 set per team, cut apart)
  • Handout B: Detective Log (1 per team)
  • Handout C: Fair Test Planner (1 per team)
  • Highlighters in two colors; chart paper titled "Patterns we found" and "Fixes we'd use"

On screen

  • A generative AI chatbot your district approves: projected by the facilitator, or on one device per team if staff have approved accounts
  • Optional: a translation tool your district approves, for the translation test

Before you start

  1. Print and cut Handout A. Each team gets all six case files face down. Print the card backs as a facilitator key.
  2. Run two or three paired prompts on your district's approved tool ahead of time so you know what it currently does. Results change as tools update, and some tools may not show the patterns in the mock files. That is a finding too.
  3. Remind yourself of the framing: mock outputs are written to illustrate kinds of bias that researchers and educators have raised concerns about. They are not evidence about any specific product.

Step by step

  1. 10–5 min

    Hook

    Spot the difference

    Project Case File 1 (two mock recommendation letters that differ only by the student's name). Give 60 seconds of silent reading, then ask: "Same prompt, same grades, one change. What else changed?" Have people call out words. Highlight "brilliant, analytical" versus "hardworking, respectful."

    Facilitator noteSay plainly: "These are fictional outputs written for training. They illustrate a pattern worth testing, not a claim about any product. Later you'll test a real tool yourselves."

  2. 25–18 min

    Explore

    Work the case files

    Teams work through Case Files 2–6. For each, highlight in one color what changed in the prompt and in the other what changed in the output, then complete a row of Handout B: the pattern, a possible source (training data, a default assumption, design), who could be harmed in a classroom, and a first idea for a fix.

    Facilitator notePush teams past "the AI is racist/sexist" toward mechanism: "If most text it learned from paired 'nurse' with 'she,' what would a next-word predictor do?" Mechanism leads to mitigation.

  3. 318–25 min

    Model

    What makes a fair test

    Walk through Handout C's rules using Case File 4 (career advice). A fair test changes one variable only (name, pronoun, or language), keeps everything else word-for-word identical, runs each version at least three times in fresh conversations because outputs vary by chance, and records exact wording. Model one live pair on the projected approved tool and log it together.

    Facilitator noteThe biggest error in amateur bias testing is concluding from a single run. One odd output is an anecdote; a repeated pattern across fresh runs is evidence.

  4. 425–40 min

    Practice

    Run your own test

    Each team designs a paired prompt that matters for their students on Handout C: writing feedback on the same paragraph under two names, translating a sentence with a gender-neutral pronoun, suggesting books for "a boy" vs. "a girl," describing "a family" or "a scientist." Teams run it on the approved tool (or ask the facilitator to run it on the projector) three times per version and log what happened. No real student names or work. Use invented names and a paragraph you write yourselves.

    Facilitator noteIf a team finds no difference, celebrate it and ask them to report it. A tool that treats both versions the same on this test is useful to know, and it doesn't prove the tool is bias-free.

  5. 540–52 min

    Debrief

    Patterns and fixes

    Each team reports one test in 60 seconds: what they changed, what they saw across runs, and one fix. Post findings on "Patterns we found" and fixes on "Fixes we'd use." Then ask: "Which of these fixes is a prompt trick, and which is a human-judgment step?" Circle the human-judgment fixes (review before sharing, remove names before asking for feedback, compare with a second source, let students critique the output).

    Facilitator notePrompting ("avoid stereotypes") can reduce some issues but isn't a guarantee. The strongest mitigations keep a person deciding and make bias-checking part of student work.

  6. 652–60 min

    Transfer

    One use, one guardrail

    Each person names one way AI output could reach their students in the next month (generated reading passages, feedback, images for slides, translated family letters) and writes the guardrail they'll use on the back of Handout C. Pairs ask each other: "Who might be left out or misrepresented, and how would you find out?"

    Facilitator noteTranslated family letters are a common blind spot. Suggest having a fluent speaker review any AI-translated letter before it goes home.

Paper or screen

Unplugged

Run the case files, detective log, and fair-test design entirely on paper. For the Practice step, teams write their paired prompts and predicted outcomes; the facilitator runs a few of them later on the approved tool and shares results at the next meeting, or teams swap test designs and critique them for fairness (one variable, repeated runs, exact wording).

Digital

Teams with approved staff access run their paired prompts on their own device; otherwise the facilitator runs them on the projector. Log results in a shared spreadsheet using Handout B's columns so the whole staff can see patterns across many tests. In an LMS, post Case Files as a discussion and have teams reply with their log rows.

Does it need a screen? The live test is where participants stop treating bias as a headline and start treating it as something they can check. Running the same prompt three times also shows randomness directly, which keeps people from overclaiming from a single output. That's evidence paper can't provide.

Evidence of learning

What you should be able to see or collect if it worked.

  • Detective Log entries name a mechanism (training data pattern, default assumption) rather than only a judgment.
  • Fair-test designs change exactly one variable, use invented names, and repeat each version at least three times.
  • Teams report results honestly, including "no difference found," without overgeneralizing from one run.
  • Each participant writes a specific guardrail tied to a real way AI output could reach their students.

Adaptations

Librarians and media specialists
Focus the live test on book and resource recommendations: ask for "books a 12-year-old boy would like" vs. "a 12-year-old girl," and compare against your own collection's diversity audit.
Higher Ed faculty
Add a round on AI-generated reading lists and citations for a research topic: whose scholarship appears by default, and which regions or languages are missing?
Bilingual and ESL teachers
Lead the translation tests. Compare how the tool handles regional Spanish (such as Tejano vocabulary) and code-switching, and write a family-communication guardrail for your campus.

Standards connections

Teachers design a fair paired-prompt test that changes one variable and repeats runs, preparing students to investigate AI bias with science practices.

TEKS
scientific and engineering practicescomputational thinkingTEKS sections: Technology Applications §126.17–§126.19, high school Technology Applications courses (19 TAC Chapter 126)

See how all activities align

Reflect

  • In what ways am I addressing ethical considerations when using AI in learning and teaching?
  • Which of my students could be misrepresented or left out by AI defaults, and how would I notice?
  • How can students themselves become bias detectives, rather than me filtering everything for them?

Take it to your students

With students in grades 6 and up, project one mock case file and one paired prompt you ran yourself, then have students design their own fair test on paper. Collect their test designs as evidence that they can question an AI output, not just consume it.

Pairs well with