A school district piloted an AI tutoring tool in three classrooms before deciding whether to expand it district-wide. For the pilot's small scale, the strategy team chose to track weekly usage rates and teacher-reported engagement rather than standardized test score changes. Why is this metric choice appropriate for a pilot of this size?
Select an answer to reveal the explanation.
Short Explanation
Think about judging a new gym membership by whether you've lost weight after one week versus whether you actually showed up and liked it — one needs time and a big enough sample to mean anything, the other tells you something real right away. Three classrooms over a short pilot window is too small a sample for test scores to move in a statistically meaningful way. Usage and engagement, though, are exactly the kind of signal that small scale can actually capture.
Full Explanation
Metric selection should match what a given sample size and timeframe can actually support: a three-classroom pilot lacks the statistical power and duration needed for standardized test-score changes to be meaningful, since test scores respond to many confounding factors and typically require a larger sample and longer observation window to isolate an AI tool's effect. Usage rates and teacher-reported engagement, by contrast, are directly observable at small scale and meaningfully indicate whether the tool is being adopted and valued early on. Claiming test scores are never appropriate for evaluating AI tutoring tools overstates the case — test scores become a legitimate lagging indicator once a pilot scales to a size and duration that can support statistical analysis. Treating usage and engagement as the pilot's default metrics because they're the only vendor-reported figures gets the causality backwards; the team should choose metrics based on what the pilot's scale can support, not settle for whatever the vendor happens to report. Attributing the choice to an accreditation requirement invents a compliance mandate that isn't part of the scenario; the actual reasoning is statistical and evidentiary, not regulatory. Caveat: usage and engagement metrics are a reasonable early substitute, not a permanent stand-in — a district-wide expansion should still eventually track outcome measures like test performance. Operational check: confirm the pilot's evaluation plan specifies at what future scale or duration test-score tracking should be added.