Who tests the AI system before your students do?

An institutional pilot should not turn students into an unacknowledged test population. Assurance must establish what the system can safely do before its outputs affect learning or opportunity.

In today’s governance note, I want you to ask who will test an AI system under real educational conditions before students are expected to rely on it.

A supplier may demonstrate that the product works. A project team may run a short pilot. Staff may experiment with it using straightforward examples. None of this necessarily establishes that the system is suitable for your institution, your students or the decisions it may influence.

Supplier testing tells you how the product performed against the supplier’s chosen measures. Institutional assurance must tell you what happens when the system encounters your environment. That question should arise before procurement and deployment, while you can still define the evidence required for approval. Begin by naming the decision or activity the system will support. Generating optional revision questions carries a different level of consequence than identifying students considered at risk, evaluating written work, or recommending access to support.

You should then define unacceptable failures before seeing the test results. Otherwise, a promising demonstration can encourage the institution to reinterpret weaknesses as manageable after the fact. Testing must extend beyond average accuracy. Use realistic cases involving incomplete records, ambiguous language, unusual learning pathways and conflicting information. Examine whether performance changes across subjects, student groups and modes of study.

Accessibility must be part of the test environment, not a later compliance check. Does the system work with assistive technology? Can a student understand their feedback without relying on colour, audio or complex language? What happens with dictated text, speech differences or unconventional sentence structures? Does poor connectivity prevent some students from receiving the same service? A system can function technically while failing institutionally.

I would use a staged assurance process:

The four-stage assurance process. Define → Test → Decide → Monitor

Students and teaching staff should contribute to this process because they understand how the system will be encountered in practice. But participation must be bounded. Students should not discover high-consequence failures through live decisions that affect their progression.

Use synthetic, historical or carefully controlled data where possible. Begin in a sandbox. Compare the system’s outputs with existing practice, not with an imagined standard of perfection. If a live pilot is necessary, limit its scope, retain a reliable non-AI route, and ensure errors can be corrected without disadvantaging the student. Someone independent of the delivery target should own the final assurance decision. The team responsible for launching the system may be well placed to explain its benefits, but it should not be solely responsible for deciding whether its evidence is sufficient. Testing must also continue after approval. Models, interfaces, data sources and vendor configurations can change. A system that passed testing six months ago is not automatically the same system operating today. Define which changes trigger reassessment and who has the authority to suspend use.

NIST’s AI Risk Management Framework treats testing, evaluation, verification and validation as continuing activities. Its Measure guidance recommends testing systems against their intended purpose, deployment context, known risks and acceptable performance limits, including pre-deployment and post-deployment comparisons. NIST AI RMF Measure guidance

UNESCO’s guidance on generative AI in education similarly warns that educational institutions may be unprepared to validate rapidly changing tools and calls for human-centred institutional policies before adoption. UNESCO guidance for generative AI in education

A pilot should answer whether the institution is ready to proceed. It should not quietly shift the responsibility for testing the system to the students who must live with its failures.

Next
Next

Notes on Strategy, Governance and the AI Rush