🤖 Artificial Intelligence (AI) in Assessment
How artificial intelligence is transforming assessment design, delivery, marking, and analysis across education and professional certification.
Why It Matters
Item development is expensive and often limited by subject-matter expert time. If AI can help produce usable first drafts, assessment organisations may expand content pipelines and reduce bottlenecks. But item quality is not just a throughput problem: weak prompts, weak source control, or weak review can create items that miss the construct, leak unintended cues, or fail defensibility tests.
Options or Comparison
A practical comparison usually sits on a spectrum rather than a binary choice: | Option | What it looks like | Main benefit | Main risk | |---|---|---|---| | Prohibit AI generation | All items written by humans only | Maximum control and simplicity | Slower development and higher SME burden | | Permit AI for drafting only | AI creates first drafts, humans approve everything | Faster production without full automation | Weak review can still let poor items through | | Integrate AI into a governe…
Risks
- Stronger enforcement without clearer recourse for candidates. - New digital dependencies that create operational or procurement risk. - AI-driven workflows adopted before their impact on fairness or validity is properly understood. - Reform that promises better integrity but does not demonstrate it. - Governance changes that make accountability less transparent, not more.
What Is Contested
What remains unsettled is how far AI can go beyond support functions without losing trust. Supplier material often frames the issue in terms of efficiency, scale, feedback quality, or headline accuracy metrics, but those claims do not by themselves settle the reliability question. The open question is whether the system holds up under the exact assessment conditions that matter to the buyer. A related tension is between automation and interpretability. The field appears to be moving towards mor…
TLDR
Assessment procurement and governance is about choosing AI-enabled assessment tools in a way that protects validity, reliability, fairness, data protection, auditability, and human accountability. The central question is not simply whether a product has AI features, but whether it is suitable for the exact assessment purpose and can be governed safely over time. Stronger guidance points towards asking for evidence before purchase, not after deployment. Vendor pages show a crowded market and a cl…
Key Concepts
- **Suitability for purpose**: whether the AI supports the specific assessment use case, rather than simply offering useful features. - **Human accountability**: keeping clear ownership for review, approval, appeal, and final decisions. - **Validation evidence**: proof that the tool performs as claimed in the relevant context. - **Local governance**: controls tailored to the service, cohort, stakes, and risk profile rather than assumed from the platform name. - **Exit and change control**: plann…
Risks
- Weak validation can lead to inappropriate use in high-stakes assessment. - Opaque systems can undermine auditability and appeal routes. - Poor procurement can create hidden data protection, accessibility, or retention problems. - Over-reliance on supplier claims can leave organisations unable to explain decisions to learners, regulators, or auditors. - Bundled AI inside core platforms can make it harder to see what is governed, what is optional, and what can be switched off. - Award or recogni…
TLDR
Automated item generation uses AI to draft, transform, or help produce assessment items, but the real question is whether it can do so without weakening validity, security, or defensibility. The strongest sources point to value in speeding up item drafting and easing subject-matter bottlenecks, provided human review remains central. The open question is not whether AI can produce item-like text, but how far that output can be trusted across subjects, levels, and assessment formats. In practice, …
TLDR
AI marking reliability is about whether automated or AI-assisted scoring can support a real assessment decision without weakening validity, consistency, fairness, transparency, or the right to challenge a mark. The central issue is not whether a system can produce a score, but whether that score is dependable in the specific qualification context where it is used. The strongest evidence points towards cautious, hybrid use with human oversight rather than opaque full automation in high-stakes set…
Vendor Landscape
The market is crowded with tools promising automated essay marking, handwritten script handling, rubric-based feedback, workflow integration, and audit support. This is a useful signal that demand is real, but the evidential weight is uneven: vendor pages describe capability, while the stronger question is whether a particular product has been independently validated in the intended assessment setting. Learnosity’s reported QWK figure, for example, is a performance claim that needs context and i…