Quick answer: AI that scores an exam is high-risk under the EU AI Act, and that compliance deadline just moved from August 2026 to 2 December 2027. Two things did not move. Transparency duties started on 2 August 2026, and emotion detection on students has been banned outright since February 2025. So a smart online exam still has homework due this month.
The Digital Omnibus on AI was signed on 8 July 2026 and came into force on 27 July. It gave education technology roughly sixteen extra months on the heaviest tier of obligations. If you run assessments, that’s real relief, and I’d rather say so plainly than pretend the sky is falling. But the relief is narrower than the headlines suggested, and the parts left standing are the parts your students can actually see.

Most exam platforms sit in all three buckets at once, which is why a product-level answer is never quite right.
The bucket that never had a deadline extension
Start here, because it’s the one that carries no compliance path at all. Using AI to infer emotional state in an educational setting is a prohibited practice under Article 5. Not high-risk. Prohibited. It has been since 2 February 2025, and the Omnibus didn’t touch it.
In practice that rules out a specific and once-popular category of proctoring feature: reading a webcam feed for stress, confusion, nervousness or so-called suspicious affect, and treating the result as a signal about academic honesty. Some vendors still market this. A few have quietly renamed it “behavioural confidence scoring” and carried on.
My view, and we’ve written about this before in the context of AI proctoring and exam integrity, is that the ban did the sector a favour. Emotion inference from video was never well-evidenced. It performed differently across neurodivergent candidates and across cultures, and it produced a category of accusation that a student could not meaningfully rebut. Losing it costs less than vendors claimed.
Worth being precise about the boundary, though. Detecting that a second face appeared in frame, or that a browser lost focus, is not emotion recognition. Those are behavioural observations, and they belong in the high-risk bucket with everything else about cheating detection. It’s the inference of an internal emotional state that’s off the table.
What became due on 2 August 2026
Article 50 is short and it’s mostly about honesty. Three obligations reach assessment platforms.
First, if a student interacts with an AI system, tell them. That covers a revision chatbot, an AI tutor, an assistant embedded in the exam interface, anything conversational. The disclosure has to be clear at the point of interaction, which means interface copy rather than a clause on page nine of a policy.
Second, AI-generated content has to be disclosed as artificial. For an assessment platform that lands squarely on generated question banks, generated distractors and AI-written feedback comments. If your item bank is partly synthetic, students and reviewers should be able to tell which items those are.
Third, machine-readable marking of generated output follows on 2 December 2026. That one is an engineering task rather than a copywriting task, so four months is less generous than it sounds.
None of this is expensive. That’s exactly why leaving it undone looks bad. A regulator or a procurement officer can check your disclosure notices in a couple of minutes without ever seeing your architecture.
The delayed bucket, and the trap inside it
Annex III explicitly lists education and vocational training. AI that determines access to a programme, evaluates learning outcomes, assigns a learner to a level, or monitors for prohibited behaviour during a test all sit inside it. That means automated grading and cheating detection are high-risk, and the full obligations now bite on 2 December 2027.
The trap is treating that as sixteen months of nothing to do. High-risk compliance is mostly evidence: risk management records, data governance for training sets, technical documentation, event logging, human oversight, accuracy and bias monitoring. Almost all of it is easier to gather as you go than to reconstruct afterwards from a system that wasn’t recording it.

Logging and appeals are the two that take longest, so they’re the two to start on.
If I had to pick one thing to begin this quarter, it would be logging. Every automated grading decision should record the model version, the timestamp, the input it saw, and the score it produced. Without that, you cannot demonstrate accuracy, you cannot investigate a complaint, and you cannot show a human reviewed anything. Retrofitting it across a year of historical results is unpleasant work that nobody budgets for.
Human oversight is a process problem, not a feature
The requirement that gets underestimated most is appeals. A learner whose result was influenced by an AI system can ask for human review, and the platform has to make that meaningful.
Meaningful has three parts. There has to be a route to request it that a student can actually find. There has to be a named person with authority to change the mark, not just to explain it. And there has to be a record of what was reviewed and what was decided. Most platforms have built the first part and stopped there.
This is the argument for hybrid designs generally, which is the position we took in AI flags, humans decide. Keep the model in the business of flagging and drafting. Keep a person in the business of concluding. It’s better assessment practice on its own terms, and it happens to be the shape the regulation is pushing everyone toward anyway.
Institutions should push on this during procurement. Ask a vendor to demonstrate the override path on a live result, not to describe it on a slide. The gap between those two answers tells you most of what you need to know.
A practical order of work
Audit feature by feature rather than product by product. A single platform can hold banned features, transparency-only features and high-risk features at the same time, and a product-level compliance claim hides exactly the detail that matters.
Then handle the buckets in the order the deadlines fall. Remove anything that infers emotional state, because that is already unlawful in an education setting. Ship your disclosure notices this month. Scope machine-readable marking for December. Start logging and oversight records now so that the December 2027 file writes itself from data you already hold.
One closing thought on positioning. Assessment vendors have spent two years marketing AI capability. The next two will reward the ones who can explain their AI clearly to a student, a registrar and an auditor. That’s a different skill, and the platforms that have kept humans in the decision loop are starting from a much better place. Our own approach to ICTExam features keeps AI grading in an assist role with human sign-off, and ICTExam does not infer emotional state from candidates.
Frequently asked questions
Is AI exam grading banned in the EU?
No. Automated grading that affects a result or a qualification is classed as high-risk under Annex III, which means it is permitted with conditions. Those conditions now apply from 2 December 2027 following the Digital Omnibus, rather than from August 2026.
What exactly is banned in an education setting?
Using AI to infer emotions in educational institutions is a prohibited practice and has been since 2 February 2025. Social scoring is also prohibited. There is no compliance route for either, and neither was affected by the delay.
Do we have to tell students when AI is used in an exam?
Yes, from 2 August 2026 for anything conversational or generative. Students must be told when they are interacting with an AI system, and AI-generated content has to be disclosed as artificial. Disclosure of high-risk grading systems to affected people also comes with the wider high-risk obligations in December 2027.
Is remote proctoring high-risk under the AI Act?
AI monitoring for prohibited behaviour during a test is listed in Annex III, so yes. Biometric identification carries its own high-risk classification on top. Emotion inference within proctoring is a separate matter and is prohibited outright in education.
Does this apply to a university outside the EU?
It applies if you assess students located in the EU, regardless of where your institution or your vendor sits. Both providers and deployers carry obligations, so a university running a third-party exam platform has duties of its own rather than inheriting the vendor’s compliance wholesale.
What should we ask an exam platform vendor before buying?
Ask which features they classify as high-risk and why, whether they log model version and reviewer identity per grading decision, how a student appeals an AI-influenced mark, and whether any component infers emotional state. Ask for a demonstration rather than a description.
Related resources
- AI Flags, Humans Decide: Why Hybrid Proctoring Won the Online Exam Debate
- AI Proctoring and Exam Integrity in 2026: Signals, Fairness, and Privacy
- How a Smart Online Exam Stays Credible Without Watching Over Shoulders
- Top 10 AI Graded Exam Software Tools in 2026
- ICTExam Tenant Admin Guide
Where to go next
If you’re mapping your own assessment stack against these deadlines, the fastest useful step is a feature-level audit rather than a vendor questionnaire. See how ICTExam handles AI grading with human sign-off, or review plans and pricing if you’re comparing platforms this term. Questions about how this maps to your institution’s setup can go to service.ictvision.net.