The online proctoring argument is settling, and neither extreme won. Fully automated AI proctoring kept losing appeals when false positives punished anxious test takers. Pure human proctoring never scaled past a few dozen simultaneous candidates. The model that’s actually spreading, reflected in this month’s industry coverage, is hybrid: AI watches every exam and flags events, humans review the flags and make every judgment call. Roughly 78% of institutions now run some form of it for high-stakes assessments, and programs pairing the two layers report up to 60% fewer confirmed integrity incidents.

Why AI-Only Proctoring Lost
The automated-only approach had a seductive pitch: no scheduling, no proctor payroll, infinite scale. Its failure mode was equally simple. An algorithm that flags a gaze shift can’t tell cheating from a candidate glancing at their crying toddler, and when the flag itself becomes the verdict, the false positives land on the people least equipped to fight them: anxious students, disabled test takers, anyone whose test environment isn’t a silent private office.
Appeals boards noticed. So did courts and regulators in several countries. The pattern across integrity disputes has been consistent: automated evidence with no human judgment attached doesn’t hold up. An unreviewed AI accusation is a liability, not a safeguard.
Human-only proctoring failed in the opposite direction. Live proctors watching video walls miss things after twenty minutes, cost real money per session, and cap how many candidates you can examine at once. A certification body running 5,000 exams a quarter can’t staff that honestly.
What the Hybrid Split Looks Like in Practice
The working model gives each layer the job it’s good at. During the exam, AI monitors everyone simultaneously: gaze patterns, additional faces in frame, tab switches, audio anomalies, copy-paste attempts. It records flagged moments with timestamps and context. Crucially, it decides nothing.
After the exam, a human reviewer works through the flagged clips. Most get dismissed in seconds: a pet walked past, a candidate stretched, someone read the question aloud to themselves. The few that survive review become integrity cases with actual evidence attached: the clip, the timestamp, the reviewer’s reasoning. Cases built that way survive appeals, which is the entire point.
This is the architecture behind ICT Exam’s AI online exam platform: automated flagging during the session, a review queue afterward, and evidence bundled with every decision. We wrote about the integrity mechanics in more depth in our AI proctoring and exam integrity piece; the short version is that the AI’s job is attention, not judgment.

Running Hybrid Well: Four Decisions
First, classify exams by stakes. Weekly quizzes don’t need proctoring at all, and pretending they do wastes reviewer hours while irritating students. Save the full pipeline for finals, certifications, and admissions tests.
Second, tune flag sensitivity to your review capacity, not to some ideal of total coverage. If your team can review 200 flags within 48 hours, configure thresholds that produce roughly 200 flags. An unreviewed flag is worse than no flag: it’s an accusation nobody examined, sitting in a log, discoverable later.
Third, publish exactly what’s monitored before exam day. Candidates who know the system flags second voices will warn their families; candidates surprised by it generate false positives and file appeals. Transparency is cheap and it works. My own view: this step matters more than any threshold tuning, because most “integrity events” in anxious cohorts are really communication failures.
Fourth, store evidence with verdicts. Clip, timestamp, reviewer decision, in one record. When an appeal arrives eight months later, the institution that kept the bundle wins in a week; the one that kept only a “flagged: yes” boolean settles.
The market context says this gets bigger, not smaller. Online proctoring is heading toward $1.8 billion as certification and hiring assessments keep moving online. Institutions choosing platforms now should be asking one question above the feature list: when the AI flags something, who decides what happens next? If the answer isn’t “a person, with the evidence in front of them,” keep looking. The ICT Exam feature set was built around that answer.
FAQ
What is hybrid proctoring?
A two-layer model for online exam monitoring: AI observes all candidates during the exam and flags suspicious events, then human reviewers examine each flag and decide whether it’s a genuine integrity incident. The AI provides attention at scale; humans provide judgment.
Why not use AI-only proctoring?
Automated flags misread ordinary behavior, glances, background noise, nervous habits, as cheating, and those false positives disproportionately hit anxious and disabled test takers. Integrity decisions based on unreviewed AI output also fare poorly in appeals and legal challenges.
Does hybrid proctoring actually reduce cheating?
Programs pairing AI flags with human review report up to 60% fewer confirmed incidents. Part is deterrence: candidates take monitoring more seriously when they know flags are actually reviewed. Part is precision: reviewed cases stick, so consequences are real.
What should candidates be told before a proctored exam?
Exactly what is monitored (camera, audio, screen activity), what triggers flags, who reviews them, and how to appeal. Publishing this before exam day reduces anxiety-driven false positives and cuts appeal volume.
How much human review capacity do we need?
Plan for reviewing every flag within 48 hours, and tune AI sensitivity to match that capacity. A typical setup produces a handful of flags per hundred exam sessions once thresholds are calibrated to your cohort and exam format.