You're viewing documentation for a non-production ProctorSafe environment. The URLs below reflect this environment — in production, use https://www.proctorsafe.eu.

Reviewer Tooling Guide

This guide walks through the reviewer-side of ProctorSafe: how a reviewer reads a session's results, how they start and close a review, how they exclude events that the detector fired on by mistake, and what the audit trail records about the sitting. It is written for the person who has to size internal staffing — operations leads, compliance officers, DPOs — rather than for the engineers building the integration.

The page is structured in two halves. Presentation of a Session Data covers the surfaces a reviewer looks at — the session header, the trust score and breakdown, the face-likeness and voice-match trajectories, the timeline, and the exclusion interaction that lets a reviewer push back on a detector call. The reviewing process covers the surfaces a reviewer acts on — starting and closing a review, the audit trail that records what was done, the staffing numbers behind it, and the cross-links to the broader compliance conversation.

A session in ProctorSafe is one proctored attempt. The system collects a small, fixed set of signals during the attempt and turns them into a trust score (0–100) and a per-type breakdown of how the score was reached. The reviewer's job is to confirm the system got it right, or to correct it where it didn't. Everything on this page is about that confirmation step.

Table of Contents

Presentation of a Session Data

  1. What the reviewer sees
  2. Reading the trust score breakdown
  3. Face likeness over time
  4. Voice match over time
  5. The timeline

The reviewing process

  1. Starting and closing a review
  2. Excluding events
  3. The audit trail
  4. Staffing cheat-sheet
  5. Where this fits

Presentation of a Session Data

The sections in this half describe what a reviewer sees on the session detail page. The order follows a reviewer's first-pass reading path: identify the session, read the headline number, drill into the signals that drove it. Excluding events — the reviewer's main interaction with the presented data — sits in The reviewing process alongside the rest of the actions a reviewer takes, so the data half of this page is purely descriptive.

What the reviewer sees

Session
sess_demo_01
COMPLETED
Application
ACME-CERT-2026-Q4
Total events
47
Duration
42.5 min
Score version
v4 (current)Compare all
62
62%Trust Score
62%No exclusions applied yet

A reviewer opens the session detail page for a single session. The page has three regions: a header with the session's identifying facts and current status; a centre column with the trust score, the breakdown, the AI summary, and the timeline; and a right column with reviewer-specific controls.

The header carries the session ID, the applicationRef your integration supplied, the current status, the start and end times, the duration, the total event count, the trust score, and the trust-score version the session was scored under.

There is no candidate name, e-mail, or identifier anywhere on the page. ProctorSafe is never sent one: applicationRef is the only reference back to your own records, and what you put in it is your choice. A reviewer who needs to know who sat the exam looks it up in your system, not in ours.

Status is one of ACTIVE, INACTIVE, COMPLETED, ABORTED, or TERMINATED. A session becomes reviewable once it has finished — that is COMPLETED, ABORTED, or TERMINATED. ABORTED sessions are reviewable too: a session that dropped out mid-attempt is often exactly the one worth a look.

The trust score is a single number from 0 to 100. The colour band is the dashboard's standard one: green at 80 and above, amber between 60 and 80, red below 60.

Which scores warrant a human look is your policy, and you set it outside ProctorSafe. There is no threshold setting in the product and no automatic review queue: every finished session is reviewable, and none is marked as needing review. Teams typically pick a cut-off — 80 is a common starting point — and work the session list filtered to scores below it, or drive their own queue from the session.completed webhook. The colour bands above are a reading aid, not that cut-off.

Below the score, two further regions matter to a reviewer:

  • The AI summary is a two-sentence narrative generated after the session ends — one sentence on the behavioural baseline, one justifying its recommendation — alongside a list of top risk factors and a recommended action: APPROVE_AUTOMATICALLY, PERFORM_SPOT_CHECK, CONDUCT_INVESTIGATION, or VOID_AUTOMATICALLY. These are triage instructions and are not the same vocabulary as the four review outcomes a reviewer closes with. It is decision-support only — the reviewer's outcome is the binding one. AI Insights is off unless your tenant enables it, and can be set to run automatically on every session or only on demand.
  • The timeline is the raw event stream the system recorded during the attempt, in chronological order. Each event has a type, a timestamp, a severity (informational / minor / major / critical), and a one-line label. Reviewers who exclude events do it from here.

Reading the trust score breakdown

Trust score contribution
Each segment’s width is proportional to the points it removed from the 100-point budget. Colour encodes severity, not type.
Score version: v4 (current)
38.0 points lost across 5 violation typesRemaining 62
Tab Blurminor · 4 incidents · 168s total
13.4
Multiple Facesmajor · 2 incidents · 21s total
9.1
Look Awayminor · 6 incidents · 73s total
7.8
Copy Pasteinformational · 3 incidents · 0s total
2.7
Second Speaker Suspectedmajor · 2 incidents · 38s total
5.0

The breakdown answers the only question the trust score itself does not: which behaviour cost the most points?

A score of 62 looks identical to a score of 62 with a different cause, so the breakdown groups all events by type and shows each type's total cost as a horizontal segment on a shared scale. The bar's full length is the 100-point integrity budget, spent downwards from 100: a perfect score is an empty bar, and a score of 62 means 38 points were spent.

A badly compromised session can be penalised past 100 while the score itself floors at 0. The scale then widens beyond the budget and a marker shows where the budget ran out, so a session that overshot stays visibly distinct from one that spent its budget exactly.

The colour encodes severity, not type. There are more scoreable types than a categorical palette can separate, so two types at the same severity look similar; identity rides on segment order (which matches the row list below the bar exactly) and on the 2-pixel gap between segments. Hovering a row in the list highlights the matching segment and dims the others.

Each row carries:

  • The type's display name (e.g. Tab blur, Multiple faces, Look away).
  • The severity, one of INFORMATIONAL, MINOR, MAJOR, CRITICAL, with the swatch that matches the bar.
  • The incident count — how many distinct episodes of this type the system recorded.
  • The raw event count, shown as "from N events" when it is higher than the incident count.
  • The total duration — how long across the attempt.
  • The total cost — the points this type removed from the score, with the matching bar segment.

The gap between the incident count and the event count is the one to read: four incidents "from nine events" is one candidate leaving the tab four times, not nine separate violations. A type can likewise show a high incident count but a low total cost (many short blips), or a low incident count with a high total cost (one sustained episode). Those distinctions are what tell a reviewer they are looking at a one-off technical blip rather than something more deliberate.

Face likeness over time

Face likeness distance per 106-second window. Detector threshold 0.70. Enrollment baseline 0.28.
ElapsedMedianMinimumMaximumReadingsAssessment
0:000.270.210.3424Consistent
1:460.330.230.4224Consistent
3:320.320.230.4124Consistent
5:180.310.220.4024Consistent
7:040.300.220.3824Consistent
8:500.290.210.3624Consistent
10:360.280.210.3524Consistent
12:220.270.210.3424Consistent
14:080.330.230.4224Consistent
15:540.320.230.4124Consistent
17:400.310.220.4024Consistent
19:260.300.220.3824Consistent
21:120.290.210.3624Consistent
22:580.280.210.3524Consistent
24:440.270.210.3424Consistent
26:300.330.230.4224Consistent
28:160.320.230.4124Consistent
30:020.740.441.0424Diverging
31:480.300.220.3824Consistent
33:340.290.210.3624Consistent
35:200.280.210.3524Consistent
37:060.270.210.3424Consistent
38:520.330.230.4224Consistent
40:380.320.230.4124Consistent

A DIFFERENT_PERSON event by itself cannot tell a reviewer whether the camera saw a real face swap or just a bad frame. Both produce detections. The shape of the distance curve over the session can. A swap walks upward and stays up with a tight spread; camera and lighting noise sprays widely with no trajectory.

The face-likeness chart plots a distance reading from the reference photo, summarised over a sampling window that defaults to 30 seconds and is configurable per tenant, on an axis that starts at 0 and runs to at least 1.4 — further if the session actually read higher. The dashed line at the tenant's DIFFERENT_PERSON threshold splits the chart into a green "Consistent" band below an amber "Inconclusive" band, with a red "Diverging" band above. The chart's per-window median is the line; the min–max spread is the ribbon around it. A single wide spike with a wide ribbon reads as noise. Two minutes of tight, high readings reads as a real swap.

Two important details:

  • The chart is collapsed by default. A reviewer triaging a queue only opens the trajectory for sessions where the shape of the curve is the question, not for every session that scored above 80. The summary line above the chart says "Consistent", "Inconclusive", or "Diverging" so the open-this-or-not decision is one glance.
  • The bands anchor on the threshold recorded for this session, not a literal. A tenant running a non-default threshold gets honest shading; the chart does not lie about its scale.

When onboarding's enrollment self-check ran (the candidate's brief self-photo at the start of the session), the chart also shows the enrollment baseline so the trajectory can be read against this session's setup rather than against an absolute number. A clean enrollment that ended around 0.28 and a session median around 0.31 is a strong "no swap" signal even before you read the curve.

Voice match over time

Voice match similarity per window, covering 42m 24s of measurable speech across 24 windows. Similarity threshold 0.30, pitch threshold 4.0 semitones.
ElapsedMedian similarityMinimumMaximumReadingsPitch divergentAssessment
0:000.760.670.85240 of 24Consistent
1:460.780.700.86240 of 24Consistent
3:320.800.730.87240 of 24Consistent
5:180.770.690.85240 of 24Consistent
7:040.790.720.86240 of 24Consistent
8:500.760.670.85240 of 24Consistent
10:360.780.700.86240 of 24Consistent
12:220.800.730.87240 of 24Consistent
14:080.770.690.85240 of 24Consistent
15:540.790.720.86240 of 24Consistent
17:400.760.670.85240 of 24Consistent
19:260.220.000.58244 of 24Diverging
21:120.800.730.87240 of 24Consistent
22:580.180.000.56244 of 24Diverging
24:440.790.720.86240 of 24Consistent
26:300.760.670.85240 of 24Consistent
28:160.780.700.86240 of 24Consistent
30:020.800.730.87240 of 24Consistent
31:480.770.690.85240 of 24Consistent
33:340.790.720.86240 of 24Consistent
35:200.760.670.85240 of 24Consistent
37:060.780.700.86240 of 24Consistent
38:520.800.730.87240 of 24Consistent
40:380.770.690.85240 of 24Consistent

The audio counterpart of the face-likeness chart. The voice-match chart plots how closely the audio matched the candidate's voice profile over the session, on a 0–1 scale where 1 is a perfect match. Unlike the face chart, higher is better: the green "Consistent" band is at the top, "Inconclusive" is in the middle, and "Diverging" is at the bottom.

The chart is what tells a reviewer whether a SECOND_SPEAKER_SUSPECTED event was the real thing or an environmental blip. A single dip with a wide ribbon is the audio equivalent of a noisy frame. Two adjacent minutes of tight, low readings is a different story — and the chart will show pitch divergence on the same window, which is a second independent signal that the audio was genuinely different.

The voice-match chart is also collapsed by default for the same reason. When the chart is open, hover on a point to see the window's offset into the session, the median similarity, the min–max range, the number of readings behind it, and the pitch line — how many of the window's pitch readings diverged, and by how many semitones on median when the detector measured one. Windows where pitch could not be measured say so rather than reading as healthy.

The timeline

100%
0:00
4:15
8:30
12:45
17:01
21:16
25:31
29:47
34:02
38:17
42:33
Sections(2)
Q1-Q326m 15s
9:12:04 AM - 9:38:19 AM
Q416m 18s
9:38:19 AM - 9:54:37 AM
Violations(5)
Tab focus(2)
Audio(1)
System Events(2)
Excluded by reviewer(2)

The timeline is the raw event stream the system recorded during the attempt, in chronological order. It is the surface a reviewer works from to understand what happened on the session, in the order it happened, with section boundaries visible. The trust score says how the system scored the session; the timeline says why.

The timeline is laid out across the session's full duration and stacked in lanes:

  • Sections at the top: the exam's question boundaries, drawn as labelled vertical dividers. A reviewer's mental model is "what part of the exam was the candidate working on when this event fired?"
  • Violations: the detections that contribute to the trust score, in chronological order. Each event is a coloured dot, with the colour carrying the severity (Critical / Major / Minor / Informational).
  • Excluded by reviewer: events a reviewer has marked as excluded. They stay on the timeline so a reviewer can see what was excluded and undo it, but they no longer cost points. This is a separate lane from Bypassed, which carries events silenced by the tenant's section-override config — only one of the two is the reviewer's decision to act on, and the lanes stay separate so that distinction is visible.

Three interactions a reviewer uses the timeline for:

  • Click an event to open its detail, and — when the session was recorded and the recording is playing — to seek the scrubber to it. Recording is captured in whole frames at an interval, so the seek lands on the nearest captured frame to the event rather than on an exact instant. Sessions run with recording disabled have no footage to seek; the timeline still works, it just has nothing to drive.
  • Drag across a lane to select every event in that time window for exclusion. The gesture is a marquee, not a shift-click: on a time axis the unit a reviewer thinks in is "that noisy forty seconds", which is a span rather than a row range. A selection tray appears at the bottom of the page, showing the running selection and the projected score after the exclusion, and the reviewer commits only when the projected score matches their judgement. (The list view, which is an ordered list, still supports shift-click to extend a range.)
  • Click a grouped burst to select the whole burst at once. Consecutive events of the same type that fall inside that type's clean-period window group into a single visual — the window is the same one the trust score uses to decide what counts as one incident, so it varies by event type and defaults to 2 seconds. Selecting the group excludes the whole incident, not just the visible dot.

The timeline also has lanes for Tab focus, Audio, System Events, Connection Status, Network Activity, and Custom Marks. Custom Marks are your own marks, emitted by your page through the SDK's proctor.mark() — not reviewer annotations. The remaining lanes are not all review-critical; they surface for diagnostic sessions where the detector or the network was the question.


The reviewing process

The sections in this half describe what a reviewer does once the surfaces above are understood. Excluding events is the headline action, but it sits inside a wider process: a review has a lifecycle (open, work, close), every action is recorded, and the volume of those actions is what staffing models have to account for.

Starting and closing a review

In review
Open since 8/26/2026, 10:02:11 AM · Demo Reviewer
Score at start
62
Stays inside ProctorSafe. Never sent to the integrator.

A review is a bounded sitting on one session. Starting one is optional: committing any event exclusion opens a review by itself, so a reviewer is never blocked from acting and no commit escapes the record. The explicit Start review button is for reviewers who want to pin the starting score before they touch anything. Nothing else opens a review — the system never opens one on a reviewer's behalf.

When a review is open, two facts are captured automatically:

  • The score at start — the trust score the session had the moment the review opened, whether the reviewer pressed Start review or the review was opened implicitly by their first exclusion commit.
  • The score at close — the trust score the session has when the review closes.

The difference between the two is the net effect of the reviewer's changes, not the trust score itself. A review can close with a higher score than it opened (because the reviewer excluded events that the detector fired on by mistake) or a lower one (because the reviewer's exclusion tightened the breakdown). Both are valid.

Closing a review requires an outcome and optionally an internal note:

OutcomeUse when
ApprovedNothing in the session undermines the result.
FlaggedConcerns remain; another reviewer should look.
VoidedThe attempt cannot stand — the issues are unambiguous enough that the candidate needs to retake.
InconclusiveNot enough evidence either way.

The outcome is a verdict you record, not an action the system takes. Closing a review marks the session as reviewed and stores the outcome on the review, and it emits the review.completed webhook carrying that outcome. It does not change the session's status, void the attempt, release the applicationRef for a retake, or move the session onto any list. Voided in particular does not terminate anything — acting on a void is your system's job, driven off the webhook or off the outcome shown on the session. ProctorSafe records the human decision; enforcing it belongs where the exam result lives.

The internal note is optional, capped at 1000 characters, and stays inside ProctorSafe. It is never sent to the candidate, and it is never included in a webhook payload. It exists so that a second reviewer (or a compliance reviewer during an audit) can read why the first reviewer decided what they decided.

A review left untouched for 30 minutes is closed automatically as Inconclusive, and marked as auto-closed so it is distinguishable from a reviewer who chose that outcome. That exists so an integrator waiting on the review.completed callback is never left hanging, but it does mean a sitting interrupted for longer than half an hour needs reopening rather than resuming.

Excluding events

Step 1 · Select1 event
Second speaker09:31:47
Projected score: 67 (+5 from current)
Step 2 · Commit

Noise, lighting or a third party outside the candidate’s control.

Step 3 · Result
Event excluded
6267
The event moves to the Excluded lane on the timeline. An audit-trail row is recorded with the before/after score, category, and comment.

The detector is allowed to be wrong. A reviewer can mark an event or a run of events as excluded from the trust-score calculation. Excluded events stay on the timeline — they do not disappear — but they move to a separate Excluded lane and no longer cost points.

The flow is select → preview → commit, in that order:

  1. Select. In the timeline, the reviewer drags across a lane to sweep a time window, clicks a grouped burst to take it whole, or clicks individual events. A sticky tray at the bottom of the page shows the running selection and the score-projection as the selection changes.
  2. Preview. The projected score — what the trust score would be after this exclusion — updates with each selection change. Reviewers can see the effect of their change before they commit.
  3. Commit. A modal asks for the category and the comment. The category is one of six:
CategoryWhen to use it
False positiveThe detection fired but the footage shows no violation.
Technical faultA camera, microphone, or network problem caused the detection.
EnvironmentalNoise, lighting, or a third party outside the candidate's control.
Approved accommodationThe behaviour was agreed in advance (e.g. a scribe).
DuplicateThe same incident is already counted by another event.
OtherNone of the above; the comment explains.

The comment is mandatory and capped at 1000 characters. It is the justification a future reader — a second reviewer, a compliance officer, a candidate's appeal — will read to understand what the reviewer saw. Asking for the comment after the reviewer has seen the projected score (rather than before) is deliberate: a required comment field asked first turns into boilerplate, and the comment is the most useful artefact the system records.

When the exclusion commits, three things happen at once: the events move to the Excluded lane in the timeline, the trust score recomputes, and an audit-trail row is recorded with the before/after score, the category, the comment, and the reviewer as actor.

A reviewer can also restore excluded events — bring them back into the counted timeline. The same comment discipline applies, with a RESTORE_EVENTS action recorded in the audit trail rather than EXCLUDE_EVENTS.

The audit trail

ACTIVECOMPLETED8/26/2026, 9:54:37 AM
By Demo Reviewer

Candidate submitted exam.

INACTIVEACTIVE8/26/2026, 9:12:04 AM
By Demo Reviewer

Every status change and every reviewer action on a session is logged. The log is append-only and cannot be edited, even by a tenant admin. It is the operational record of who did what, when, and why — and it is what a compliance reviewer reads during an audit.

The audit trail merges two kinds of row:

  • Status changes — when the session moved from one status to another (e.g. ACTIVECOMPLETED, or an operator terminating a session). Each row carries the from-status, the to-status, the reason (when given), the actor, and the timestamp. Note these are changes to the session, which a review outcome does not cause — see above.
  • Scoring overridesEXCLUDE_EVENTS and RESTORE_EVENTS from a reviewer. Each row carries the action type, the category, the comment, the score before, the score after, the actor, the timestamp, and the events touched. Collapsed, a row names the event types and the score movement; expanded, it shows the comment and lists the individual events by clock time, each one a link that jumps the timeline to it. A row that reverses an earlier exclusion says so.

Rows are newest first. The default filter is All actions; reviewers can narrow to Scoring overrides or Status changes when they want one stream, and to a single reviewer when more than one person has touched the session. Scoring-override rows are expandable — individually, or all at once — so the comment text and the event list do not dominate the list by default.

Note that the review sitting itself — who opened it, when it closed, and with what outcome — is shown by the review controls beside the audit trail, not as a row within it. The audit trail is the record of changes made; the review is the container they were made in.

Review activity is also pushed to your systems via the ProctorSafe webhook system rather than requiring you to poll:

  • review.started — a sitting opened, with the score it opened at.
  • review.completed — a sitting closed, with the outcome, the score at open and close, the categories used, and the net set of events excluded and restored. This is the one to build on: it fires once per sitting, whether the reviewer closed it or the idle timer did.
  • score.updated — fires once per individual reviewer commit. It is not in the default subscription, deliberately: a reviewer working through a session in five selections produces five callbacks, none of which is the final answer. Subscribe to it only if you genuinely want the intermediate deltas.

The comment on an exclusion and the note on a review are never carried in any of these payloads; the category is.

Staffing cheat-sheet

The numbers below are informal estimates offered as a starting point for a first sizing conversation. They are not measured from your data, not benchmarks, and not commitments — ProctorSafe does not instrument how long a reviewer spends on a session, so nothing here is something we can verify on your behalf. Replace them with your own figures after your first sitting.

Time per sitting. A reviewer who is approving a clean session spends about 2–3 minutes on it: glance at the trust score, scan the breakdown, confirm the AI summary, close with Approved. A reviewer who is flagging spends about 6–10 minutes: read the timeline, decide whether the issues warrant exclusion, often exclude one or two events, then close with Flagged and a comment. A reviewer who is voiding spends about 10–15 minutes: the void requires confidence, so reviewers typically re-read the full audit trail and the AI summary before committing. Inconclusive sits at the high end of the flagging range because the reviewer has to write down what they could not decide.

Reviewer-to-candidate ratio. This has no clean answer, because it depends almost entirely on where you draw the line for what gets looked at — and that line is your policy, not a product setting. A team that reviews everything scoring below 80 will see several times the queue of one that reviews below 60, on identical exams. Most of the variance is cut-off-driven rather than volume-driven, so the honest advice is to run your first sitting with a deliberately low cut-off, measure what fraction of sessions it produces and how many of those turned out to be worth reviewing, and size from your own numbers.

Authentication and audit binding. Every reviewer action is bound to a real authenticated user account, and a tenant cannot disable the audit trail. The one action not attributable to a person is the automatic close of an abandoned sitting, which is marked as auto-closed and attributed to whoever opened it. Deleting a staff account does not orphan their history: the name and e-mail are snapshotted onto each row at the time it is written, so the trail stays attributed. The integration you set up determines who counts as a reviewer (your staff), not whether every action is logged — the logging is on by default and cannot be turned off at the tenant level.

Working hours. Review does not have to be real-time. Sessions become eligible the moment they finish — COMPLETED, ABORTED, or TERMINATED — and the audit trail records when the reviewer opened and closed the sitting, not when the session ended. A team that reviews in batches once or twice a day is functionally equivalent to one that reviews continuously, with a backlog that clears overnight. The one thing that does not survive a long break is an open sitting: leave one idle for 30 minutes and it closes itself as Inconclusive, so batch reviewers should close each session before stepping away rather than leaving several open at once.

Where this fits

This page covers the operational view of human review: what a reviewer sees, what they can do, and what the system records. The legal and ethical framing — the requirement that a human must meaningfully oversee an automated decision, the conditions under which ProctorSafe qualifies, and the limits of what a human reviewer can do — lives in two adjacent articles: