Reviewer Tooling Guide
This guide walks through the reviewer-side of ProctorSafe: how a reviewer reads a session's results, how they start and close a review, how they exclude events that the detector fired on by mistake, and what the audit trail records about the sitting. It is written for the person who has to size internal staffing — operations leads, compliance officers, DPOs — rather than for the engineers building the integration.
The page is structured in two halves. Presentation of a Session Data covers the surfaces a reviewer looks at — the session header, the trust score and breakdown, the face-likeness and voice-match trajectories, the timeline, and the exclusion interaction that lets a reviewer push back on a detector call. The reviewing process covers the surfaces a reviewer acts on — starting and closing a review, the audit trail that records what was done, the staffing numbers behind it, and the cross-links to the broader compliance conversation.
A session in ProctorSafe is one proctored attempt. The system collects a small, fixed set of signals during the attempt and turns them into a trust score (0–100) and a per-type breakdown of how the score was reached. The reviewer's job is to confirm the system got it right, or to correct it where it didn't. Everything on this page is about that confirmation step.
Table of Contents
Presentation of a Session Data
- What the reviewer sees
- Reading the trust score breakdown
- Face likeness over time
- Voice match over time
- The timeline
The reviewing process
Presentation of a Session Data
The sections in this half describe what a reviewer sees on the session detail page. The order follows a reviewer's first-pass reading path: identify the session, read the headline number, drill into the signals that drove it. Excluding events — the reviewer's main interaction with the presented data — sits in The reviewing process alongside the rest of the actions a reviewer takes, so the data half of this page is purely descriptive.
What the reviewer sees
- Application
- ACME-CERT-2026-Q4
- Total events
- 47
- Duration
- 42.5 min
- Score version
- v4 (current)Compare all
A reviewer opens the session detail page for a single session. The page has three regions: a header with the session's identifying facts and current status; a centre column with the trust score, the breakdown, the AI summary, and the timeline; and a right column with reviewer-specific controls.
The header carries the session ID, the applicationRef your integration supplied, the current status, the start and end times, the duration, the total event count, the trust score, and the trust-score version the session was scored under.
There is no candidate name, e-mail, or identifier anywhere on the page. ProctorSafe is never sent one: applicationRef is the only reference back to your own records, and what you put in it is your choice. A reviewer who needs to know who sat the exam looks it up in your system, not in ours.
Status is one of ACTIVE, INACTIVE, COMPLETED, ABORTED, or TERMINATED. A session becomes reviewable once it has finished — that is COMPLETED, ABORTED, or TERMINATED. ABORTED sessions are reviewable too: a session that dropped out mid-attempt is often exactly the one worth a look.
The trust score is a single number from 0 to 100. The colour band is the dashboard's standard one: green at 80 and above, amber between 60 and 80, red below 60.
Which scores warrant a human look is your policy, and you set it outside ProctorSafe. There is no threshold setting in the product and no automatic review queue: every finished session is reviewable, and none is marked as needing review. Teams typically pick a cut-off — 80 is a common starting point — and work the session list filtered to scores below it, or drive their own queue from the session.completed webhook. The colour bands above are a reading aid, not that cut-off.
Below the score, two further regions matter to a reviewer:
- The AI summary is a two-sentence narrative generated after the session ends — one sentence on the behavioural baseline, one justifying its recommendation — alongside a list of top risk factors and a recommended action:
APPROVE_AUTOMATICALLY,PERFORM_SPOT_CHECK,CONDUCT_INVESTIGATION, orVOID_AUTOMATICALLY. These are triage instructions and are not the same vocabulary as the four review outcomes a reviewer closes with. It is decision-support only — the reviewer's outcome is the binding one. AI Insights is off unless your tenant enables it, and can be set to run automatically on every session or only on demand. - The timeline is the raw event stream the system recorded during the attempt, in chronological order. Each event has a type, a timestamp, a severity (informational / minor / major / critical), and a one-line label. Reviewers who exclude events do it from here.
Reading the trust score breakdown
The breakdown answers the only question the trust score itself does not: which behaviour cost the most points?
A score of 62 looks identical to a score of 62 with a different cause, so the breakdown groups all events by type and shows each type's total cost as a horizontal segment on a shared scale. The bar's full length is the 100-point integrity budget, spent downwards from 100: a perfect score is an empty bar, and a score of 62 means 38 points were spent.
A badly compromised session can be penalised past 100 while the score itself floors at 0. The scale then widens beyond the budget and a marker shows where the budget ran out, so a session that overshot stays visibly distinct from one that spent its budget exactly.
The colour encodes severity, not type. There are more scoreable types than a categorical palette can separate, so two types at the same severity look similar; identity rides on segment order (which matches the row list below the bar exactly) and on the 2-pixel gap between segments. Hovering a row in the list highlights the matching segment and dims the others.
Each row carries:
- The type's display name (e.g.
Tab blur,Multiple faces,Look away). - The severity, one of
INFORMATIONAL,MINOR,MAJOR,CRITICAL, with the swatch that matches the bar. - The incident count — how many distinct episodes of this type the system recorded.
- The raw event count, shown as "from N events" when it is higher than the incident count.
- The total duration — how long across the attempt.
- The total cost — the points this type removed from the score, with the matching bar segment.
The gap between the incident count and the event count is the one to read: four incidents "from nine events" is one candidate leaving the tab four times, not nine separate violations. A type can likewise show a high incident count but a low total cost (many short blips), or a low incident count with a high total cost (one sustained episode). Those distinctions are what tell a reviewer they are looking at a one-off technical blip rather than something more deliberate.
Face likeness over time
A higher distance means the face matched the reference photo less closely. The shaded band shows the spread within each window: a narrow band that sits high is a steady difference, a wide band that swings is usually camera angle or lighting. A high reading is a prompt to review the footage, not a finding on its own. The enrollment baseline is what this session’s own reference photo measured before the exam began: readings near it are this candidate matching themselves.
| Elapsed | Median | Minimum | Maximum | Readings | Assessment |
|---|---|---|---|---|---|
| 0:00 | 0.27 | 0.21 | 0.34 | 24 | Consistent |
| 1:46 | 0.33 | 0.23 | 0.42 | 24 | Consistent |
| 3:32 | 0.32 | 0.23 | 0.41 | 24 | Consistent |
| 5:18 | 0.31 | 0.22 | 0.40 | 24 | Consistent |
| 7:04 | 0.30 | 0.22 | 0.38 | 24 | Consistent |
| 8:50 | 0.29 | 0.21 | 0.36 | 24 | Consistent |
| 10:36 | 0.28 | 0.21 | 0.35 | 24 | Consistent |
| 12:22 | 0.27 | 0.21 | 0.34 | 24 | Consistent |
| 14:08 | 0.33 | 0.23 | 0.42 | 24 | Consistent |
| 15:54 | 0.32 | 0.23 | 0.41 | 24 | Consistent |
| 17:40 | 0.31 | 0.22 | 0.40 | 24 | Consistent |
| 19:26 | 0.30 | 0.22 | 0.38 | 24 | Consistent |
| 21:12 | 0.29 | 0.21 | 0.36 | 24 | Consistent |
| 22:58 | 0.28 | 0.21 | 0.35 | 24 | Consistent |
| 24:44 | 0.27 | 0.21 | 0.34 | 24 | Consistent |
| 26:30 | 0.33 | 0.23 | 0.42 | 24 | Consistent |
| 28:16 | 0.32 | 0.23 | 0.41 | 24 | Consistent |
| 30:02 | 0.74 | 0.44 | 1.04 | 24 | Diverging |
| 31:48 | 0.30 | 0.22 | 0.38 | 24 | Consistent |
| 33:34 | 0.29 | 0.21 | 0.36 | 24 | Consistent |
| 35:20 | 0.28 | 0.21 | 0.35 | 24 | Consistent |
| 37:06 | 0.27 | 0.21 | 0.34 | 24 | Consistent |
| 38:52 | 0.33 | 0.23 | 0.42 | 24 | Consistent |
| 40:38 | 0.32 | 0.23 | 0.41 | 24 | Consistent |
A DIFFERENT_PERSON event by itself cannot tell a reviewer whether the camera saw a real face swap or just a bad frame. Both produce detections. The shape of the distance curve over the session can. A swap walks upward and stays up with a tight spread; camera and lighting noise sprays widely with no trajectory.
The face-likeness chart plots a distance reading from the reference photo, summarised over a sampling window that defaults to 30 seconds and is configurable per tenant, on an axis that starts at 0 and runs to at least 1.4 — further if the session actually read higher. The dashed line at the tenant's DIFFERENT_PERSON threshold splits the chart into a green "Consistent" band below an amber "Inconclusive" band, with a red "Diverging" band above. The chart's per-window median is the line; the min–max spread is the ribbon around it. A single wide spike with a wide ribbon reads as noise. Two minutes of tight, high readings reads as a real swap.
Two important details:
- The chart is collapsed by default. A reviewer triaging a queue only opens the trajectory for sessions where the shape of the curve is the question, not for every session that scored above 80. The summary line above the chart says "Consistent", "Inconclusive", or "Diverging" so the open-this-or-not decision is one glance.
- The bands anchor on the threshold recorded for this session, not a literal. A tenant running a non-default threshold gets honest shading; the chart does not lie about its scale.
When onboarding's enrollment self-check ran (the candidate's brief self-photo at the start of the session), the chart also shows the enrollment baseline so the trajectory can be read against this session's setup rather than against an absolute number. A clean enrollment that ended around 0.28 and a session median around 0.31 is a strong "no swap" signal even before you read the curve.
Voice match over time
A lower similarity means the microphone matched the enrolled voice less closely. The line is only drawn where the candidate was speaking — gaps are silence, not a steady match. The shaded band shows the spread within each window: a narrow band sitting low is a steady difference, a wide band that swings is usually room noise or a distant microphone. The pitch strip below flags windows where the voice pitch alone diverged. A low reading is a prompt to review the recording, not a finding on its own.
| Elapsed | Median similarity | Minimum | Maximum | Readings | Pitch divergent | Assessment |
|---|---|---|---|---|---|---|
| 0:00 | 0.76 | 0.67 | 0.85 | 24 | 0 of 24 | Consistent |
| 1:46 | 0.78 | 0.70 | 0.86 | 24 | 0 of 24 | Consistent |
| 3:32 | 0.80 | 0.73 | 0.87 | 24 | 0 of 24 | Consistent |
| 5:18 | 0.77 | 0.69 | 0.85 | 24 | 0 of 24 | Consistent |
| 7:04 | 0.79 | 0.72 | 0.86 | 24 | 0 of 24 | Consistent |
| 8:50 | 0.76 | 0.67 | 0.85 | 24 | 0 of 24 | Consistent |
| 10:36 | 0.78 | 0.70 | 0.86 | 24 | 0 of 24 | Consistent |
| 12:22 | 0.80 | 0.73 | 0.87 | 24 | 0 of 24 | Consistent |
| 14:08 | 0.77 | 0.69 | 0.85 | 24 | 0 of 24 | Consistent |
| 15:54 | 0.79 | 0.72 | 0.86 | 24 | 0 of 24 | Consistent |
| 17:40 | 0.76 | 0.67 | 0.85 | 24 | 0 of 24 | Consistent |
| 19:26 | 0.22 | 0.00 | 0.58 | 24 | 4 of 24 | Diverging |
| 21:12 | 0.80 | 0.73 | 0.87 | 24 | 0 of 24 | Consistent |
| 22:58 | 0.18 | 0.00 | 0.56 | 24 | 4 of 24 | Diverging |
| 24:44 | 0.79 | 0.72 | 0.86 | 24 | 0 of 24 | Consistent |
| 26:30 | 0.76 | 0.67 | 0.85 | 24 | 0 of 24 | Consistent |
| 28:16 | 0.78 | 0.70 | 0.86 | 24 | 0 of 24 | Consistent |
| 30:02 | 0.80 | 0.73 | 0.87 | 24 | 0 of 24 | Consistent |
| 31:48 | 0.77 | 0.69 | 0.85 | 24 | 0 of 24 | Consistent |
| 33:34 | 0.79 | 0.72 | 0.86 | 24 | 0 of 24 | Consistent |
| 35:20 | 0.76 | 0.67 | 0.85 | 24 | 0 of 24 | Consistent |
| 37:06 | 0.78 | 0.70 | 0.86 | 24 | 0 of 24 | Consistent |
| 38:52 | 0.80 | 0.73 | 0.87 | 24 | 0 of 24 | Consistent |
| 40:38 | 0.77 | 0.69 | 0.85 | 24 | 0 of 24 | Consistent |
The audio counterpart of the face-likeness chart. The voice-match chart plots how closely the audio matched the candidate's voice profile over the session, on a 0–1 scale where 1 is a perfect match. Unlike the face chart, higher is better: the green "Consistent" band is at the top, "Inconclusive" is in the middle, and "Diverging" is at the bottom.
The chart is what tells a reviewer whether a SECOND_SPEAKER_SUSPECTED event was the real thing or an environmental blip. A single dip with a wide ribbon is the audio equivalent of a noisy frame. Two adjacent minutes of tight, low readings is a different story — and the chart will show pitch divergence on the same window, which is a second independent signal that the audio was genuinely different.
The voice-match chart is also collapsed by default for the same reason. When the chart is open, hover on a point to see the window's offset into the session, the median similarity, the min–max range, the number of readings behind it, and the pitch line — how many of the window's pitch readings diverged, and by how many semitones on median when the detector measured one. Windows where pitch could not be measured say so rather than reading as healthy.
The timeline
The timeline is the raw event stream the system recorded during the attempt, in chronological order. It is the surface a reviewer works from to understand what happened on the session, in the order it happened, with section boundaries visible. The trust score says how the system scored the session; the timeline says why.
The timeline is laid out across the session's full duration and stacked in lanes:
- Sections at the top: the exam's question boundaries, drawn as labelled vertical dividers. A reviewer's mental model is "what part of the exam was the candidate working on when this event fired?"
- Violations: the detections that contribute to the trust score, in chronological order. Each event is a coloured dot, with the colour carrying the severity (Critical / Major / Minor / Informational).
- Excluded by reviewer: events a reviewer has marked as excluded. They stay on the timeline so a reviewer can see what was excluded and undo it, but they no longer cost points. This is a separate lane from Bypassed, which carries events silenced by the tenant's section-override config — only one of the two is the reviewer's decision to act on, and the lanes stay separate so that distinction is visible.
Three interactions a reviewer uses the timeline for:
- Click an event to open its detail, and — when the session was recorded and the recording is playing — to seek the scrubber to it. Recording is captured in whole frames at an interval, so the seek lands on the nearest captured frame to the event rather than on an exact instant. Sessions run with recording disabled have no footage to seek; the timeline still works, it just has nothing to drive.
- Drag across a lane to select every event in that time window for exclusion. The gesture is a marquee, not a shift-click: on a time axis the unit a reviewer thinks in is "that noisy forty seconds", which is a span rather than a row range. A selection tray appears at the bottom of the page, showing the running selection and the projected score after the exclusion, and the reviewer commits only when the projected score matches their judgement. (The list view, which is an ordered list, still supports shift-click to extend a range.)
- Click a grouped burst to select the whole burst at once. Consecutive events of the same type that fall inside that type's clean-period window group into a single visual — the window is the same one the trust score uses to decide what counts as one incident, so it varies by event type and defaults to 2 seconds. Selecting the group excludes the whole incident, not just the visible dot.
The timeline also has lanes for Tab focus, Audio, System Events, Connection Status, Network Activity, and Custom Marks. Custom Marks are your own marks, emitted by your page through the SDK's proctor.mark() — not reviewer annotations. The remaining lanes are not all review-critical; they surface for diagnostic sessions where the detector or the network was the question.
The reviewing process
The sections in this half describe what a reviewer does once the surfaces above are understood. Excluding events is the headline action, but it sits inside a wider process: a review has a lifecycle (open, work, close), every action is recorded, and the volume of those actions is what staffing models have to account for.
Starting and closing a review
A review is a bounded sitting on one session. Starting one is optional: committing any event exclusion opens a review by itself, so a reviewer is never blocked from acting and no commit escapes the record. The explicit Start review button is for reviewers who want to pin the starting score before they touch anything. Nothing else opens a review — the system never opens one on a reviewer's behalf.
When a review is open, two facts are captured automatically:
- The score at start — the trust score the session had the moment the review opened, whether the reviewer pressed Start review or the review was opened implicitly by their first exclusion commit.
- The score at close — the trust score the session has when the review closes.
The difference between the two is the net effect of the reviewer's changes, not the trust score itself. A review can close with a higher score than it opened (because the reviewer excluded events that the detector fired on by mistake) or a lower one (because the reviewer's exclusion tightened the breakdown). Both are valid.
Closing a review requires an outcome and optionally an internal note:
| Outcome | Use when |
|---|---|
Approved | Nothing in the session undermines the result. |
Flagged | Concerns remain; another reviewer should look. |
Voided | The attempt cannot stand — the issues are unambiguous enough that the candidate needs to retake. |
Inconclusive | Not enough evidence either way. |
The outcome is a verdict you record, not an action the system takes. Closing a review marks the session as reviewed and stores the outcome on the review, and it emits the review.completed webhook carrying that outcome. It does not change the session's status, void the attempt, release the applicationRef for a retake, or move the session onto any list. Voided in particular does not terminate anything — acting on a void is your system's job, driven off the webhook or off the outcome shown on the session. ProctorSafe records the human decision; enforcing it belongs where the exam result lives.
The internal note is optional, capped at 1000 characters, and stays inside ProctorSafe. It is never sent to the candidate, and it is never included in a webhook payload. It exists so that a second reviewer (or a compliance reviewer during an audit) can read why the first reviewer decided what they decided.
A review left untouched for 30 minutes is closed automatically as Inconclusive, and marked as auto-closed so it is distinguishable from a reviewer who chose that outcome. That exists so an integrator waiting on the review.completed callback is never left hanging, but it does mean a sitting interrupted for longer than half an hour needs reopening rather than resuming.
Excluding events
Noise, lighting or a third party outside the candidate’s control.
The detector is allowed to be wrong. A reviewer can mark an event or a run of events as excluded from the trust-score calculation. Excluded events stay on the timeline — they do not disappear — but they move to a separate Excluded lane and no longer cost points.
The flow is select → preview → commit, in that order:
- Select. In the timeline, the reviewer drags across a lane to sweep a time window, clicks a grouped burst to take it whole, or clicks individual events. A sticky tray at the bottom of the page shows the running selection and the score-projection as the selection changes.
- Preview. The projected score — what the trust score would be after this exclusion — updates with each selection change. Reviewers can see the effect of their change before they commit.
- Commit. A modal asks for the category and the comment. The category is one of six:
| Category | When to use it |
|---|---|
False positive | The detection fired but the footage shows no violation. |
Technical fault | A camera, microphone, or network problem caused the detection. |
Environmental | Noise, lighting, or a third party outside the candidate's control. |
Approved accommodation | The behaviour was agreed in advance (e.g. a scribe). |
Duplicate | The same incident is already counted by another event. |
Other | None of the above; the comment explains. |
The comment is mandatory and capped at 1000 characters. It is the justification a future reader — a second reviewer, a compliance officer, a candidate's appeal — will read to understand what the reviewer saw. Asking for the comment after the reviewer has seen the projected score (rather than before) is deliberate: a required comment field asked first turns into boilerplate, and the comment is the most useful artefact the system records.
When the exclusion commits, three things happen at once: the events move to the Excluded lane in the timeline, the trust score recomputes, and an audit-trail row is recorded with the before/after score, the category, the comment, and the reviewer as actor.
A reviewer can also restore excluded events — bring them back into the counted timeline. The same comment discipline applies, with a RESTORE_EVENTS action recorded in the audit trail rather than EXCLUDE_EVENTS.
The audit trail
“Candidate submitted exam.”
Every status change and every reviewer action on a session is logged. The log is append-only and cannot be edited, even by a tenant admin. It is the operational record of who did what, when, and why — and it is what a compliance reviewer reads during an audit.
The audit trail merges two kinds of row:
- Status changes — when the session moved from one status to another (e.g.
ACTIVE→COMPLETED, or an operator terminating a session). Each row carries the from-status, the to-status, the reason (when given), the actor, and the timestamp. Note these are changes to the session, which a review outcome does not cause — see above. - Scoring overrides —
EXCLUDE_EVENTSandRESTORE_EVENTSfrom a reviewer. Each row carries the action type, the category, the comment, the score before, the score after, the actor, the timestamp, and the events touched. Collapsed, a row names the event types and the score movement; expanded, it shows the comment and lists the individual events by clock time, each one a link that jumps the timeline to it. A row that reverses an earlier exclusion says so.
Rows are newest first. The default filter is All actions; reviewers can narrow to Scoring overrides or Status changes when they want one stream, and to a single reviewer when more than one person has touched the session. Scoring-override rows are expandable — individually, or all at once — so the comment text and the event list do not dominate the list by default.
Note that the review sitting itself — who opened it, when it closed, and with what outcome — is shown by the review controls beside the audit trail, not as a row within it. The audit trail is the record of changes made; the review is the container they were made in.
Review activity is also pushed to your systems via the ProctorSafe webhook system rather than requiring you to poll:
review.started— a sitting opened, with the score it opened at.review.completed— a sitting closed, with the outcome, the score at open and close, the categories used, and the net set of events excluded and restored. This is the one to build on: it fires once per sitting, whether the reviewer closed it or the idle timer did.score.updated— fires once per individual reviewer commit. It is not in the default subscription, deliberately: a reviewer working through a session in five selections produces five callbacks, none of which is the final answer. Subscribe to it only if you genuinely want the intermediate deltas.
The comment on an exclusion and the note on a review are never carried in any of these payloads; the category is.
Staffing cheat-sheet
The numbers below are informal estimates offered as a starting point for a first sizing conversation. They are not measured from your data, not benchmarks, and not commitments — ProctorSafe does not instrument how long a reviewer spends on a session, so nothing here is something we can verify on your behalf. Replace them with your own figures after your first sitting.
Time per sitting. A reviewer who is approving a clean session spends about 2–3 minutes on it: glance at the trust score, scan the breakdown, confirm the AI summary, close with Approved. A reviewer who is flagging spends about 6–10 minutes: read the timeline, decide whether the issues warrant exclusion, often exclude one or two events, then close with Flagged and a comment. A reviewer who is voiding spends about 10–15 minutes: the void requires confidence, so reviewers typically re-read the full audit trail and the AI summary before committing. Inconclusive sits at the high end of the flagging range because the reviewer has to write down what they could not decide.
Reviewer-to-candidate ratio. This has no clean answer, because it depends almost entirely on where you draw the line for what gets looked at — and that line is your policy, not a product setting. A team that reviews everything scoring below 80 will see several times the queue of one that reviews below 60, on identical exams. Most of the variance is cut-off-driven rather than volume-driven, so the honest advice is to run your first sitting with a deliberately low cut-off, measure what fraction of sessions it produces and how many of those turned out to be worth reviewing, and size from your own numbers.
Authentication and audit binding. Every reviewer action is bound to a real authenticated user account, and a tenant cannot disable the audit trail. The one action not attributable to a person is the automatic close of an abandoned sitting, which is marked as auto-closed and attributed to whoever opened it. Deleting a staff account does not orphan their history: the name and e-mail are snapshotted onto each row at the time it is written, so the trail stays attributed. The integration you set up determines who counts as a reviewer (your staff), not whether every action is logged — the logging is on by default and cannot be turned off at the tenant level.
Working hours. Review does not have to be real-time. Sessions become eligible the moment they finish — COMPLETED, ABORTED, or TERMINATED — and the audit trail records when the reviewer opened and closed the sitting, not when the session ended. A team that reviews in batches once or twice a day is functionally equivalent to one that reviews continuously, with a backlog that clears overnight. The one thing that does not survive a long break is an open sitting: leave one idle for 30 minutes and it closes itself as Inconclusive, so batch reviewers should close each session before stepping away rather than leaving several open at once.
Where this fits
This page covers the operational view of human review: what a reviewer sees, what they can do, and what the system records. The legal and ethical framing — the requirement that a human must meaningfully oversee an automated decision, the conditions under which ProctorSafe qualifies, and the limits of what a human reviewer can do — lives in two adjacent articles:
- EU AI Act compliance for online proctoring — what the regulation asks of a deployer and how ProctorSafe's architecture answers it.
- Human-in-the-loop: bias, fairness, and the role of the reviewer — the ethics of the reviewer sitting itself: what training a reviewer needs, what second-reviewer structures look like, and what ProctorSafe's audit trail is and is not designed to support.