AI-simulated exams · certification practice · published research

An exam is a room,
not a quiz.
We simulate the room.

We build AI that simulates real examinations — virtual patients who never see the mark scheme, virtual examiners whose questions are chosen by code, verdicts you can audit line by line. And we publish what we measure.

For most candidates, the first exam that behaves like the real one is the real one — and getting it wrong has a price list.

Research
#001 published
Products
4 live, 1 in build
Specialisms
CAT · adaptive · AI examiners
Founded
By product people
01 — Flagship

The exam room, simulated —
with marking you can audit.

Station · Abdominal pain07:42
Patient

“It started last night, low on the right. Walking makes it worse.”

Candidate

“Does the pain move anywhere? Any fever or vomiting?”

Examiner · question chosen by code

“What is your differential diagnosis?”

mark scheme — withheld from patientverdict — computed in codeIllustrative excerpt

Our flagship work is a simulated clinical examination: an AI patient and an AI examiner in one room, on a clock, while the candidate types. One model call plays both roles.

The engineering that matters is what each persona is not allowed to know — three design decisions the conversation never reveals.

01

One call, two personas

The patient and the examiner are two roles inside a single model call. Faster — 0.7s mean turn latency, measured — and the two can never drift out of sync with each other.

02

The patient never sees the mark scheme

Criteria, essential items and red flags are not parameters of the actor prompt. There is nothing to forget to strip out — and no way for the patient to steer the candidate around a failing point.

03

The model judges. The code decides.

Pass/fail is arithmetic over individual judgements, not a model’s opinion of how the conversation went. Every mark can be audited line by line, and the marking’s stability is measured, not assumed.

30
Published stations
609
Marking criteria
209
Red flags
0.7s
Mean turn latency
03 — Products

Five products.
Five different design choices.

A home health aide, a nurse registering with the NMC, a physician entering the Canadian match, a Malaysian SPM student and a CDL career-changer each need the test to behave differently. These are five deliberately different systems — not one engine in five skins.

04 — Specialisms

The craft beneath
the products.

Five areas where the work goes deeper than a screenshot can show.

01

AI clinical examinations you can audit.

A simulated patient and a simulated examiner in one room, on a clock, while the candidate types. The interesting engineering isn’t the conversation — it’s that the patient is never shown the marking criteria, the examiner’s questions are chosen by code rather than by the model, and the pass/fail verdict is arithmetic over individual judgements instead of a model’s opinion of how it went.

Proof — NMCMATE’s OSCE simulator: 30 published stations, 609 criteria, 209 red flags. Marking stability measured, not assumed. Read Research #001

02

Computer-adaptive testing engines.

The engine picks the next question based on what the candidate just answered. Done well, it shortens the exam without losing accuracy — and the routing logic itself is something we can explain in plain English.

Proof — Practice & mock exam platform for AMC candidates in Australia, with a CAT engine choosing the next question.

03

Adaptive practice tuned to small-sample reality.

Most practice sessions are 5–50 questions — too few for the heavy statistical models most platforms claim. We use threshold-based methods that answer "you’re at 62% on air brakes" in plain English.

Proof — QuizSprint and PassHHA both run on this.

04

Real exam blueprints, to the percentage point.

If practice doesn’t mirror the actual exam’s category mix, candidates walk in surprised. Our PassHHA mock exam matches the HHA blueprint (Personal Care 24%, Safety 20%, …) exactly, and NMCMATE’s papers are generated from the NMC’s own published test specification.

Proof — PassHHA and NMCMATE.

05

Accessibility for learners who aren’t the default user.

Low-vision support, simplified interaction patterns, alternative answer flows. Practice systems usually optimise for the median candidate; we’ve built for the ones the median ignores.

05 — Contact

We also build for education
& certification businesses.

If you run a testing or training operation and want systems like the ones above — tell us what you’re running today and what’s breaking. Honest read within two business days, including the case where the answer is no.

hello@vibeserve.dev or talk to the founder directly: kent@vibeserve.dev