AI clinical examinations you can audit.
A simulated patient and a simulated examiner in one room, on a clock, while the candidate types. The interesting engineering isn’t the conversation — it’s that the patient is never shown the marking criteria, the examiner’s questions are chosen by code rather than by the model, and the pass/fail verdict is arithmetic over individual judgements instead of a model’s opinion of how it went.
Proof — NMCMATE’s OSCE simulator: 30 published stations, 609 criteria, 209 red flags. Marking stability measured, not assumed. Read Research #001



