Research

What we measure
when we build.

Every report here publishes first-hand measured data from systems we run in production — marking stability, model comparisons, latency, and the experiments that killed our own ideas. No opinion pieces. No industry roundups. If a number appears, we measured it.

  1. #001An AI clinical examiner you can auditOne model call plays both patient and examiner, the mark scheme stays withheld, and the verdict is computed in code — plus the measurements that killed our own ideas.September 2026 · VibeServe Research · Kent Tan · 17 min read