Original data on how QA candidates interview
Every page here states the row count behind it in its first paragraph. Where the sample is too small to support a claim, the page says so and publishes the method instead of a result.
Where QA candidates lose points in mock interviews
Across 498 scored mock interviews from 247 QA candidates, technical accuracy is almost never the weakest dimension. Candidates lose points on concrete examples and on depth: examples is the lowest-scoring dimension in 46 percent of reports, depth in 39 percent, technical accuracy in 3 percent. The pattern holds in every interview type with enough reports to say.
Testing a non-deterministic system: what 65 golden cases and three eval runs showed
Six findings from running 65 golden cases through a real LLM interview-scoring prompt three times, with a judge model and consistency sampling, in a public repository anyone can rerun. The model was far more deterministic than the eval budgeted for, the authored expectations were harsher than the prompt, and the judge needed the same context as the system under test.
How the interview readiness diagnostic adapts and scores
This page explains how the free interview readiness diagnostic works: 55 seeded questions across eight QA categories, ten served per session, difficulty adapting after every answer, and a difficulty-weighted score. It publishes no results yet. As of 2026-09-13 the diagnostic has 38 sessions and 24 completed reports, too few to support claims about candidates.