Educave research
For government, inspectorates and regulators

Evidence, ungated

Nothing on this page asks for an email. Public bodies can read the findings, the method and the limits, then ask us for the rest.

1
expert tutor whose practice this describes
48
recurring students
1,418
recorded one-to-one expert STEM lessons, Aug 2024 to Jun 2026, speaker-attributed
7,546
student errors classified as slip, misconception or missing prerequisite
74,258
tutor questions classified by type
308
held-out expert decision-points used as a benchmark for AI tutors

Ofsted Areas of Research Interest

Area one: gaps between advantaged and disadvantaged learners

We cannot answer this question. The corpus is fee-paying one-to-one tutoring, structurally an advantaged setting, with no pupil premium, SEND or EAL indicator linked at source.

What we do hold is a precondition finding. Trust executives responsible for around 50,000 students could not evidence centrally which of their edtech works, with data across 12 or more disconnected systems and about a term of lag before drift is visible. The question cannot currently be answered by schools themselves.

We also hold an instrument rather than a conclusion: a coding scheme with double-coded agreement of 0.95 on question type and 0.98 on scaffolding, and a held-out benchmark design. Both transfer to a disadvantaged cohort.

Area two: cognitive and socioemotional development

We hold process measures, not developmental outcomes.

These measure how problem-solving support and productive struggle are managed. We hold no measure of resilience or socioemotional outcomes.

We hold no evidence on online harms, their third area.

Limits, stated up front

Questions this evidence answers

Can a tutoring product's judgement be scored against an expert?

The corpus gives a measurable benchmark rather than an opinion. A generic AI tutor matches the expert's decision to give or withhold an answer 46% of the time; a state-aware policy derived from the corpus reaches 72%. Any product claim about tutoring quality can be scored the same way.

Human tutor intelligence

What are schools actually able to measure?

Around 45% of corrections carry a diagnosis — a wrong model or a missing prerequisite — that no management information system records. Attainment data cannot answer questions about mechanism because the mechanism is never captured.

Finding one

How do experts adapt to the learner?

Withholding an answer is conditioned on the student's state and rises as the relationship matures. Systems with no memory across sessions cannot reproduce that move, whatever a single-turn demo shows.

Finding four

What do trust executives report about edtech impact?

Interviews and survey responses with UK multi-academy trust executives on what they can evidence, what they cannot, and what they buy anyway.

Edtech impact survey

What we can provide

Coding scheme and codebook
The full classification scheme, definitions and coder agreement statistics.
Held-out benchmark set
The 308 expert decision-points, for scoring any tutoring product against expert practice.
Aggregate data tables
Figures behind every published chart, in machine-readable form.
Briefings and evidence sessions
A one-hour walkthrough for officials, with the method and the limits in the room.

Raw transcripts are never shared. Learners are minors and anonymisation precedes any analysis.

Contact

Saif Sarwar and Carl Morris, directors · hello@educave.ai

State the question you need answered and the deadline. We reply with what the corpus can and cannot support.

Methods, ethics and interests · All research

Methods v1.0, August 2026