Evidence, ungated
Nothing on this page asks for an email. Public bodies can read the findings, the method and the limits, then ask us for the rest.
Ofsted Areas of Research Interest
Area one: gaps between advantaged and disadvantaged learners
We cannot answer this question. The corpus is fee-paying one-to-one tutoring, structurally an advantaged setting, with no pupil premium, SEND or EAL indicator linked at source.
What we do hold is a precondition finding. Trust executives responsible for around 50,000 students could not evidence centrally which of their edtech works, with data across 12 or more disconnected systems and about a term of lag before drift is visible. The question cannot currently be answered by schools themselves.
We also hold an instrument rather than a conclusion: a coding scheme with double-coded agreement of 0.95 on question type and 0.98 on scaffolding, and a held-out benchmark design. Both transfer to a disadvantaged cohort.
Area two: cognitive and socioemotional development
We hold process measures, not developmental outcomes.
- 69% of 74,258 expert questions are closed or procedural. Metacognitive prompting, the move most guidance recommends, is 2.4%.
- Withholding tracks student state: 77% partially right, 60% wrong, 30% truly stuck, 11% on a direct question.
- A generic always-explain policy agrees with the expert on 46% of held-out decisions, against 72% for a state-aware policy.
These measure how problem-solving support and productive struggle are managed. We hold no measure of resilience or socioemotional outcomes.
We hold no evidence on online harms, their third area.
Limits, stated up front
- Descriptive, not causal. The corpus records practice; it does not measure attainment.
- Purposive sample. Findings describe these lessons and these organisations, not the sector.
- Learners predominantly in England, with a minority in the UAE.
- No disadvantage breakdown yet. Pupil-premium and SEND indicators are not linked at source, so we make no gap-narrowing claim.
- Commercial interest declared: Educave sells education infrastructure and the corpus is collected with a tutoring provider whose leaders are directors here. Methods and interests.
Questions this evidence answers
The corpus gives a measurable benchmark rather than an opinion. A generic AI tutor matches the expert's decision to give or withhold an answer 46% of the time; a state-aware policy derived from the corpus reaches 72%. Any product claim about tutoring quality can be scored the same way.
Around 45% of corrections carry a diagnosis — a wrong model or a missing prerequisite — that no management information system records. Attainment data cannot answer questions about mechanism because the mechanism is never captured.
Withholding an answer is conditioned on the student's state and rises as the relationship matures. Systems with no memory across sessions cannot reproduce that move, whatever a single-turn demo shows.
Interviews and survey responses with UK multi-academy trust executives on what they can evidence, what they cannot, and what they buy anyway.
What we can provide
Raw transcripts are never shared. Learners are minors and anonymisation precedes any analysis.
Saif Sarwar and Carl Morris, directors · hello@educave.ai
State the question you need answered and the deadline. We reply with what the corpus can and cannot support.
Methods, ethics and interests · All research
Methods v1.0, August 2026

