Educave research
Field study · August 2026

Human tutor intelligence

An analysis of expert one-to-one STEM lessons over two years. 

Carl Morris · Executive summary · 7 min
1,4181
expert STEM lessons recorded, around 1,404 hours
9432
analysable coded lessons across 48 recurring students
7,5463
student errors classified by type
3084
expert decision-points in the held-out eval set
01

Executive summary

The study draws upon 1,418 recorded one-to-one STEM lessons between August 2024 and June 2026. We analysed who spoke, for how long, what was asked, what went wrong, and what the tutor did next. The result is a description of expert practice as a contribution to the ongoing debates around human vs AI tutoring.

Measure738 lessons943 lessons
Careless slips56%55%
Misconceptions29%30%
Missing prerequisite15%15%
Median student talk share0.1500.155
Headline proportions across a near-doubling of coded lessons. Educave and Phoenix Tutors, coded corpus, August 2026.

02

Five findings

One. Grades can't signal 45% of corrections 

Of 7,546 classified error-corrections3, 55% are slips. The remaining 45% are misconceptions or a missing prerequisite, which is a diagnostic signal . Coding leaned conservatively toward slip, so the deep share is a floor.

Careless slip: 55%Misconception: 30%Missing prerequisite: 15%
Hover or tab through the chart to read each figure.
Error type, 7,546 coded corrections across 943 lessons. Exploratory; proportions reliable, absolute counts not yet.

Two. The expert asks concrete questions, not exploratory ones

Across 74,258 classified tutor questions5, closed and procedural questions account for 69%. Metacognitive prompting, the move most often recommended in guidance, is 2.4% of the total. An AI tutor imitating this expert would be far more concrete than the default instruction to always ask open questions.

Hover or tab through the chart to read each figure.
Tutor question type, 74,258 coded questions. Educave and Phoenix Tutors, August 2026.

Three. Subject and level change the behaviour, not just the content

Biology8 is the most lectured and the most misconception-heavy. Maths is the most student-led, with guided scaffolding overtaking direct instruction and errors that are mostly slips. Metacognitive questioning scales with level, from about 0.6 markers a lesson at GCSE to about 1.8 at A-level and IB. One generic tutoring policy cannot be right for all of these at once.

Median student talk share by subject, 738-lesson sample. Directions stable across every refresh.

Four. Withholding is a property of the relationship

Across students with long histories, the expert holds space for productive struggle more as the relationship matures, in 24 of 30 students7. The withholding index rises in 18 of 30 and student talk in 20 of 30. Within a single lesson there is no such arc. We report this as a consistent majority trend rather than a confirmed population effect, because the mixed-effects significance is still sample-dependent.

At the five-student pilot stage this finding was a null. That is the clearest argument in the study for building benchmarks on a full corpus.

Five. Generic AI tutors get the give-or-withhold call wrong half the time

Across 308 held-out decision-points6, the expert's decision to withhold the answer depends almost entirely on the student's state: 77% when the student is partially right, 30% when truly stuck, 11% when the student asks a direct question.

Hover or tab through the chart to read each figure.
Rate at which the expert withholds the answer, by student state. 308 held-out decision-points.

Scored against that ground truth, a generic tutor that always explains agrees with the expert 46% of the time. A state-aware policy derived from this corpus reaches 72%.6

Hover or tab through the chart to read each figure.
Agreement with the expert's withhold decision, 308 decision-points. Benchmark set, held out from policy derivation.

03

The policy the corpus produces

For every correction, we recorded the expert's next move. The map is conditioned on the error, and it is specific enough to run.

Slip, careless
Correct it outright (51%) or locate it for the student (39%).
Misconception, wrong model
Guide or reteach (83%). The answer is handed over in 17% of cases.
Missing prerequisite
Reteach the underlying concept (91%).

From 121 coded correction episodes; scales as coding completes.9


04

How to read these numbers

A deterministic pre-processor anonymises each transcript, attributes tutor and student speech, and computes every word, turn and talk-share count. The language model never counts; it classifies discourse only. A quality gate sets aside no-shows, aborted connections and planning calls filed as lessons.

Two independent coders double-coded ten lessons. They agree strongly on the discourse mix (question type 0.95, scaffolding 0.98, error type 0.80) and disagree on raw counts, which are a segmentation judgement. So the proportions and the directions of change quoted here are reliable; the absolute counts are not yet. Every figure is exploratory until a coder-calibration round lifts item-level agreement. Coding stands at 81% complete, 1,144 of 1,418 lessons processed.

Notes and sources
  1. 1.1,418 lessons, about 1,404 hours: all one-to-one STEM sessions recorded August 2024 to June 2026. Counted deterministically from transcript metadata before any model step. Educave and Phoenix Tutors, coded corpus, August 2026.
  2. 2.943 analysable lessons across 48 recurring students: the 1,418 recordings less no-shows, aborted connections and planning calls removed by the quality gate. Coding stands at 81% complete, 1,144 of 1,418 lessons processed. Same source.
  3. 3.7,546 error-corrections classified as careless slip, misconception or missing prerequisite across the 943 coded lessons. Double-coded agreement on error type is 0.80; proportions are reliable, absolute counts are not yet. Same source.
  4. 4.308 decision-points in the held-out evaluation set: moments where the expert either gave or withheld the answer, reserved from policy derivation so agreement scores are not fitted on their own data. Same source.
  5. 5.74,258 tutor questions classified as closed or factual, procedural, Socratic, open-exploratory or metacognitive. Double-coded agreement on question type is 0.95. Same source.
  6. 6.Withhold rates by student state, and 46% versus 72% agreement, are scored against the 308 held-out decision-points in note 4. The 46% baseline is a tutor that always explains; 72% is a state-aware policy derived from this corpus.
  7. 7.Relationship trend across 30 students with long histories: withholding rises in 18 of 30 and student talk in 20 of 30. Reported as a majority trend, not a population effect; mixed-effects significance remains sample-dependent.
  8. 8.Subject signatures use the 738-lesson sample available at the previous coding refresh, retained because directions held stable when the corpus near-doubled to 943.
  9. 9.Error-to-response mapping comes from 121 coded correction episodes and scales as coding completes.
What this means for you
  • School and trust leaders
    Ask any tutoring supplier, human or AI, for its withhold policy by student state. A tool that always explains is teaching to the 46% baseline.
  • Policy and governance
    Grades record the outcome and discard the reason. Around 45% of the errors in this corpus carry a diagnosis no management information system holds.
  • Edtech builders
    The model is not the differentiator. A conditioned error-to-response policy and a held-out eval set are what let you prove a tutor beats the generic baseline.
  • Researchers
    Findings stabilised at roughly 230 coded sessions, and one headline result was a null at five students. Pilot-scale tutoring studies should be read with that in mind.
Full paper · twelve studies

Request the full research paper

Methods, all twelve studies, the error-to-response policy map, the eval set and the reliability appendix. One email when it publishes.

Your details stay with the research team and are never passed to a supplier.

Contribute evidence

If your school, trust or product has evidence that tests this piece, send it in. Submissions are anonymised before anything is published.

Contribute evidence
Method and anonymisation

Naturalistic, longitudinal, speaker-attributed one-to-one expert STEM lessons recorded August 2024 to June 2026 across biology, chemistry, physics and mathematics. Learners are minors: anonymisation precedes any model step, an automated scrub verifies that no real names remain in coded output, and only aggregate, paraphrased results leave the repository. Figures are exploratory pending coder calibration.

Back to research · Peer review