Human tutor intelligence
An analysis of expert one-to-one STEM lessons over two years.
Executive summary
The study draws upon 1,418 recorded one-to-one STEM lessons between August 2024 and June 2026. We analysed who spoke, for how long, what was asked, what went wrong, and what the tutor did next. The result is a description of expert practice as a contribution to the ongoing debates around human vs AI tutoring.
| Measure | 738 lessons | 943 lessons |
|---|---|---|
| Careless slips | 56% | 55% |
| Misconceptions | 29% | 30% |
| Missing prerequisite | 15% | 15% |
| Median student talk share | 0.150 | 0.155 |
Five findings
One. Grades can't signal 45% of corrections
Of 7,546 classified error-corrections3, 55% are slips. The remaining 45% are misconceptions or a missing prerequisite, which is a diagnostic signal . Coding leaned conservatively toward slip, so the deep share is a floor.
Two. The expert asks concrete questions, not exploratory ones
Across 74,258 classified tutor questions5, closed and procedural questions account for 69%. Metacognitive prompting, the move most often recommended in guidance, is 2.4% of the total. An AI tutor imitating this expert would be far more concrete than the default instruction to always ask open questions.
Three. Subject and level change the behaviour, not just the content
Biology8 is the most lectured and the most misconception-heavy. Maths is the most student-led, with guided scaffolding overtaking direct instruction and errors that are mostly slips. Metacognitive questioning scales with level, from about 0.6 markers a lesson at GCSE to about 1.8 at A-level and IB. One generic tutoring policy cannot be right for all of these at once.
Four. Withholding is a property of the relationship
Across students with long histories, the expert holds space for productive struggle more as the relationship matures, in 24 of 30 students7. The withholding index rises in 18 of 30 and student talk in 20 of 30. Within a single lesson there is no such arc. We report this as a consistent majority trend rather than a confirmed population effect, because the mixed-effects significance is still sample-dependent.
At the five-student pilot stage this finding was a null. That is the clearest argument in the study for building benchmarks on a full corpus.
Five. Generic AI tutors get the give-or-withhold call wrong half the time
Across 308 held-out decision-points6, the expert's decision to withhold the answer depends almost entirely on the student's state: 77% when the student is partially right, 30% when truly stuck, 11% when the student asks a direct question.
Scored against that ground truth, a generic tutor that always explains agrees with the expert 46% of the time. A state-aware policy derived from this corpus reaches 72%.6
The policy the corpus produces
For every correction, we recorded the expert's next move. The map is conditioned on the error, and it is specific enough to run.
From 121 coded correction episodes; scales as coding completes.9
How to read these numbers
A deterministic pre-processor anonymises each transcript, attributes tutor and student speech, and computes every word, turn and talk-share count. The language model never counts; it classifies discourse only. A quality gate sets aside no-shows, aborted connections and planning calls filed as lessons.
Two independent coders double-coded ten lessons. They agree strongly on the discourse mix (question type 0.95, scaffolding 0.98, error type 0.80) and disagree on raw counts, which are a segmentation judgement. So the proportions and the directions of change quoted here are reliable; the absolute counts are not yet. Every figure is exploratory until a coder-calibration round lifts item-level agreement. Coding stands at 81% complete, 1,144 of 1,418 lessons processed.
- 1.1,418 lessons, about 1,404 hours: all one-to-one STEM sessions recorded August 2024 to June 2026. Counted deterministically from transcript metadata before any model step. Educave and Phoenix Tutors, coded corpus, August 2026.
- 2.943 analysable lessons across 48 recurring students: the 1,418 recordings less no-shows, aborted connections and planning calls removed by the quality gate. Coding stands at 81% complete, 1,144 of 1,418 lessons processed. Same source.
- 3.7,546 error-corrections classified as careless slip, misconception or missing prerequisite across the 943 coded lessons. Double-coded agreement on error type is 0.80; proportions are reliable, absolute counts are not yet. Same source.
- 4.308 decision-points in the held-out evaluation set: moments where the expert either gave or withheld the answer, reserved from policy derivation so agreement scores are not fitted on their own data. Same source.
- 5.74,258 tutor questions classified as closed or factual, procedural, Socratic, open-exploratory or metacognitive. Double-coded agreement on question type is 0.95. Same source.
- 6.Withhold rates by student state, and 46% versus 72% agreement, are scored against the 308 held-out decision-points in note 4. The 46% baseline is a tutor that always explains; 72% is a state-aware policy derived from this corpus.
- 7.Relationship trend across 30 students with long histories: withholding rises in 18 of 30 and student talk in 20 of 30. Reported as a majority trend, not a population effect; mixed-effects significance remains sample-dependent.
- 8.Subject signatures use the 738-lesson sample available at the previous coding refresh, retained because directions held stable when the corpus near-doubled to 943.
- 9.Error-to-response mapping comes from 121 coded correction episodes and scales as coding completes.
- School and trust leadersAsk any tutoring supplier, human or AI, for its withhold policy by student state. A tool that always explains is teaching to the 46% baseline.
- Policy and governanceGrades record the outcome and discard the reason. Around 45% of the errors in this corpus carry a diagnosis no management information system holds.
- Edtech buildersThe model is not the differentiator. A conditioned error-to-response policy and a held-out eval set are what let you prove a tutor beats the generic baseline.
- ResearchersFindings stabilised at roughly 230 coded sessions, and one headline result was a null at five students. Pilot-scale tutoring studies should be read with that in mind.
Request the full research paper
Methods, all twelve studies, the error-to-response policy map, the eval set and the reliability appendix. One email when it publishes.
Your details stay with the research team and are never passed to a supplier.
If your school, trust or product has evidence that tests this piece, send it in. Submissions are anonymised before anything is published.
Contribute evidenceNaturalistic, longitudinal, speaker-attributed one-to-one expert STEM lessons recorded August 2024 to June 2026 across biology, chemistry, physics and mathematics. Learners are minors: anonymisation precedes any model step, an automated scrub verifies that no real names remain in coded output, and only aggregate, paraphrased results leave the repository. Figures are exploratory pending coder calibration.

