Educave research
Standards

Methods, ethics and interests

What we count, what a model does, what we will not claim, and who pays us.

Unit of analysis

One expert tutor. 48 recurring students. 943 coded lessons of 1,418 recorded, August 2024 to June 2026, across biology, chemistry, physics and mathematics, GCSE through A-level and IB. Learners are predominantly in England with a minority in the UAE.

Findings describe one expert's practice at high resolution. They are not a description of expert tutoring in general.

How the evidence is made

Sources
Recorded one-to-one lessons, executive interviews and structured survey responses from schools, trusts and edtech teams in the UK and UAE.
Counting
A deterministic pre-processor attributes speech and computes every word, turn and talk-share figure. A language model classifies discourse only; it never counts.
Quality gate
No-shows, aborted connections and planning calls filed as lessons are set aside before analysis.
Agreement
Two independent coders double-coded 10 lessons, about 1% of the coded corpus, so the agreement statistics carry wide intervals. Current agreement: question type 0.95, scaffolding 0.98, error type 0.80. Proportions and directions are reportable; absolute counts are not, and are labelled exploratory until calibration lifts item-level agreement.
Held-out evaluation
Benchmark decision-points are reserved before any policy is derived, so agreement scores are never fitted on their own data. The held-out set shares its tutor and its students with the derivation set, so it tests fidelity to this expert and not external validity.

Ethics

Learners in the tutoring corpus are minors. Anonymisation runs before any model step, an automated scrub verifies that no real names survive into coded output, and only aggregate, paraphrased results leave the repository. Raw transcripts are never shared, published or used for training. Survey and interview participants are anonymised at organisation level; quotes are cleared before publication.

Consent and data protection

Recording and research use of lessons are covered by the Phoenix Tutors service terms agreed by the fee-paying parent or guardian at the point of engagement. No separate research-specific consent was sought. No external ethics review has been conducted. Learners are minors. Anonymisation precedes any model step and raw transcripts never leave the repository.

We regard the absence of a research-specific consent process and an independent ethics review as limitations of the current corpus, and we are addressing both before the full paper.

Data availability

Coding scheme, codebook, agreement statistics and aggregate data tables are available on request. The 308-point benchmark set is available to public bodies and researchers on request. Raw transcripts are not shareable under the consent terms.

Peer review status

This work is not peer reviewed. It is published as working evidence and is open to challenge through peer review.

What we do not claim

Interests

Educave Research is the research arm of Educave, which builds and sells education infrastructure to schools, edtechs and governments. The tutoring corpus is collected with Phoenix Tutors, a commercial tutoring provider whose leaders are directors here. We therefore publish the method, the coder agreement and the held-out benchmark alongside every claim, and we state where a finding would advantage a product we sell. Findings are not withheld or amended for commercial reasons.

Corrections and versioning

Every published piece carries a version number and a date. Substantive changes increment the version and are noted on the page. Corrections can be sent to hello@educave.ai or through peer review.

Methods v1.0, August 2026. Research contact: Carl Morris, hello@educave.ai.

For policymakers · Read the research