Educave research
For anyone reading a supplier's evidence

The evidence benchmark

Five tests, nine kinds of evidence, and the score each one can reach.

Every claim on this site carries a source and a grade. The grade comes from the rubric below, which we apply to our own findings on the same terms as a supplier's. The purpose is to make the ceiling of each evidence type visible before an argument is built on it: a case study cannot become a trial by being repeated, and a randomised trial cannot tell you what happened in the room.

Scoring is two points where the test is met, one where it is partly met, none where it is absent. Ten is the maximum. Nothing available in this market currently scores ten.

The five tests

Comparison
Is there a group that did not receive the thing, chosen by someone other than the supplier?
Independence
Who paid for the study, and could the result have gone against them?
Method on record
Are the sample, the measure and the analysis published in enough detail to be checked?
Attrition
Are pupils who dropped out counted in the result, or defined out of it?
Transfer
Does the setting resemble yours closely enough for the finding to carry?

How evidence types score

A filled circle is a test met, a grey circle partly met, an empty circle not met.

Evidence types scored against five tests, out of ten
Evidence typeComparisonIndependenceMethod on recordAttritionTransferScore
Supplier case studyA named school, a warm quotation, no counterfactual. Useful for understanding what the product does, not whether it works.1 / 10
Supplier-run trialComparison groups are usually completers against non-starters, which measures who finished rather than what the product did.3 / 10
Peer recommendationHigh transfer and no method. The recommending peer may hold a commercial relationship that neither of you has named.3 / 10
Independent pilot, single siteHonest and small. Treat the effect size as a hypothesis for your own trial rather than a finding.7 / 10
Randomised trial, multi-siteThe strongest single design available, and usually two to four years behind the product being sold.9 / 10
Meta-analysis or toolkit entryAverages across settings that differ from yours. Read the range and the conditions, not the headline months of progress.7 / 10
Government or NGO evaluationMethod is usually published and slow. Commissioning bodies sometimes hold a stake in the policy being evaluated.8 / 10
Your own comparison termTwo schools adopt, one holds, both record the same measure. The cheapest strong evidence available to a trust.9 / 10
Structured field observationNo counterfactual by design. It answers what happens and why, which is the question a trial cannot reach.7 / 10

Compare two claims side by side

Pick the evidence type a supplier has offered you and the one you would rather have. The bars show where each design is strong, and the panel underneath names the test that holds it back.

Compare up to four at a time. 3 selected.

Comparison
Supplier case studyNot met
Randomised trial, multi-siteMet
Your own comparison termMet
Independence
Supplier case studyNot met
Randomised trial, multi-siteMet
Your own comparison termMet
Method on record
Supplier case studyNot met
Randomised trial, multi-siteMet
Your own comparison termPartly met
Attrition
Supplier case studyNot met
Randomised trial, multi-siteMet
Your own comparison termMet
Transfer
Supplier case studyPartly met
Randomised trial, multi-sitePartly met
Your own comparison termMet
What limits each one
Supplier case study
Weakest on comparison, independence, method on record, attrition. A named school, a warm quotation, no counterfactual. Useful for understanding what the product does, not whether it works.
Randomised trial, multi-site
Weakest on transfer. The strongest single design available, and usually two to four years behind the product being sold.
Your own comparison term
Weakest on method on record. Two schools adopt, one holds, both record the same measure. The cheapest strong evidence available to a trust.
Evidence types scored against each test, out of two
Evidence typeComparisonIndependenceMethod on recordAttritionTransfer
Supplier case studyNot metNot metNot metNot metPartly met
Supplier-run trialPartly metNot metPartly metNot metPartly met
Peer recommendationNot metPartly metNot metNot metMet
Independent pilot, single sitePartly metMetMetPartly metPartly met
Randomised trial, multi-siteMetMetMetMetPartly met
Meta-analysis or toolkit entryMetMetMetPartly metNot met
Government or NGO evaluationMetPartly metMetPartly metMet
Your own comparison termMetMetPartly metMetMet
Structured field observationNot metMetMetPartly metMet

Our strongest findings, graded

Each finding below carries its basis, its limits and its score on the same rubric.

Field observation · 7 of 10
0 of 3

No trust in our field research could evidence edtech impact centrally

Structured discovery calls with executive leaders across trusts responsible for around 50,000 students, coded from transcripts.

Source: What trust leaders cannot see, Educave Research, June 2026

Practitioner report · 4 of 10
1

An AI tutoring intervention was discontinued after costing more and producing worse outcomes than human tutoring

A single trust's internal comparison, reported to us rather than independently verified. One case, named as one case.

Source: How trust leaders read AI claims, Educave Research, June 2026

Use it on a live bid

Score the next supplier claim you receive against the five tests, then rehearse the drafting that follows in The Procurement Room. If you hold evidence that scores well and has not been published, send it through the call for evidence.