A 7-cent AI tutor matched a $75/hour human on GRE learning gains
Handshake's StudentBench ran 2,383 students through an hour of GRE tutoring from 12 LLMs, expert humans, or nobody. Pooled AI tutoring was statistically equivalent to human tutoring, and the cheapest model, Gemma 4 31B at $0.067 a session, matched humans at 918x lower cost per point gained.
An hour of tutoring from an LLM raised GRE scores about as much as an hour with an expert human tutor. Handshake AI Research ran 2,383 students through a pre-test, one hour of tutoring, and a post-test, with 12 different models, former ETS and Kaplan tutors, or no tutoring at all. Pooled across every model, AI tutoring landed within 0.58 points of human tutoring, and the equivalence test passed (p=.015). The cheapest model in the study, open-weight Gemma 4 31B, cost $0.067 for the whole session.
- Learning gain: AI 6.15 points over control, human 7.04. Difference: -0.58 points (90% CI [-2.18, 1.03]), inside the ±4.09-point equivalence bound
- The 12 AI tutors were statistically indistinguishable from each other on learning gain (p=.755), from Opus 4.8 at 16.0 points down to Gemini 3.6 Flash at 12.1
- Gemma 4 31B: $0.0052 per percentage point gained vs. $4.81 for a human, a 918x gap. It passed its own equivalence test against humans (p=.044)
- Latency drove engagement: reply times ranged from 1.9s (GPT-5.4 mini) to 31.0s (GPT-5.5 Pro), and faster tutors got more student messages (Spearman Ï=-0.81)
- Only immediate gains were measured; nobody was retested weeks later
The setup
Each student took a 27-question GRE Quantitative or Verbal test at real exam timing, spent five minutes reviewing their mistakes, then got one of three things for an hour: a randomly assigned AI tutor (the student wasnât told which), a live one-on-one video session with a human tutor, or unrelated educational videos (the control). Then a second 27-question test. All questions were newly written by former GRE exam writers so no model could have trained on them, and the two test forms were randomly swapped between pre and post so a harder form couldnât pass for learning.
Both human and AI tutors saw the studentâs pre-test mistakes. Neither saw the post-test. The AI side was deliberately thin: two prompts, one to plan a lesson and write practice problems, one to tutor live in chat. The authors wanted to measure the models, not their own scaffolding.
Hereâs where everyone ended up. Scores are raw points gained from pre-test to post-test, adjusted for starting score:
Learning gain after one hour, percentage points
Combined Quantitative + Verbal, adjusted for pre-test score. ~160 sessions per AI tutor, 140 human, 190 control. Reasoning setting in parentheses.
Every tutorâs 95% interval spans roughly ±2.5 points. The 12 AI tutors are not significantly different from one another (p=.755).
Two things jump out. First, the control group gained 7.8 points just from reviewing their mistakes and retaking a test, so roughly half of everyoneâs raw gain is practice effect. The tutoring effect is the gap above that: about 6 to 7 points, or 1.5 to 2 more correct answers out of 27. Second, the spread among models is narrow and statistically flat. A 31B open-weight model and GPT-5.5 Pro produced learning gains the study canât tell apart.
On a per-section basis, AI was equivalent in Quantitative (+0.88 points vs. human) but not in Verbal (-2.19, p=.085). Human tutors kept the highest mean in all three Verbal domains against pooled AI, though the single best AI tutor in each domain beat the human mean in five of seven.
Cost is where the models actually separate
If learning gain is a wash, the price tag decides. The paper proposes a unit for this: dollars per percentage point of learning gain, i.e. mean session cost divided by mean gain.
Cost of one tutoring session
AI: measured inference cost for lesson plan, practice problems, and one hour of chat. Human: a $75/hour market reference rate.
Human, per point gained
$4.81
Gemma 4 31B, per point gained
$0.0052
Ratio
918x
Across the 12 models, session cost spanned more than 300x, from $0.067 to $21.24. Every model on the cost/learning Pareto frontier cost under $5 a session. The same four models were cheapest per point in both sections: Gemma 4 31B, GPT-5.4 mini, Kimi K2.6, and Gemini 3.7 Flash. Some expensive models were strictly dominated: Gemini 3.1 Pro had a higher mean gain, lower cost, and faster replies than GPT-5.5 Pro.
Six of the 12 tutors passed individual equivalence tests against human tutoring. Gemma 4 31B was one (mean gain 12.7 vs. 15.6 raw for humans; difference after adjustment -1.09 points, p=.044).
Why a cheap, fast model can keep up
The study also built three âteaching qualityâ leaderboards: 51 expert tutors did 2,028 blind pairwise comparisons of AI lesson plans and practice problems, and six rule-based indicators scored 1,971 AI transcripts for things like asking students to explain their reasoning and leaving room to attempt a problem before showing the solution.
Anthropic models, especially Opus, won the expert rankings for lesson planning and practice-problem design. The conversational-pedagogy scores clustered cleanly by provider, with no overlap between model families, which the authors read as teaching style coming from company-wide training choices. Student answer disputes tracked expert ratings too: GPT-5.4 mini had the highest dispute rate, with students flagging about 11% of its Quantitative practice answers.
But those quality rankings didnât translate into significantly different learning gains. The paperâs best candidate explanation is latency. In 1,137 Quantitative sessions, there was a chain: faster replies went with more student messages, more messages with more correct practice problems, and more correct practice with larger gains (all p under .002). A model that answers in 2 seconds gets more reps in an hour than one that thinks for 30. The chain didnât hold in Verbal, where students sent fewer, longer messages.
For anyone building a tutoring product, the takeaway is that the hour is the budget, and the studentâs practice volume inside it matters. A slower, smarter model has to be much better per turn to make up for fewer turns.
Caveats
- Immediate gains only. The post-test came right after the session. Nothing says the 6 points survive a week.
- No âpractice problems aloneâ control. Control students watched unrelated videos. So the study canât separate what an AI tutor adds over just giving a motivated student a stack of targeted practice problems.
- The human arm is small and clustered. 140 human sessions vs. 2,139 AI. Because a few tutors taught many students, the combined comparison has just 1.59 effective degrees of freedom, with 102 of the 140 human sessions in one connected cluster. Dropping repeat students or each tutor in turn kept equivalence, but a simpler gain-only analysis clustered by tutor didnât establish it in either section.
- Equivalence means âwithin about 4 points,â not âidentical.â The bound was ±0.25 standard deviations (±4.09 points), against a tutoring effect of only ~6 to 7 points. Verbal alone didnât pass.
- Per-model tests werenât corrected for multiple comparisons, so âsix models passed individuallyâ should be read loosely.
- Different modalities. Humans tutored over live video; AI tutored in text chat.
- Study population: paid volunteers (median age 21) recruited through Handshake, which also funded the study. The $75/hour human rate is a published market reference, not what these tutors were paid.
The platform is live at studentbench.org, with data on Hugging Face and code on GitHub. The strongest version of the result is narrower than the headline: for a one-hour GRE session, which model you pick barely changed how much students learned, so the cheap, fast one is the rational default.
Liked this? Engineer's Codex sends one deep dive and a link roundup every week.
No spam. Unsubscribe anytime.