All articles · Published 2026-08-28

How Accurate Is a TOEIC Score Estimate? Four Test-Takers Sent Us Their Real Scores, So Let's Do the Math

The day the August 16 TOEIC results came out, four test-takers reported their real scores back to us — the first time Handy 990's score estimate has faced a real, graded exam in batch. The headline: listening was startlingly accurate (two of the four were off by exactly 5 points), reading carries three systematic biases that each have a name, and the estimate ranked all four people in exactly the right order. This post lays the four daily estimate trajectories against the four real score reports, and answers a harder question along the way: what does it actually mean when an estimate "converges"?

At the end of How Handy 990 estimates your TOEIC score, we made a promise: this estimation method is honest about its own uncertainty. It was an easy promise to write, because at the time nobody could check it. That has changed — after your exam date passes, the app asks a simple question, "how did it go?" — and on August 28, the morning results were released, four real score reports arrived within two hours of each other. All four are founding members; all four sat the August 16 administration. The sample is four people, and this article will keep remembering that fact from start to finish. But four people turn out to be enough for certain patterns to be too clear to hide.

Four score reports, one table

"Pre-exam estimate" is each person's last daily estimate snapshot before test day — the number on their screen when they walked into the exam room.

Test-taker Practice volume Pre-exam estimate (L/R) Real score (L/R) Error
A 1,477 questions 540 (310/230) 495 (315/180) 45 too high
B 731 questions 685 (395/290) 760 (400/360) 75 too low
C 307 questions 795 (395/400) 855 (430/425) 60 too low
D 2,697 questions 815 (460/355) 895 (495/400) 80 too low

The overall picture first: the mean absolute error on the total was 65 points (on the 990 scale), and all four landed within 80. And the ordering — who scored higher than whom, and by roughly how much — came out exactly right. Getting four people in the right order by guessing has a probability of 1/24; we're not popping champagne over it, but it does say the estimate at least measured relative position correctly, and relative position is precisely what the question "how far am I from my target?" actually needs.

The interesting part appears when you split the totals open.

Why was listening accurate to within 5 points?

The four listening errors were +5, +5, +35, and +35 — a mean absolute error of 20 points. For scale: ETS's own TOEIC Listening and Reading Score User Guide states that the standard error of measurement (SEM) of the official test is about 25 points per section — meaning the same person retaking the official exam a week later would see fluctuations of roughly that size anyway. Our listening error is the same order of magnitude as the official test's error about itself.

Test-taker A deserves a special mention: estimated listening 310, real listening 315. And in their report, they told us the previous sitting — July 28 — was the same story: "the listening score was close." Same person, two exams, two direct hits — that does not look much like luck.

Why can listening be this accurate? We think the answer is a bit boring: because practice conditions match exam conditions. Listening is paced by the audio, not by you — doing a Part 3 set in the app and doing one in the exam room puts you under identical time pressure. The estimate measures "your accuracy under these conditions," and the exam happens to be exactly these conditions.

[Update, Aug 29] The day after this article went live, a fifth score report arrived: pre-exam estimate 810 (Listening 435 / Reading 375), real score 805 (Listening 405 / Reading 400). The total was off by just 5 points, and the ranking across all five people is still exactly right — but hold the listening victory lap: this is our first over-estimate of listening, by a full 30 points. The test-taker had already guessed the cause, writing in their report that "I sometimes replay the audio several times, which might make the score inaccurate" — exactly right. Replaying audio is to listening what the missing clock is to reading: when practice conditions are more forgiving than the exam room's, the estimate drifts toward the forgiving side. The total only looks perfect because reading was simultaneously under-estimated by 25 points (this test-taker's estimate climbed from 475 to 830 over the three pre-exam weeks — the lagging indicator couldn't keep up, as usual), and the two biases happened to cancel. So this section's claim that "practice conditions match exam conditions" needs one rider: they match only if you play each clip once.

Reading is the opposite.

The three kinds of reading error, each with a name

The four reading errors were −50, +70, +25, and +45 — a mean absolute error of 48 points, nearly two and a half times listening's, and in both directions. Spread the four cases out and three systematic biases each show themselves.

First: untimed ability is not timed performance. Test-taker A's reading was over-estimated by 50 points — and their July 28 sitting followed the same script ("reading was off by a lot," in their words). Everyday practice is untimed; you can take your time with a Part 7 passage and read the whole thing properly. The real reading section is 100 questions in 75 minutes, and what you can't finish, you can't finish. Test-taker D also wrote that reading felt rushed — "almost didn't finish" — and D scored 895. The estimate measures untimed accuracy; the exam grades timed accuracy; between those two sits an independent skill called speed.

Second: drill your weaknesses, and the estimate follows your weaknesses. Test-taker B was under-estimated by 75 points, and B wrote the diagnosis into the report personally: "I focus my practice on my weaker parts (4, 5, 6), so the score comes out lower." Exactly right. Their pre-exam practice concentrated on their weakest question types, so the answer history feeding the estimator was a sample deliberately biased toward weakness — but the real exam doesn't only test your weaknesses. It's a textbook case of selection bias, except this time the test-taker wrote the textbook. (Incidentally: B also used another score-estimating app, which put them at 640. Both apps under-estimated them; we were 75 low, it was 120 low. Under-estimating people who prepare seriously appears to be an industry-wide condition.)

Third: the estimate is a lagging indicator by design, and you are improving. Test-taker C practiced only 307 questions, the fewest of the four, yet got the second-smallest error. More interesting are C's two real-world anchors: a 720 from a no-preparation June sitting, and an 855 in August — 135 points of improvement in two months. Our daily estimate climbed from 735 to 795 over that stretch: it started accurate (735 against a two-month-old real 720), moved in the right direction, and couldn't keep up with the slope. That's not an excuse — it's the design's necessary consequence. The estimator decays old evidence over time, so it measures "the average you of the past few weeks"; the person who walks into the exam room is "the best you of exam day." The faster you improve, the wider that gap. Test-taker D adds a fourth, smaller bias on top — ceiling compression: D's listening estimate sat pinned at 460 for the entire final week (blinking down to 455 exactly once), because above roughly 90% accuracy, sampling noise is larger than the score increments and the estimate can't climb. D's real listening score was 495. A perfect section.

So — did the estimate converge?

This is the question we most wanted to answer, and the honest answer is: "converge" means two different things. We achieved the first; the second, only halfway.

The first meaning is statistical convergence: the estimate stops jittering. All four achieved it. Over the final two pre-exam weeks, the daily total-estimate standard deviations were 19, 34, 20, and 22 points — and the 34 (test-taker B) wasn't noise: their listening estimate genuinely climbed from 320 to 395 in two weeks, and the trend got counted into the standard deviation. For contrast, the older method described in How Handy 990 estimates your TOEIC score — a fixed EMA — jitters by about ±11 percentage points forever; the current method, a few hundred questions in, moves only a few points day to day. The displayed uncertainty intervals narrowed honestly too: B's pre-exam screen read Listening 395±28, Reading 290±20.

The second meaning is convergence to the true value. Listening got there: all four errors fall within a reasonable multiple of the interval. Reading did not: B's real reading of 360 landed a full 50 points outside 290±20. This is the part that needs saying plainly — the ± we display measures sampling uncertainty: "if the exam questions resemble your practice questions, and your ability resembles the past few weeks, the score lands in this range." It knows nothing about timed pressure, nothing about practice deliberately skewed toward your weaknesses, nothing about the fact that you're improving. Convergence is not correctness; it only guarantees correctness when its premises hold — and these four test-takers each broke one premise apiece.

What we'll change (and what we won't)

Listening stays as it is. Conditions match, and the error is the same size as the official SEM; it doesn't get better than that.

Of reading's three biases, two already have tools — we just hadn't said so plainly enough: for the timing problem, practice under the clock with timed mock exams, and train Part 7 speed directly (Can learning speed reading save my reading section?); for improvement lag, enter your real score after the exam and the estimate recalibrates against the official score report (Your TOEIC score report is more than two numbers). The weakness-drilling bias is the hardest: punishing users for the single most correct exam-prep behavior — practicing their weaknesses — would be absurd, so the correction has to happen on the estimation side, not the practice side. We're working out how to correct it without pretending to have data we don't have; until then, the plain version goes here in print: if your final pre-exam weeks were spent drilling weaknesses, your real level is most likely above your estimate.

Hold on — before you quote these numbers

First, n=4. This is a stack of case studies, not a validation study; we have not changed a single parameter off the back of these four data points. Second, reporting is voluntary, so the sample skews toward people with something to say — people who beat their estimate may well be happier to come back and report (three of four here did), which means the "estimates run conservative" pattern may itself be a reflection of reporting bias. Third, all four sat the same administration; whatever equating error exists between administrations, this data cannot see it. Fourth, even a perfect model runs into ETS's own per-section SEM of 25 points — predicting the official test more precisely than the official test predicts itself is not a mathematically winnable game. The goal was never prophecy. It's an honest answer to "roughly where am I, how far to my target, and what should I practice next." By that standard, these four score reports gave us one passing grade, two clear pieces of homework, and a reason to keep collecting score reports.

If you also took the August 16 sitting — or any after it — the app will ask how it went when your results arrive. Please tell it. The sequel to this article gets written by your score report.


The score estimate lives on the Handy 990 home screen; tap through to "how the estimate is calculated" for each section's uncertainty interval and daily trend. The full method is in How Handy 990 estimates your TOEIC score; the post-exam calibration mechanism is in Your TOEIC score report is more than two numbers; timed reading training is in Can learning speed reading save my reading section?.

← All articles