All articles · Published 2026-08-27

Your TOEIC Score Report Is More Than Two Numbers: What ABILITIES MEASURED Means, and How It Calibrates Your Estimated Score

Below the two big numbers on a TOEIC score report sits a row of bar charts most people glance at once and never again — ABILITIES MEASURED, ten percent-correct figures across ten skills. This article explains why those bars are the most information-dense part of the whole report, and what Handy 990 does when you enter your real result: it calibrates your estimated score — keeping the shape of strengths and weaknesses your months of practice established, and shifting only the overall level onto the official truth. We'll also own up to something: our first version got this wrong.

At the end of How your TOEIC score is estimated, we made a promise: as real exam results come in, the estimate will get sharper. This is the sequel that promise was waiting for — the score report arrived. Now what?

Let's make the scene concrete. You've practiced with the app for three months, hundreds of questions across the seven Parts; the app says 85% on Part 5, 55% on Part 7, and draws you a radar chart. Then you sit the real TOEIC, and two or three weeks later the report shows up. You now hold two accounts of the same thing — your English ability. One is a fine-grained profile built from hundreds of answers; the other is an official number with ETS's stamp on it. They will almost certainly disagree.

Which one should the app believe?

Two obvious approaches, both wrong

Approach one: the report wins — overwrite everything. The official score is obviously more authoritative than our estimate, so wipe the estimate and restart from the report. The problem: the report is two numbers — one for Listening, one for Reading. It doesn't know your Part 5 runs thirty points ahead of your Part 7, let alone that you handle single passages fine and fall apart on triple passages. Those facts took three months and hundreds of questions to establish, and two numbers cannot hold them. Overwriting throws away the most valuable thing the app has.

Approach two: the report is "for reference only" — keep the estimate as is. The opposite instinct: trust your own model. But we admitted in the previous article that the model's constants are reasoned, not fitted, and that our question bank's difficulty doesn't line up perfectly with the real exam — and the gap differs from person to person. If the app says 750 and the report says 680, continuing to display 750 isn't confidence. It's plugging your ears.

Each approach discards half the information. The right answer hides in one observation: these two sources are good at different things.

The report knows the altitude; your practice history knows the ridgeline

Picture your ability as a mountain range: seven Parts, seven peaks of different heights.

Your practice history — hundreds of questions, spread over weeks, recorded one answer at a time — has traced the ridgeline of that range: which peaks are high, which valleys are low, and by how much. The tracing is fine enough to tell apart Part 7's single, double, and triple passages as three different slopes. But it has one built-in weakness: the whole ridgeline's altitude may be marked wrong. Our questions aren't written by ETS, so the difficulty calibration carries some systematic offset — your ridgeline's shape is right, but the entire range may be drawn a few dozen meters too high or too low.

The score report is exactly the opposite. ETS's scores go through large-scale equating; they are the truth about altitude. But about the ridgeline the report is nearly silent: one number for Listening, one for Reading — two "average peaks" and nothing more.

So there's only one correct move: use the report to calibrate the altitude, and your practice history to keep the ridgeline.

What is ABILITIES MEASURED?

Except the report isn't entirely silent about the ridgeline — which is where those overlooked bars come in.

The lower half of an official TOEIC score report (the paper one — under Taiwan's official schedule it's mailed on the 16th working day after the test, while online score lookup opens on the 10th working day but shows only the scores) prints ABILITIES MEASURED: five Listening skills and five Reading skills, each with a percent-correct. They're organized by skill, not by Part — the Listening five are roughly "gist of short spoken texts," "gist of extended spoken texts," "details in short spoken texts," "details in extended spoken texts," and "a speaker's implied meaning"; the Reading five are "inference," "locating specific information," "connecting information across sentences and texts," "vocabulary," and "grammar."

ETS computes those ten numbers from your actual test — the hundred Listening questions and hundred Reading questions you answered. In other words, this is official, coarse-grained ridgeline information. Handy 990's score-entry sheet now includes an expandable panel where you can copy those ten percentages in. The ten rows in the app deliberately use the report's own printed wording, in the report's printed order, so you can transcribe top to bottom with the report in your other hand. All ten fields are optional — enter as many or as few as you like, including none.

How the calibration works

Three steps.

Step one: keep the ridgeline. Each Part's starting point is the accuracy your own practice established — not the report's average. Your Part 5 is strong and your Part 7 weak? After calibration, that's still true.

Step two: fold the ability figures in — as evidence, not as truth. Here's a detail that's easy to miss: the ten percentages rest on very different amounts of data. A full test has only 6 Part 1 questions but 39 Part 3 questions — so the bar about "details in short spoken texts" is backed by far less evidence than it appears. We fold each ability figure into the matching Parts' estimates as if it were that many fresh observations, weighted by exactly its question count on the test: Part 3's evidence folds in with the weight of 39 questions, Part 1's with the weight of 6. Your own practice record sits on the other side of the scale — if a Part already carries dozens of questions of effective evidence, the ability figure nudges the estimate rather than replacing it. A brand-new user with no record, by contrast, has nothing on that side of the scale, so the ability figures essentially decide. One formula covers both people, with no "new user vs. veteran" branch anywhere.

Step three: shift the whole section to match the altitude. With the first two steps done, every Part in a section is shifted by the same amount, until the section score computed back from the estimates is exactly the official score you entered. This step is mathematically exact — the Parts' question counts in each section sum to precisely one hundred, so raising every Part by one percentage point raises the section by exactly one question — which means that no matter how the shape moved, the score the app derives afterwards is the one printed on your report. The official score always has final authority; the ability figures only distribute it, never change its total.

The radar chart follows the same logic: ETS's ten abilities map almost one-to-one onto the six axes of the app's radar ("grammar" to grammar, "connecting information" to inference, and so on), so an entered panel pushes the radar toward the report — pushes, never overwrites.

Our first version got this wrong

Full disclosure: when this feature first shipped, it used exactly "approach one" above. Enter a score, and the whole estimate was wiped and reseeded. For a brand-new user that's harmless — there is no ridgeline to preserve — but for someone with months of practice it traded hundreds of recorded answers for two numbers. Shortly after shipping, we walked through the "practice three months, then take the test" scenario ourselves, saw the problem, and replaced it with the calibration described here within days. We're leaving this paragraph in as a record: authoritative data can still destroy valuable information. "More authoritative" does not mean "should overwrite."

Two weeks after your exam, the app will remind you

There was a more basic problem, too: most people simply don't know the report is worth entering. The app's old behavior — the home screen's exam countdown quietly disappears the day after the test, and then… nothing. As if the exam had no aftermath.

Now, starting two weeks after your registered exam date (when online scores are already up and the paper report is nearly in the mail), if you haven't entered a result for that sitting, a card appears on the home screen: "Has your score report arrived?" Tapping it opens score entry directly — the scores, the abilities panel, and then, in the same flow, a fresh per-Part target plan for your next exam. You can dismiss it; the same sitting will never ask again, and your next registered exam re-arms the reminder naturally. It also stops appearing three months after the exam — asking "has your report arrived?" a quarter of a year later would just prove the app hadn't been paying attention.

The honest caveats

As always, the limits, stated plainly.

  • The ability-to-Part mapping is our judgment, not an ETS formula. ETS doesn't publish which questions feed which ability, so how much of a 70% "grammar" bar belongs to Part 5 versus Part 6 is a weighting we reasoned out. But step three's shift guarantees the damage is bounded: however imperfect the mapping, it can only affect the distribution's shape — never the score derived back.
  • One test is one sample. The same person sitting the exam two weeks apart will score differently — the day's condition, the form they're handed, luck on guesses. A score report is the best single data point available, but it is still a single data point; that's one more reason the abilities are folded in rather than swapped in.
  • Reports go stale. The score you enter may be from a test three weeks ago, and you've kept practicing since. Calibration treats it as evidence about the present, which is slightly off — but far better than ignoring it or copying it wholesale, and your subsequent practice keeps updating the estimate as usual.
  • Your entered score also grades us. Every time you enter an official result, the app privately records — before calibrating anything — the gap between what it had been estimating and the real number. The previous article's promise that "the constants will be refined against real data" runs on exactly these gaps. While you calibrate the app, you're calibrating us.

Score entry lives in Handy 990's goal-consultation flow, reached from the home screen's goal card (this path runs the calibration described here immediately; from two weeks after your exam, the home screen also offers it proactively). "Settings → Enter official TOEIC score" records a result for later — it's applied at your next consultation. For the model underneath — the decayed Bayesian posterior, the uncertainty bands, why not IRT — read How your TOEIC score is estimated.

← All articles