All articles · Published 2026-08-01
Difficulty Overview: what does "as hard as the real exam" actually mean? — how we measure it, what we compare against, and what we don't claim
A technical explainer for the curious. This is the overview of our difficulty series; each Part then gets its own post. This one covers what all the later posts share: the formal definition of difficulty in testing, why nobody — us included — can measure it directly, what we use as an external reference, and the difference between what we are willing to claim and what we are not. Once you've read this, each Part's post can go straight to where that Part's difficulty actually comes from. No statistics background needed.
Why "difficulty" can't be left to instinct
Every practice question in the bank carries a label: easy, medium, hard. The field looks unremarkable, and it is the most failure-prone one in the whole bank — because nothing makes a sound when it is wrong.
If an answer key is wrong, someone reports it. If an audio file is broken, you hear it the moment you press play. But a question labelled "hard" that is actually easy gives off no signal at all: it still appears, still scores, still looks like a hard question.
And that label does two important jobs. First, it decides the mix of questions you actually practise — a practice set or a mock is only meaningful if it has a gradient from easy to hard, and that only holds when "hard" really means hard. Second, it is what makes "how to prepare for hard questions" sayable at all: if we can't say why a question is hard, all we can offer is "practise more", which is true and useless. Break difficulty into a few concrete, recognisable sources and those sources are themselves a study list.
The formal definition, and its brutal prerequisite
Psychometrics defines difficulty without ambiguity:
A question's difficulty is the fraction of real examinees who answer it correctly.
The technical term is the p-value. A question 85% of people get right is easy; one 40% get right is hard. That direct. It isn't whether an expert finds it hard — it's how many people actually got it wrong.
But the definition has a brutal prerequisite: a large number of examinees must already have answered that exact question. ETS has that data — every live item is pre-tested before it counts. No third party has it: not because we're being coy, but because no publisher, no cram school, and no app has it either.
So everyone who prints "difficulty: high" beside a question is using some stand-in. There's nothing wrong with that; stand-ins are the norm in this field. The only difference that matters is:
Is that stand-in one person's judgment, or a set of rules that can be re-run, checked, and proven wrong?
We chose the latter. The method: ignore everyone's difficulty labels, and compute from the question's own text a set of structural features that make a question harder, measured item by item. Because the features come from the text, the same program can measure our bank and any set of questions printed on paper — which is the key to the next section.
"As hard as the real exam" can honestly mean three things
The phrase gets used as if it were one claim. It's three, and only the first two are checkable by anyone today.
| Claim | Checkable? | |
|---|---|---|
| 1 | Same blueprint — same question types, same proportions, same structure per form | Yes, objectively |
| 2 | Same difficulty distribution — same easy/medium/hard split, per an external reference | Yes, but with a "judged by whom" caveat |
| 3 | Same accuracy rates — real examinees get each question right at the same rate | No, without official data or a large answer sample |
We can support (1) strongly, (2) with explicit caveats, and we state plainly that (3) is not established.
Anyone claiming (3) without a large body of answer data is guessing. That includes us — every post in this series returns to the point.
What we compare against: a six-form commercial mock set
You can't invent the markings on a ruler; they have to be aligned to something outside yourself. Ours is 全新!新制多益 TOEIC 題庫解析 狠準 6 回 (by Hackers Academia, published in Taiwan by 國際學村), a complete six-form set of practice exams widely regarded as one of the most faithful on the market.
There was one very practical reason to choose it: the book prints the publisher's own difficulty rating (high / medium / low) beside every question. That handed us over a thousand ready-made, professionally edited per-question judgments to compare against — which is exactly what you need in order to check whether your own instrument works.
How we read those ratings is worth a line. Each is printed as three dots (high = ●●●). Transcribing a thousand-plus of them by hand is too error-prone, so we didn't use text recognition at all; we used a more reliable signal: only the filled rating dots use the book's accent ink. Scan each page, count the discs in that colour, and the ratings come back. With two sanity checks (a genuine rating disc fills a specific fraction of its box; the neighbouring slots must be grey outline rings), we recovered 1,203 per-question ratings across the six forms, and verified the detector against a hand-read page.
Then the step the whole method rests on:
The same feature-extraction program runs over our bank and over the reference book's questions.
One instrument, two corpora. That's the precondition for the comparison meaning anything — we aren't grading ourselves against our own standard, we're putting both sides on one ruler.
This ruler measures groups well and individuals badly
This is the most important point in the method, and the easiest to skip, so it's worth a comparison.
Imagine a bathroom scale accurate to ±5 kg. You cannot use it to decide "did I gain a kilo this week" — the error is larger than the change you're trying to see, so the reading is close to meaningless. But if you weigh a thousand people on it and then weigh a different thousand, the two averages are still comparable: the errors run high sometimes and low sometimes, and across a thousand readings they largely cancel, so a remaining gap is real.
Our structural classifier is that scale. Judging whether one individual question is hard, it is unreliable — and that isn't us being modest, it's something we measured; the numbers and the method are in the Part 5 post (coming soon). But apply the same rules across a thousand-plus questions and the individual misjudgements scatter both ways and cancel, leaving proportions you can trust.
So the line falls here:
- We can say: the "four distinct content words" pattern is 42.3% of our Part 5 and 33.1% of the reference's. ✓
- We cannot say: this question is hard, because the classifier put it in the hardest bucket. ✗
The practical consequence is direct: we never overwrite any question's difficulty label with the classifier's output. The ruler is used where it holds up — gating the overall composition of new content — not to score single items. Every difficulty comparison in this series should be read under that limit.
The reference's limits, not papered over
- It's a publisher's mock set, not ETS material. That makes it strong evidence on format (near-zero variance across its six forms — exactly what faithfully copying the official blueprint looks like) and weaker evidence on difficulty (the questions are the publisher's own, and the ratings are its editors' judgment). We deliberately report the format comparison confidently and the difficulty comparison with reservations.
- It's only six forms. For structural features that appear on every single form, six is plenty to confirm. For fine differences in proportion, six is not.
Where things stand: our questions are not easier than the reference's
Start with the numbers most likely to be misread. Comparing the two sides' labels for "hard", it looks like we're easier nearly everywhere:
| Part | Reference rated "high" | Ours labelled "hard" |
|---|---|---|
| 2 | 20% | 18.9% |
| 3 | 12% | 21.7% |
| 4 | 18% | 16.2% |
| 5 | 33% | 10.3% |
| 6 | 28% | 25.8% |
| 7 | 29% | 21.6% |
Four of the six Parts land within a few points. That 23-point Part 5 gap looks alarming — and it's the most important lesson in the series, because when we measured by structure instead of comparing the two label sets, the conclusion flipped:
| Part 5 option pattern | Ours | Reference |
|---|---|---|
| one word family, with a local cue (easiest) | 3.6% | 2.9% |
| one word family, no cue | 29.4% | 28.4% |
| function-word / collocation choice | 24.7% | 25.6% |
| four distinct content words (hardest) | 42.3% | 33.1% |
We carry about 9 points more of the hardest pattern than the reference, not less. The two middle buckets overlap almost exactly.
In other words: the 23-point gap is an artifact of labelling, not a difference in content. Our authors tend to mark grammatically tricky items as hard, while the reference's rubric treats vocabulary and collocation load as hardest — the two sides are measuring the same thing with different rulers. Our labels are conservative; our content is not.
That's the conclusion available today, and it applies to the whole bank rather than any single item: on the one axis we can measure objectively, our questions are at least equivalent to the reference's, and carry more of the hardest pattern. We won't extend that to "we're harder than the real TOEIC" — that would need claim (3), and nobody can prove (3) yet.
What we don't claim
- We have no accuracy rates. Nothing here establishes that any question is as hard for a real examinee as a real question.
- The reference is a publisher's mock, not ETS. Format conclusions are strong; difficulty conclusions inherit an editor's judgment.
- Our listening audio is cleaner than the real test's. It's generated with neural TTS; speech rate and accent mix we control and measure, but the elision, overlapping speech, and room noise of a live recording we can't yet reproduce. That makes our listening slightly easier than the real exam in a way our own metrics cannot see — we'd rather say so than have you discover it on test day.
- We don't judge any single question's difficulty. The reason is the bathroom scale above: this ruler is trustworthy on groups, untrustworthy on individuals. Every comparison here is about the composition of two banks, never about how hard one question is.
- Two scales are still not one scale. The reference's high/medium/low is an editor's judgment; ours is an author's. Only large gaps are worth acting on; small ones are noise.
What would make this more certain
In order of value:
- Accuracy rates from learners. That's the real measurement. Per-question outcomes are already accumulating anonymously; what's missing is volume. When it arrives, this structural ruler gets recalibrated against real data — or replaced by it.
- Estimated score versus official score. Every time a learner enters an official score, we can check how closely our estimate tracked it — and that claim is worth quoting after only a few dozen samples.
What's next
With the shared groundwork in place, each Part's post can go straight to its own subject: where that Part's difficulty actually comes from, and what to do about it when you're answering.
- Part 1 — photographs; difficulty lives in verb aspect, the minimal contrast between a state and an action in progress.
- Part 2 — question–response; difficulty lives in the prompt's form and how indirectly the answer replies.
- Part 3 / Part 4 — conversations and talks; difficulty lives in reasoning questions, graphics, and speaker-intent items.
- Part 5 — single-sentence completion; difficulty lives almost entirely in how the four options relate to each other.
- Part 6 — text completion; difficulty lives in cohesion across sentences and in the inserted-sentence item.
- Part 7 — reading; difficulty lives in paraphrase distance — the correct answer almost never repeats the passage's wording.
The back half of every post reads that Part's difficulty sources backwards into an answering strategy: every device that makes a question harder also announces where the answer is hiding.