All articles · Published 2026-08-04
Part 5's difficulty hides in the four options — the four measured patterns, and how to answer each
A technical explainer for the curious. The difficulty series runs one post per Part, published in order; this is the fifth. In the Part 7 post the difficulty was spread across a whole passage; Part 5 is the exact opposite — one sentence, four options, and the difficulty sits almost entirely in <strong>how the options are arranged</strong>. This post explains how we turned that into something measurable, how we calibrate it against a commercial mock set, how we measured our own instrument, found it wasn't accurate enough, and abandoned a change we had planned — and, most useful to you, what to do with each pattern once you can see it, and how many seconds it deserves. No statistics background needed.
A question type that looks structureless
Part 5 is the first 30 questions of the Reading section: one sentence, one blank, four options. No passage, no context, no paragraph to go back to. Because of that, most people prepare for Part 5 by "learning more words and more grammar" — treating it as one undifferentiated mass.
But it does have structure, and the structure is right in front of you: the four options themselves.
Compare these two:
The new inventory system will allow staff to track shipments more ___. (A) accurate (B) accuracy (C) accurately (D) accurateable
Employees who wish to take unpaid leave must ___ a formal request at least two weeks in advance. (A) comply (B) submit (C) respond (D) apply
In the first, the four options are four forms of the same word. You don't even have to understand the sentence — the more sitting before the blank has all but settled it: an adverb is required.
In the second, the four options are four independent verbs. No grammatical clue helps you at all; you have to understand what the sentence means, and know that submit a request is the idiomatic pairing while apply a request is not.
One question each, the same 20-second budget — and they ask you to do completely different things. That is where Part 5's difficulty lives, and it can be measured question by question.
How we measure it (the one-minute version)
Difficulty has a formal definition in testing: a question's difficulty is the fraction of real examinees who answer it correctly. The problem is that this requires thousands of people to have answered that exact question first, and nobody outside ETS has that data. So everyone who labels difficulty is using a stand-in; the only difference is whether that stand-in is one person's judgment or a set of rules that can be re-run and proven wrong.
We chose the latter, and calibrated it against a commercial six-form mock set — a book that prints the publisher's own difficulty rating beside every question, which handed us over a thousand professional editorial judgments to compare against.
The full method, why that reference was chosen, and its limits are in the difficulty overview. This post covers only Part 5's own part: ignore everyone's difficulty labels, and look only at how the four options relate to each other.
The four measured patterns
Here is the classification actually produced across all 1,216 of our Part 5 questions. The list doubles as the answering strategy in the second half: each arrangement corresponds to a completely different thing you should do.
1. One word family, with a decisive cue beside the blank (easiest, 3.6%)
The four options are different parts of speech built on one root, and a word near the blank settles the answer directly — than, more, most, very, or a be-verb, auxiliary or to in front of it.
It is ___ than before for small companies to profit. (A) hard (B) hardly (C) harder (D) hardest
than is right there, so the answer is the comparative. This type requires no understanding of the sentence's meaning; three words either side of the blank are enough.
2. One word family, no cue (29.4%)
Still forms of one word, but nothing nearby helps you.
All staff must complete the mandatory ___ training before handling personal data. (A) comply (B) complied (C) compliance (D) compliant
You have to work out what the slot needs yourself: training is a noun, it wants a modifier in front, and compliance training is a noun-modifying-noun compound. That takes structural analysis — but you still don't need to know what "compliance" means.
3. Function-word / collocation choice (24.7%)
All four options are prepositions, conjunctions, conjunctive adverbs, or the particle after a verb.
Employees are entitled ___ three weeks of paid vacation per year. (A) for (B) at (C) to (D) of
This type isn't testing sentence meaning, it's testing a fixed pairing: be entitled to. What you're looking for isn't what the sentence says, it's which preposition the word before the blank (here, entitled) habitually takes.
Phrasal verbs belong here too, and they're the hardest members of the group:
(A) carry on (B) carry over (C) carry out (D) carry off
All four share one verb; only the particle changes. (An aside: our own classifier used to misfile these as the easiest type, because all four options share a root — we fixed that in July 2026. A phrasal verb isn't word form, it's collocation.)
4. Four independent content words (hardest, 42.3%)
Four unrelated words of the same part of speech, decided only by meaning and collocation.
The sales team exceeded its quarterly ___ by fifteen percent. (A) limit (B) ceiling (C) target (D) boundary
All four are nouns, all four are about limits or ceilings, and all four are grammatical. Only exceed a target is natural business English. This type has no shortcut — it tests vocabulary and collocational feel at once, and it's the most time-consuming of the four.
Calibrating the ruler on a commercial mock set
The classification alone isn't enough — you also need to know whether the proportions are right. The key step: the same program runs over our bank and over the reference book's questions. One instrument, two corpora; that's what makes the comparison mean anything. (Which reference, why it was chosen, and its limits: the difficulty overview.)
| Option pattern | Ours | Reference |
|---|---|---|
| one word family, with a cue (easiest) | 3.6% | 2.9% |
| one word family, no cue | 29.4% | 28.4% |
| function word / collocation | 24.7% | 25.6% |
| four independent content words (hardest) | 42.3% | 33.1% |
Look at that last row, because it may be the opposite of what you expect:
In the hardest of the four patterns, we are at 42.3% and the reference is at 33.1% — we carry 9 points more.
The two middle buckets overlap almost exactly (29.4% vs 28.4%, 24.7% vs 25.6%), and the easiest sits near 3% on both sides. In other words, on the one axis we can measure objectively, our Part 5 is not easier than this commercial mock set; in the pattern that leans hardest on vocabulary and collocation, we carry more of it.
That's worth stressing, because looking at labels alone gives exactly the opposite impression — the reference rates 33% of its Part 5 at the top difficulty, and we label only 10% "hard". A 23-point gap. The next section is the resolution of that contradiction.
We measured our own instrument, found it wanting, and dropped a planned change
This deserves its own section, because it's the most important failure in the whole method.
The resolution of that contradiction — labels saying we're much easier, structure saying we're equal or harder — is that the gap is in the labelling, not the content. Our authors tend to mark grammatically tricky items as hard (subjunctives, participle clauses), while the reference's rubric treats vocabulary and collocation load as hardest. The two sides were measuring the same thing with different rulers.
Following that conclusion, the obvious next step was: since the structural classification is more objective than hand-applied labels, use it to overwrite all 1,216 difficulty labels.
We didn't. Because before touching anything, we did the thing that ought to be done first: run the classifier against the reference's printed per-question ratings and see whether it reproduces a professional editor's judgment.
The result: only half right.
Splitting the reference's rated questions into our four buckets, and looking at the average rating an editor gave each bucket (1 = easiest, 3 = hardest):
| Our bucket | Editor's average rating |
|---|---|
| one word family, with a cue (easiest) | 1.75 |
| one word family, no cue | 2.12 |
| function word / collocation | 2.37 |
| four independent content words (hardest) | 2.39 |
The good news is that the ordering is exactly right — our easy-to-hard sequence runs the same direction as the editor's judgment, and the easiest bucket really is markedly easiest (1.75 against a book-wide average of 2.26).
The bad news is the last two rows: 2.37 and 2.39, a gap of 0.02. Concretely: take one question from "function word / collocation" and one from "four independent content words", and the chance the editor rated the second one harder is about a coin flip. We thought we had separated "medium" from "hardest"; in the editor's eyes, those two piles are near-identical in difficulty.
That's why the change was rejected. Compared in bulk, errors across a thousand questions cancel each other out, so using it to compare the overall composition of two banks holds up; but for judging whether one individual question is hard, it can't tell. Using it to overwrite 1,216 labels would mean taking a ruler that separates only "easiest" from "everything else" and using it to assign three grades to every item.
So the classifier was redeployed where it does hold up: as a gate on new content — before new material enters the bank, its composition must fall within tolerance of the reference profile. Same tool, used at the resolution where it actually works.
We're writing this down because a measurement method that has never overturned its own conclusion usually means it has never really been tested.
The honest caveats
- The reference is a publisher's mock set, not ETS material. On format it's highly credible; on difficulty it carries an editor's judgment, not real examinees' accuracy rates.
- The proportion table above is for comparing two banks, not for judging a single question. The reason is the previous section: this ruler separates "easiest" from the rest, but can't separate the two middle buckets from the hardest one. Across a thousand items the individual misjudgements cancel and the proportions stay trustworthy; on one item they don't. So we never overwrite any question's difficulty label with it.
- The real endpoint is accuracy rates. As anonymous answer statistics accumulate, each question gradually earns its true p-value, and this structural ruler gets recalibrated against real data. Until then, this is the most honest measurement we can make.
Once you can see the pattern: what to do with each
Now read that classification backwards. The single most important Part 5 habit is: look at the four options first, then decide whether you need to read the sentence at all.
That's worth repeating, because it runs against most people's instinct. Most people start Part 5 at the beginning of the sentence and look at the options afterwards. But one glance at the options tells you which type this is, and the type decides whether you need to finish the sentence.
You see four forms of one word → find the part of speech, not the meaning
The four options look alike (accurate / accuracy / accurately / accurateable), so it's a word-form item. What to do:
- Look at the three words either side of the blank first.
than,more,most,as … as→ comparative or superlative; answer on the spot. A be-verb or auxiliary before it → verb form.the,a, or an adjective before it → a noun is needed. - If there's no cue, work out the blank's role in the sentence: subject? verb? modifier?
- Don't stop to think about what the word means. You can answer this type without recognising the root at all — it's the only Part 5 type where vocabulary size is irrelevant, and it's a third of all questions.
Target time: under 10 seconds.
You see four prepositions / conjunctions / particles → find what it attaches to
All four options are function words like by / with / among / during, so it's a collocation item. What to do: don't read the sentence for meaning — find the content word before the blank, usually a verb, adjective or noun, and ask which one it habitually takes.
be entitled ___ → to. comply ___ → with. account ___ → for. These are memorised as pairs, not reasoned out.
Phrasal verbs (carry on / over / out / off) are the hardest of this group, because all four share a verb and their meanings diverge wildly. There is no exam-room technique for these — only accumulated exposure. The good news is that they don't appear often.
Target time: 15 seconds. You either know it or you don't; if you don't, pick one and move on rather than stalling.
You see four unrelated content words → this is where your time belongs
Four different words (limit / ceiling / target / boundary), and grammar is no help. What to do:
- Read the whole sentence. This is the only type that requires it.
- Find what the blank has to pair with — usually the verb before it or the noun after. In the example above the key is
exceeded, not the rest of the sentence. - Substitute all four options back in and say them. Collocational wrongness is most audible when spoken.
- If two options both work, choose the one more common in a business context. TOEIC's answer is almost always the thing people actually say in an office.
Target time: 30 seconds; if it hasn't come, mark it and skip.
Pacing: Part 5 is where you save time for Part 7
The Reading section is 100 questions in 75 minutes, and Part 5 is 30 of them. The usual advice is to hold Part 5 to under 10 minutes — an average of 20 seconds per question.
That looks tight, but it's reasonable once you break it down: a third of the questions are word-form items that should be done in 10 seconds, and a quarter are collocation items that take seconds when you know them. Every second you save belongs to Part 7. Spending an extra minute on one Part 5 question, wrestling with a word you don't know, can cost you a whole four-question set at the end of Part 7 — and a question you never reach scores exactly like a guess.
So the real Part 5 discipline isn't accuracy, it's knowing when to let go. Four independent content words, two of which you don't recognise, and 30 seconds gone with no instinct forming — that's the moment to pick one, flag it, and move.
In one sentence
Part 5's difficulty isn't in the sentence, it's in how the four options are arranged: one word family with or without a cue, a function-word collocation, or four independent content words. Those four arrangements can be measured question by question, and we calibrate them against a commercial mock set that prints its publisher's own per-question ratings — the same measurement that showed the ruler is too coarse to judge any single question, which is why a planned bank-wide relabel was abandoned. Read that same classification backwards and it becomes your answering order: look at the options first, then decide whether to read the sentence; and give every second you save to Part 7.