All articles · Published 2026-08-04
Part 7 difficulty is measured, not guessed — and once you see the structure, you know how to answer
A technical explainer for the curious. The difficulty series runs one post per Part, published in order; this is the seventh and last. It takes on the hardest part of TOEIC to pin down — Part 7, reading comprehension — and explains how we turned its difficulty into something measurable, how we calibrate that ruler against a commercially published mock-exam set, and (most useful to you) how seeing exactly *how* a question was made hard tells you where to spend your time. No statistics background needed.
Why Part 7 difficulty is the hardest to pin down
When a Part 5 question is hard, you can usually point at one grammar point and say "that's it." When a Part 1 question is hard, you can point at a sentence pattern. Part 7 has no such single point: one passage, three or four questions, and the difficulty is smeared across the length of the text, the wording of the options, and where the evidence is hiding. Ask ten test-takers why a passage was hard and you'll get ten answers that are all correct — and all only partly so.
For a practice app, this is a problem that has to be solved, for the reason this series has argued all along: the difficulty label decides the mix of questions you actually practise, and something that can't be articulated can't be checked or calibrated either. If "hard" is just how the question writer felt that day, then the so-called difficulty distribution is just a distribution of feelings.
So we took Part 7 difficulty apart and measured it.
The definition of difficulty, and its stand-in
Psychometrics has an unambiguous definition of difficulty: a question's difficulty is the fraction of real examinees who answer it correctly (the term is p-value). A question 85% of people get right is easy; one 40% get right is hard. That direct.
But the definition has a brutal prerequisite: you need a large number of examinees to have answered that exact question. ETS has that data; no third party — not us, not any publisher — does. So everyone who prints "difficulty: high" next to a question is using some stand-in. The only difference is whether that stand-in is one person's judgment, or a set of rules that can be re-run and verified.
We chose the latter. The approach: identify a set of structural features that make a question harder, and measure them question by question. Because the features are computed from the question's text, the same ruler can measure our bank — and any set of questions printed on paper. That point becomes important below.
The directions a Part 7 question can be made hard from
Here are the main features we actually compute for every Part 7 question. This list is also a preview of the second half of this article: every direction a question can be made hard from corresponds to a way of spending your time in the right place.
1. Question type — what kind of work you're being asked to do
Part 7 questions are not one kind of question but a spectrum. Ordered by how much processing the answer requires, shallow to deep:
| Type | What you have to do |
|---|---|
| Detail (What is indicated about…?) | Locate one fact in the text |
| Main idea / purpose (What is the purpose of…?) | Summarise the whole passage, not any single line |
| Vocabulary (The word "…" is closest in meaning to) | Choose a sense from context, not from the dictionary |
| Sentence insertion (choose position [1]–[4]) | Judge how one sentence connects to what surrounds it |
| NOT questions (What is NOT mentioned…?) | Verify three options are true; the leftover is the answer |
| Inference (What is suggested/implied…?) | Combine two stated facts into a conclusion that isn't written |
| Cross-text (double/triple passages) | Apply a rule from one passage to a case in another |
The further down the list, the more seconds a question eats. We didn't invent this ordering — it is simply "how far the answer sits from the surface of the text."
2. Paraphrase distance — the real heart of Part 7
This is the single most important feature in all of Part 7. The correct answer almost never repeats the passage's wording.
The passage says: Refunds are issued within five business days of our receiving the returned item. An easy question lets the key read: Within five business days (verbatim). The real exam's key reads: About a week after the store gets the product back — "five business days" has become "about a week," "our receiving the returned item" has become "the store gets the product back." Every keyword has been swapped out; not a shade of the meaning has changed.
For every question we compute the word overlap between the options and the passage's sentences. The lower the overlap — the more thoroughly the key has been "translated" — the harder the question, because you cannot find it by scanning for keywords.
3. Trap density — wrong options that look like answers
The mirror image of paraphrase distance: how many words that really do appear in the passage show up in the wrong options. An option that lifts "five business days" verbatim but attaches it to the wrong event (say, the shipping time instead of the refund time) is lethal to anyone reading by scan. The more wrong options carry the passage's own words, the more dangerous the question.
4. Evidence position — where the answer is buried
Evidence in the first paragraph and evidence in the last paragraph cost different amounts; evidence concentrated in one sentence and evidence scattered across two paragraphs cost more different still. We record the paragraph position of each question's evidence. Incidentally, you can see this data in the app: after you answer, the evidence sentence is highlighted right in the passage.
5. Set shape — how many words you read per question
Single, double, or triple passage, plus total text length. Measuring real exam forms gives a benchmark that's easy to remember: in Part 7, each question carries roughly 84 words of reading behind it. A 300-word passage with three questions is a standard load; the same three questions on 500 words, and the reading cost alone has pushed the difficulty up.
To be concrete about what "long" means: on the real exam, a single passage's length tracks its question count — a 2-question document (a text chain, a receipt) runs about 110–200 words, a 3-question one about 180–300, and a 4-question article or formal letter 250–400. Our bank enforces the same ranges: a new set whose word count falls outside them doesn't get in.
While we're here, one thing you may already have noticed in the app deserves saying plainly: the bank also keeps a number of deliberately shorter Part 7 sets, used in everyday practice only — friendlier when your commute comes in five-minute pieces, or while you're still building reading stamina. Mock tests draw exam-length sets only, so your mock experience and score estimate are never diluted by short passages. Down the road we may also add "baby TOEIC" starter questions for learners still fighting at a lower level — first finish the reading, then race it.
6. Document kind — what sort of text this is
Notice, e-mail, advertisement, article, online chat. They are not interchangeable: an expository article is naturally harder than a text-message chain, while chat passages carry their own signature hard question in return (the intent question: "What does she mean when she writes…?").
Put these features together and every question gets a difficulty score, which is then cut into easy / medium / hard. Only one problem remains: where do the cut points belong?
Calibrating the ruler on a commercially published mock-exam set
You can't invent the scale's markings yourself; they have to be aligned to an external reference. Ours is Hackers TOEIC Reading/Listening Mock Tests — 《全新!新制多益 TOEIC 題庫解析 狠準 6 回》 (by Hackers Academia, published in Taiwan by 國際學村) — a complete six-form mock set widely regarded as one of the most faithful on the market. Choosing it had one very practical reason: the book prints the publisher's own difficulty rating (high / medium / low) next to every question, which gave us 1,200 ready-made, professionally edited per-question judgments to compare against.
Then the key step: the same feature-extraction program runs over our bank and over the reference book's questions. One instrument, two corpora — that is what makes the comparison meaningful. The scale's cut points are calibrated on the proportions the reference book itself prints: first the ruler is made to read out the publisher's own answer on real mock items, and only then is it used to read our bank. And before new content enters the bank, it must pass automated gates: its length must land on the real exam's reading-per-question budget, and its difficulty profile must land within tolerance of the reference profile.
The reference also tells us something useful to any test-taker along the way: the publisher itself rates Reading as far harder than Listening. Across the Listening parts, 34–45% of questions are rated "low" (easy); in Reading, only 10–12% are — and 29% of Part 7 questions are rated "high." If Reading has always felt like the harder grind to you, that isn't you. It's the shape of the exam.
The honest caveats
- The reference is a publisher's mock exam, not ETS material. On format it is extremely credible (near-zero variance across its six forms — exactly what faithful copying of the official blueprint looks like); on difficulty it carries an editor's judgment, not real examinees' accuracy rates.
- We make claims at the distribution level, not the single-question level. We ran a per-question validation on Part 5 once: structural features explain very little of the difficulty of one individual question. So the correct use of this instrument is comparing and aligning the difficulty composition of two banks — not certifying any single question's score as precise. Per-question scores carry error bars, and are never the sole basis for judging a question.
- The real endpoint is accuracy rates. As anonymous answer statistics accumulate, every question gradually earns its true p-value, and this structural ruler will be recalibrated against real data. Until then, this is the most honest measurement we know how to make.
Once you see the structure: how to answer efficiently
Now read that feature list backwards. Every direction a question can be made hard from is an instruction about where your time should go.
Read the stems first, classify first, then decide how to read the passage
When you get a question set, spend five seconds sorting each question into the type table above. That step decides your reading plan:
- Main idea / purpose — the evidence is almost always in the first paragraph (for letters, the opening line or two) plus the last. You can answer without close-reading the middle, so this should be the question you knock off the moment you finish paragraph one.
- Detail — pull anchor words from the stem (dates, amounts, and names are the most reliable, because they can't be paraphrased), scan to that neighbourhood, then slow down and read closely. Remember paraphrase distance: you are looking for where the stem's meaning appears, not the stem's words.
- Vocabulary — substitute each of the four options back into the original sentence. The key is the one that holds up in this sentence, and it's often not the word's most common meaning. This type doesn't require understanding the whole passage — it has the best return per second in all of Part 7.
- Inference — the answer is a combination of two facts, and those two facts are often a paragraph apart. Watch for the trap in the opposite direction too: the passage says customers are invited to attend, the option says customers will attend — an option that outruns what the passage supports is wrong no matter how plausible it sounds.
Price NOT questions at triple
"Which of the following is NOT mentioned?" costs not one question but three — you must find evidence in the text for each of three options; the leftover is the answer. It's worth doing, but do it last: once you've already gotten to know the passage answering the set's other questions, verifying three "is mentioned" checks goes much faster. Get the order wrong and you'll read the same passage twice for nothing.
Treat verbatim-copy options with suspicion
This is what the trap-density feature teaches, and it deserves its own rule: when an option matches the passage word for word, suspect it first. Question writers know that people in a hurry scan for keywords, so the cheapest trap is to attach the passage's own words to the wrong subject, time, or condition. Conversely, an option that thoroughly rephrases the passage deserves your serious attention — the key usually looks like that.
Sentence insertion: let the pronouns and connectives do the work
The sentence you're given almost always contains a hook: These changes (the preceding text must have mentioned plural changes), However (the previous sentence must point the opposite way), He then (there must already be this person and a prior action). Find the hook in the sentence first, then look for the position in the passage it can latch onto — two of the four candidates usually eliminate themselves on the spot.
Cross-text questions: the rule is in one passage, the case in another
In double and triple passages, some questions are unanswerable from any single passage. Their structure is remarkably fixed: one passage states a rule (free shipping over NT$800, 10% member discount, prices change after March), another gives a concrete case (an invoice for NT$740, one member's order, an inquiry sent in April). Your move: while reading the first passage, mark every number, date, and threshold you see — cross-text answers hang on those points almost every time. When a question asks "how much will this customer actually pay" and the amount isn't written in any one passage, that's the test asking you to connect two of them.
And finally, time — because Part 7 is first of all a speed test
The Reading section is 100 questions in 75 minutes, and Part 7 is 54 of them. The usual pacing advice is to hold Part 5 under 10 minutes and Part 6 under 8, leaving about 55 minutes for Part 7 — an average of one minute per question, reading included. Combined with the ~84 words of reading per question, that demands a steady reading speed plus all of the "never read it twice" discipline above.
This is also why, in our score-estimation post, the app simulates the exam clock using your actual seconds-per-question: a question you never reach scores like a guess. In Part 7, skipping a whole four-question set costs more than spending thirty extra seconds on one hard question — if you must surrender something, surrender single questions (NOT questions especially), never a whole set.
In one sentence
Part 7 difficulty is not a fog. It is the sum of features that can be measured question by question — question type, paraphrase distance, trap density, evidence position, reading load, document kind — and we calibrate that measurement against a six-form commercial mock set that prints its publisher's own per-question ratings, so that our bank's difficulty composition is aligned to an external, checkable reference rather than to our own feelings. Read the same feature list in reverse and it becomes your answering strategy: every device that makes a question harder also announces where the answer is hiding, and where your next minute should go.