All articles · Published 2026-08-02
Part 2 difficulty: the hard part isn't understanding it, it's that understanding it isn't enough — three ways a question-response item is made hard
A technical note for the curious. One post per Part in the difficulty series; this is the Part 2 one. Question-Response is only 25 items, and the question plus all three options exist only as sound — not one word is printed. This post explains where the difficulty actually sits: not in whether you understand the English, but in the step between "I understood that" and "that counts as an answer". It ends with the part most useful to you: what to listen for in the first word, and when to raise your guard. No statistics background needed.
25 items, not one word on paper
Part 2 is questions 7 through 31 of the listening test: a line is spoken, then three responses (A)(B)(C), and you pick the one that fits best.
Compared with every other section, Part 2 has one unique feature: the question itself isn't printed either. Part 1 at least gives you a photograph. Parts 3 and 4 print their questions and options. Part 2 gives you nothing — the prompt is read once, each option is read once, and then the next item starts.
That has a consequence many people never notice: you have to hold the prompt and all three options in your head at the same time. By the time you hear (C), the prompt is seven or eight seconds in the past. Hesitate on (A) and even what the question asked starts to blur.
And because each item is only two short lines, many people treat Part 2 as the simplest stretch of the listening test and prepare by drilling vocabulary and dictation. But when we pulled the bank apart and measured it, we found that what actually defeats people in Part 2 has almost nothing to do with whether you understand the words.
Look at the difference between these two items. The English in both is very plain; not one word is above high-school level.
The first:
Who approved the travel budget? (A) It starts next Monday. (B) The flight was delayed. (C) Ms. Alvarez from accounting did.
The moment you hear the first word, Who, you know the answer has to be a person. (C) gives a name and a department; the other two give a time and a flight — the wrong class of thing entirely, no close listening required. This item is over in about two seconds.
The second:
Have we settled on a date for the vendor walkthrough yet? (A) We're still waiting to hear back from their office. (B) Yes, the walkthrough went smoothly last quarter. (C) The vendors parked in the visitor lot.
Every word here is simpler than in the first item. Yet it is far harder — and it uses all three of Part 2's difficulty levers at once.
This is a Yes/No question, and the correct answer contains neither Yes nor No. (A) says they are still waiting to hear back — it never answers "is the date settled", it states a fact and leaves you to derive "not yet".
Meanwhile (B), the only option containing Yes, is wrong.
One more thing: (B) repeats walkthrough from the prompt, and (C) repeats vendor. Both wrong options engage the prompt and both sound familiar. The correct answer, (A), repeats not a single word from it.
That is the core of Part 2's difficulty, and it is not a listening problem. It is an inference problem.
Three things that make Part 2 hard
We took the whole bank (1,225 items) apart item by item and measured the structure of each. Only three things actually track difficulty.
1. The correct answer doesn't answer directly
This is the strongest lever by a wide margin — stronger than the other two combined.
On the original labels, before we reclassified anything, the "answer doesn't respond directly" feature appeared in about 5% of easy items and about 56% of hard ones — more than a tenfold gap. In other words, the people writing those items never wrote the rule down, but what they meant by "hard" was in practice exactly this.
Answering indirectly takes several forms:
- Implying the answer with a fact: "Has the order arrived?" → "The courier just dropped it off."
- Deferring: "When does this shipment arrive?" → "I'll go ask the scheduler."
- Professing ignorance: "Who approved the overtime?" → "You should ask the floor supervisor."
- Correcting the premise: "The auditor wanted the March ledger, right?" → "She asked for the reconciliations."
- Answering with a question: responding with a counter-question outright
What they share: you understood every word, and the sentence still never stated the answer. You have to walk the last step yourself.
2. The prompt has no question word
Who, When, Where come with a large built-in advantage: they announce the class of the answer. Hear When and you know to expect a time; hear Where and you expect a place. You can sketch the shape of the answer before the options arrive, then match against it.
Some prompts deny you that:
- Yes/No questions:
Has the order arrived? - Tag questions:
You reserved the suite, didn't you? - Negative questions:
Haven't the samples arrived? - Statements: not a question at all, just an assertion ("The loading ramp is icy again this morning.")
These add up to under a fifth of our bank, but their density in hard items is several times what it is in easy ones. The reason is simple: you have lost the tool that narrows the field in advance.
Tag and negative questions add one more wrinkle. The Haven't in Haven't the samples arrived? is negative, but the thing being asked is positive — did the samples arrive. Plenty of people flip Yes and No right here.
3. The wrong options are about the same thing
For an option to interfere with you, it has to relate to the prompt. If two of the three options are about something else entirely, the item is only testing whether you caught a keyword.
We measured, for every item, how many wrong options genuinely engage the prompt's content. Hard items average more than easy ones — and this measure has a special role: it is the gate we use to block fake hard items.
We found 27 items in our own bank labelled "hard" in which not one of the three options offered any interference at all. They wore a hard label and were in practice two-second items. The current rule demotes every one of them automatically.
A misconception worth dismantling: sound-alike traps
Almost every Part 2 study guide teaches the sound-alike trap — the prompt says hire, a wrong option says higher; the prompt says contract, a wrong option says contact. Nearly identical in sound, completely different in meaning.
It is a real technique, and an effective one.
But when we went to measure our own bank, we found something unflattering: not one item in 1,015 had one.
Not "few" — zero. More precisely: 303 items had a trappable word in the prompt (hire, fare, personnel…), and not one of them put the matching sound-alike into a wrong option. The material was there the whole time; the technique was never once used.
That is our own oversight, not a discovery. Both later batches deliberately used the technique, and there are now 76 items carrying a properly wired sound-alike trap.
Worth noting: this trap has one easy way to get it wrong. The sound-alike has to sit in a wrong option. Drop it into the correct answer by accident and the trap becomes a gift — the listener who mishears is rewarded instead. Our checking script blocks that case specifically.
Checking against a commercial mock test: 150 items
Everything above is our own ruler. Whether that ruler points in the right direction has to be checked against something outside.
We have a commercially published six-test mock set whose explanation volume prints a difficulty rating (low / mid / high) for every single item. Part 2 is 25 items per test, so six tests is 150 items — a sample four times larger than the 36 our Part 1 post could draw on.
Reading all 150 printed ratings off the page and setting them beside our bank — the first comparison did not look good. Our bank was 1,075 items then, and one band was plainly off: medium sat at 37.1% against the reference's 48.0%, the only row outside the confidence interval. Our bank was more bimodal than the reference: items leaned easy or leaned hard, and the middle ran thin. Hard, meanwhile, was 25.9%, just 0.5 points below the interval's upper bound of 26.4%.
So we wrote 150 medium items in response. After that:
| Reference (150) | Ours (1,225) | |
|---|---|---|
| low / easy | 32.7% | 32.5% |
| mid / medium | 48.0% | 44.8% |
| high / hard | 19.3% | 22.7% |
All three bands now land inside the reference's 95% confidence interval. Easy is near-identical (reference 32.7%, ours 32.5%); medium moved from 37.1% outside the interval to 44.8%, inside the 40.2%–55.9% band; and hard fell from 25.9% to 22.7% purely because the denominator grew, so it is no longer sitting half a point from the ceiling.
The order of those events is the part worth stating plainly: we did not tune the ratios and then go looking for a reference. We compared first, found the middle too thin, and wrote to fill it. All 150 were medium; not one hard item was added.
Incidentally, when the reference's explanations discuss wrong options, they routinely name "this pair of similar-sounding words" outright — several items in every one of the six tests are marked that way. That is precisely the sound-alike trap we had used zero times. It is standard equipment in a real commercial paper, not a theory.
Read it backwards: how to listen to Part 2
Everything so far has been about how items are made hard. This section is the part most useful to you: what to change once you know.
The first word decides everything
The first word of a Part 2 prompt is worth far more than any other word in it.
Hear Who / When / Where / Which / Why / How — fix the class of the answer in your head immediately. Person, time, place, thing, reason, method. Once it's fixed, any option of the wrong class needs no close listening at all. These are the largest group in our bank and almost all of them sit in easy and medium.
Hear something that is not a question word — Has / Are / Didn't / Haven't, or no interrogative shape at all — and that is your signal to raise your guard. These items give no advance warning, and they cluster heavily in the hard band.
Don't wait for Yes or No
This is the single habit most worth breaking in Part 2.
Hearing a Yes/No question and going hunting for Yes or No among the options is a natural reflex, but on hard items that reflex walks you straight into a wrong option — because the writer knows you will hunt that way, and puts the Yes in a wrong option.
The right move: when you hear a Yes/No question, ask yourself not "which option has Yes" but "which option told me the answer". "We're still waiting to hear back from their office" has no Yes, and it told you.
Treat a repeated word as a trap first
If the prompt says budget and one option also says budget, that repetition is usually a trap rather than a clue.
This is Part 2's most common interference technique: use a word you are certain you heard, so the option sounds familiar. The real answer often repeats nothing at all — as in the item above, where the two wrong options repeated walkthrough and vendor while the correct We're still waiting to hear back from their office repeated nothing.
The same logic covers sound-alikes: hear a word that is close to one in the prompt but not quite the same (higher against hire), and raise your guard.
Tag questions: answer the fact, not the negative
With sentences like You reserved the suite, didn't you? and Haven't the samples arrived?, don't get tangled in the didn't and the Haven't.
Rewrite them in your head in the plainest possible form — "was the suite reserved", "did the samples arrive" — then handle them exactly like a Yes/No question. The negative auxiliary does not change what is being asked.
Hear all three before deciding
Because neither the prompt nor the options are printed, many people commit as soon as an option "sounds right", and then relax.
Hard Part 2 items often place the most answer-like wrong option at (A). By the time you reach (C) and realise that is the correct one, the impression (A) left is hard to shake. Hearing all three before deciding costs only a few seconds.
Honest caveats
- The reference is a publisher's mock test, not official ETS material, and its difficulty rating is an editor's judgement, not the proportion of real candidates who answered correctly. The difficulty overview covers this fully.
- 150 items only supports coarse conclusions. What the table above can say is "all three bands are broadly comparable"; it cannot support finer differences. The intervals are not narrow either — the medium cell spans 40.2%–55.9%, over 15 points wide.
- Today's three levers align perfectly with difficulty because the current classification is computed from those three levers. The real evidence is the set of numbers from before the reclassification, when labels were assigned by feel: "answer doesn't respond directly" appeared in about 5% of the items they called easy and about 56% of the ones they called hard. That tenfold gap is the finding; today's perfect alignment is just a definition.
- We did not count the reference's sound-alikes item by item. "Several in every test" is an observation noted in passing while reading, not a figure obtained by scanning all 150 explanations, so no percentage is given here.
- These proportions describe our bank, not the real exam. They explain how we construct hard items; they make no claim that real ETS papers are composed this way.
- Our listening audio is cleaner than the real test. The difficulty overview covers this more fully: synthesised speech lacks the elision and background noise of live recording, so our listening is slightly easier than the real exam along an axis we cannot measure ourselves.
- "Does a wrong option engage the prompt" is computed from literal word overlap, so it misses options that share no word yet still compete semantically. It is a coarse ruler; we use it only to block obvious fake hard items, never to reject any individual item.
In one sentence
Part 2's difficulty is not in the vocabulary; it is in the step between "I understood that" and "that can serve as the answer". The two most effective adjustments: if the first word is a question word, fix the answer class immediately; if it isn't, know the item will be harder — and stop waiting for Yes or No to appear. Go find the option that told you the answer, even if it repeats not a single word.