All articles · Published 2026-08-04
What makes Part 3 difficult is “who said what” plus “the answer being reworded” — what the measurements showed, and the two things we changed because of them
A technical note for the curious. This difficulty series covers one Part at a time, in order; this is the third article. Part 5 concentrates its difficulty among four options, while Part 7 spreads it across an entire text. Part 3 dialogues sit between the two: a dialogue lasting less than 90 seconds comes with three questions, and the difficulty comes both from what happens while you listen and from how the options are written when you read them. Here, we measure our question bank item by item against all six tests in a commercial practice-test book — dialogue length, number of turns, number of speakers, and question-type mix — to explain where we match and where we are deliberately harder. The measurements exposed two gaps — our questions were too shallow, and the answers were too easy to guess — and we have now closed both: we rewrote 575 questions and 434 sets of options, and wrote another 108 new dialogues. No background in statistics required.
First, what shape is Part 3?
In the real test, Part 3 has 13 dialogues, with 3 questions each, for a total of 39 questions — the largest section of the Listening test (39 of its 100 questions). Each dialogue is played only once. The three questions are printed in the test book, and you have to answer while listening.
This part of our question bank currently contains 694 dialogues and 2,082 questions (it grew from 586 dialogues and 1,758 questions while this article was being written, for reasons explained below). But how much we have is not the point. What matters is whether it has the right shape. So we extracted every Part 3 transcript from all six tests in a commercial practice-test book (there is a full section later on what this reference is and what its limitations are), then measured both banks with the same ruler:
| Measure | Reference tests (6 tests, 69 dialogues) | Our question bank (586 dialogues) |
|---|---|---|
| Words per dialogue | Mean 85 (median 87, range 45–110) | Mean 87 (median 89, range 43–148) |
| Turns per dialogue | Mean 4.5 turns | Mean 5.4 turns |
| Three-speaker dialogues | 13.0% (1–2 of the 13 dialogues per test) | 25.6% (150 dialogues) |
The first row is reassuring: the lengths are almost identical. 85 versus 87 words, with medians of 87 and 89. We have not misjudged how long a Part 3 dialogue should be. (The only difference is at the upper limit: our longest is 148 words, while the reference's longest is 110. A few unusually long dialogues form a tail of our own; they are not part of the test design.)
The second and third rows show the real differences, and both point to the same thing:
- With the same number of words, we divide our dialogues into more turns. Nearly one extra turn on average. More turns without more words means each turn is shorter and the speaker changes more often.
- We have almost twice as many three-speaker dialogues as the reference. The real test design includes 1–2 three-speaker dialogues among the 13 in each test; they make up a quarter of our bank.
Together, those two points create a “tracking cost”. And three-speaker dialogues really are harder; we measured this rather than guessed it: 29.1% of questions attached to three-speaker dialogues are labelled hard, compared with just 19.1% for two-speaker dialogues. Adding one person adds another task. You not only have to understand the content, but remember who said it, because a question may ask directly what the man suggests or what the second woman is worried about.
In other words, our Part 3 is difficult because of tracking, not because of the amount of information. Keep that conclusion in mind. The figure below showing that we are 8 percentage points harder than the reference will then cease to be a mystery and start to have an explanation.
A caveat about extraction: the reference-test transcripts were printed in the answer book without a text layer, so we reconstructed them using optical character recognition (the two-column layout had to be rebuilt from text coordinates, or both columns would be read as one line). There were 78 dialogues across the six tests, and we successfully reconstructed 69 dialogues (88%); recognition errors prevented the rest from being separated reliably. The reference figures in the table above are statistics for those 69 dialogues.
The different ways a dialogue can be made difficult
These are the main features we actually calculate or label for every Part 3 question. As in the other articles in this series, the list will later double as a set of answering strategies.
1. Number of speakers and turns — the tracking cost
Two speakers versus three, and the number of turns into which a dialogue is divided, as discussed in the first section. The difficulty of a three-speaker dialogue lies not in the amount of information but in attribution: who says a given line can change the answer to at least one of the three questions. The number of turns is the continuous version of the same problem — every time the speaker changes, your attention has to reattach.
2. Question type — what you are being asked to do
We wrote a program that infers what each question is actually testing from the wording of its stem. It ran across our entire bank, and across all 230 Part 3 questions in the six reference tests — one program, two question banks, so the comparison means something.
| Question type | Reference tests | Before rewriting | After rewriting | What you have to do |
|---|---|---|---|---|
| Detail | 35.7% | 52.8% | 35.7% | Locate a stated fact |
| Request / suggestion | 11.7% | 4.3% | 11.6% | Identify who asks whom to do what |
| Graphic question | 7.8% | 2.4% | 7.8% | Cross-reference the audio with a printed graphic to find the answer |
| Reason (Why…?) | 7.4% | 16.4% | 7.4% | Identify cause and effect, usually across two turns |
| Location | 6.5% | 2.7% | 6.4% | Identify where the scene takes place |
| Next action (What will… do next?) | 6.5% | 4.4% | 6.5% | Identify what happens after the dialogue ends |
| Inference | 6.5% | 1.1% | 6.5% | Combine two facts into an unstated conclusion |
| Speaker intention | 5.7% | 0.9% | 5.9% | Work out “What does the speaker mean by this?” |
| Problem (What problem…?) | 4.3% | 6.3% | 4.2% | Identify the obstacle being raised or complained about |
| Speaker identity | 3.5% | 1.1% | 3.5% | Identify who the speaker is |
| Main idea | 3.0% | 1.9% | 2.9% | Summarise the whole dialogue, not any one sentence |
| Purpose | 1.3% | 1.8% | 1.3% | Identify why this call or conversation is taking place |
| Time / quantity | 0% | 3.8% | 0.3% | Retrieve a single piece of information |
| Reasoning questions, total | 30.9% | 12.6% | 31.9% |
This table matters much more than the one in the previous section, because it describes not the scale but the content.
Before the rewrite, only a third of the reference-test questions were detail questions; more than half of ours were. Conversely, question types requiring an extra step — inference, intention, main idea, purpose, next action, and graphic questions — made up 30.9% of the reference but only 12.6% of ours, less than half as much. The conclusion at the time was that the real test design asked you to understand what you heard and then take another step more than twice as often as ours did.
That version has now been fixed. The fourth column in the table is the current question bank: all fourteen question types are within two percentage points of the reference, and reasoning questions total 31.9% versus 30.9%.
Closing that gap required two entirely different methods, worth spelling out:
- Eleven question types could be added by “asking something different”. The dialogue and audio stayed the same; we rewrote the question itself, asking about something different and providing a completely new set of options. Not one word of the dialogue audio changed, because each question's audio is a separate clip. The result is a genuinely different question, not a paraphrase of “Why…?”. The latter would make the figures look better without making the question harder.
- Graphic questions and speaker-intention questions could not be added that way; they had to be written from scratch. A graphic question needs the dialogue to supply only one half of information that can be cross-referenced. An intention question needs a line whose literal and intended meanings differ — we found only 2 such lines in the existing 490,000 words of dialogue. These two question types therefore arrived through 108 newly written dialogues (with newly recorded audio), which was also the slowest part of the entire project.
Two other things were fixed along the way: time / quantity questions fell from 3.8% to 0.3% (none appeared in the six reference tests, and their absence across all six is unlikely to be chance — single-point dictation such as “What time will they meet?” is more like Part 2 work); and the position of the correct option, which at one point was A in 70% of new questions, is now split equally among the four letters.
Finally, this classifier has one useful property that makes it worth taking seriously: it never saw the difficulty labels, yet it ranked the question types in the same order as those labels — 96.7% of speaker-intention questions were labelled hard, as were 58.6% of inference questions and 33.6% of reason questions, while only 0.8% of purpose questions and 0% of quantity questions were. A classifier that has never seen the answers has to reproduce this ordering before “question type” can be said to capture something real.
(The reference column also came from optical character recognition: of 234 questions across the six tests, 230 questions (98%) were reconstructed successfully. Two independent checks give us confidence that the extraction was not distorted — the program counted exactly 3.0 graphic questions and about 2.2 intention questions per test, matching the test-design quotas we had previously measured by an entirely different method.)
3. Paraphrase distance — how far the correct answer's wording is from the transcript
This is the most consequential factor in Part 3, and the central character in the second gap below.
The dialogue says: I'll revise the interview guide before tomorrow's sessions. An option that is too easy looks like this: Revise the interview guide (copied verbatim). An option on the real test looks like this: Update some materials — “revise” becomes “update”, and “interview guide” becomes “materials”. The meaning stays the same; every word changes.
The difference is that the first can be answered by matching keywords, while the second requires understanding. We calculate the wording overlap between each option and the transcript.
4. Distractor density — wrong options that look like the answer
This is the mirror image of paraphrase distance. Words that really appear in the dialogue are planted in wrong options, specifically to catch people scanning for keywords. The current averages in our bank are a wording overlap of 0.72 between the correct option and the transcript, and 0.21 for wrong options. That gap is too large — that is precisely the problem, as the next section explains.
5. Evidence position and question order
We labelled the evidence sentence for every question (the line of dialogue on which the answer is based). Every question has one, for 100% coverage; this is also why, after you answer a question in the app, that sentence is highlighted directly in the transcript.
Evidence positions let us measure something very practical: in 94% of cases, the evidence for the three questions appears in the same order as it does in the audio. The answer to the first question comes early, and the answer to the third comes late. The real test is designed this way too, so it is not a simplification on our part — it is a pattern you can use. The strategy section explains how.
6. Graphic and intention questions — fixed quotas in the test design
There was zero variation across the six reference tests: every Part 3 had 3 graphic questions and 2 speaker-intention questions. That is test design, not coincidence.
Our question bank originally had none of either type. It now has both and, more importantly than the number of questions, practice tests draw them according to that quota rather than leaving them to chance. If a question type exists in the bank but never appears in your practice tests, it may as well not exist.
What the reference comparison showed: we are harder, and we are keeping it that way
A scale cannot be invented in isolation; it has to be aligned with an external reference. We used 《全新!新制多益 TOEIC 題庫解析 狠準 6 回》 by Hackers Academia, published by International Learning Village — all six complete tests, with the publisher's own difficulty rating (high / medium / low) printed beside every question, giving us 1,200 ready-made expert judgements for comparison.
The Part 3 results:
| Low (easy) | Medium | High (hard) | |
|---|---|---|---|
| Reference tests (205 questions) | 41% | 47% | 12% |
| Our question bank | 23.8% | 54.5% | 21.7% |
We label about 8 percentage points more questions hard than the reference does.
When a gap like this appears, the first question is not “Should we adjust it?” but “Is this a difference in the content or in the ruler?” The Part 5 article in this series already demonstrated this once: its labels looked very different from the reference, but closer investigation found that our questions were not structurally easier at all. The difference lay in what each side meant by “hard”, not in the questions themselves. Distinguishing the two before deciding whether to change the content is the most useful lesson from this entire series.
The direction of these 8 points in Part 3 is the opposite of Part 5, and our response is to keep them.
There are two reasons. First, the two scales are not directly commensurable: one reflects editorial ratings and the other question writers' judgements, and a gap of 8 points lies within a range that can plausibly be explained by either noise or a real difference. Second, and more practically, if a practice bank is going to deviate from the standard in one direction, the harder side is safer for learners. Practising on harder material gives you a margin in the test room; the reverse gives you a nasty surprise.
As for why we are harder, the table in the first section has already given an answer far more solid than saying that we are simply stricter: our dialogues are no longer than the reference's, but they contain more turns, and we have twice as many three-speaker dialogues. Three-speaker dialogues themselves carry 10 percentage points more hard questions (29.1% versus 19.1%), and they account for 25.6% of our bank but 13.0% of the reference. The difference in that one proportion alone explains most of the 8-point gap in the labels.
In other words, we have not made every question more devious; our bank simply contains more of a kind of dialogue that is inherently harder. The discrepancy is not properly understood until we know that its source is the “mix”, not the “method”. If we ever want to tune it back, the thing to change is the share of three-speaker dialogues, not individual questions.
But at the time, there was a contradiction that had to be stated, and it leads into the next section: the labels said we were harder, while the question-type mix said we were shallower. In the same bank, the listening side (three speakers, many turns) skewed hard, while the asking side (more than half were detail questions) skewed easy. Both facts were measured, and both were true. This is exactly why we do not treat the difficulty labels as a conclusion: they reflect question writers' sense of “how complicated this sounds”, not “how difficult a task this question requires”.
Honest caveats
- The reference book contains a publisher's practice tests, not official ETS material. Its format is highly credible (near-zero variation across six tests is exactly what faithful replication of the official structure looks like); its difficulty ratings are editorial judgements, not real test-takers' correct-response rates.
- Part 3 difficulty labels are assigned by question writers, not calculated by a rule. This is the crucial difference between Part 3 and Part 1: every Part 1 label is recalculated from the question itself by a program and matches all 335 questions in the bank; Part 3 has no equivalent yet. Some of the features above (number of speakers, length, question type, evidence position, wording overlap) are already measured, but they have not yet been combined into a difficulty score that overrides the labels. So when we discuss the Part 3 difficulty distribution, we are discussing the distribution of a group of question writers' judgements.
- We tried to find a more objective referee and failed. The idea was to have AI models act as “simulated test-takers” on the reference tests, to see whether their correct-response rates could reproduce the publisher's printed difficulty ordering. If so, we would have a ruler capable of measuring both banks. Instead, two models an order of magnitude apart in capability scored 89.8% / 89.9%, and showed no response at all to difficulty (the correlation coefficient was approximately 0). The reason is not hard to understand: learners are held back by their command of vocabulary and collocations, both of which any serviceable model already possesses. We discarded that ruler, but think the failure is worth recording.
- The real destination is still the correct-response rate. As anonymous response statistics accumulate, every question will gradually acquire a real p-value. Until then, this is the most honest measurement we can make.
The two gaps we found ourselves have now both been closed
This section is for readers who want the whole picture. The 8-percentage-point gap in labels above is actually the least important figure in this article — it is a question of calibration. There were two real problems, both problems with the content and both found through our own measurements.
Gap one: the questions were too shallow — now fixed
This is the question-type table. Originally, 52.8% of our questions were detail questions versus 35.7% in the reference, while reasoning questions totalled 12.6% versus 30.9%. This meant that someone who completed all of our Part 3 material received only about 40% as much practice in “understanding what you hear and then taking another step” as the real test demands. They would become very good at locating a stated fact in a dialogue, but would be seriously underprepared for inference, intention, and integrative question types such as what the man asks the woman to do.
All fourteen question types are now within two percentage points of the reference, and reasoning questions total 31.9% versus 30.9%. Section 2 explained the method: 575 rewritten questions, plus 108 new dialogues and new audio. This also expanded the bank from 1,758 questions to 2,082.
One point that needs to be acknowledged honestly: this line is measured using our own classifier. The same program ran on the reference tests, so the definition is consistent between the two sides, but it ultimately works backwards from the wording of the question stem, not from “how difficult this question really is”. Matching the question types does not mean the difficulty is matched. We will know that only when actual user-response data comes in.
Gap two: the answers were too easy to guess — now fixed too
At the time, 653 Part 3 questions (37%) had correct options that could be found through keyword matching.
That was the result from the scripts/option_source_overlap.py scanner. For every question, it compares wording overlap between the correct option and the audio transcript, and flags cases where the correct option overlaps conspicuously more than the three wrong options. This was not a new discovery: a user first reported the problem, pointing out that some questions could be answered by matching keywords without understanding the audio. In response, we carried out one large rewrite, replacing 576 Part 3 options (and 1,511 across Part 3, Part 4, and Part 7 together).
Now the figure is 28 questions (1.3%), and all 28 have been retained deliberately (as explained below). Not one question that genuinely needed changing remains.
The figure came down in two stages:
- The 575 rewritten questions from the previous section removed more than half as a by-product. Every rewritten question had to pass the same scanner before it was admitted to the database, so aligning the question types simultaneously fixed this problem in a fifth of the bank. By the start of this section, the figure was already down to 462 questions (22.2%).
- The remaining 434 questions were edited individually. There were two methods, and the second is the key: rewrite the correct option as a paraphrase; when a paraphrase was still not enough, give words from the transcript to a wrong option. That is exactly the shape of a real test question — the option copied from the source is the trap, not the answer.
Here is a real example. At the end of a dialogue about a catering order, the man says, “Send me the room number before three o'clock, and I can change the order.” The question asks, “What is the woman asked to send?”
Before the rewrite, the four options were: the room number / the guest list / a payment receipt / the menu prices. Only the words “room number” appeared in the audio; not one word from any of the other three did. You did not need to understand what the sentence meant, or even know who was speaking to whom — you could simply match a phrase you heard to an option and be done.
After the rewrite, the correct option is “the location where the event will be held”: the same thing, put differently. One of the wrong options, meanwhile, became “the number of vegetarian meals” — every one of those words appears in the audio (the woman had just asked to add five vegetarian meals), and it sounds entirely plausible. The option whose words now match is the wrong one. To answer correctly, you have to understand that he wants the room number, not the number of meals.
Why not change those 28 questions? Because their answers cannot be paraphrased. The four options for a graphic question are the labels printed on the graphic (“Cone Six”, “Loading Area Four”), so they have to be reproduced as printed; place names, people's names, and amounts of money are the same. Real tests ask questions this way too. We register these cases in accepted.json, with an item-by-item reason, so the scanner's figure remains meaningful and does not force worse rewrites merely to silence the program.
In measurement terms, wording overlap between the correct option and the transcript fell from 0.72 to 0.39; for wrong options, it is 0.25. The two figures are now close, which is exactly what we said at the outset an ideal Part 3 should look like.
One more honest qualification: passing the scanner proves only that “shared vocabulary no longer gives away the answer”; it does not mean the question has become difficult. The real validation is still users' correct-response rates.
The two gaps were originally opposite sides of the same problem: we were not asking deeply enough, and we were making the answers too easy to pick out. Together, they produced a Part 3 that sounded difficult but was relatively easy to answer — exactly the contradiction from the previous section. With both sides fixed, the contradiction no longer exists.
One harder limitation remains
The register of our question bank is too uniform. Across all 1,018 dialogues and monologues (nearly 490,000 words), there are only 2 idioms, while conversational markers such as “I guess” and “to be honest” are entirely absent. We discovered this while adding speaker-intention questions. Such questions need a quotable line whose literal meaning differs from its practical meaning; we could not find one, so we had to write new material. What was missing was never merely the question type, but the looseness of natural speech itself.
This also explains why closing the first gap took so long. Intention and inference questions need more than a different way of asking; the dialogue itself must first contain something that can be inferred. Those final 108 dialogues were written from scratch, not rewritten. This remains a constraint on every new batch of content.
Once you know the structure: how to listen and answer in Part 3
Read the feature list above in reverse.
The three questions are your listening outline — always read them first
The single most valuable habit in Part 3 is to read all three question stems before the audio begins. The questions are printed in the test book; the rules allow this, and the section is designed around it. Once you have read all three, you know whether to listen for a name, a time, or one person's suggestion. Instead of “trying to remember the whole dialogue”, you are now “waiting for three specific things to appear”.
When the explanation for the previous question is playing and you still have time, read ahead. This is the most important rhythm in Part 3.
Trust the order: 94% of questions follow the audio
The answer to the first question is almost always near the beginning of the dialogue, and the third near the end. So keep moving down the page while you listen: once you hear the answer to the first question, select it and immediately turn your attention to the second. Do not wait until the dialogue is over and then return to the first question. At that point, you are fighting your memory rather than the question.
This also gives you a signal to cut your losses: if the dialogue is nearing its end and you still have no answer to the first question, let it go and save the second and third. Losing one question from a dialogue is far better than losing all three.
Treat “who said it” as information worth remembering
Question stems say the man / the woman / the second speaker. This matters especially in three-speaker dialogues. In practice: when you hear a suggestion, complaint, promise, or similar remark, make a mental note of who said it (man / woman, or which speaker). Remembering the content but assigning it to the wrong person still produces the wrong answer — and this is precisely where those extra 10 percentage points of hard questions in three-speaker dialogues come from.
Be wary of an option that is “exactly what you heard”
This is the lesson of distractor density: when the words in an option match exactly what you have just heard, be suspicious. Question writers know that people under time pressure rely on keywords, so the easiest trap is to attach words from the transcript to the wrong person, time, or event. Conversely, an option that thoroughly paraphrases the source is often the correct one — that is what correct answers on the real test usually look like.
(This is also why those 653 questions in the previous section genuinely mattered to us: on this point, our question bank is not yet fully aligned with the behaviour of the real test.)
Three question types that need special treatment
- Next-action questions (What will the man do next?) — the answer is almost always in the last one or two lines. When you hear “I'll…”, “Let me…”, or “I'd better…”, that is it. You can relax through the rest of the dialogue, but you must be awake for the last two lines.
- Intention questions (What does the woman mean when she says, "…"?) — the quoted line is usually harmless on its face; the real answer is in the line immediately before or after it. Do not stare at the words inside the quotation marks. Ask what she is responding to by saying them.
- Graphic questions — look at the graphic before the audio begins and understand what its columns represent (times matched to sessions? prices matched to plans?). The answer always comes from the audio giving you one field and the graphic giving you another; your job is to connect them. Reading the graphic first gets half the work done before you begin.
Finally: there are no second chances here
Part 3 is not played twice, and there is no going back. The entire Listening test moves relentlessly through its fixed 45 minutes; fall behind for one dialogue and you are behind for the whole dialogue. Every strategy above therefore serves just one shared purpose: to keep you on the next question, not the previous one.
In one sentence
Part 3 difficulty comes from three measurable places: how many people and how many turns you have to track, how difficult a task the question asks you to perform, and how extensively the correct answer has been reworded.
Running the same program across both our question bank and all six tests in a commercial practice-test book produced three answers. The length of our dialogues is almost identical to the real test design, but our turns are denser and we have twice as many three-speaker dialogues — explaining why we have 8 percentage points more hard questions, a difference we have chosen to keep. Our question-type mix was too shallow, with reasoning questions at only 40% of the real proportion. And 37% of our options could still be answered by keyword matching. The last two were content debt, and we wrote about them here rather than waiting until they were fixed.
Read the same list in reverse and it becomes your answering strategy: read the three questions first, follow the audio in order, remember who said what, and distrust an option that matches word for word.