All articles · Published 2026-08-04

The Difficulty in Part 4 Is Not in the Monologue, but in the Questions: Four Measured Question Types, and How to Answer Each One

A technical note for the curious. This difficulty series has one article for each Part; this is the Part 4 article. In Part 5, difficulty centers on how the four options are arranged. Part 4 is almost the opposite: the same monologue, paired with different questions, can differ by several difficulty levels. This article explains how we turned “what makes a question difficult” into something measurable, how the answer we originally expected measured out <strong>backwards</strong>, what we saw after comparing against a commercial practice-test set, and, most useful to you, once you understand the four question types, how to listen for each one and when to let go. No statistics background required.

A Question Type That Looks Like Pure Listening Endurance

Part 4 is questions 71 to 100 of the Listening test: ten monologues, three questions each. The content includes announcements, voice messages, broadcasts, advertisements, and tour introductions.

It differs from Part 3 in only one way, but that one thing matters: there is no second person. In Part 3, if you miss one sentence, the other speaker often confirms the same idea again in different words; Part 4 has no such safety net.

That is exactly why many people prepare for Part 4 by “listening more, and listening faster”: treating it as a mass of sound that can only be muscled through by ear.

But it has structure, and the structure is not in the monologue; it is in the questions. Look at two questions paired with the same monologue:

Monologue: Welcome to the Fenwick Harbor walking tour. We'll start at the old customs house and finish out at the light station. I hope you all brought your walking shoes. The path out to the point is uneven in places, and there's no shuttle back.

Question A: Where does the tour end? (A) At the customs house (B) At the light station (C) At a photography studio (D) At the water cart

Question B: Why does the speaker say, "I hope you all brought your walking shoes"? (A) To warn that the route is long and rough (B) To recommend a nearby shoe shop (C) To explain a change to the dress code (D) To apologize for a cancelled shuttle

What Question A asks you to do is: find one sentence in the monologue. finish out at the light station is right there; once you hear it, you are done. It does not matter if you do not fully understand what a customs house is.

Question B asks you to do something completely different: you can see that sentence, and you know every word in it, but the test was never asking for its literal meaning. It is asking why he says that sentence in that position: and the answer is in the next sentence (uneven path, no shuttle), not in the sentence itself.

The same audio file, the same 30 seconds, but the two questions demand tasks of a completely different order. That is where Part 4 difficulty lives, and it can be measured question by question.


How We Measure It (The One-Minute Version)

The formal definition of difficulty, why no one except ETS can obtain it, and why everyone therefore has to use proxy metrics are already covered in “Difficulty Overview.” Here we will cover only the Part 4-specific part: what proxy metric we use, and why it is defensible.

The method is to break “what extra work this question asks you to do” into seven features that can be judged question by question:

Feature How It Is Judged
Evidence falls late The sentence containing the answer appears in the final third of the monologue
Semantic turn The answer sentence itself contains markers such as however / actually / instead that overturn what came before
Requires synthesis The answer requires information from two or more sentences, not a single sentence alone
Paraphrase distance More than half of the content words in the correct option do not appear in the answer sentence: the answer has been paraphrased, not copied
Distractor density Two or more incorrect options use words that really appear in the monologue
Inference-type question The question type is quoted meaning, inference, graphic comparison, or next action
Requires calculation The answer is a calculated number, and it is never spoken from beginning to end

A question is labelled difficult only if it has three or more features, and at least one of them is one of these three: semantic turn, inference-type question, requires calculation. That condition is necessary. Without it, this tier gets filled with questions that are only time-consuming, not truly difficult, such as “the monologue is long and the answer is late.”

This rule has a minimum test it has to pass: the order it produces must line up with the difficulty that item writers marked before they had ever seen this rule. The actual result:

Number of features Share labelled difficult by item writers
0 7.9%
1 14.6%
2 35.6%
3 51.3%
4 100%

Monotonically increasing. The two sides are independent, so the line itself is evidence.


The Four Question Types We Measured

Among the seven features, the strongest determinant of difficulty is question type: the action the question asks you to perform. Below is the classification we ran across all Part 4 questions. The percentage in parentheses is its share of the question bank; “share labelled hard” is the item writers’ independent label.

1. Gist questions (18.9%, share labelled hard 0%)

These ask for the purpose, main idea, who the speaker is, or where the speaker is.

Question: What is the purpose of the message? (A) To advertise a new service (B) To announce a planned power interruption (C) To collect a payment (D) To report a street light fault

The answer is almost always in the first one or two sentences. The opening of a monologue has to establish “what this is and what it is for”; otherwise the listener cannot even enter the situation.

Share labelled hard: 0%: purpose 0/88, main-idea 0/49, location 0/80. This is not coincidence. It is the nature of the type: the evidence is always at the very front, always requires only one sentence, and is almost never paraphrased. None of the three features applies.

2. Detail questions (72.9%, share labelled hard 14%)

These ask about something that was stated: time, place, quantity, what listeners are asked to do, or what the next step is. This is the main workhorse question type in Part 4.

Question: What are listeners asked to do? (A) Wait until Thursday (B) Leave the unit unlocked (C) Telephone the shop today (D) Collect the part in person

The answer is in one sentence in the audio. The hard part is not understanding; it is still being awake at the right moment.

The reason this category’s share labelled hard is not 0 is that it hides two mechanisms that really can raise the difficulty:

Mechanism one: semantic turn. The answer sentence contains however, actually, or instead.

……originally asked for a plated salmon course. However, they switched to a vegetarian menu yesterday.

The first half of the turn is the version everyone remembers; the second half is the real one. The wrong option is prepared for that first half, so this kind of question naturally has high distractor density.

Mechanism two: making you calculate.

Monologue: …confirming your cleaning appointment for Wednesday at nine fifteen A M. Please arrive ten minutes early… Question: What time should the listener arrive? (A) At nine oh five (B) At nine fifteen (C) At nine twenty-five (D) At four

nine fifteen will definitely appear in the options, and you just heard it. The answer is the number that was not spoken.

3. Graphic questions (2.7%, share labelled hard 100%)

The question booklet prints a table, schedule, or list, and the stem begins with Look at the graphic.

Graphic: North Depot — Morning Departures | 7:20 Via Ash Lane | 7:40 Direct | 8:00 Direct | 8:20 Via Ash Lane Monologue: From Monday the seven forty service is withdrawn… Question: Which departure will replace the withdrawn one? → 8:00

The answer is not in the audio, and not in the table; it exists only at the intersection of the two. The audio gives you one condition (the seven forty service is withdrawn), the table gives you the rest, and you have to connect them yourself. Share labelled hard: 100%, all 29 questions.

4. Intent & inference questions (5.5%, share labelled hard 100% / 50%)

These ask about something that was not said out loud. There are two forms:

Quoted meaning (share labelled hard 100%): the stem quotes a sentence from the monologue and asks what the speaker means. This is Question B from the opening of this article.

Inference (share labelled hard 50%): asks “which of the following is probably true.”

Monologue: The seat belt sign remains on while we work our way up through this weather… Question: What is probably true about the flight at the moment? (A) It has not yet reached a smooth altitude (B) It is about to land (C) It is delayed on the ground (D) It has changed its destination

No sentence says “the plane is still climbing.” You have to infer it from “the seat belt sign remains on” and “working up through weather.” This type necessarily has maximum paraphrase distance: none of the words in the correct option appear in the monologue.


One Feature We Measured Backwards

Before arriving at the list above, we first got one thing wrong, and that mistake deserves its own section.

Our original assumption was the same as most people’s: Part 4 is hard because of information density. An announcement packed with times, flight numbers, and floor numbers should be scarier. We even wrote this assumption into our own design document, listing it as the feature that best represents Part 4.

So we counted how many numbers, times, and proper nouns appear per 100 words in each monologue, then compared that against difficulty:

Item-writer difficulty label Numbers / times / proper nouns per 100 words
Easy 3.57
Medium 3.12
Hard 2.82

Exactly the opposite. The questions we labelled difficult were actually sparser than the easy ones.

Once you think it through, this is not surprising. In an announcement that densely lists numbers, the speaker slows down, segments the information, uses signposts like “first / next,” and the question is usually asking for one of the numbers that was spoken. You only need to catch that one; the rest is background. Conversely, sparse, loosely connected monologues are often sparse precisely because the answer was not stated directly.

This feature is still computed and printed in the report, but it does not count toward difficulty. The point is to keep that inverse relationship visible, so no one adds it back next time just because intuition says so.

Two other features were excluded for a different reason. The question’s position within the set (first question 4.5% hard, second question 18.4%, third question 26.1%) and the evidence position within the monologue both align strongly with the difficulty labels: strongly enough that you could build a very accurate model using only them.

But they explain nothing. “The third question is harder” describes the test-writing convention of where hard questions are placed, not what makes a question hard. Evidence position is subtler: once you look only at third questions, the evidence-position difference among easy, medium, and hard groups disappears completely (0.719 / 0.727 / 0.757). It was never an independent thing; it was just another way of saying “the evidence for third questions is normally later.”

A metric that can predict but cannot explain looks the most like an achievement, and is actually the least useful.


Calibrating the Ruler Against a Commercial Practice-Test Set

Classification alone is not enough: you also have to know whether the proportions are right. The key step was this: the same program runs on our question bank and on the reference book’s questions. One instrument, two corpora; only then is the comparison meaningful. (For which reference book we used, why we chose it, and its limitations, see “Difficulty Overview.”)

We extracted all 180 questions from the reference corpus’s six Part 4 tests and ran the same classifier.

First, Check That the Ruler Is Not Broken

The extraction can validate itself. Across the reference book’s six tests, two question types have zero variation in count:

Per test Graphic questions Quoted-meaning questions
Tests 1–6, each test 2 3

All six tests are the same. And months earlier, using a completely different method: counting the difficulty-rating dots printed for each question in the answer book, we measured the same test-design structure and got the same two numbers. Two independent measurements match exactly, which means the extraction was not wrong, and also means these two counts are not accidental; they are part of the test-design structure.

Comparing Question-Type Composition

Question type Ours Reference
Gist questions 18.9% 22.2%
Detail questions 70.8% 52.8%
Graphic questions 2.1% 6.7%
Intent & inference questions 1.6% 13.3%

Pay attention to the second and fourth rows:

We ask “find a sentence that was said” nearly twenty percentage points more often, while we ask “something that was not said out loud” at one-eighth the frequency of the reference book.

The most extreme case is quoted-meaning questions: the reference tests have 3 in every test, while our entire 1,350-question bank has only 14 questions.

A Reason Hidden Behind the Numbers

Quoted-meaning questions need a sentence whose literal meaning and actual meaning differ. We scanned the full text of all 1,018 dialogues and monologues, 491,912 characters in total, and counted how often this kind of language appeared.

The result: two idioms. I guess, I suppose, and to be honest appeared zero times.

Our speakers are literal and transactional from beginning to end. They announce things, give instructions, and report times, but they do not imply, hedge, or use tone to communicate what is not said out loud. What is missing is not a question type; it is a way of speaking. The question type is only the exposed corner of it.

Comparing Difficulty Labels, and Why That Is Less Useful Than It Looks

The reference book prints its own publisher difficulty rating for every question. Put it beside our labels:

Low Medium High
Reference book Part 4 (n=157) 36% 46% 18%
Labelled “hard” by us 16.3%

A difference of 1.7 percentage points looks almost identical. But this number should not be treated as a conclusion, for two reasons.

First, the two sides are two rulers that have never been calibrated against each other: the reference book’s high / medium / low labels are the publisher editor’s judgment, while ours are the item writers’ judgment. They may happen to land on the same number while measuring completely different things.

Second, and more importantly: when we recompute with our own seven-feature rule, only 2.2% of old questions count as difficult, while item writers labelled 16.3%. Even under the loosest algorithm (counting a question as difficult if it requires any cognitively harder action at all), the figure is only 9.1%.

In other words, that “almost identical” 16.3% rests on overly loose labels. What can truly be compared is question-type composition, because the same program measured two corpora; the difficulty-label comparison can only be a reference point, not a conclusion.


Honest Caveats

  • There is no external claim here about “difficulty.” Question-type composition is measurable and reproducible; difficulty is not. We deliberately keep those two things separate.
  • The reference is a commercial practice-test set, not official ETS questions. Question-type structure is the part publishers are best at copying faithfully (zero variation across six tests is the evidence), so the type comparison is credible; the difficulty rating carries editorial judgment, not real test-taker accuracy.
  • Our audio is cleaner than the real test. Synthetic speech does not have the reductions, overlap, or background noise of live recording, so our listening material is slightly easier than the real test on an axis we cannot measure ourselves.
  • Question type is derived from stem text. That makes it reproducible and still correct for content written later, but what it captures is the way the question is asked. A question with an ordinary-looking stem that is actually tricky will not be detected by this ruler.
  • The real endpoint is accuracy rate. As anonymous answer statistics accumulate, each question will slowly get its real p-value, and this ruler will be recalibrated against real data. Until then, this is the most honest measurement method available to us.

Once You Know the Question Type: How to Listen for Each One

Now read the classification above in reverse. The single most important Part 4 habit is this: before the audio starts, scan the three question stems.

Part 4 gives you the time to do this: there are eight seconds between sets, and three stems take only five seconds to scan. Once you scan them, you know what this set is asking you to do, and that directly determines how you allocate attention.

Another rule makes this workable: the answers to the three questions almost always appear in the order of the monologue. You can “move along with it”: once you hear the answer to the first question, shift your attention to the second.

If You See “purpose / main idea / who is the speaker” → Solve It in the First Two Sentences, Then Drop It

The answer is in the opening. Your action: by the end of the first two sentences, this question should already have an answer. Do not come back to it afterward.

If you find yourself still hesitating over the first question in the second half of the monologue, the second and third questions are basically already gone. This is the most common way to lose points in Part 4, and the points lost are the ones that should have been easiest to take.

Target time: within the first two sentences.

If You See “when / where / how much / what are listeners asked to do” → Lock Onto a Keyword and Wait for It

This is a detail question. Your action:

  1. Pull one keyword from the stem: a place, time, or action.
  2. Wait for the sentence near that word in the audio.
  3. When you hear however, actually, or instead, raise your guard immediately: the previous sentence has probably been invalidated, and the answer is after the turn.
  4. If one number in the options is a number you just heard, it is usually not the answer. The answer is the one you have to calculate.

Target time: the moment the answer appears. If you miss it, let it go; do not use the third question’s time to chase the second.

If You See “Look at the graphic.” → Read the Table First, Then Listen to the Audio

This is a graphic question, and it is the only type where half the answer is already in the question booklet. Your action:

  1. Understand the table before the audio starts: how many columns, what units, and which column is not listed in the options.
  2. The options usually list one column of the table (for example, times), so the audio will give you another column (for example, route or room number).
  3. While listening, catch only that one condition; once you hear it, go back to the table and match it.

Do not start reading the table only after the audio begins. This is the only Part 4 type where you can complete half the work in advance; wasting that is costly.

Target time: read the table in the 8 seconds before the audio starts; while listening, wait for only one condition.

If You See Quotation Marks in the Stem → This Tests Tone; the Answer Is Around That Sentence

The quotation marks are the signal; you can spot them at a glance. Your action:

  1. Do not try to understand that sentence from its literal meaning. It is printed in the question, and you can read it; that is not what is being tested.
  2. Wait for that sentence to appear in the audio, then remember the sentence before it and the sentence after it. The answer is almost certainly there.
  3. Ask yourself: why does the speaker say this here? Is it a warning, a complaint, a polite refusal, or irony?

The same action applies to inference questions such as What is probably true…: the answer was not said out loud; what you are looking for is the sentence that supports it. And because these questions necessarily have maximum paraphrase distance, options that contain the audio’s exact words are usually traps.

Target time: this is the only question type worth spending extra effort on, and the one whose options most need to be compared one by one.

Pacing: The Real Discipline in Part 4 Is Letting Go

Part 4 has a fixed rhythm; you cannot go back. There are eight seconds between sets and eight seconds between the three questions. There is only one correct use of that space: look ahead to the next set’s stems, not back at the previous question.

Once a monologue finishes, it is gone forever. Spending ten more seconds on the second question costs you the whole third question: and a question you did not hear is, for scoring purposes, the same as a guess.

So the discipline in Part 4, as in Part 5, is not getting everything right; it is knowing when to let go. If you did not catch the answer, choose the option least likely to be a trap (especially not the number you just heard), then hand your attention immediately to the next question.


One-Sentence Summary

Part 4 difficulty is not about how dense the monologue is or how many numbers it contains: that intuition measured out backwards, and it is still printed in our reports as a reminder. Difficulty can be broken into seven question-level features, and the strongest determinant among them is what the question asks you to do: find one sentence in the opening, wait for one fact in the middle, connect the table to the audio, or infer what the speaker did not say out loud. When we measured a commercial practice-test set with the same ruler, what we saw was a difference in question-type composition; and when you read the same classification in reverse, it becomes your answer order: scan the three stems before the audio starts, solve the first question in the first two sentences, assume the previous sentence has been invalidated when you hear a turn, listen around the sentence when you see quotation marks, and then, the moment each question ends, let go.

← All articles