All articles · Published 2026-08-04

Part 6 difficulty comes down to "did you actually need the passage?" — a test that covers up the text

A technical explainer for the curious. The difficulty series runs one post per Part, published in order; this is the sixth. Part 5 difficulty focuses on how the four options are arranged, while Part 7 difficulty is scattered across the entire passage. Part 6 sits in between, and its difficulty comes from a place that is easily overlooked: <strong>does this blank actually require you to read the passage?</strong> This post explains how we turned that question into a measurable number, the method we used so this measurement wouldn't just be "one person's feeling," what the results were (the first measurement looked very bad for us, so we rewrote the bank based on it), and—most useful to you—how knowing that Part 6 contains two completely different types of blanks tells you how to manage your time on the exam. No statistics background needed.

What Part 6 ought to test

Part 6 makes up Questions 131 to 146 of the Reading section: four short passages, each with four blanks, and four options per blank.

On the surface, it looks a lot like "stuffing Part 5 into a passage." But it shouldn't be. Part 5 gives you an isolated sentence, so the answer can only ever come from inside that single sentence. Part 6 gives you an entire passage—which means the answer can be hidden in surrounding sentences.

That is the true reason Part 6 exists. If every blank could be solved using only its own sentence, the text would be mere decoration, and there would be no reason for this Part to exist on its own.


A simple test: cover up the passage

To judge whether a blank is truly testing the passage, there is a very straightforward method:

Cover up the entire passage, leaving only the sentence containing the blank and the four options. Then ask: how many options still make sense?

If only one option still makes sense, then the blank doesn't need the passage at all—it's just a Part 5 question hiding inside a text.

If two or more options still make sense, then only the passage can decide the answer. This blank is genuinely testing Part 6.

We call blanks that fit the latter description "context-dependent blanks." Every number in the rest of this post is a count of these.

Three concrete examples

Example 1: Unaffected by covering the text

Before entering, all staff must put on a full protective suit, gloves, and a face mask in the ------- room. (A) checking (B) changing (C) charging (D) chilling

No need to read the passage. The place where you put on protective suits, gloves, and masks is a changing room. The other three options will not work no matter what the surrounding context says. Even if you deleted the rest of the text entirely, this question's difficulty would remain completely unchanged.

Example 2: Impossible to answer if covered

Bring your friends. ------- groups of six or more should book ahead. (A) Because of this, (B) That said, (C) For that reason, (D) In the meantime,

All four options are grammatically correct and coherent ways to open a sentence. Looking at this sentence alone, you have zero reason to pick B over A.

Only the preceding line, "Bring your friends," lets you know that the following sentence introduces a qualification—group visits are welcome, but groups of six or more must book in advance. That qualification is conveyed by "That said."

For this question, the passage is a strict requirement for finding the answer.

Example 3: Impossible to tell what the text is even about when covered

Even stubborn ------- don't stand a chance against our gentle, modern process. (A) sprains (B) strains (C) stains (D) streams

In isolation, you cannot even tell whether this is an ad for a laundromat or a physical therapy clinic. "Sprains" and "stains" make equal sense in this isolated sentence.

Only the passage can tell you what is being discussed. This is context dependence at its most extreme.


Turning "makes sense" into a credible number

Up to this point, it's still just "how I feel." And "how I feel" is not a measurement.

Psychometrics has a formal definition of difficulty: a question's difficulty is the proportion of real test-takers who answer it correctly. That requires thousands of examinees to take the test first—data nobody outside of ETS possesses. The complete explanation is in the Overview. What we need to handle here is a smaller, trickier problem: how do we evaluate "how many options still make sense when covered" without letting it turn into one person's subjective judgment?

We did three things.

1. Isolate each question down to "only its sentence"

We used a script to extract the target sentence containing each blank. Other blanks in the same passage were filled in with their correct answers so that the surrounding sentence structure remained complete. All 490 blanks were successfully extracted this way.

This step immediately set aside an entire category of questions: sentence insertion questions (where the four options are complete sentences). By definition, that question type is always context-dependent; including it would simply pad the numbers. Furthermore, our sentence insertion ratio already matches the real test blueprint—one in every four questions. So what we are measuring here is the remaining three blanks.

2. Hand them to two raters who cannot see each other's work

We presented "isolated sentence + four options" to two independent AI raters. Both received the exact same written grading rubric, and neither was told which option was the correct answer (knowing the answer severely biases judgment). They only needed to answer one thing: looking at the sentence alone, how many of the four options make sense?

We used two raters instead of one because a single rater cannot be audited. With two, we can calculate how often they agree:

Agreement rate Cohen's κ
Our item bank (490 items) 80.4% 0.51
Reference mock set (72 items) 90.3% 0.81

κ (kappa) measures agreement beyond what would occur by chance. The 0.81 on the reference set qualifies as very high; our score of 0.51 is moderate—meaning our item bank contained a higher number of ambiguous edge cases, which is informative data in itself.

3. Count only what both raters agree on

Where the two raters disagreed, we inspected the items manually. Randomly sampling 16 items for item-by-item analysis revealed that one rater was noticeably overly permissive, counting items where only one answer actually worked as context-dependent (for example, on the changing room item above, it claimed two options were viable). Crucially, its permissiveness was unequal between the two sets—it over-flagged by 18% on our bank, but by only 10% on the reference set.

This matters: relying solely on its judgments would unfairly flatter one side. Therefore, the final numbers count only items where both raters agreed.

(We also calculated the permissive version—where an item counts if either rater flags it. The conclusion is identical, as detailed below.)


What the measurement revealed

The first time we measured, the results looked very bad for us:

Context-dependent blanks 95% confidence interval
Our item bank (first measurement) 82 / 490 = 16.7% 13.7%–20.3%
Reference mock set 35 / 72 = 48.6% 37.4%–59.9%

In the reference set, roughly one out of every two blanks genuinely requires the passage. In our bank, only about one out of six did.

The two 95% confidence intervals do not overlap (our upper bound of 20.3% falls below their lower bound of 37.4%), meaning this gap cannot be explained away by sampling error. Re-running the calculation under the permissive definition yields 36.3% for us and 58.3% for them—again with no overlap. Changing the calculation method does not change the conclusion.

As a sanity check: when we deliberately rewrote blanks to be context-dependent, both raters flagged every single one without knowing what had been changed. If a ruler cannot even register items built specifically for it, nothing it measures can be trusted.

Then we fixed it

After measuring this, we didn't tuck the number away; we edited our item bank according to what it showed. The fix wasn't "writing more passages"—that wouldn't help, because adding a passage adds three new non-insertion blanks, growing the denominator alongside the numerator. To raise the proportion, the only way was to take existing blanks one by one and rewrite them so they genuinely require the passage.

After completing the rewrites, we re-measured using the exact same ruler, the same two blind raters, and the same policy of counting only unanimous agreements:

Context-dependent blanks 95% confidence interval Average per passage
Our item bank (first measurement) 82 / 490 = 16.7% 13.7%–20.3% 0.50
Our item bank (current) 239 / 490 = 48.8% 44.4%–53.2% 1.46
Reference mock set 35 / 72 = 48.6% 37.4%–59.9% 1.46

The point estimates differ by 0.2 percentage points, while the average density per passage matches at an identical 1.46. The confidence intervals overlap heavily.

To put it plainly: on the specific question of whether a blank requires the passage, there is now no measurable difference between our bank and the reference set.

We must be careful not to overstate this. This does not mean our questions are "as hard as the real exam"—difficulty has many dimensions, and this post measures only one of them. Nor does it mean we are better than the reference set: the gap between 48.8% and 48.6% sits well inside measurement error, so neither side wins. The only claim we can make is this: a gap on this dimension that was once large enough to detect statistically has now vanished.

An honest footnote: while spot-checking the examples printed in this article against our actual item bank after the rewrite, we discovered four items broke during the revision process—options had been attached to the wrong blank within the same passage, leaving those four items without a correct answer. They have since been fixed, and we added an automated check so that this category of error will be caught before reaching production in the future. We caught this precisely because we reconciled the printed examples in this post against our live bank.

Why this happened in the first place

It didn't happen because anyone was being lazy, but because the natural tendency of question writing points this way.

Writing a blank that "can be solved by its sentence alone" is easier and safer: you are certain it has a single solution, certain it won't trigger disputes, and certain it is grammatically sound. Writing a blank that "requires the preceding sentence to decide" forces you to manage the relationship between two sentences simultaneously while ensuring all four options make sense on their own—otherwise, it degrades right back into a Part 5 question.

Without measuring it explicitly, any item bank naturally drifts toward "Part 5 stuffed into a passage." And we drifted. Before covering up the passages, this was completely invisible—viewed individually, every question looked like a great item.

This is why we turned this measurement into a permanent step that runs every time we update questions: drift doesn't stop on its own; only measurement stops it.


The most useful part for you: how to tackle Part 6

Knowing that Part 6 contains two completely different kinds of blanks makes your exam strategy clear: classify first, then decide whether to read back.

Step 1: Look at the options to classify the blank

When you encounter a question, check what the four options look like:

  • Four forms of the same word (apply / applied / applying / application) → Almost certainly a single-sentence blank. Look around the blank; do not read back into the passage.
  • Four unrelated content words (nouns, verbs, adjectives) → Usually a single-sentence blank, but watch out for cases like Example 3 where you "can't tell what the text is even about." Try solving it using the sentence alone first; only read back if that fails.
  • Four sentence-opening transitions (However, / As a result, / Even so, / In the meantime,) → You must read back to the preceding sentence. This type can never be solved by looking at the target sentence alone; staring at it is just wasting time.
  • Four complete sentences → Sentence insertion question. Look at both the preceding and following sentences, and usually leave this for last.

Step 2: For transition questions, you only need to read back one sentence

This is the single most valuable tip to remember. Transition questions do not require re-reading the entire passage; you only need to read the single sentence right before it, and then ask yourself one question:

What is the relationship between the preceding sentence and this sentence?

  • The second sentence is the result of the first → As a result, / Therefore, / Accordingly,
  • The second sentence qualifies or refutes the first → However, / That said, / Even so,
  • The second sentence is an example of the first → For example,
  • The second sentence is the effect brought by the means in the first → That way, / In this way,
  • The first sentence states a condition, and the second is the consequence of not meeting it → Otherwise,

Once you identify the relationship, the answer appears immediately. This step should take under 10 seconds; if you're still hesitating past 20 seconds, you are usually trying to hunt for clues inside the target sentence—where no clues exist.

Step 3: Put saved time toward sentence insertion questions

Every Part 6 passage contains four blanks, typically including one sentence insertion question—the most time-consuming item in Part 6, because you must verify it connects logically with both the sentence before it and the sentence after it.

The main takeaway of these steps is: don't waste time reading back into the passage for single-sentence blanks; focus your time on the one or two questions per passage that genuinely require reading the text.


Limitations of this measurement

Three honest caveats:

The reference set contains only 72 blanks. This represents every non-insertion blank across all six mock forms—the complete set of data available to us—but 72 is still a modest sample size. That is why the 48.6% estimate carries a confidence interval spanning 22 percentage points. Our 490 items represent a complete census of our bank, yielding a much tighter interval.

The raters are AI models, not human test-takers. They evaluate whether an option "makes sense in isolation"—a stand-in metric, not actual examinee accuracy rates. Its relationship to true difficulty is logical, but they are not the same thing.

We make claims at the distribution level, not the individual item level. Much like a bathroom scale with a ±5 kg margin of error, weighing one person yields little insight, but taking the average of a thousand people reveals clear population differences. This measurement can say "the distributions of the two banks differ significantly on this dimension"; it cannot certify that "Item A is more context-dependent than Item B."


Difficulty series: Overview · Part 5 · Part 6 (this post) · Part 7

← All articles