All articles · Published 2026-08-02

Part 1 Difficulty: it lives in a single instant of the verb — five ways to make photo questions hard, and how to listen for each

A technical explainer for the curious. Each Part gets its own post in this series; this is the Part 1 post. At just six questions, Photographs is the shortest part of the test, and the one most often dismissed as "just looking at a picture and listening to a sentence." This post explains where its difficulty actually lives — not in vocabulary, but in whether the verb describes a completed state or an action in progress — then compares our bank against 36 questions from a commercial mock set (yielding a confirmation we didn't expect), and ends with the most useful part: what to look at before the audio starts, and which word to listen for first once it plays. No statistics background needed.

Six questions, ninety seconds, played once

Part 1 is the first six questions of the listening test: a photograph on screen, four English sentences in your headphones, pick the one that best fits the image. The four options are not printed in the booklet — they are read aloud once, and the test moves immediately to the next question.

That format has a consequence that is easily underestimated: there is no room to look back. In Part 5 you can read the four options three times over; in Part 7 you can return to the passage to check an answer. In Part 1, once the four sentences pass through your ears, they are gone.

Because the questions are short and the photos straightforward, many test-takers treat Part 1 as a warmup and prepare by memorising nouns — tables, ladders, handcarts, scaffolding. Nouns matter, of course, but what actually trips people up in Part 1 is almost never the nouns.

Consider the difference between these two questions.

In the first, the photo shows two workers in a warehouse stacking boxes onto a cart:

(A) A customer is trying on a jacket. (B) Two employees are stacking boxes on a cart. (C) The boxes are being loaded into a car. (D) A worker is sweeping an empty hallway.

The subjects of the four sentences are a customer, two employees, boxes, and a worker — four completely distinct entities. As soon as you hear "there are two people moving boxes," the other three options are eliminated automatically. If you know the nouns, you get the question right; it takes about three seconds.

In the second, the photo shows an empty staff locker room with a row of lockers, one door standing wide open:

(A) One of the lockers is being opened by a staff member. (B) One of the lockers is open. (C) The lockers are being repainted along the wall. (D) The lockers are being emptied of their contents.

In this question, the subjects of all four options are the lockers. The subject offers zero help; you have to hear the verb, and specifically distinguish (A) from (B):

  • is being opened — someone is in the middle of opening it
  • is open — it is open

There is no person in the photo. So "someone is opening it" is wrong, and "it is open" is right. The two options use the exact same root word open; the entire difference lies in that single grammatical form. Memorising the word locker inside out will not help you answer this question correctly.

That is where Part 1's difficulty actually lives.


Five things that make Part 1 hard

The five items below are the classification extracted across all 335 questions in our bank. Each represents a concrete, recognisable technique; the two columns on the right show its occurrence rate in easy versus hard questions — the larger the gap, the more effective the technique is at making a question hard.

Technique Easy questions Hard questions
1. No people in the photo 0% 65.0%
2. All four sentences share a subject 0% 100%
3. State / action minimal pair 0% 70.0%
4. Sound-alike trap 0% 32.5%
5. Longer sentences 0% 52.5%

1. No people in the photo

A photo showing only objects or scenes — an empty meeting room, stacked pallets, a parked van. The difficulty is that you have no "main subject" to anchor on.

When people are present, you automatically focus on what they are doing; any sentence describing a different action can be eliminated immediately. When no people are present, every single sentence must be cross-checked against the image individually, and English naturally uses passive constructions and state verbs to describe objects, stacking the difficulty straight onto Item 3.

2. All four sentences share a subject

The locker example above is a case in point. Once the subject is repeated across all options, it loses all filtering power, pushing the entire burden of judgment onto the verb and the trailing details.

In our hard questions, the occurrence rate for this item is 100% — zero exceptions. It is the most fundamental foundation: as long as the four options feature distinct subjects, the question cannot be particularly hard.

3. State / action minimal pair (the core of Part 1)

Two sentences feature the same subject, differing only in the verb's grammatical aspect: one describes a completed state, the other an action in progress.

State (completed) Action (in progress)
The gate is open. The gate is being opened.
A van is parked. A van is being parked.
The crates are stacked. The crates are being stacked.
The man is wearing a helmet. The man is putting on a helmet.
A boat is moored. A boat is being moored.
The trays have been arranged. The trays are being arranged.

This is the one source of difficulty in Part 1 that is completely unavoidable: it doesn't test whether you know the vocabulary — the words in the two options are often identical. It tests whether you can hear the difference between two grammatical forms within a fraction of a second and immediately cross-reference it against the photo.

Among our hard questions, 70.0% carry a contrast of this kind.

4. Sound-alike trap

An option contains a word that sounds very similar to something actually present in the photo. Common pairs include:

writing / riding, cart / card, walking / working, copy / coffee, glasses / classes, sweeping / sleeping, pouring / pointing, mail / male, parked / packed

In an actual question, the photo shows a woman standing beside a bicycle, writing on a clipboard:

(A) A woman is riding a bicycle. (B) A woman is writing on a clipboard. (C) A woman is repairing a bicycle wheel. (D) A woman is loading a box onto a bicycle.

There is indeed a bicycle in the photo, and riding differs from writing by only a single sound. If you listen too fast and see a bicycle in the image, (A) becomes extremely persuasive.

5. Longer sentences

The longer the sentence, the more information you must hold in working memory — and you only hear it once. On its own, sentence length rarely makes a question hard, but it amplifies the preceding four techniques — a state/action minimal pair buried in an eleven-word sentence is far harder to catch than one in a six-word sentence.


A misconception to dismantle first: "is being + past participle" is not a difficulty signal

Many test-takers hear passive progressive constructions like is being handed or are being stacked and instinctively assume "this must be a trap."

In our question bank, 76.7% of questions feature this structure in at least one option — across easy, medium, and hard items alike. It is not a marker of difficulty; it is simply one of the most natural ways to describe a photo in English.

So please do not use "which sentence is in the passive progressive" to guess the answer — in either direction: it neither implies the sentence is a trap nor that it is the correct answer. What you actually need to listen for is not "whether there is a being," but the question introduced in the next section.


Comparing against a commercial mock set: 36 questions

The classification above is computed by our own tools, which leaves one question unanswered: is this difficulty composition reasonable?

Our external reference is a commercial six-form mock exam set, with the publisher's own printed difficulty rating (low ●○○ / mid ●●○ / high ●●●) beside every question. With six questions per form across six forms, that makes 36 questions in total — far fewer than other Parts, but enough for a rough comparison.

Difficulty Reference book (36 questions) Ours (335 questions)
Low / Easy 50.0% 29.0%
Mid / Medium 36.1% 59.1%
High / Hard 13.9% 11.9%

The bottom row is the most revealing: in the hardest tier, the reference sits at 13.9% while ours is 11.9% — a difference of less than two percentage points. For a sample of 36 questions, that gap lies well within the margin of error (the 95% confidence interval for the reference's 13.9% spans 6.1%–28.7%, which is very wide).

The real divergence shows up in the top two rows: the reference skews "easy" (50.0%), whereas we skew "medium" (59.1%). We place many questions into "medium" that the reference would label "low", but the proportion of hard questions remains consistent between the two.

An unexpected confirmation

More persuasive than the overall proportions is which questions the reference book rates as high difficulty.

Across all six forms, Part 1 in every single form features exactly one object/scene photo (photos with no people) — six questions out of six forms, or 16.7%, which closely matches our bank's 20.6%. And of those six questions, three are rated at the highest difficulty. While high-difficulty items account for only 13.9% of the reference as a whole, they represent 50.0% of the photos without people.

Look further at how the options for those questions are constructed: in the questions rated high-difficulty, the four options consist almost entirely of state vs. action contrasts — "has already been..." versus "is being...". In other words, without any knowledge of our classification scheme, the publisher's editors assigned their highest difficulty ratings to photos without people and state/action minimal pairs.

These are precisely Items 1 and 3 from our framework. Two different rulers pointing to the exact same place.

But 36 questions are still just 36 questions

  • The sample size is small. Each individual question represents 2.8 percentage points, so the proportions above can only confirm whether the general direction is sound, not support fine-grained comparisons.
  • Variance across forms is high. The count of high-difficulty questions across the six forms is 2, 0, 0, 1, 1, 1 respectively — two of the forms have none at all. Looking at a single form would yield a completely different impression.
  • The reference book is a publisher's mock set, not ETS material. Those difficulty ratings reflect an editor's judgment, not real examinee accuracy rates. This limitation is explained in full in the Overview post.

Reading it backwards: how to listen in Part 1

Each of the five techniques above reveals where the answer is hiding. Read them backwards and they become a concrete set of actions.

Before the audio plays: examine the photo and ask yourself three questions

There is a multi-second gap between questions in Part 1, and the photo is visible first. Don't let those seconds go to waste; use them to ask:

  1. Are there people in the photo? If there are no people, any subsequent sentence describing "someone is doing something" can be eliminated right away — this is the highest-ROI judgment in all of Part 1.
  2. If there are people, what are their hands touching? This prepares you for the rule below.
  3. How is this photo most likely to be described? Formulate a single sentence in English in your head. Once you've pre-framed it, processing a similar sentence when the audio plays will be much faster.

While playing: Subject → Aspect → Details

Listen in this exact order, and drop a sentence the moment something doesn't match without waiting for it to finish:

  1. Subject — Person or object? Singular or plural? If all four sentences have different subjects, the question is usually decided right here.
  2. Verb aspect — A completed state or an action in progress? This is the watershed moment for hard questions.
  3. Details — Location, quantity, spatial arrangement. You only need to listen this far if the sentence passes the first two checks.

The single most practical rule: look at the hands

The state vs. action contrast in a photo is almost always determined by the hands:

If no human hand is touching the object → it is a "state", not "someone is doing something".

The door is open but no hand is on the door → is open, not is being opened. The boxes are stacked but no hand is on the boxes → are stacked, not are being stacked. The vehicle is parked and nobody is in the driver's seat → is parked, not is being parked.

The inverse also holds: if a person is making no active movement with their hands — a jacket is fully on their body, arms hanging by their side → is wearing, not is putting on.

What makes this rule so effective is that it converts an abstract grammatical judgment into something you can spot instantly in the photo.

When encountering sound-alike traps: trust the photo, not what sounded familiar

When you hear a word that matches an object in the image, pause for half a second and ask: is someone in the image actually performing that action?

In the bicycle question above, there is indeed a bicycle in the image, but nobody is riding it. The sound riding appeared, but the corresponding action did not. Almost all traps in Part 1 take this exact shape: the object is in the picture, but the action is not.

Pacing: commit your answer and move on

With only six questions, Part 1 offers plenty of time; the real risk is not running out of time, but getting stuck.

If all four sentences have been read and you are still torn between two choices, pick one and move on. The photo for the next question is already on screen; every extra second spent dwelling is stolen from asking yourself those "three pre-audio questions" for the next item — and the value of those three questions far outweighs five extra seconds of hesitating over the previous one.


Honest caveats

  • The external reference for Part 1 consists of only 36 questions. While other Parts have reference samples numbering in the hundreds, Part 1 has only 36, meaning the comparison table above supports only broad conclusions like "the proportion of hard questions is roughly comparable," rather than fine-grained differences.
  • The five-category classification is computed from the structural features of the questions themselves, rather than copying ratings from external sources. The reference book's ratings serve to verify whether our ruler points in the right direction, not to define difficulty.
  • The proportions above describe our question bank, not the official exam. They illustrate how we construct difficult questions; we make no claim that live ETS exam forms share this exact distribution.
  • Our listening audio is cleaner than the real test's. This point is covered more fully in the Overview: synthesized speech lacks the elision and background ambient noise of live recordings, making our listening items slightly easier than the real exam along a dimension our metrics cannot measure.

The bottom line in one sentence

Part 1's difficulty lives not in the nouns, but in a single instant of the verb: whether something is already completed or currently happening. Photos without people, shared subjects across all four options, and longer sentences are all designed to force your attention onto that exact instant. Armed with that knowledge, your answering procedure boils down to two steps: before the audio plays, confirm whether there are people in the photo and what their hands are touching; once the audio begins, listen first for the subject, then for the verb aspect.

← All articles