All articles · Published 2026-08-04

In Defense of Drilling: Wittgenstein, Linguists, Memory Science, and AI Take the Stand

In East Asia it has a name — 刷題, "grinding through practice questions" — and a reputation to match: cram-school culture, exam-factory learning, the reason people "score high but can't hold a conversation." Real language learning, the story goes, happens somewhere else. This essay is a serious defense of drilling. Taking the stand: one of the twentieth century's most important philosophers of language, several linguists who actually counted, a century of memory science — and one surprise witness, the large language model itself. Verdict up front: drilling is not a detour around language learning; short of moving to an English-speaking country, it is the closest thing there is to the direct route.

The Charge

Let's hear the accusation in full — it deserves to be taken seriously.

It goes like this: language is for communicating, not for testing. Doing multiple-choice questions all day trains "test technique" — elimination, keyword matching, question-type instincts — not the language itself. That's why East Asian test-takers famously "score high but can't speak"; that's why "drilling" sounds like an unhealthy obsession with exams rather than a way of learning anything.

The accusation sounds forceful because of a premise hiding underneath it, one that looks self-evident:

A language is a set of rules plus a stock of words. Master the rules and you "know" the language; questions merely measure whether you do. Repeating a measurement doesn't improve the thing being measured.

If that premise holds, drilling really is suspect — a hundred medical checkups won't make you healthier.

The problem is that the premise is wrong. And its demolition is one of the most famous reversals in the history of philosophy.


A Philosopher Changes His Mind

Ludwig Wittgenstein wrote two books, embodying two ways of looking at language — and the second was written to tear down the first.

The first, the Tractatus Logico-Philosophicus, holds that the essence of language is logic: sentences are pictures of facts, and once the logical structure is analyzed, language becomes transparent. On publishing it he considered philosophy essentially finished, and went off to teach primary school in rural Austria. Notice the family resemblance between this and the premise in the previous section: if language really reduces to a compact system of rules, then "learn the rules" is the highway and "mass exposure" is the long way around.

Then he changed his mind. A philosopher publicly refuting himself is a rare event — most spend a lifetime defending their early work — and what Wittgenstein refuted was his own masterpiece. The later Philosophical Investigations dismantles the early position almost clause by clause, and what emerges is this: the meaning of a word lies not in a definition but in its actual use in the language. Language is not a logical system floating in mid-air; it is countless "language games" embedded in daily life. To imagine a language is to imagine a form of life.

He also left an argument that matters especially to learners, known since as the rule-following paradox: no explicitly stated rule can fully determine how it is to be applied — you always also need the knack of applying it, and that knack is not another rule; it is a feel acquired through practice. Translated into exam-prep terms: every rule in a grammar book presupposes a reader who has already seen enough examples. Rules are after-the-fact tidying, not generators.

Most interesting of all is his choice of word for how children acquire their mother tongue: Abrichtung — a German word ordinarily used for training animals; "drill" is about as close as English gets. In his account, the teaching of language "is not explaining, but training." The century's deepest philosopher of language, describing how language is learned, reached for a metaphor far closer to "drill" than to "explain."

This is also why our About page says: you can learn the words "give" and "in" perfectly well and still never guess that "give in" means "to yield" — if language were really that reasonable, it would have given in to you a long time ago.


The Linguists Counted: Language Is Mostly Convention

Philosophy supplied the direction; linguistics supplied the numbers.

In 1983, Pawley and Syder published an observation the field has never gotten around since: grammar rules mostly generate sentences that are perfectly grammatical and that nobody says. To ask the time, grammar happily licenses "What is the hour?" and "How late is it?" (which happens to be exactly how Dutch does it) — all well-formed, but only "What time is it?" is English. Native fluency rests not on computing from rules in real time but on retrieving ready-made expressions from an enormous conventional stock.

Corpus linguistics then quantified it. Sinclair, after staring at real corpora, proposed the "idiom principle": the first principle of language use is not free assembly by rule but selection from prefabricated parts. Erman and Warren (2000) actually counted: pick up a stretch of authentic English, and roughly half of it is laid down from such ready-made combinations, large and small. Wray (2002) reached the same conclusion at book length: formulaic sequences are not the scraps of language — they are its substance.

And these conventions are exactly what logic cannot derive:

  • heavy rain can't become strong rain, yet strong coffee can't become heavy coffee;
  • you make a decision, but you do homework;
  • the way to say yes to "Would you mind closing the window?" is "Not at all" — by logic, "I don't mind in the least"; by convention, "sure, I'll close it";
  • to "You haven't sent the report, have you?", English answers with the facts — "No" (I haven't) — while Chinese, Japanese, and Korean instinct says "Yes" (right, I haven't). Neither side is more logical; they are two different conventions.

Now look at what TOEIC tests: Part 5 tests collocation and usage, Part 2 tests response conventions, Parts 3 and 4 test the fixed routines of workplace conversation, Parts 6 and 7 test the set phrases of business correspondence — "Please find attached," "effective immediately." This is not exam trivia; this is the body of the language — the very half the linguists counted. A good TOEIC question bank is, at bottom, an anthology of workplace-English conventions, typeset as multiple choice.


Brains Already Learn This Way: Frequency and Statistics

If conventions can't be derived, they can only be absorbed through contact. Fortunately, the brain is built for exactly that.

The classic experiment of Saffran, Aslin, and Newport (1996): eight-month-old infants, after two minutes of a continuous artificial syllable stream, can carve out "word" boundaries purely from the statistics of which syllables follow which. The brain is a statistics engine — and a statistics engine runs on one fuel only: sample size.

Tomasello's usage-based theory connects this to language acquisition: children do not receive a grammar and then apply it; they induce ever more abstract constructions out of one concrete usage event after another. Nick Ellis (2002) extends it to second languages: adult language learning, too, is largely implicit statistical tallying over input — every encounter adds a tick next to some collocation, some construction. Frequency is not a supporting factor. It is the main one.

This, incidentally, explains the advice everyone already agrees on: "the best way to learn a language is to move there." Immersion works not because English is in the air, but because of dosage — living inside the language means passively meeting thousands of convention specimens a day, the statistics engine running at full speed. Wittgenstein would say: you have joined the form of life.

If you can't move, you have to engineer the dosage yourself. Extensive reading is one arrangement; drilling is another — and, as the next section shows, one with a measured advantage.


Memory Science: Testing Beats Rereading, and Getting It Wrong Beats Never Being Wrong

Suppose you accept that mass exposure is necessary. One question remains: does the format of the exposure matter? Are ten articles read and ten questions answered (with explanations) the same thing?

They are not — and the direction may be the opposite of your intuition. One of the most robustly replicated findings in memory science is the "testing effect": Roediger and Karpicke (2006) had subjects study the same passage, one group rereading it repeatedly, the other taking tests on it instead; a week later, the tested group won decisively. Rereading breeds the fluent illusion of "I know this"; retrieval is what leaves the trace. Karpicke and Blunt (2011), in Science, went further: retrieval practice beat even concept mapping, that most sophisticated of study techniques.

Bjork summed up this family of phenomena as "desirable difficulties": the smoother the studying feels, the less it leaves behind; what the brain has to strain to fetch is what sticks. Answering questions turns every encounter into a retrieval — you can't swipe past; you have to commit to a judgment.

Even wrong answers earn their keep. Kornell and colleagues showed that guessing wrong and then seeing the correct answer beats being shown the answer directly. So your wrong answers are not a record of failure — they are your inventory of convention gaps, one per item, with coordinates attached. Miss "meet a deadline" once in Part 5 and that convention is flagged for you; the next time it appears in a real email, it is no longer background noise.

Line up what learning science prescribes against what drilling is, and the columns match almost embarrassingly well:

What learning science prescribes What drilling delivers
Mass exposure to real conventions Every item is a condensed specimen: a business letter, a customer call, an idiomatic reply
Retrieval, not rereading Every item forces a judgment — passive skimming doesn't count
Immediate feedback You know at once whether you were right, and the explanation tells you why
Learning from errors Wrong answers pinpoint the exact convention you haven't internalized
Spacing Five questions on the train, three in a queue — naturally spread across days, beating any week-before binge

Unpack the act of "drilling" and it reads: dense convention exposure × forced retrieval × immediate feedback × natural spacing. Not one of the four is "test technique." All four are on learning science's shopping list.


Why the "Large" in Large Language Model

The human witnesses have spoken. The last witness is unusual: it isn't human, and it only learned to talk a few years ago.

Computer scientists first tried to teach machines language on exactly the premise this essay opened with: write the grammar as rules, build the vocabulary into a dictionary, and have programs parse and generate from there. That research program was pursued seriously for decades, and the outcome is well known — fine on toy sentences, shattered on real language, with rules multiplying and exceptions patching exceptions, forever chasing what the language actually does. Fred Jelinek, the speech-recognition pioneer, reportedly joked that every time he fired a linguist, recognition accuracy went up. The joke became a signpost for the whole field.

The real breakthrough came from abandoning rules-first altogether: statistical methods, then deep learning, then large language models. Computers handle natural language passably for the first time not because someone finally finished writing out the grammar, but because the models were soaked in trillions of words of real text and tallied the conventions one by one — the same road Wittgenstein pointed down. Rich Sutton compressed this history into the famous essay "The Bitter Lesson": approaches that build in human expert knowledge lose, again and again, to approaches that learn it themselves from computation and data.

The word "large" is the point. If language really reduced to a few rules, models wouldn't need to be large at all — install the rules and be done; a few thousand lines of code would suffice. Models must be enormous, and corpora oceanic, precisely because the substance of language is an irreducible inventory of conventions that can only be memorized item by item. A large model's parameter count can be read as an engineering answer to the question "how many conventions does a language actually contain?" — and the answer runs to billions.

Better still, look at how they're trained; it will feel familiar. Hide the next word; make the model guess; check the answer; adjust; next question. Billions of items a day. BERT's training objective is literally called a cloze task in the paper — the same question format as TOEIC Parts 5 and 6. In other words, the only successful large-scale project in history to teach a natural language to "someone" other than a human being used drilling — merely at astronomical volume.

Of course, brains are not transformers, and the analogy shouldn't be pushed too far. But the direction of the result is hard to dodge. Machines hold no philosophy and carry no opinions about East Asian study culture; they report only what works. The derive-it-from-rules route failed. The absorb-the-conventions-at-scale route produced the thing you can now talk to.


What to Keep Separate, and What to Concede

A defense is only credible if it draws its boundaries. Two distinctions first:

Drilling is not pattern drill. The behaviorist "pattern practice" of the 1960s — mechanically transforming sentence frames with no meaning attached — was rightly discarded by language pedagogy, and if that is what the critics mean, they are correct. But TOEIC items are not that: every item makes you process meaning — read a passage or follow a conversation, then commit to a judgment about what it meant. In Krashen's terms, this is comprehensible input; a good question bank is a stack of comprehensible input, sorted by difficulty, with explanations stapled on.

In a comprehension test, "test prep" and "learning" nearly coincide. TOEIC L&R has no essay to template, no free response to game; two hundred items all ask the same two questions — did you understand what you read, did you understand what you heard. In a test of comprehension, the best test technique is comprehension itself. This is the structural reason TOEIC rewards drilling: the road to the score and the road to the language are, unusually, the same road.

Then three concessions:

  1. Drilling doesn't train production. Swain's output hypothesis and Long's interaction hypothesis both stand: to speak and write, you must actually speak and write. Drilling takes care of the input side; arrange the output side separately. (In fairness: TOEIC L&R only tests the input side anyway.)
  2. Memorizing answer keys isn't drilling. If "drilling" means memorizing "this one is B," the critics win outright. The learning lives inside the item: the passage read, the audio heard, the explanation studied, the missed questions revisited. What you drill is the English in the questions, not the letters beside them.
  3. Question quality bounds exposure quality. What you are mass-exposed to is the language inside the items — so the language inside the items had better be real English. A shoddy bank exposes you, at volume, to conventions that don't exist, which is worse than not drilling at all. This is why we treat difficulty and originality as engineering problems (see the Difficulty Overview for how).

Closing: Grind in Peace

Back to the original charge: "drilling is test prep, not learning." It can now be answered in full.

If language really were a rule set, the charge would stick — measuring doesn't improve the measured. But a philosopher spent the second half of his life showing that language is no such thing; linguists counted conventions at half of all real text; psychologists proved that answering questions stores more than rereading and that errors teach more than error-free study; and the engineers supplied the living proof — the only student ever successfully taught a natural language was trained by drilling at astronomical scale. Four testimonies, one conclusion: the central work of learning a language is meeting the way it is actually used — at volume, repeatedly, with feedback — and that is precisely the job drilling does, cheaply and well.

Language doesn't listen to reason, so your only option is to keep meeting it. Drilling is how you schedule the appointments — on the train, in a queue, with one hand free.

Our app's Chinese name, 單手刷990, means "grind to 990 with one hand." The grind is right there in the title. It isn't self-mockery; it's a thesis.

← All articles