All articles · Published 2026-08-30

TOEIC study plans that plan themselves: how Handy 990 decides what you practice today

Open Handy 990 and the home screen lists today's work: a 23-minute smart session, 3 mistakes to review, 14 vocabulary cards, a 1/3-length listening mock. None of those numbers is random, and none of them is a black-box AI hunch — every one comes from a rule we can write down. This article opens the machinery: why the plan never asks for more time than you said you have, why the mock tests are a third of a full form, why your weakest Part gets the most minutes, and why the Parts you've already mastered don't disappear. Our position is simple: a plan you can read is a plan you can trust with three months of your life.

Why do self-made study plans die in week two?

Because when people plan, they plan for an idealized version of themselves.

Psychology has a formal name for this: the planning fallacy. The classic experiment is Buehler, Griffin and Ross (1994, Journal of Personality and Social Psychology): students were asked to estimate how many days they'd need to finish their senior theses. The average estimate was 33.9 days — the average reality was 55.5, and only about a third finished within their own prediction. Note what these students were estimating: the task they knew best, with the most at stake. They still missed by sixty percent.

TOEIC study plans are the same story. Sunday-evening you generously schedules "two hours a day, plus a full mock test on Saturday" for the whole week. Wednesday-at-9pm you, just home from overtime, looks at that schedule and closes the app. The plan wasn't wrong. It was written for a different person.

So the first design principle of this system is: a small plan that gets finished beats a grand plan that gets abandoned. A finished third is worth more than an abandoned whole — nearly every rule below is some version of that sentence.

It asks how much time you have, not how much you "should" spend

When you set an exam date, the app asks one question: how long can you study per day? From then on, the entire plan is capped by that number — it will never ask for more time than you yourself said you have.

Within that ceiling, the closer the exam, the larger the share the plan takes. Days-to-exam splits the run-up into four phases, each drawing a fixed fraction of your stated maximum:

Phase Days to exam Share of your ceiling
Foundation over 42 65%
Build 15–42 80%
Sprint 4–14 100%
Eve 3 or fewer light review only

Foundation taking only 65% isn't politeness — it's strategy. Run you at full throttle from day one and the plan won't survive its first month. Sprint takes everything, because those two weeks come with an exam date for fuel. And the last three days flip the other way: no new practice, no mocks — just a pre-exam cram deck of your collected mistakes and vocabulary, so you walk in with everything sorted.

Two more rules from the same family: rest days you configure are respected, not argued with — the whole list collapses to a single "rest" row; and on a day with a scheduled mock, the practice session is automatically halved — asking you to sit an exam and do a full session in the same evening is how you teach someone to skip both.

Why are the mock tests only 1/3-length?

A full TOEIC mock runs about two hours; with review, over two and a half. For most working test-takers that only ever happens on a weekend — which is why the mocks that apps schedule (including an earlier version of ours) mostly share one fate: postponed indefinitely.

An abandoned mock is the worst possible outcome. It occupies a whole evening's worth of mental budget and produces no data at all. A short, finished mock is the opposite — fewer questions means more noise per sitting, but it's real evidence. So the current rules:

  • Routine mocks are fractional: 1/3 of a form in Foundation and Build, 1/2 in Sprint.
  • Listening and reading are split into two separate tasks: a 1/3 listening mock is about 15 minutes, a 1/3 reading mock about 25 — do listening on the morning commute and reading at night, tick them off independently.
  • Mock answers carry more weight in the score estimate: ordinary practice counts as 1× evidence, a mock 1.5×, a serious-mode mock 2× — performance under exam conditions says more about real ability than practice you can pause. This is also why the plan insists on regular mocks at all: it's collecting the highest-grade samples for the estimation engine described in How Handy 990 estimates your TOEIC score.

There is exactly one exception: your last mock before the exam is a full-length dress rehearsal, recommended in serious mode. Fractional mocks train the questions; only the full two hours trains the two hours — that specific heaviness in your head around question 150 is something no shortened version can preview for you. The whole arrangement is "most sittings shortened, one rehearsal complete," and the dates are listed right on the home screen under "Coming up."

Writing the dates down is itself backed by the literature, incidentally. Gollwitzer and Sheeran (2006, Advances in Experimental Social Psychology) meta-analyzed 94 independent tests and found that upgrading a goal from "I will do X" to "I will do X at this time and place" improves goal attainment with an average effect size of d = 0.65 — large, by behavioral-science standards. "I should take a mock sometime" is a wish. "1/3 listening mock by 9/26" is a plan.

How does Smart Practice split your 30 minutes?

Tap Smart Practice, give it a time budget, and it hands back a per-Part allocation. Four rules in there are worth knowing.

First: the further behind a Part is, the more time it gets. Each Part's weight is its share of questions on the real exam multiplied by your gap to target. The neediest Part can receive over eight times the minutes of a Part you've already mastered. This is the everyday version of what What Moneyball Teaches Us argues: outsource the "what should I practice" decision to the data.

Second: a Part you've mastered doesn't disappear. An early version got this wrong — the weight of an at-target Part wasn't cut deeply enough, so someone already strong at Part 3 kept getting Part 3 every day, because Part 3's large exam share out-multiplied a genuinely lagging Part 5. We later cut a mastered Part's weight to roughly forty percent of its old value: it stays in the rotation to keep your touch, but it never again outbids a Part that's actually behind. We tell this mistake in public because it demonstrates the point of the whole system: rules you can write down are rules you can catch being wrong, and fix.

Third: every Part has a minimum worthwhile portion, and a Part that can't be funded to it sits out. Cram seven Parts into thirty minutes and each gets four — deep enough to practice nothing. So a Part that can't reach its minimum (at least 5 questions of Part 5, at least one whole passage of Part 7) is set aside — but not forever: each new session rotates one squeezed-out Part back in, so a short budget never permanently neglects the same Parts.

Fourth — the one that surprises people most: half of your time budget is reserved for reading the explanations. Every question's time estimate is multiplied by 2 — one share for answering, one share for the explanation, the translations, and understanding why you were wrong. Without that ×2, the same thirty minutes would fit twice as many questions, every one of them wasted. The value of drilling is in the review — we made that argument in In Defense of Drilling; here we just wrote it into the algorithm.

The hard parts in practice

As usual, we finish by saying plainly what this system cannot do.

The plan is rules, not mind-reading. It knows your answer history; it doesn't know your mood, last night's sleep, or that it's quarter-end at work. What it gives you is a sensible default, and you can always adjust — mock length can go up (not down; that's deliberate), per-Part question counts can be edited by hand, rest days are yours to set.

It is also deliberately not AI. The same state in always produces the same plan out — no black box, no surprises; the rules you see today are the rules you'll see in three months. Why we think boring rules beat clever models for this particular job is argued in full in I Already Have ChatGPT. Why Would I Need a TOEIC App?

Fractional mocks are noisier per sitting. A 1/3 listening mock is thirty-odd questions, so any single result swings more than a full mock's would. That's the price of "actually gets finished," and we think the trade is good — but you should know its terms. Whether the estimation engine's multi-sample smoothing actually holds up against real scores is a matter of public record: see Four test-takers sent us their real scores.

It only sees the you inside the app. The workbook from your prep class, the podcast on your commute — the plan knows nothing about them. If the app is one part of your preparation, treat its plan as your minimum effective dose, not the whole prescription.


The daily plan sits at the top of the Handy 990 home screen and activates once you set an exam date; if you prefer minimalism, switch the home to Mission Mode and the screen reduces to just today's work. For how the score estimate is computed, read How Handy 990 estimates your TOEIC score; for how it performed against real score reports, read Four test-takers sent us their real scores.

← All articles