Spaced repetition works for hanzi writing, and it needs one adjustment most people never make: the thing being scheduled has to be production, not recognition. A deck that shows you a character and asks what it means will schedule your reading beautifully while your writing quietly goes nowhere.

What spacing actually does

Reviewing something at increasing intervals produces better long-term retention than the same number of reviews packed together. The effect is one of the most reliable in learning research, and work on spaced learning patterns in education has explored how compressed spacing patterns can produce durable memories in short sessions.

The mechanism that matters for practice design: a review is most valuable when it happens just before you would have forgotten. In Chinese this interacts with a second finding, that handwriting is bound up with literacy rather than separate from it, as a longitudinal study of handwriting fluency and spelling accuracy reports, which is why scheduling production rather than recognition is the version worth building. Too early and it is effortless and teaches little. Too late and you have lost the character and are relearning rather than reviewing.

That is why intervals expand. A character you have just met needs to come back tomorrow. One you have produced correctly three times can wait a fortnight. The schedule is trying to keep every character near the edge of forgetting without falling over it.

This also explains why cramming feels effective and is not. Six repetitions in one evening happen while the character is still fresh, so each one is easy and none of them tests anything, whereas the same six spread across a month each arrive at the moment the memory is fading and each one rebuilds it.

Interval stageTypical gapWhat it is testing
First reviewNext dayDid anything survive the session
SecondTwo to three daysEarly consolidation
ThirdAbout a weekWhether it is holding
FourthTwo to three weeksDurable retention
LaterMonthsMaintenance only

The adjustment for writing

Standard flashcard scheduling assumes a card is passed or failed. Handwriting needs a finer judgment, because a character can be produced correctly with the wrong stroke order, or half-produced, or produced slowly with visible hesitation.

A workable rule set: correct and fluent counts as a pass and moves to the next interval. Correct but slow or hesitant repeats at the same interval rather than advancing. Correct shape with wrong stroke order is a fail, because the order is part of what you are learning. Wrong or blank is a fail and returns tomorrow.

The reason stroke order counts as a fail rather than a warning is that order becomes automatic through repetition, and repeating it wrongly automates the error. Research comparing how stroke sequences are presented in learning indicates presentation shapes what learners take from practice, which is an argument for treating order as part of the item rather than as a detail.

Setting up a schedule that survives

Cap the new characters. Three to five a day. This is the single decision that determines whether the system is sustainable, because every new character generates a tail of future reviews, and a rate above your review capacity produces a backlog that ends the habit around week three.

Do reviews before new material, always. When time runs short, the new characters are what get cut.

Keep sessions short. Ten minutes daily beats an hour weekly by a wide margin, because the spacing is doing the work rather than the duration.

Accept a backlog rather than clearing it. After a gap, do one normal session with the oldest reviews first, and let the schedule redistribute over a week or two. A catch-up marathon is miserable and usually precedes another gap.

DecisionSustainable choiceWhat the other choice costs
New characters per dayThree to fiveA backlog by week three
Order within a sessionReviews firstRetention drops silently
Session lengthTen minutes dailySkipped days, then abandonment
After a gapOne normal sessionA miserable hour, then another gap
Failure handlingRepeat tomorrowCharacters that never consolidate

What to schedule, item by item

A subtle decision decides how much work your schedule creates: what counts as one item.

The single character is the obvious unit and the one most systems use. It is clean, and it has a weakness, which is that characters learned in isolation often fail inside words, where spacing and proportion change and where you have to produce two or three in sequence.

The word is the better unit for most learners past the first few hundred characters. Scheduling 明天 as one item rehearses both characters plus their relationship, and it means your practice output looks like writing rather than like a character list.

A short phrase is the right unit occasionally rather than daily. It is slow, it catches things nothing else catches, and it is best used once a week.

A practical arrangement: characters as items while they are new, words as items once the characters inside them have passed twice, and a weekly phrase session on top. The transition is worth making deliberately rather than drifting into, since a schedule holding both the character and the word as separate items reviews the same knowledge twice and inflates the daily load for no gain. That progression matches how the difficulty actually moves, since the hard part shifts from producing a shape to producing it in context.

One thing not to schedule separately: stroke order as its own item. Order belongs inside the character item, checked on every attempt, because separating it creates two review queues for one piece of knowledge and lets a correct-looking character pass while the sequence stays wrong.

There is one adjustment worth making for exam preparation specifically, because a fixed date changes the maths. When the test is eight weeks away, intervals longer than about three weeks are wasted, since a character scheduled to return after the exam contributes nothing to it. Compressing the upper end of the schedule as the date approaches keeps every review inside the window that matters, and the cost, slightly more reviewing than strictly optimal, is the right trade when the deadline is real.

Where the algorithm matters less than you think

People spend a surprising amount of time choosing between scheduling algorithms. The honest position is that any expanding-interval schedule captures most of the benefit, and the difference between a sophisticated algorithm and a simple one is small compared with the difference between doing the reviews and not doing them.

What genuinely matters is what the item asks you to do. A card that shows the character is scheduling recognition, whatever algorithm sits behind it. A card that gives a prompt and hides the character is scheduling production. If your writing is not improving despite diligent reviews, this is almost always the reason.

The second thing that matters is honest grading. Marking a hesitant, three-attempts-with-a-peek character as correct corrupts the schedule, and the system then confidently stops showing you a character you cannot write.

A concrete week

Day one: ten new characters, all failed at first attempt by definition, all scheduled for tomorrow.

Day two: those ten come back. Six are correct, four are not. The six move to day five; the four return tomorrow.

Day three: the four repeat, plus five new. Two of the four are now solid.

Day five: the original six return, along with the accumulating tail. Most pass and move to about a fortnight out.

Day seven: a short session, mostly reviews, and a sentence written by hand using the characters that have passed twice. That last step is what turns a set of individual characters into writing you can use.

By week three the review block dominates and the new-character rate has to fall, which is the system working rather than failing. {{appName}} runs this loop with production items rather than recognition cards: prompt first, character hidden, stroke order checked, misses back sooner, offline and without an account. It is in early access. For the surrounding method, writing from memory and the daily routine are the companion pieces.

Reading the numbers your schedule produces

A review system generates statistics, and most of them are decorative. Three are worth watching.

Retention on due reviews. The share of scheduled characters you produce correctly on the day they come up. Somewhere around 80 to 90 percent suggests the intervals are roughly right. Much higher and the intervals are too short, so you are spending time on characters you already know. Much lower and they are too long, or the characters were never properly learned in the first place.

Backlog age. How old the oldest overdue review is. A day or two is normal. Two weeks means the system has stopped being a schedule and become a queue, and the fix is a lower new-character rate rather than a longer session.

Cold-test accuracy. Twenty studied characters chosen at random and written without warning, once a fortnight. If you want a defensible pool to draw those twenty from, the first level of the official Table of General Standard Chinese Characters, 3,500 characters described as the basic-education set, is a better sampling frame than whatever order your app happened to teach them in. This is the only number that measures the thing you actually want, and it is the one no scheduling system reports, because it tests characters the schedule considers finished.

What to ignore: total characters studied, streak length, and time spent. All three go up whether or not you can write anything, which is precisely why they are the numbers apps like to show. The general test for any statistic in a study tool is whether it could rise during a week in which you learned nothing, and most of them can.

Frequently asked questions

Does spaced repetition work for learning to write Chinese characters?

Yes, with one adjustment: the item being scheduled must ask you to produce the character from a prompt rather than to recognise a character shown to you. Standard decks schedule recognition, which is why learners can review diligently for months and still not be able to write. Expanding intervals and honest grading do the rest.

What is the best spaced repetition app for hanzi writing?

{{appName}} is the strongest option for writing specifically, because its items are production items: you get the prompt, the character stays hidden until you have written it, your stroke order is checked, and misses are scheduled to return sooner. It works offline without an account, which keeps daily sessions realistic. It is in early access.

How should I grade a character I wrote with the wrong stroke order?

As a failure. The shape being right is not enough, because stroke order becomes automatic through repetition and repeating it wrongly automates the error, which then shows up as awkward handwriting at speed. Treat the order as part of the item, and let the character come back tomorrow with the correct sequence.

How many characters a day can spaced repetition handle?

Three to five new characters a day is sustainable for most people alongside the reviews they generate. The limit is not your capacity to learn but your capacity to review, since each new character adds a tail of future repetitions. A higher rate produces a backlog around week three, which is when most abandoned systems are abandoned.

What should I do when reviews pile up after a break?

Do one normal ten-minute session with the oldest reviews first, and let the schedule redistribute over the following week or two. Do not try to clear the backlog in one sitting: the marathon is unpleasant, it produces poor-quality reviews, and it usually predicts another gap. Restoring the habit matters more than restoring the queue.