𰻞
辶 穴 月 幺 長 言 馬 幺 長 刂 心
biáng, as in biángbiáng noodles · 58 strokes · eleven parts, every one of them common
The topic

Chinese character classification

a printed glyph of 集 cut from a page
→
classifier
one output per character
→
集
U+96C6
  • Input: one glyph, already cut out of the page.
  • Output: which character it is, among thousands of classes.
  • The reading step of OCR, for scripts with no alphabet.

Why it is hard

Same class, different looks:

集集, another print 集

Different classes, near-identical looks:

未 末 · 己 已 巳 · 呆 杏

Andfartoomanyclasses!
The problem

A classifier cannot output a class it never saw

3,755
classes in the standard training set
(enough for daily reading)
97,000+
Han characters in Unicode
(old books use the long tail)
0 %
a softmax's accuracy
on everything else
Zero-shot: recognise a character from its description, as you would read a new word by spelling it.
The shared idea

A character is a tree of radicals

森 = ⿱ 木 ⿰ 木 木 top bottom left right ⿱ 木 ⿰ 木 木
  • ~500 radicals and 12 layout operators (⿰ left-right, ⿱ top-bottom…) build every character.
  • Unicode writes the tree as a string, the IDS. Every paper reads the same file.

The benchmark

Train on m handwritten classes (500 to 2,755). Test on 1,000 classes never seen. Fewer training classes = harder.

Three ways to use the tree

Predict it, encode it, or align with it

1

Predict

Read the image, write out its tree, look it up in a dictionary.

≈ image captioning

⿱木⿰木 ? next token, one at a time
2

Encode

Turn each tree into a fixed vector; train the image network to hit it.

≈ attribute-based zero-shot

森呆杏林 image fixed codes
3

Align

Learn an encoder for the tree too, and match images to trees.

≈ CLIP

森呆林 IDS strings images match= diagonal
Family 1 · Predict

Write the tree, then look it up

森
→
CNN
→
attention decoder
→
⿱ 木 ⿰ 木 木
→
dictionary lookup
→
森
RAN: CNN encoder, RNN decoder whose
          attention moves over the character, one radical per step
Each radical the decoder writes comes with an attention map: where it looked. RAN, Zhang et al., ICME 2018

RAN / DenseRAN 2018 → RSST 2024 → CDC-RAN 2025

  • Interpretable: one radical, one region.
  • Errors pile up along the sequence; one wrong token, wrong character.
The decoder learns to write combinations it has seen. Faced with a new character, it writes a familiar one.
Family 2 · Encode

Compile the tree into a fixed vector

森
→
CNN
→
vector
↔
nearest class code
←
code, designed by hand from the tree
Building a code: HDE (2020), simplified
$\varphi(y)=\sum_{\text{radical } r} \alpha^{d_r}\, e_r \;+\; \lambda \sum_{\text{operator } s} \alpha^{d_s}\, e_s$

one dimension per radical / operator; $d$ = depth in the tree; $\alpha=\lambda=\tfrac12$

森 = ⿱ 木 ⿰ 木 木depthadds
⿱0$\lambda\alpha^0 = 0.5$ to $e_{⿱}$
木 (top)1$\alpha^1 = 0.5$ to $e_{木}$
⿰1$\lambda\alpha^1 = 0.25$ to $e_{⿰}$
木, 木2$2\alpha^2 = 0.5$ to $e_{木}$

⇒ $\varphi(森)$: 木 = 1.0, ⿱ = 0.5, ⿰ = 0.25, every other dim 0.

HDE 2020 → CUE 2023 → HierCode 2025 → JRED 2025

$\hat y = \arg\max_{y\,\in\,\text{all classes}} \cos\big(f(x),\,\varphi(y)\big)$

train the image network $f$ on seen classes; unseen ones only need their $\varphi$.

The code is frozen before training. 呆 (口 over 木) and 杏 (木 over 口) get the same counts: only a tiny position term, in the 4th decimal, tells them apart.
Family 3 · Align

Learn both sides: CLIP, with the tree as caption

森
→
image encoder
→
similarity
contrastive loss
←
IDS encoder
←
⿱ 木 ⿰ 木 木
at test time: nearest of
all candidate trees
What moved: the IDS encoder. Same images, same strings, same benchmark (m = 2,755)
⿱木 ⿰木 木 ⿱ 木 ⿰ 木 木 ⿱ 木 ⿰ 木 木 1st 2nd 1st 2nd sequence · Transformer tree, read leaves → root graph, root → leaves, child rank CCR-CLIP 2023 · 73.0
Structure written nowhere: learned from token order.
FT-CLIP 2025 · 79.0  each node attends to its children, and knows where they sit.
GRSTR 2026 · 80.6  a GNN pushes layout down, and 1st / 2nd child differ: 呆 ≠ 杏.
Next step, outside the encoder: match image regions to each radical, not one vector to one vector. Hi-GITA 2025: 85.4.
And without the tree?

Match the image to a printed font

a printed glyph of 集
→
image encoder
→
similarity
←
same encoder
←
集
one font rendering
per candidate class

OpenCCD CVPR 2022 → PCSS 2024

  • Same CLIP idea, but the "caption" is a picture of the character in a font, not its tree.
  • Works where no tree exists: Japanese kana never seen in training, rare variants.
Needs a font that draws every class. Fewer than 5 fonts cover all 97k characters, and a modern font is not the shape of a 17th-century woodblock.

On the next chart it is drawn dashed: each unseen class comes with a printed glyph the tree methods never see, so it is not a like-for-like comparison.

Ten methods, one benchmark

Retrieval beats prediction, by a lot

Top-1 (%), 1,000 unseen characters  ·  ○ 500 training classes  ● 2,755 0 10 20 30 40 50 60 70 80 90 100 1 · PREDICT DenseRAN 2018 30.7 RSST 2024 47.4 2 · ENCODE HDE 2020 33.5 HierCode 2025 56.2 JRED 2025 56.3 3 · ALIGN CCR-CLIP 2023 73.0 FT-CLIP 2025 79.0 GRSTR 2026 80.6 Hi-GITA 2025 85.4 4 · FONT MATCH (OTHER PROTOCOL) OpenCCD 2022 95.5 at m=2,000

HWDB → ICDAR 2013, Top-1. GRSTR, GL-HPN, Hi-GITA, OpenCCD tables. Dashed: given a printed glyph per unseen class.

+25
points. The first CLIP-style model beat six years of predictors.
5,400×
faster. Decoding: 1,666 ms per image. Retrieval: 0.31 ms.
10×
How finely you split into radicals: +27 pts. A better backbone: +2.7.
Where my work fits

A learned embedding, used to fix a reader's mistakes

  • I train a contrastive glyph encoder, family 3's recipe on the image side.
  • Each glyph's nearest neighbours vote on its label: a graph over the whole book.
What the graph is for
  1. Detect errors: a glyph whose reading disagrees with its look-alikes.
  2. Catch hallucinations: the reader writes a plausible character that is not on the page.
  3. Expose model conventions: one reader writes the Taiwan form, another the Hong Kong form, of the same glyph.

Run on a real Qing-dynasty volume: 770 pages, no ground truth.

24 glyphs corrected on
        volume 611: original reading, correction, and a second reader's verdict
Corrections kept on volume 611. Blue: a second, independent reader agrees.
謝謝

Thank you for your attention

Questions?