이 블로그의 코드 글꼴은 한글을 정확히 두 칸으로 그립니다. 왜 그래야 하는지는 터미널이 글자의 폭을 정하는 방식에서 나옵니다. 아래 속성 값은 Unicode 18.0의 EastAsianWidth.txt에서 확인했고, 코드 출력은 Node.js(ICU 78.3)에서 실행한 결과입니다.
폭은 문자에 붙은 속성입니다
Unicode는 모든 문자에 East_Asian_Width라는 속성을 붙여 둡니다(UAX #11). 터미널은 대개 이 값을 보고 글자 하나에 몇 칸을 줄지 정합니다.
값
뜻
칸
예
W
Wide
2
한ㄱ漢
F
Fullwidth
2
! (U+FF01)
Na, H
Narrow, Halfwidth
1
a1
A
Ambiguous
1 또는 2
§·→─
N
Neutral
대개 1
먼저 W나 F인지 봅니다. 한글 완성형 음절 11,172자(U+AC00–U+D7A3)는 모두 W라서 두 칸입니다.
W도 F도 아니면 A인지 봅니다. A는 원래의 문자 집합에 따라 폭이 달랐던 문자들입니다.
A는 동아시아 환경으로 설정된 터미널에서는 두 칸, 그 밖에서는 한 칸입니다. 같은 글이 터미널에 따라 어긋나는 이유가 대개 여기에 있습니다.
조합형의 중성과 종성은 앞 글자에 붙어 한 글자를 이룹니다. 그래서 대부분의 wcwidth 구현은 이들을 0칸으로 셉니다. 한을 NFD로 풀면 코드 포인트 세 개가 되지만, 초성의 두 칸만 남아 여전히 두 칸입니다.
한글 (NFD) → 1112 1161 11AB 1100 1173 11AF
모호한 폭
§·→, 그리고 상자를 그리는 ─│(U+2500–U+254B)는 모두 A입니다. EUC-KR 같은 예전 동아시아 문자 집합에서는 이 문자들이 두 칸이었기 때문입니다. 터미널마다 모호한 폭을 두 칸으로 볼지 정하는 설정이 있고, 그 설정과 글꼴이 실제로 그리는 폭이 다르면 상자의 선이 어긋납니다.
이 블로그의 코드 글꼴 Monoplex KR은 상자 문자를 반각으로 그립니다. 그래서 모호한 폭을 한 칸으로 보는 일반적인 환경에서 한글이 섞인 상자도 맞습니다.
┌────────┬──────┐
│ 이름 │ 칸 │
├────────┼──────┤
│ 한글 │ 4 │
│ abc │ 3 │
└────────┴──────┘
문자열 길이는 칸 수가 아닙니다
JavaScript의 length는 UTF-16 코드 단위의 수입니다. 같은 한글도 완성형이면 2, 조합형이면 6입니다. 칸을 세려면 글자를 사용자가 보는 단위(grapheme)로 나눈 다음, 글자마다 폭을 정해야 합니다.
Why one Hangul syllable is as wide as two Latin letters, why arrows and box-drawing characters change width from one setup to the next, and why string length is not a column count.
The code font on this blog draws each Hangul syllable exactly two columns wide. Why that matters comes from how a terminal decides how wide a character is. The property values below were checked against EastAsianWidth.txt from Unicode 18.0, and the code output was produced by Node.js (ICU 78.3).
Width is a property of the character
Unicode gives every character an East_Asian_Width property (UAX #11). Terminals mostly read that value to decide how many columns a character gets.
Value
Meaning
Columns
Examples
W
Wide
2
한ㄱ漢
F
Fullwidth
2
! (U+FF01)
Na, H
Narrow, Halfwidth
1
a1
A
Ambiguous
1 or 2
§·→─
N
Neutral
usually 1
First, is it W or F? All 11,172 precomposed Hangul syllables (U+AC00–U+D7A3) are W, so two columns.
If it is neither, is it A? Ambiguous characters are the ones whose width depended on the character set they came from.
An A character is two columns in a terminal set up for East Asian text and one column everywhere else. When the same text lines up in one terminal and not in another, this is usually why.
Hangul
Hangul can be encoded more than one way, and each way has its own property values.
Precomposed syllables, U+AC00–U+D7A3, are W.
Compatibility jamo such as ㄱ (U+3131–U+318E) are W too.
Decomposed (NFD), the leading consonants U+1100–U+115F are W, while the vowels and trailing consonants U+1160–U+11FF are N.
Decomposed vowels and trailing consonants attach to the letter before them to form one syllable, so most wcwidth implementations count them as zero columns. Decomposing 한 gives three code points, yet only the leading consonant’s two columns remain: still two columns.
한글 (NFD) → 1112 1161 11AB 1100 1173 11AF
Ambiguous width
§, ·, → and the box-drawing ─│ (U+2500–U+254B) are all A, because older East Asian character sets such as EUC-KR made them two columns. Terminals have a setting for whether ambiguous characters are double width, and when that setting and the width the font actually draws disagree, the lines of a box no longer meet.
Monoplex KR, the code font here, draws box-drawing characters half width. So in the usual setup, where ambiguous means one column, a box with Hangul inside still lines up.
┌────────┬──────┐
│ Name │ Cols │
├────────┼──────┤
│ 한글 │ 4 │
│ abc │ 3 │
└────────┴──────┘
String length is not a column count
In JavaScript, length counts UTF-16 code units. The same 한글 is 2 precomposed and 6 decomposed. To count columns, split the string into the characters a reader sees (graphemes), then give each one its width.
Intl.Segmenter splits the string into the characters a reader sees. The three code points of a decomposed syllable come out as one.
The W and F ranges, and the A range. Hangul, CJK ideographs, kana and fullwidth forms count as wide; for A, only the characters this post mentions.
Each character’s width comes from its first code point. With cjk on, A counts as two columns.
What it returns:
Input
columns
length
한글
4
2
한글 (NFD)
4
6
Redis 키
8
7
─→·
3
3
─→·, cjk: true
6
3
ㄱㄴ
4
2
!
2
1
This function does not cover emoji or every combining sequence. For real use, an implementation that follows the full Unicode tables is safer: string-width in JavaScript, go-runewidth in Go.