~/blog/hangul-width/ko.mdx

터미널에서 한글은 왜 두 칸인가

한글 한 글자가 영문 두 칸을 차지하는 이유, 화살표와 상자 문자의 폭이 환경마다 다른 이유, 그리고 문자열 길이로 칸을 세면 안 되는 이유.

이 블로그의 코드 글꼴은 한글을 정확히 두 칸으로 그립니다. 왜 그래야 하는지는 터미널이 글자의 폭을 정하는 방식에서 나옵니다. 아래 속성 값은 Unicode 18.0의 EastAsianWidth.txt에서 확인했고, 코드 출력은 Node.js(ICU 78.3)에서 실행한 결과입니다.

폭은 문자에 붙은 속성입니다

Unicode는 모든 문자에 East_Asian_Width라는 속성을 붙여 둡니다(UAX #11). 터미널은 대개 이 값을 보고 글자 하나에 몇 칸을 줄지 정합니다.

값뜻칸예
WWide2한 ㄱ 漢
FFullwidth2! (U+FF01)
Na, HNarrow, Halfwidth1a 1
AAmbiguous1 또는 2§ · → ─
NNeutral대개 1
예 아니오 동아시아 설정 그 밖 글자 하나 W 또는 F 두 칸 A 한 칸

먼저 W나 F인지 봅니다. 한글 완성형 음절 11,172자(U+AC00–U+D7A3)는 모두 W라서 두 칸입니다.

W도 F도 아니면 A인지 봅니다. A는 원래의 문자 집합에 따라 폭이 달랐던 문자들입니다.

A는 동아시아 환경으로 설정된 터미널에서는 두 칸, 그 밖에서는 한 칸입니다. 같은 글이 터미널에 따라 어긋나는 이유가 대개 여기에 있습니다.

한글

한글은 표현하는 방법이 여러 가지이고, 방법마다 속성이 다릅니다.

  • 완성형 음절 U+AC00–U+D7A3은 W입니다.
  • 호환 자모 ㄱ(U+3131–U+318E)도 W입니다.
  • 조합형(NFD)으로 풀면 초성 U+1100–U+115F는 W, 중성과 종성 U+1160–U+11FF는 N입니다.

조합형의 중성과 종성은 앞 글자에 붙어 한 글자를 이룹니다. 그래서 대부분의 wcwidth 구현은 이들을 0칸으로 셉니다. 한을 NFD로 풀면 코드 포인트 세 개가 되지만, 초성의 두 칸만 남아 여전히 두 칸입니다.

한글 (NFD) → 1112 1161 11AB 1100 1173 11AF

모호한 폭

§ · →, 그리고 상자를 그리는 ─ │(U+2500–U+254B)는 모두 A입니다. EUC-KR 같은 예전 동아시아 문자 집합에서는 이 문자들이 두 칸이었기 때문입니다. 터미널마다 모호한 폭을 두 칸으로 볼지 정하는 설정이 있고, 그 설정과 글꼴이 실제로 그리는 폭이 다르면 상자의 선이 어긋납니다.

이 블로그의 코드 글꼴 Monoplex KR은 상자 문자를 반각으로 그립니다. 그래서 모호한 폭을 한 칸으로 보는 일반적인 환경에서 한글이 섞인 상자도 맞습니다.

┌────────┬──────┐
│ 이름 │ 칸 │
├────────┼──────┤
│ 한글 │ 4 │
│ abc │ 3 │
└────────┴──────┘

문자열 길이는 칸 수가 아닙니다

JavaScript의 length는 UTF-16 코드 단위의 수입니다. 같은 한글도 완성형이면 2, 조합형이면 6입니다. 칸을 세려면 글자를 사용자가 보는 단위(grapheme)로 나눈 다음, 글자마다 폭을 정해야 합니다.

columns.js
const graphemes = new Intl.Segmenter('ko', { granularity: 'grapheme' })
const WIDE = /^[\u1100-\u115F\u2E80-\u303E\u3041-\u33FF\u3400-\u4DBF\u4E00-\u9FFF\uA960-\uA97F\uAC00-\uD7A3\uF900-\uFAFF\uFE30-\uFE4F\uFF00-\uFF60\uFFE0-\uFFE6]/
const AMBIGUOUS = /^[\u00A7\u00B7\u2190-\u2199\u2500-\u254B]/
function columns(text, { cjk = false } = {}) {
let n = 0
for (const { segment } of graphemes.segment(text)) {
if (WIDE.test(segment)) n += 2
else if (AMBIGUOUS.test(segment)) n += cjk ? 2 : 1
else n += 1
}
return n
}

Intl.Segmenter가 문자열을 사용자가 보는 글자 단위로 나눕니다. 조합형 한의 코드 포인트 세 개는 글자 하나로 묶입니다.

W·F 범위와 A 범위입니다. 한글, 한자, 가나, 전각 문자를 넓게 보고, A는 이 글에 나온 것만 골랐습니다.

글자의 첫 코드 포인트로 폭을 정합니다. cjk를 켜면 A를 두 칸으로 셉니다.

실행한 결과입니다.

입력columnslength
한글42
한글 (NFD)46
Redis 키87
─→·33
─→·, cjk: true63
ㄱㄴ42
!21

이 함수는 이모지나 결합 문자를 모두 다루지는 않습니다. 실제로 쓸 때는 JavaScript의 string-width나 Go의 go-runewidth처럼 Unicode 표 전체를 따르는 구현을 쓰는 편이 안전합니다.

~/blog/hangul-width/en.mdx

Why Hangul takes two columns in a terminal

Why one Hangul syllable is as wide as two Latin letters, why arrows and box-drawing characters change width from one setup to the next, and why string length is not a column count.

The code font on this blog draws each Hangul syllable exactly two columns wide. Why that matters comes from how a terminal decides how wide a character is. The property values below were checked against EastAsianWidth.txt from Unicode 18.0, and the code output was produced by Node.js (ICU 78.3).

Width is a property of the character

Unicode gives every character an East_Asian_Width property (UAX #11). Terminals mostly read that value to decide how many columns a character gets.

ValueMeaningColumnsExamples
WWide2한 ㄱ 漢
FFullwidth2! (U+FF01)
Na, HNarrow, Halfwidth1a 1
AAmbiguous1 or 2§ · → ─
NNeutralusually 1
yes no East Asian setting anything else one character W or F two columns A one column

First, is it W or F? All 11,172 precomposed Hangul syllables (U+AC00–U+D7A3) are W, so two columns.

If it is neither, is it A? Ambiguous characters are the ones whose width depended on the character set they came from.

An A character is two columns in a terminal set up for East Asian text and one column everywhere else. When the same text lines up in one terminal and not in another, this is usually why.

Hangul

Hangul can be encoded more than one way, and each way has its own property values.

  • Precomposed syllables, U+AC00–U+D7A3, are W.
  • Compatibility jamo such as ㄱ (U+3131–U+318E) are W too.
  • Decomposed (NFD), the leading consonants U+1100–U+115F are W, while the vowels and trailing consonants U+1160–U+11FF are N.

Decomposed vowels and trailing consonants attach to the letter before them to form one syllable, so most wcwidth implementations count them as zero columns. Decomposing 한 gives three code points, yet only the leading consonant’s two columns remain: still two columns.

한글 (NFD) → 1112 1161 11AB 1100 1173 11AF

Ambiguous width

§, ·, → and the box-drawing ─ │ (U+2500–U+254B) are all A, because older East Asian character sets such as EUC-KR made them two columns. Terminals have a setting for whether ambiguous characters are double width, and when that setting and the width the font actually draws disagree, the lines of a box no longer meet.

Monoplex KR, the code font here, draws box-drawing characters half width. So in the usual setup, where ambiguous means one column, a box with Hangul inside still lines up.

┌────────┬──────┐
│ Name │ Cols │
├────────┼──────┤
│ 한글 │ 4 │
│ abc │ 3 │
└────────┴──────┘

String length is not a column count

In JavaScript, length counts UTF-16 code units. The same 한글 is 2 precomposed and 6 decomposed. To count columns, split the string into the characters a reader sees (graphemes), then give each one its width.

columns.js
const graphemes = new Intl.Segmenter('ko', { granularity: 'grapheme' })
const WIDE = /^[\u1100-\u115F\u2E80-\u303E\u3041-\u33FF\u3400-\u4DBF\u4E00-\u9FFF\uA960-\uA97F\uAC00-\uD7A3\uF900-\uFAFF\uFE30-\uFE4F\uFF00-\uFF60\uFFE0-\uFFE6]/
const AMBIGUOUS = /^[\u00A7\u00B7\u2190-\u2199\u2500-\u254B]/
function columns(text, { cjk = false } = {}) {
let n = 0
for (const { segment } of graphemes.segment(text)) {
if (WIDE.test(segment)) n += 2
else if (AMBIGUOUS.test(segment)) n += cjk ? 2 : 1
else n += 1
}
return n
}

Intl.Segmenter splits the string into the characters a reader sees. The three code points of a decomposed syllable come out as one.

The W and F ranges, and the A range. Hangul, CJK ideographs, kana and fullwidth forms count as wide; for A, only the characters this post mentions.

Each character’s width comes from its first code point. With cjk on, A counts as two columns.

What it returns:

Inputcolumnslength
한글42
한글 (NFD)46
Redis 키87
─→·33
─→·, cjk: true63
ㄱㄴ42
!21

This function does not cover emoji or every combining sequence. For real use, an implementation that follows the full Unicode tables is safer: string-width in JavaScript, go-runewidth in Go.

ko.mdx

단축키Keyboard

/ ⌘K
글 검색Search posts
t
밝은/어두운 테마Light or dark theme
l
한국어/영어, 읽던 자리 유지Korean or English, keeping your place
[ ]
이전/다음 절Previous or next section
?
이 목록This list

↑↓ 이동move↵ 열기openesc 닫기close