Fast Hangul text utilities for Python.
JamoLib helps you decompose Hangul syllables into jamo, compose jamo back into syllables, and convert two-set Korean keyboard input into readable Korean text.
- Linear-time text scanning for large Hangul strings.
- Small API surface that is easy to drop into search, normalization, keyboard-input, and NLP preprocessing pipelines.
- Preserves mixed text such as English, numbers, punctuation, and whitespace.
- Handles common batchim boundary cases like
값이,닭이, and읽어. - Supports common compound-medial combinations such as
ㄱㅗㅏ -> 과.
pip install jamolibimport jamolib
text = "한글과 English 123"
decomposed = jamolib.decomposeHangulText(text)
print(decomposed)
# ㅎㅏㄴㄱㅡㄹㄱㅘ English 123
print(jamolib.composeHangulText("ㄱㅏㅂㅅㅇㅣ"))
# 값이
print(jamolib.translateEngToKor("dkssudgktpdy"))
# 안녕하세요| Function | Input | Output | Use case |
|---|---|---|---|
decomposeHangul |
Single Hangul syllable | Compatibility jamo string | Token-level preprocessing |
decomposeHangulText |
Mixed text | Text with Hangul syllables decomposed | Search normalization, phonetic indexing |
composeHangul |
초성 + 중성 [+ 종성] |
Single Hangul syllable | Rebuilding syllables |
composeHangulText |
Jamo text | Re-composed Hangul text | UI input handling, postprocessing |
translateEngToKor |
Two-set English keyboard input | Hangul text | Keyboard typo correction |
getCharset |
None | Supported compatibility jamo list | Validation and custom pipelines |
examples/quickstart.py: decomposition, composition, and keyboard conversion in one scriptexamples/mixed_text.py: preserving non-Hangul text while processing Hangulexamples/batchim_boundaries.py: tricky batchim and syllable-boundary cases
Run an example from the repository root:
python examples/quickstart.pydecomposeHangulexpects a single Hangul syllable.composeHangulexpects compatibility jamo in the order초성 + 중성 [+ 종성].composeHangulTextalso combines common compound medials likeㅗㅏ,ㅜㅓ, andㅡㅣ.translateEngToKoruses the standard two-set Korean keyboard mapping.- Mixed strings are preserved as-is outside Hangul processing.
The current implementation uses a single-pass scanner instead of repeated global string replacement. Local measurements on Python 3.12 in this repository produced the following averages:
| Operation | Input shape | Average time |
|---|---|---|
decomposeHangulText |
Repeated Hangul sentence x500 | 0.0041s |
composeHangulText |
Recompose decomposed sentence x500 | 0.0205s |
translateEngToKor |
Keyboard string x2000 | 0.0123s |
These numbers are environment-dependent, but they reflect the optimized code currently in the repository.
python -m pip install -e .[test]
pytest
python scripts/benchmark.py
python -m build
