Python and HWP files: extract text and tables, convert to PDF (tested on Hangul 2024)
HWP is the word-processor format of Hancom's Hangul, and in Korea it is everywhere: government notices, school forms, contracts. Sooner or later someone needs to pull the text or a table out of hundreds of them, or turn them into PDFs. This guide does those three jobs from Python and shows what actually happened on a Windows 11 PC with Hancom Office 2024 (Hwp.exe 13.0.0.3380): which calls worked, the security prompt that freezes unattended scripts, why the officially documented fix did not take effect on this version, and the workaround that removed the prompt in every test. There are also two routes that need no Hangul at all, for .hwpx and for .hwp.
Three routes, and what each can do
There are two file formats to deal with. .hwp is the older binary format: an OLE compound file, the same container old .doc files used. .hwpx is the newer one: a ZIP archive of XML files, much like .docx. Hangul saves both. What you can do from Python depends on which file you have and whether Hangul is installed on the machine running the script.
| Route | Needs | Text | Tables | |
|---|---|---|---|---|
| Hangul + pywin32 | Windows, Hangul, pip install pywin32 | Yes | Yes (rows and columns) | Yes |
| .hwpx + standard library | Python only | Yes | Yes (rows and columns) | No |
| .hwp + olefile | pip install olefile | Yes | One cell per line | No |
Test machine for everything below: Windows 11, Hancom Office 2024 (Hwp.exe 13.0.0.3380, a 32-bit program), Python 3.11, pywin32 312, olefile 0.47. Hancom's developer site states that its automation interface is free for personal, non-commercial use; commercial products need a separate licence from Hancom.
Starting Hangul from Python
Hangul registers a COM object called HWPFrame.HwpObject. With pip install pywin32 you can start it like any other COM server. Here it took 2.3 to 2.8 seconds to come up, and the window stayed hidden (Visible was False) unless the script turned it on.
import win32com.client as win32
hwp = win32.gencache.EnsureDispatch("HWPFrame.HwpObject")
print(hwp.Version) # 13, 0, 0, 3380 on the test PC
print(hwp.XHwpWindows.Item(0).Visible) # False - it runs hidden
hwp.Quit()The method names below come straight from the type library that ships with Hangul (HwpAutomation.tlb) and from Hancom's API manual (HwpAutomation, April 2025 edition): Open(filename, Format, arg), SaveAs(Path, Format, arg), GetTextFile(Format, option), Clear(option) and Quit().
The prompt that freezes your script
The first time a script opened or saved a file in an ordinary folder, Hangul showed this dialog and waited:
D:\WebSite\GUI_Tools\hwp_automation_test\sample.hwp
νκΈμ μ΄μ©νμ¬ μ νμΌμ μ κ·Όνλ €λ μλ(νμΌμ μμ λλ μ μΆμ μν λ±)κ° μμ΅λλ€.
μ μμ μΈ μμ
κ³Όμ μλ§ μ κ·Όμ νμ©νμμμ€.
[μ κ·Ό νμ©(Y)] [λͺ¨λ νμ©(N)] [νμ© μ ν¨(A)] [λͺ¨λ μ ν¨(C)]In English: someone is trying to access this file through Hangul, which could damage or leak it; allow access only for normal work. The buttons are Allow, Allow all, Deny and Deny all. Nothing else happens until a person clicks. For a script meant to run unattended over hundreds of files, that is the end of it: the test script sat waiting until it was killed after six seconds.
The same steps in the user temp folder (%TEMP%) never showed the prompt. Saving as HWP took 0.33 seconds, saving as PDF 0.32 seconds and reopening 0.02 seconds, with no window at all. Hancom's own sample checker module has an IsTempPath rule that allows the temp folder, which is consistent with what I saw, although Hancom does not document the built-in behaviour.
The official fix, and why it did not work here
Hancom's documented answer is a security module: a DLL in the 'security module (Automation)' download, registered under HKEY_CURRENT_USER\Software\HNC\HwpAutomation\Modules with its full path, then enabled in code with hwp.RegisterModule("FilePathCheckDLL", "FilePathCheckerModuleExample"). I followed that exactly on the test PC. RegisterModule returned False, and the prompt still appeared.
To rule out my own mistakes I checked several things. The DLL is 32-bit like Hwp.exe and loads fine into a 32-bit Python. The value name was tried as both FilePathCheckerModuleExample and FilePathCheckerModule (the name most tutorials use), with the window hidden and shown. Smart App Control was off and no code-integrity block was logged. Every combination still returned False. The decisive check was the list of modules loaded into Hwp.exe: 183 before RegisterModule and the same 183 after, so Hangul never loaded the DLL at all. Tutorials report this working on earlier versions; on Hangul 2024 13.0.0.3380 it did not.
If you do register it, know what it does. In the source code that ships in the same ZIP, the exported IsAccessiblePath function begins with return TRUE;. It approves every path for any program on your account that asks for that module name. That is acceptable on a personal machine running your own scripts, but it is a real loosening of a security check.
Pull out the text and tables
This script takes a .hwp or .hwpx from any folder, copies it into the temp folder, prints the text and then every table as a list of rows. Passing an empty format string to Open lets Hangul detect the type. forceopen:true stops the read-only dialog from appearing when a file has to open read-only.
import shutil, sys, tempfile
from html.parser import HTMLParser
from pathlib import Path
import win32com.client as win32
class TableParser(HTMLParser):
"""Collect tables from GetTextFile("HTML") as [[row], [row], ...]."""
def __init__(self):
super().__init__()
self.tables, self.row, self.cell = [], None, None
def handle_starttag(self, tag, attrs):
if tag == "table":
self.tables.append([])
elif tag == "tr":
self.row = []
elif tag in ("td", "th"):
self.cell = []
def handle_endtag(self, tag):
if tag in ("td", "th") and self.cell is not None:
self.row.append("".join(self.cell).strip())
self.cell = None
elif tag == "tr" and self.row is not None:
self.tables[-1].append(self.row)
self.row = None
def handle_data(self, data):
if self.cell is not None:
self.cell.append(data)
src = Path(sys.argv[1])
work = Path(tempfile.gettempdir()) / "hwp_work"
work.mkdir(exist_ok=True)
tmp = work / src.name
shutil.copy2(src, tmp) # Hangul only ever opens the copy in %TEMP%
hwp = win32.gencache.EnsureDispatch("HWPFrame.HwpObject")
hwp.Open(str(tmp), "", "forceopen:true") # "" = detect .hwp or .hwpx
text = hwp.GetTextFile("UNICODE", "")
html = hwp.GetTextFile("HTML", "")
hwp.Clear(1)
hwp.Quit()
tmp.unlink()
print(text)
parser = TableParser()
parser.feed(html)
for i, table in enumerate(parser.tables, 1):
print(f"Table {i}:", table)On the test document, a title, a three-by-three inspection table and a sign-off line, the plain text put every table cell on its own line (μ₯λΉ, μν, μ κ²μΌ - device, status, checked - then the next row's cells). The HTML pass gave the table back with its structure intact: [['μ₯λΉ', 'μν', 'μ κ²μΌ'], ['ESP32 보λ', 'μ μ', '10/01'], ['USB νλΈ', 'κ΅μ²΄ νμ', '10/02']]. Run against a file in an ordinary folder on D:, it finished with no prompt.
Use UNICODE, not TEXT. Hancom's manual says TEXT drops anything that only exists in Unicode, such as hanja, archaic Hangul and special symbols. The manual also notes that GetTextFile copies the document three or four times in memory; for very large files, saving to disk with SaveAs is the recommended route.
Convert a whole folder to PDF
Same idea in a loop: copy each .hwp into the temp folder, open it, save it as PDF there, then move the PDF next to the original. versionwarning:false suppresses the 'written in a newer version' message that would otherwise block on newer files.
import shutil, sys, tempfile, time
from pathlib import Path
import win32com.client as win32
src_dir = Path(sys.argv[1]) # folder with .hwp files
work = Path(tempfile.gettempdir()) / "hwp_work" # Hangul only works inside %TEMP%
work.mkdir(exist_ok=True)
hwp = win32.gencache.EnsureDispatch("HWPFrame.HwpObject")
t0 = time.perf_counter()
done = 0
for src in sorted(src_dir.glob("*.hwp")):
tmp_in = work / src.name
tmp_out = tmp_in.with_suffix(".pdf")
shutil.copy2(src, tmp_in)
if hwp.Open(str(tmp_in), "HWP", "forceopen:true;versionwarning:false"):
if hwp.SaveAs(str(tmp_out), "PDF", ""):
shutil.move(tmp_out, src.with_suffix(".pdf"))
done += 1
hwp.Clear(1) # close without saving
tmp_in.unlink(missing_ok=True)
hwp.Quit()
print(f"{done} converted in {time.perf_counter() - t0:.1f}s")Ten one-page documents in a folder on D: converted in 2.4 seconds, not counting the 2 to 3 seconds Hangul takes to start, and every output started with %PDF-1.6. The same SaveAs call with "HWPX" converts old .hwp files to .hwpx (both worked in testing), which is handy if you later want to use the no-Hangul route below.
No Hangul: .hwpx with the standard library
A .hwpx file is a ZIP archive. The body lives in Contents/section0.xml (and section1.xml and so on for more sections), with text in hp:t elements and tables in hp:tbl > hp:tr > hp:tc. Python's zipfile and xml.etree are enough, so this runs on Linux and macOS too.
import sys, zipfile
import xml.etree.ElementTree as ET
HP = "{http://www.hancom.co.kr/hwpml/2011/paragraph}"
def cell_text(tc):
return "".join(t.text or "" for t in tc.iter(HP + "t"))
with zipfile.ZipFile(sys.argv[1]) as z:
for name in sorted(n for n in z.namelist() if n.startswith("Contents/section")):
root = ET.fromstring(z.read(name))
for p in root.findall(HP + "p"): # body paragraphs (not the ones inside tables)
line = "".join(t.text or "" for run in p.findall(HP + "run") for t in run.findall(HP + "t"))
if line:
print(line)
for tbl in p.iter(HP + "tbl"): # tables keep their rows and columns
print("Table:", [[cell_text(tc) for tc in tr.findall(HP + "tc")] for tr in tbl.findall(HP + "tr")])On the same test document this printed the title, the table as three rows of three cells, and the sign-off line, in document order.
No Hangul: .hwp with olefile
The binary format takes more work. Inside the OLE container are named streams. The body is in BodyText/Section0 (one per section), compressed with zlib and stored as a sequence of records. Each record has a 4-byte header packing a tag, a nesting level and a size, and paragraph text sits in records with tag 67.
import struct, sys, zlib
import olefile
# control characters that take 16 bytes (8 characters) inside paragraph text - placeholders for tables, images and so on
WIDE = {1, 2, 3, 4, 5, 6, 7, 8, 9, 11, 12, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23}
def records(data):
i = 0
while i < len(data):
head = struct.unpack_from("<I", data, i)[0]
i += 4
tag, level, size = head & 0x3FF, (head >> 10) & 0x3FF, head >> 20
if size == 0xFFF: # large records keep the size in the next 4 bytes
size = struct.unpack_from("<I", data, i)[0]
i += 4
yield tag, level, data[i:i + size]
i += size
def para_text(raw):
out, i = [], 0
while i < len(raw):
c = struct.unpack_from("<H", raw, i)[0]
if c in WIDE:
i += 16
continue
if c == 9:
out.append("\t")
elif c >= 32:
out.append(chr(c))
i += 2
return "".join(out)
ole = olefile.OleFileIO(sys.argv[1])
compressed = ole.openstream("FileHeader").read()[36] & 1
for entry in sorted(e for e in ole.listdir() if e[0] == "BodyText"):
data = ole.openstream(entry).read()
if compressed:
data = zlib.decompress(data, -15)
for tag, level, body in records(data):
if tag == 67: # HWPTAG_PARA_TEXT - paragraph text
print(para_text(body))It printed every paragraph of the test file, with each table cell on its own line (cell paragraphs sit at nesting level 3, body paragraphs at level 1). Rebuilding the rows would mean also parsing the table records, which is where converting to .hwpx once with Hangul becomes the easier path.
For a quick peek there is a shortcut: the PrvText stream holds plain, uncompressed preview text, and it even marks table rows as <cell><cell><cell>. It is only a preview, though. On a 12,798-character test document it held the first 1,022 characters, about 8% of the text.
What did not work on this PC
Two things are worth knowing before you build on this. First, the security module, covered above. Second, every action called through HAction.Run returned False on the test PC: moving the cursor (MoveRight, MoveDocBegin), SelectAll, starting a new paragraph, moving between table cells. Calling them from PowerShell's COM support gave the same result, so it is not a pywin32 issue. Actions run with HAction.Execute and a parameter set, such as inserting text or creating a table, did work, and so did MovePos.
That is why this guide sticks to reading and converting, which never needed Run. If you need to fill in forms or edit documents, test those actions on your own Hangul version first; many tutorials rely on them, and I could not confirm them on Hangul 2024 13.0.0.3380.
FAQ
Q. Does this work on macOS or Linux?
The two no-Hangul routes do: zipfile for .hwpx and olefile for .hwp are plain Python. Driving Hangul itself uses COM, which is Windows-only.
Q. Do I need pyhwpx?
No. pyhwpx is a convenience wrapper around the same HWPFrame.HwpObject COM object; everything here uses pywin32 directly. I did not test pyhwpx for this guide.
Q. Why does my script freeze with no error?
Almost certainly the file-access prompt. Hangul runs hidden, so the dialog may appear on screen while the script waits. Work on copies in the temp folder as shown above.
Q. Can I use this in a commercial product?
Hancom's developer site says its automation interface is free for personal, non-commercial use and requires a licence from Hancom for commercial use. The no-Hangul routes read the file format directly and do not involve that interface.
Q. HWP or HWPX - which should I ask people to send?
HWPX if you have the choice. It is a ZIP of XML that any language can read, and tables keep their structure without extra parsing.