If you have ever tried to pull bold or italic text out of a PDF, you probably started with a regex. It feels like the obvious move. Search the font name for the word “Bold,” flag the text, and move on. That strategy works just long enough to give you false confidence. Then someone exports the same document from a different version of Acrobat, or from LibreOffice, or from a print-to-PDF driver, and every assumption you coded falls apart.

Why Font Names Lie

PDF.js will hand you strings like ABCDEF+TimesNewRomanPS-BoldMT. The first six characters are a random prefix injected during export, and they change every time the file is regenerated. Betting your parser on that prefix is betting on noise. Other exporters are even less helpful. Some emit bare identifiers like Font12 or F1. These labels carry no semantic meaning at all; they are internal resource tags that happened to be nearby when the file was written.

Because the Portable Document Format was never designed to make downstream text extraction easy, font names are simply references to embedded or subsetted resources. They were never intended to act as a stable API for style detection. When you write a regex that looks for the substring Bold or Italic, you are scraping a label that the creator application was free to format however it pleased. You are not reading the actual typographic properties. You are reading a file-naming convention, and conventions are not contracts.

Read the Descriptor, Not the Label

The ground truth lives elsewhere. In PDF.js, every page object exposes commonObjs, a map that holds the real font descriptors needed to render the page. When the library parses a page, it populates this map with genuine font objects. Those objects expose boolean properties for bold and italic. Those booleans are not inferred from the name string. They originate from the font descriptor embedded inside the PDF, derived from glyph metrics, OS/2 table flags, and the symbolic descriptors that the typesetter wrote into the document.

That means you can stop guessing. Before you walk the text items on a page, create a fontStyleMap by iterating over page.commonObjs. For each font ID, store an object that records the actual .bold and .italic properties provided by the font object. Later, when you process each text item, look up its font reference in your map and read the precomputed flags. You suddenly have a deterministic answer that does not change between exports.

You can keep a fallback. If the descriptor is somehow absent or incomplete, clean the font name strip the prefix, drop the random tags, and run a conservative regex against what remains. But this should be your last resort, not your primary logic. The difference in reliability is dramatic. Where name-based parsing fractures across exporters, descriptor-based parsing holds steady because it asks the file what it actually contains.

When the Font Lies Too: Synthetic Styles

Even the descriptor can miss a trick. Some PDF creators do not bother embedding a separate italic typeface. Instead, they take the upright roman font and slant it with a transform matrix. This is common in files generated by design tools or older word processors that favor file size over typographic purity.

Every text item in PDF.js carries a transform array, a six-element affine matrix that maps the glyph coordinate system into the page coordinate system. The third element of that array controls horizontal shear. When that value is non-zero, the text is being mechanically slanted by the renderer. If you only trust the font descriptor, you will classify this text as upright roman. If you inspect the matrix, you catch the synthetic italic and mark it correctly. The same logic applies to synthetic bold created by overprinting, though that is harder to detect from geometry alone. For slanted text, the shear value is your smoking gun.

Underlines Are Drawn, Not Declared

Bold and italic are font properties. Underline is not. In PDF, an underline is a vector path. The renderer issues a drawing command for a thin horizontal segment positioned near the text baseline. It is a graphical element that happens to sit underneath glyphs, not a character attribute stored in a cmap.

Sự phân biệt này rất quan trọng vì việc đọc các mô tả phông chữ (font descriptor) sẽ không bao giờ phát hiện ra gạch chân. Bạn cần xem xét các toán tử vẽ thô (raw drawing operators) trên trang, hoặc nói ngắn gọn là đầu ra ở cấp độ hình học (geometry-level), để tìm các đoạn thẳng nằm ngang ngắn chạy song song với đường cơ sở (baseline) ở khoảng cách phù hợp. Khi công cụ trích xuất của bạn phát hiện một đoạn như vậy nằm bên dưới một chuỗi văn bản, bạn sẽ gắn thẻ chuỗi đó là có gạch chân. Việc coi đây là một lớp phát hiện riêng biệt giúp mô hình dữ liệu của bạn luôn chính xác: in đậm và in nghiêng là đặc tính nội tại của phông chữ, trong khi gạch chân là phần trang trí ngoại lai được trình bày bởi tài liệu.

Một Quy trình Hoạt động

Một hệ thống trích xuất chuẩn mực sẽ tách biệt các nhiệm vụ thành các lớp riêng biệt. Đầu tiên, một trình xử lý hình học (geometry worker) sẽ duyệt qua trang. Nó truy vấn page.commonObjs để xây dựng fontStyleMap của bạn, kiểm tra mảng biến đổi (transform array) của từng mục văn bản để bắt các kiểu in nghiêng nhân tạo (synthetic italics), và quét các đường dẫn vector (vector paths) lân cận để phát hiện gạch chân. Đầu ra của nó là một cấu trúc trung gian sạch sẽ, gọi là textMeta, trong đó mỗi chuỗi văn bản mang theo ba giá trị boolean đơn giản: bold, italic, và underline.

Tiếp theo, một trình tái tạo văn bản (text rebuilder) sẽ tiêu thụ cấu trúc đó và tạo ra đầu ra đã được đánh dấu (marked-up output). Thứ tự lồng nhau ở đây rất quan trọng. Thứ tự phân cấp chính xác là đặt gạch chân ở ngoài cùng, sau đó đến in nghiêng, và cuối cùng là in đậm ở trong cùng. Điều đó có nghĩa là một chuỗi văn bản được định dạng đầy đủ sẽ trở thành <u><i><b>text</b></i></u>. Thứ tự này ngăn chặn việc chồng chéo HTML không hợp lệ và giữ cho việc hiển thị nhất quán trên các trình duyệt và trình chuyển đổi tài liệu. Nó cũng phản ánh logic kiểu chữ: phần trang trí bao quanh sự nhấn mạnh về ngữ nghĩa, và sự nhấn mạnh về ngữ nghĩa bao quanh trọng lượng cấu trúc.

Sức mạnh của phương pháp này là nó chỉ sử dụng những gì chính tệp PDF đã biết về bản thân nó. Không cần đến OCR, không cần dịch vụ thị giác đám mây (cloud vision service), và không có mô hình học máy nào phải đoán các kiểu chữ từ các điểm ảnh đã được raster hóa. Bạn đang đọc lớp ngữ nghĩa của chính tệp tin, được hiển thị thông qua các thông số hình học và siêu dữ liệu (metadata) mà ứng dụng tạo tệp đã tính toán sẵn. Kết quả là tốc độ nhanh, có tính xác định và chính xác trên khắp bối cảnh hỗn loạn của các trình tạo PDF.

Bài học Thực tế

Sử dụng Regex đối với tên phông chữ là một cái bẫy. Nó có vẻ như là một lối tắt vì nó hoạt động trên tệp duy nhất mà bạn đã thử nghiệm, nhưng nó sẽ thất bại ngay khi gặp áp lực nhẹ từ một lần xuất tệp thứ hai. Thông tin thực sự đã nằm sẵn bên trong PDF, nằm trong các mô tả (descriptors), ma trận (matrices) và các đường dẫn vector (vector paths). Hãy xây dựng quy trình của bạn dựa trên những sự thật đó. Hãy hỏi đối tượng phông chữ xem nó có phải là in đậm hay không. Kiểm tra ma trận biến đổi để tìm độ nghiêng (shear). Quan sát các đoạn được vẽ để tìm gạch chân. Nếu bạn truy vấn chính các thông số kỹ thuật của tài liệu thay vì cào các nhãn bề mặt của nó, bạn sẽ có được các kiểu chữ tồn tại xuyên suốt từ trình tạo PDF này sang trình tạo PDF khác.