ਜੇਕਰ ਤੁਸੀਂ ਕਦੇ ਵੀ PDF ਵਿੱਚੋਂ ਬੋਲਡ (bold) ਜਾਂ ਇਟੈਲਿਕ (italic) ਟੈਕਸਟ ਕੱਢਣ ਦੀ ਕੋਸ਼ਿਸ਼ ਕੀਤੀ ਹੈ, ਤਾਂ ਸ਼ਾਇਦ ਤੁਸੀਂ regex ਨਾਲ ਸ਼ੁਰੂਆਤ ਕੀਤੀ ਹੋਵੇਗੀ। ਇਹ ਇੱਕ ਸਪੱਸ਼ਟ ਕਦਮ ਲੱਗਦਾ ਹੈ। ਫੌਂਟ ਦੇ ਨਾਮ ਵਿੱਚ "Bold" ਸ਼ਬਦ ਲੱਭੋ, ਟੈਕਸਟ ਨੂੰ ਫਲੈਗ ਕਰੋ, ਅਤੇ ਅੱਗੇ ਵਧੋ। ਇਹ ਰਣਨੀਤੀ ਤੁਹਾਨੂੰ ਝੂਠਾ ਭਰੋਸਾ ਦੇਣ ਲਈ ਬਹੁਤ ਕੁਝ ਸਮੇਂ ਤੱਕ ਕੰਮ ਕਰਦੀ ਹੈ। ਫਿਰ ਕੋਈ ਇੱਕ ਉਹੀ ਦਸਤਾਵੇਜ਼ Acrobat ਦੇ ਕਿਸੇ ਵੱਖਰੇ ਵਰਜ਼ਨ, ਜਾਂ LibreOffice, ਜਾਂ print-to-PDF ਡਰਾਈਵਰ ਤੋਂ ਐਕਸਪੋਰਟ ਕਰਦਾ ਹੈ, ਅਤੇ ਤੁਹਾਡੇ ਦੁਆਰਾ ਕੋਡ ਕੀਤਾ ਗਿਆ ਹਰ ਅਨੁਮਾਨ ਟੁੱਟ ਜਾਂਦਾ ਹੈ।
ਫੌਂਟ ਦੇ ਨਾਮ ਝੂਠ ਕਿਉਂ ਬੋਲਦੇ ਹਨ
PDF.js ਤੁਹਾਨੂੰ ABCDEF+TimesNewRomanPS-BoldMT ਵਰਗੇ ਸਟ੍ਰਿੰਗਸ (strings) ਦੇਵੇਗਾ। ਪਹਿਲੇ ਛੇ ਅੱਖਰ ਐਕਸਪੋਰਟ ਦੌਰਾਨ ਪਾਏ ਗਏ ਰੈਂਡਮ ਪ੍ਰੀਫਿਕਸ (prefix) ਹਨ, ਅਤੇ ਹਰ ਵਾਰ ਫਾਈਲ ਨੂੰ ਦੁਬਾਰਾ ਬਣਾਉਣ 'ਤੇ ਇਹ ਬਦਲ ਜਾਂਦੇ ਹਨ। ਆਪਣੇ ਪਾਰਸਰ (parser) ਨੂੰ ਉਸ ਪ੍ਰੀਫਿਕਸ 'ਤੇ ਨਿਰਭਰ ਕਰਨਾ ਸ਼ੋਰ (noise) 'ਤੇ ਦਾਅ ਲਗਾਉਣ ਦੇ ਬਰਾਬਰ ਹੈ। ਹੋਰ ਐਕਸਪੋਰਟਰ ਹੋਰ ਵੀ ਘੱਟ ਮਦਦਗਾਰ ਹੁੰਦੇ ਹਨ। ਕੁਝ Font12 ਜਾਂ F1 ਵਰਗੇ ਸਾਧਾਰਨ ਆਈਡੈਂਟੀਫਾਇਰ (identifiers) ਜਾਰੀ ਕਰਦੇ ਹਨ। ਇਹ ਲੇਬਲਾਂ ਦਾ ਕੋਈ ਅਰਥ ਨਹੀਂ ਹੁੰਦਾ; ਇਹ ਅੰਦਰੂਨੀ ਰਿਸੋਰਸ ਟੈਗ ਹਨ ਜੋ ਫਾਈਲ ਲਿਖਣ ਸਮੇਂ ਉੱਥੇ ਮੌਜੂਦ ਸਨ।
ਕਿਉਂਕਿ Portable Document Format ਨੂੰ ਟੈਕਸਟ ਐਕਸਟਰੈਕਸ਼ਨ (extraction) ਨੂੰ ਆਸਾਨ ਬਣਾਉਣ ਲਈ ਕਦੇ ਵੀ ਡਿਜ਼ਾਈਨ ਨਹੀਂ ਕੀਤਾ ਗਿਆ ਸੀ, ਇਸ ਲਈ ਫੌਂਟ ਦੇ ਨਾਮ ਸਿਰਫ਼ ਐਂਬੈਡਡ (embedded) ਜਾਂ ਸਬਸੈੱਟ ਕੀਤੇ ਰਿਸੋਰਸਾਂ ਦੇ ਹਵਾਲੇ ਹਨ। ਉਹਨਾਂ ਨੂੰ ਸਟਾਈਲ ਡਿਟੈਕਸ਼ਨ (style detection) ਲਈ ਇੱਕ ਸਥਿਰ API ਵਜੋਂ ਕੰਮ ਕਰਨ ਲਈ ਕਦੇ ਵੀ ਨਹੀਂ ਬਣਾਇਆ ਗਿਆ ਸੀ। ਜਦੋਂ ਤੁਸੀਂ Bold ਜਾਂ Italic ਵਰਗੇ ਸਬਸਟ੍ਰਿੰਗ (substring) ਨੂੰ ਲੱਭਣ ਲਈ regex ਲਿਖਦੇ ਹੋ, ਤਾਂ ਤੁਸੀਂ ਇੱਕ ਅਜਿਹੀ ਲੇਬਲ ਨੂੰ ਸਕ੍ਰੈਪ ਕਰ ਰਹੇ ਹੁੰਦੇ ਹੋ ਜਿਸ ਨੂੰ ਬਣਾਉਣ ਵਾਲੀ ਐਪਲੀਕੇਸ਼ਨ ਆਪਣੀ ਮਰਜ਼ੀ ਅਨੁਸਾਰ ਫਾਰਮੈਟ ਕਰਨ ਲਈ ਆਜ਼ਾਦ ਸੀ। ਤੁਸੀਂ ਅਸਲ ਟਾਈਪੋਗ੍ਰਾਫਿਕ ਗੁਣਾਂ (typographic properties) ਨੂੰ ਨਹੀਂ ਪੜ੍ਹ ਰਹੇ ਹੋ। ਤੁਸੀਂ ਫਾਈਲ-ਨਾਮ ਰੱਖਣ ਦੇ ਤਰੀਕੇ (convention) ਨੂੰ ਪੜ੍ਹ ਰਹੇ ਹੋ, ਅਤੇ ਰਵਾਇਤਾਂ ਕੋਈ ਇਕਰਾਰਨਾਮਾ ਨਹੀਂ ਹੁੰਦੀਆਂ।
ਡਿਸਕ੍ਰਿਪਟਰ (Descriptor) ਪੜ੍ਹੋ, ਲੇਬਲ ਨਹੀਂ
ਅਸਲ ਸੱਚਾਈ ਕਿਤੇ ਹੋਰ ਹੈ। PDF.js ਵਿੱਚ, ਹਰ ਪੇਜ ਆਬਜੈਕਟ commonObjs ਨੂੰ ਐਕਸਪੋਜ਼ (expose) ਕਰਦਾ ਹੈ, ਜੋ ਇੱਕ ਮੈਪ (map) ਹੈ ਜਿਸ ਵਿੱਚ ਪੇਜ ਨੂੰ ਰੈਂਡਰ ਕਰਨ ਲਈ ਲੋੜੀਂਦੇ ਅਸਲੀ ਫੌਂਟ ਡਿਸਕ੍ਰਿਪਟਰ (font descriptors) ਹੁੰਦੇ ਹਨ। ਜਦੋਂ ਲਾਇਬ੍ਰੇਰੀ ਇੱਕ ਪੇਜ ਨੂੰ ਪਾਰਸ ਕਰਦੀ ਹੈ, ਤਾਂ ਇਹ ਇਸ ਮੈਪ ਨੂੰ ਅਸਲੀ ਫੌਂਟ ਆਬਜੈਕਟਾਂ ਨਾਲ ਭਰ ਦਿੰਦੀ ਹੈ। ਉਹ ਆਬਜੈਕਟ bold ਅਤੇ italic ਲਈ ਬੂਲੀਅਨ (boolean) ਪ੍ਰੋਪਰਟੀਜ਼ ਪ੍ਰਦਾਨ ਕਰਦੇ ਹਨ। ਉਹ ਬੂਲੀਅਨ ਨਾਮ ਵਾਲੇ ਸਟ੍ਰਿੰਗ ਤੋਂ ਨਹੀਂ ਲਏ ਜਾਂਦੇ। ਉਹ PDF ਦੇ ਅੰਦਰ ਐਂਬੈਡਡ ਫੌਂਟ ਡਿਸਕ੍ਰਿਪਟਰ ਤੋਂ ਆਉਂਦੇ ਹਨ, ਜੋ ਗਲਾਈਫ ਮੈਟ੍ਰਿਕਸ (glyph metrics), OS/2 ਟੇਬਲ ਫਲੈਗਸ, ਅਤੇ ਉਹਨਾਂ ਸਿੰਬੋਲਿਕ ਡਿਸਕ੍ਰਿਪਟਰਾਂ ਤੋਂ ਪ੍ਰਾਪਤ ਕੀਤੇ ਜਾਂਦੇ ਹਨ ਜੋ ਟਾਈਪਸੈਟਰ ਨੇ ਦਸਤਾਵੇਜ਼ ਵਿੱਚ ਲਿਖੇ ਸਨ।
ਇਸਦਾ ਮਤਲਬ ਹੈ ਕਿ ਤੁਸੀਂ ਅੰਦਾਜ਼ੇ ਲਗਾਉਣਾ ਬੰਦ ਕਰ ਸਕਦੇ ਹੋ। ਪੇਜ 'ਤੇ ਟੈਕਸਟ ਆਈਟਮਾਂ ਨੂੰ ਪ੍ਰੋਸੈਸ ਕਰਨ ਤੋਂ ਪਹਿਲਾਂ, page.commonObjs 'ਤੇ ਇਟਰੇਟ (iterate) ਕਰਕੇ ਇੱਕ fontStyleMap ਬਣਾਓ। ਹਰੇਕ ਫੌਂਟ ID ਲਈ, ਇੱਕ ਅਜਿਹਾ ਆਬਜੈਕਟ ਸਟੋਰ ਕਰੋ ਜੋ ਫੌਂਟ ਆਬਜੈਕਟ ਦੁਆਰਾ ਪ੍ਰਦਾਨ ਕੀਤੀਆਂ ਗਈਆਂ ਅਸਲ .bold ਅਤੇ .italic ਪ੍ਰੋਪਰਟੀਜ਼ ਨੂੰ ਰਿਕਾਰਡ ਕਰਦਾ ਹੋਵੇ। ਬਾਅਦ ਵਿੱਚ, ਜਦੋਂ ਤੁਸੀਂ ਹਰੇਕ ਟੈਕਸਟ ਆਈਟਮ ਨੂੰ ਪ੍ਰੋਸੈਸ ਕਰਦੇ ਹੋ, ਤਾਂ ਆਪਣੇ ਮੈਪ ਵਿੱਚ ਉਸਦੇ ਫੌਂਟ ਰੈਫਰੈਂਸ ਨੂੰ ਲੱਭੋ ਅਤੇ ਪਹਿਲਾਂ ਤੋਂ ਕੈਲਕੂਲੇਟ ਕੀਤੇ ਫਲੈਗਸ ਨੂੰ ਪੜ੍ਹੋ। ਅਚਾਨਕ ਤੁਹਾਡੇ ਕੋਲ ਇੱਕ ਨਿਸ਼ਚਿਤ (deterministic) ਜਵਾਬ ਹੋਵੇਗਾ ਜੋ ਐਕਸਪੋਰਟ ਦੇ ਵਿਚਕਾਰ ਨਹੀਂ ਬਦਲਦਾ।
ਤੁਸੀਂ ਇੱਕ ਫਾਲਬੈਕ (fallback) ਰੱਖ ਸਕਦੇ ਹੋ। ਜੇਕਰ ਡਿਸਕ੍ਰਿਪਟਰ ਕਿਸੇ ਤਰ੍ਹਾਂ ਗੈਰ-ਹਾਜ਼ਰ ਜਾਂ ਅਧੂਰਾ ਹੈ, ਤਾਂ ਫੌਂਟ ਦੇ ਨਾਮ ਨੂੰ ਸਾਫ਼ ਕਰੋ, ਪ੍ਰੀਫਿਕਸ ਨੂੰ ਹਟਾਓ, ਰੈਂਡਮ ਟੈਗਸ ਨੂੰ ਸੁੱਟੋ, ਅਤੇ ਜੋ ਬਚਦਾ ਹੈ ਉਸ 'ਤੇ ਇੱਕ ਸਾਵਧਾਨੀ ਵਾਲਾ regex ਚਲਾਓ। ਪਰ ਇਹ ਤੁਹਾਡਾ ਆਖਰੀ ਵਿਕਲਪ ਹੋਣਾ ਚਾਹੀਦਾ ਹੈ, ਤੁਹਾਡਾ ਮੁੱਖ ਲੌਜਿਕ ਨਹੀਂ। ਭਰੋਸੇਯੋਗਤਾ ਵਿੱਚ ਅੰਤਰ ਬਹੁਤ ਵੱਡਾ ਹੈ। ਜਿੱਥੇ ਨਾਮ-ਅਧਾਰਤ ਪਾਰਸਿੰਗ ਵੱਖ-ਵੱਖ ਐਕਸਪੋਰਟਰਾਂ ਵਿੱਚ ਟੁੱਟ ਜਾਂਦੀ ਹੈ, ਉੱਥੇ ਡਿਸਕ੍ਰਿਪਟਰ-ਅਧਾਰਤ ਪਾਰਸਿੰਗ ਸਥਿਰ ਰਹਿੰਦੀ ਹੈ ਕਿਉਂਕਿ ਇਹ ਫਾਈਲ ਤੋਂ ਪੁੱਛਦੀ ਹੈ ਕਿ ਇਸ ਵਿੱਚ ਅਸਲ ਵਿੱਚ ਕੀ ਹੈ।
ਜਦੋਂ ਫੌਂਟ ਵੀ ਝੂਠ ਬੋਲਦਾ ਹੈ: ਸਿੰਥੈਟਿਕ ਸਟਾਈਲ (Synthetic Styles)
ਡਿਸਕ੍ਰਿਪਟਰ ਵੀ ਗਲਤੀ ਕਰ ਸਕਦਾ ਹੈ। ਕੁਝ PDF ਕ੍ਰਿਏਟਰ ਇੱਕ ਵੱਖਰੇ ਇਟੈਲਿਕ ਟਾਈਪਫੇਸ (typeface) ਨੂੰ ਐਂਬੈਡ ਕਰਨ ਦੀ ਭੈਲ ਨਹੀਂ ਕਰਦੇ। ਇਸ ਦੀ ਬਜਾਏ, ਉਹ ਸਿੱਧੇ ਰੋਮਨ ਫੌਂਟ ਨੂੰ ਲੈਂਦੇ ਹਨ ਅਤੇ ਉਸਨੂੰ ਇੱਕ ਟ੍ਰਾਂਸਫਾਰਮ ਮੈਟ੍ਰਿਕਸ (transform matrix) ਨਾਲ ਤਿਰਛਾ ਕਰ ਦਿੰਦੇ ਹਨ। ਇਹ ਡਿਜ਼ਾਈਨ ਟੂਲਸ ਜਾਂ ਪੁਰਾਣੇ ਵਰਡ ਪ੍ਰੋਸੈਸਰਾਂ ਦੁਆਰਾ ਬਣਾਈਆਂ ਗਈਆਂ ਫ
This distinction matters because no amount of font descriptor reading will reveal an underline. You need to look at the raw drawing operators on the page, or at the geometry-level output, for short horizontal line segments that run parallel to the baseline at the correct proximity. When your extraction engine spots such a segment underneath a text run, you tag that run as underlined. Treating this as a separate detection layer keeps your data model honest: bold and italic are intrinsic to the font, while underline is extrinsic decoration rendered by the document.
A Working Pipeline
A clean extraction system separates concerns into distinct layers. First, a geometry worker walks the page. It queries page.commonObjs to assemble your fontStyleMap, inspects each text item’s transform array to catch synthetic italics, and scans nearby vector paths to spot underlines. Its output is a clean intermediate structure call it textMeta where every text run carries three simple booleans: bold, italic, and underline.
Next, a text rebuilder consumes that structure and emits marked-up output. Nesting order is important here. The correct hierarchy places underline outermost, then italic, then bold innermost. That means a fully styled run becomes <u><i><b>text</b></i></u>. This ordering prevents invalid HTML overlaps and keeps rendering consistent across browsers and document converters. It also mirrors the typographic logic: decoration wraps semantic emphasis, and semantic emphasis wraps structural weight.
The power of this approach is that it uses only what the PDF already knows about itself. There is no OCR involved, no cloud vision service, and no machine learning model guessing at styles from rasterized pixels. You are reading the file’s own semantic layer, exposed through the geometry and metadata that the creator application already calculated. The result is fast, deterministic, and accurate across the chaotic landscape of PDF generators.
The Real Takeaway
Regex against font names is a trap. It feels like a shortcut because it works on the one file you tested, but it collapses under the mild pressure of a second export. The real information is already inside the PDF, sitting in descriptors, matrices, and vector paths. Build your pipeline around those facts. Ask the font object whether it is bold. Check the transform matrix for shear. Look at the drawn segments for underlines. If you query the document’s own engineering instead of scraping its surface labels, you get styles that survive from one PDF generator to the next.
