जर तुम्ही कधीही PDF मधून बोल्ड (bold) किंवा इटालिक (italic) मजकूर काढण्याचा प्रयत्न केला असेल, तर तुम्ही बहुधा regex ने सुरुवात केली असेल. हे एक स्वाभाविक पाऊल वाटते. फॉन्टच्या नावात "Bold" हा शब्द शोधा, मजकुराला फ्लॅग करा आणि पुढे जा. ही रणनीती तुम्हाला केवळ चुकीचा आत्मविश्वास देईपर्यंतच काम करते. त्यानंतर कोणीतरी तोच दस्तऐवज Acrobat च्या वेगळ्या व्हर्जनमधून, किंवा LibreOffice मधून, किंवा print-to-PDF ड्रायव्हरमधून एक्सपोर्ट करतो आणि तुम्ही कोडमध्ये केलेले सर्व गृहीतक (assumptions) कोलमडून पडतात.
फॉन्टची नावे का खोटं बोलतात
PDF.js तुम्हाला ABCDEF+TimesNewRomanPS-BoldMT सारखे स्ट्रिंग्स देईल. पहिले सहा कॅरेक्टर्स हे एक्सपोर्ट दरम्यान जोडलेले रँडम प्रीफिक्स (prefix) असतात आणि फाईल पुन्हा तयार केल्यावर ते प्रत्येक वेळी बदलतात. तुमच्या पार्सरचा (parser) आधार या प्रीफिक्सवर ठेवणे म्हणजे केवळ गोंधळावर (noise) अवलंबून राहण्यासारखे आहे. इतर एक्सपोर्टर्स तर त्याहूनही कमी उपयुक्त ठरतात. काही वेळा Font12 किंवा F1 सारखे साधे आयडेंटिफायर्स (identifiers) मिळतात. या लेबल्सना कोणताही अर्थ (semantic meaning) नसतो; ते केवळ अंतर्गत रिसोर्स टॅग्स असतात जे फाईल लिहिताना जवळ उपलब्ध होते.
Portable Document Format ची रचना मजकूर काढणे (text extraction) सोपे करण्यासाठी केलेली नव्हती, त्यामुळे फॉन्टची नावे ही केवळ एम्बेड केलेल्या (embedded) किंवा सबसेट केलेल्या (subsetted) रिसोर्सेसचे संदर्भ असतात. स्टाईल शोधण्यासाठी (style detection) एक स्थिर API म्हणून काम करण्यासाठी त्यांची कधीही कल्पना नव्हती. जेव्हा तुम्ही Bold किंवा Italic हा सबस्ट्रिंग शोधणारा regex लिहिता, तेव्हा तुम्ही अशा लेबलला स्क्रॅप करत असता जे तयार करणाऱ्या ॲप्लिकेशनला हवे तसे फॉरमॅट करण्याचे स्वातंत्र्य असते. तुम्ही प्रत्यक्ष टायपोग्राफिक गुणधर्म (typographic properties) वाचत नाही आहात. तुम्ही केवळ फाईल-नेमिंग कन्व्हेन्शन (file-naming convention) वाचत आहात आणि कन्व्हेन्शन्स म्हणजे करार (contracts) नसतात.
लेबल नाही, तर डिस्क्रिप्टर (Descriptor) वाचा
खरी माहिती दुसरीकडे असते. PDF.js मध्ये, प्रत्येक पेज ऑब्जेक्ट commonObjs प्रदर्शित करतो, जो एक मॅप (map) आहे आणि त्यात पेज रेंडर करण्यासाठी आवश्यक असलेले खरे फॉन्ट डिस्क्रिप्टर्स असतात. जेव्हा लायब्ररी एखादे पेज पार्स करते, तेव्हा ती या मॅपमध्ये अस्सल फॉन्ट ऑब्जेक्ट्स भरते. ते ऑब्जेक्ट्स bold आणि italic साठी बुलियन (boolean) प्रॉपर्टीज देतात. हे बुलियन नावाच्या स्ट्रिंगवरून काढलेले नसतात. ते PDF मध्ये एम्बेड केलेल्या फॉन्ट डिस्क्रिप्टरमधून येतात, जे ग्लिफ मेट्रिक्स (glyph metrics), OS/2 टेबल फ्लॅग्स आणि टाईपसेटरने दस्तऐवजात लिहिलेल्या सिम्बॉलिक डिस्क्रिप्टर्सवरून मिळवले जातात.
याचा अर्थ असा की तुम्ही अंदाज लावणे थांबवू शकता. पेजवरील मजकूर घटक (text items) तपासण्यापूर्वी, page.commonObjs वरून इटरेट करून एक fontStyleMap तयार करा. प्रत्येक फॉन्ट ID साठी, फॉन्ट ऑब्जेक्टद्वारे प्रदान केलेल्या प्रत्यक्ष .bold आणि .italic प्रॉपर्टीज नोंदवणारा एक ऑब्जेक्ट स्टोअर करा. नंतर, जेव्हा तुम्ही प्रत्येक मजकूर घटक प्रोसेस कराल, तेव्हा तुमच्या मॅपमध्ये त्याचा फॉन्ट संदर्भ शोधा आणि आधीच मोजलेले (precomputed) फ्लॅग्स वाचा. तुम्हाला अचानक एक निश्चित (deterministic) उत्तर मिळते जे एक्सपोर्ट्समध्ये बदलत नाही.
तुम्ही एक फॉलबॅक (fallback) ठेवू शकता. जर डिस्क्रिप्टर काही कारणास्तव अनुपस्थित किंवा अपूर्ण असेल, तर फॉन्टचे नाव स्वच्छ करा, प्रीफिक्स काढून टाका, रँडम टॅग्स काढून टाका आणि उरलेल्या भागावर एक सुरक्षित regex चालवा. पण हे तुमचा शेवटचा पर्याय असावा, प्राथमिक लॉजिक नाही. विश्वासार्हतेमधील फरक लक्षणीय आहे. जिथे नावावर आधारित पार्सिंग वेगवेगळ्या एक्सपोर्टर्समध्ये कोलमडते, तिथे डिस्क्रिप्टर-आधारित पार्सिंग स्थिर राहते कारण ते फाईलला विचारते की त्यात प्रत्यक्षात काय आहे.
जेव्हा फॉन्ट सुद्धा खोटं बोलतो: सिंथेटिक स्टाइल्स (Synthetic Styles)
डिस्क्रिप्टर सुद्धा काही वेळा चुकू शकतो. काही PDF क्रिएटर्स वेगळा इटालिक टाईपफेस (typeface) एम्बेड करण्याची कटकट करत नाहीत. त्याऐवजी, ते सरळ रोमन फॉन्ट घेतात आणि ट्रान्सफॉर्म मॅट्रिक्सद्वारे (transform matrix) त्याला तिरपे करतात. डिझाइन टूल्स किंवा जुन्या वर्ड प्रोसेसर्सद्वारे तयार केलेल्या फाईल्समध्ये हे सामान्य आहे, जे टायपोग्राफिक शुद्धतेपेक्षा फाईलच्या आकाराला प्राधान्य देतात.
PDF.js मधील प्रत्येक मजकूर घटक एक transform ॲरे वाहून नेतो, जो सहा घटकांचा अॅफिन मॅट्रिक्स (affine matrix) असतो जो ग्लिफ को-ऑर्डिनेट सिस्टमला पेज को-ऑर्डिनेट सिस्टममध्ये मॅप करतो. त्या ॲरेचा तिसरा घटक हॉरिझॉन्टल शेअर (horizontal shear) नियंत्रित करतो. जेव्हा ती व्हॅल्यू शून्य नसते, तेव्हा रेंडररद्वारे मजकूर यांत्रिकदृष्ट्या तिरपा केला जात असतो. जर तुम्ही फक्त फॉन्ट डिस्क्रिप्टरवर विश्वास ठेवला, तर तुम्ही या मजकुराचे वर्गीकरण सरळ रोमन म्हणून कराल. जर तुम्ही मॅट्रिक्सची तपासणी केली, तर तुम्हाला सिंथेटिक इटालिक समजेल आणि तुम्ही त्याला योग्यरित्या मार्क कराल. ओव्हरप्रिंटिंगद्वारे तयार केलेल्या सिंथेटिक बोल्डला देखील हेच लॉजिक लागू होते, जरी केवळ भूमितीवरून (geometry) ते शोधणे कठीण असले तरी. तिरप्या मजकुरासाठी, शेअर व्हॅल्यू (shear value) हा तुमचा निर्णायक पुरावा आहे.
अंडरलाईन्स (Underlines) काढल्या जातात, घोषित केल्या जात नाहीत
बोल्ड आणि इटालिक हे फॉन्ट गुणधर्म आहेत. अंडरलाईन तसे नाही. PDF मध्ये, अंडरलाईन हा एक वेक्टर पाथ (vector path) असतो. रेंडरर मजकुराच्या बेसलाइनजवळ स्थित एका पातळ आडव्या भागासाठी ड्रॉइंग कमांड देतो. हा एक ग्राफिकल घटक आहे जो ग्लिफ्सच्या खाली असतो, तो cmap मध्ये साठवलेला कॅरेक्टर ॲट्रिब्युट (character attribute) नसतो.
This distinction matters because no amount of font descriptor reading will reveal an underline. You need to look at the raw drawing operators on the page, or at the geometry-level output, for short horizontal line segments that run parallel to the baseline at the correct proximity. When your extraction engine spots such a segment underneath a text run, you tag that run as underlined. Treating this as a separate detection layer keeps your data model honest: bold and italic are intrinsic to the font, while underline is extrinsic decoration rendered by the document.
A Working Pipeline
A clean extraction system separates concerns into distinct layers. First, a geometry worker walks the page. It queries page.commonObjs to assemble your fontStyleMap, inspects each text item’s transform array to catch synthetic italics, and scans nearby vector paths to spot underlines. Its output is a clean intermediate structure call it textMeta where every text run carries three simple booleans: bold, italic, and underline.
Next, a text rebuilder consumes that structure and emits marked-up output. Nesting order is important here. The correct hierarchy places underline outermost, then italic, then bold innermost. That means a fully styled run becomes <u><i><b>text</b></i></u>. This ordering prevents invalid HTML overlaps and keeps rendering consistent across browsers and document converters. It also mirrors the typographic logic: decoration wraps semantic emphasis, and semantic emphasis wraps structural weight.
The power of this approach is that it uses only what the PDF already knows about itself. There is no OCR involved, no cloud vision service, and no machine learning model guessing at styles from rasterized pixels. You are reading the file’s own semantic layer, exposed through the geometry and metadata that the creator application already calculated. The result is fast, deterministic, and accurate across the chaotic landscape of PDF generators.
The Real Takeaway
Regex against font names is a trap. It feels like a shortcut because it works on the one file you tested, but it collapses under the mild pressure of a second export. The real information is already inside the PDF, sitting in descriptors, matrices, and vector paths. Build your pipeline around those facts. Ask the font object whether it is bold. Check the transform matrix for shear. Look at the drawn segments for underlines. If you query the document’s own engineering instead of scraping its surface labels, you get styles that survive from one PDF generator to the next.
