หากคุณเคยอัปโหลดเรซูเม่ไปยังระบบติดตามผู้สมัคร (applicant tracking system) และสงสัยว่าทำไมถึงไม่มีมนุษย์คนไหนได้เห็นมันเลย คุณก็คงเข้าใจถึงปัญหาแบบกล่องดำ (black-box problem) ของการจ้างงานด้วย AI แล้ว เครื่องมือให้คะแนนเรซูเม่ส่วนใหญ่ซ่อนตรรกะของตนเองไว้หลังแดชบอร์ด SaaS และอีเมลปฏิเสธที่สุภาพ แต่ HackerRank เลือกใช้เส้นทางที่ต่างออกไป เนื่องจาก Hiring Agent ของเขาเป็นโอเพนซอร์ส (open source) ซึ่งหมายความว่าใครก็สามารถแกะมันออกมา ไล่ดูโค้ด และดูได้ว่า LLM เปลี่ยนไฟล์ PDF และลิงก์ GitHub ให้กลายเป็นตัวเลขได้อย่างไร นักพัฒนาคนหนึ่งได้ทำเช่นนั้นพอดี และสิ่งที่เขาพบไม่ใช่โครงสร้างการสรรหาบุคลากรที่ขัดเกลามาอย่างดี แต่มันคือกระจกที่สะท้อนให้เราเห็นว่าระบบอัตโนมัติสามารถเปลี่ยนความคิดเห็นส่วนตัวให้กลายเป็นรหัสคำสั่งได้อย่างง่ายดายเพียงใด

Under the Hood

กระบวนการทำงานนั้นดูเรียบง่ายจนน่าประหลาดใจ เรซูเม่รูปแบบ PDF ของผู้สมัครจะถูกแปลงเป็น Markdown จากนั้นจะถูกแยกแยะเข้าสู่โครงสร้าง JSON ที่ตายตัว โดยมีฟิลด์สำหรับประวัติการทำงาน, ทักษะ, การศึกษา และโปรเจกต์เสริม สคริปต์ Python จะทำหน้าที่ส่งต่อข้อมูลจากขั้นตอนหนึ่งไปยังอีกขั้นตอนหนึ่ง แต่การ "คิด" ที่แท้จริงเกิดขึ้นภายในชุดของพรอมต์ (prompts) แต่ละส่วนจะได้รับพรอมต์ของตัวเอง LLM จะอ่านข้อมูลที่มีโครงสร้างแล้วใช้กฎการให้คะแนนที่เขียนด้วยภาษาอังกฤษธรรมดา จากนั้นจึงส่งเกรดกลับมา

สถาปัตยกรรมนี้มีความสำคัญมาก งานหนักไม่ได้เกิดขึ้นในอัลกอริทึมที่ชาญฉลาดหรือลูปการฝึกฝน (training loops) แต่มันเกิดขึ้นในการเลือกใช้คำในพรอมต์ เพียงแค่เปลี่ยนคำคุณศัพท์ไม่กี่คำในชุดคำสั่ง วิศวกรคนเดิมก็อาจเปลี่ยนจากผู้สมัครที่น่าจ้างงานอย่างยิ่ง กลายเป็นผู้สมัครที่อ่อนแอได้ทันที นั่นทำให้เครื่องมือนี้เปราะบาง แต่มันก็ทำให้เครื่องมือนี้ซื่อสัตย์ด้วยเช่นกัน ผู้ให้บริการด้านการจ้างงานด้วย AI ส่วนใหญ่จะไม่มีวันยอมให้คุณเห็นพรอมต์ แต่ต้นแบบของ HackerRank เปิดเผยความจริงที่ว่า การให้คะแนนเรซูเม่นั้นเป็นเรื่องของเกณฑ์การให้คะแนน (rubric) มาโดยตลอด ไม่ใช่เรื่องของโค้ด

อำนาจเบ็ดเสร็จของ 35 เปอร์เซ็นต์

อคติที่เด่นชัดที่สุดซ่อนอยู่ในเกณฑ์การให้คะแนน (scoring rubric) การมีส่วนร่วมในโปรเจกต์โอเพนซอร์สคิดเป็น 35 เปอร์เซ็นต์ของคะแนนทั้งหมด ซึ่งถือเป็นน้ำหนักที่มหาศาล หากจะให้เห็นภาพ ประวัติการทำงาน การศึกษา และชุดทักษะทั้งหมดของผู้สมัครต้องมาแข่งขันกับเพียงส่วนเสี้ยวหนึ่งของชีวิตการเขียนโค้ดนอกเวลา เพื่อแย่งชิงอีก 65 เปอร์เซ็นต์ที่เหลือ

กฎเกณฑ์นั้นเข้มงวดกว่าน้ำหนักที่แสดงไว้เสียอีก เรโพสิทอรี (repository) ส่วนตัวบน GitHub ไม่ถูกนำมานับ การดูแลไลบรารีของตัวเอง ไม่ว่าจะมีประโยชน์เพียงใด ก็จะได้คะแนนเป็นศูนย์ เครื่องมือนี้จะให้รางวัลเฉพาะการมีส่วนร่วมในโปรเจกต์ของผู้อื่นเท่านั้น ผู้สมัครต้องเป็น committer ใน codebase ของผู้อื่นเพื่อที่จะได้รับคะแนนเหล่านั้น

ความลำเอียงนี้ส่งผลกระทบต่อกลุ่มประชากรอย่างแท้จริง วิศวกรที่ดูแลเครื่องมือของตัวเองมักทำเช่นนั้นเพราะพวกเขาได้แก้ปัญหาที่ไม่มีใครอื่นแก้ พวกเขาอาจทำงานในตำแหน่งที่ห้ามการมีส่วนร่วมจากภายนอก ทำงานในภูมิภาคที่มีชุมชนโอเพนซอร์สขนาดใหญ่น้อยกว่า หรือเพียงแค่มีภาระทางครอบครัวที่ทำให้ไม่สามารถเขียนโค้ดโดยไม่ได้รับค่าตอบแทนหลังเลิกงานได้ ด้วยการเขียนพรอมต์ในลักษณะนี้ เครื่องมือจึงไม่ได้วัดความสามารถทางวิศวกรรมที่แท้จริง แต่มันวัดการมีส่วนร่วมในวัฒนธรรมการเขียนโค้ดเฉพาะกลุ่ม แล้วเรียกสิ่งนั้นว่าความเที่ยงธรรม

เมื่อคำสั่งไม่ได้ผล

เกณฑ์การให้คะแนนยังพยายามจะให้รางวัลกับประสบการณ์ในสตาร์ทอัพด้วย โดยพรอมต์ระบุอย่างชัดเจนว่าให้คะแนนพิเศษแก่ผู้ก่อตั้ง (founders) และวิศวกรในระยะเริ่มต้น (early-stage engineers) ซึ่งฟังดูสมเหตุสมผลในทางทฤษฎี เพราะผู้เชี่ยวชาญในสตาร์ทอัพมักต้องรับบทบาทหลายอย่างและต้องส่งมอบงานภายใต้ความกดดัน ดังนั้นผู้ทดสอบจึงลองทำการทดลอง โดยใช้เรซูเม่เพียงฉบับเดียวและไม่เปลี่ยนอะไรเลยนอกจากชื่อตำแหน่งงานล่าสุด แล้วรันผ่าน agent สามครั้งด้วยป้ายกำกับที่แตกต่างกันสามแบบ ได้แก่ Senior Java Engineer, Founding Engineer และ Co-founder / CTO

คะแนนแทบจะไม่ขยับเลย โดยพื้นฐานแล้ว LLM ได้เพิกเฉยต่อคำสั่งนั้น

นี่คือหนึ่งในการค้นพบที่สำคัญที่สุดจากการตรวจสอบทั้งหมด มันพิสูจน์ว่ากฎในพรอมต์เป็นเพียงข้อแนะนำเท่านั้น โมเดลภาษาขนาดใหญ่ (Large language models) ถูกฝึกฝน

A LinkedIn profile is worth exactly one point. Not the quality of the profile. Not the number of recommendations or the depth of the work history. Simply having a URL on the resume adds a single point to the total. Meanwhile, being a Google Summer of Code participant is worth five points. And if a candidate has forked repositories on GitHub, the agent ignores any fork with fewer than five forks of its own.

Each of these rules makes a loud value judgment disguised as a quiet coefficient. Why is LinkedIn presence worth a point at all? It signals that a candidate knows how to fill out a social network, not that they can architect a distributed system. Why is GSoC worth five times as much as a LinkedIn link? Perhaps because the prompt author respects the program. That respect is now a hiring policy. And why draw the line at five forks? A tool with ten users might solve a critical niche problem. Under this system, it might as well not exist.

These numbers do not emerge from regression analysis. They were chosen by individuals. One person decided open source participation is more than a third of an engineer’s worth. Another person decided a LinkedIn profile is worth 1 point. When you automate those guesses, you give them the authority of software.

Every Prompt Is a Prejudice

The hardest part of building a resume-scoring agent is not parsing PDFs or calling an API. It is deciding what matters. Every word in a scoring prompt is a value judgment about what makes a good engineer. Should side projects outweigh day jobs? Should public code matter more than private enterprise work? Should a social media profile matter at all? There are no mathematically correct answers to these questions. There are only cultural preferences.

When a recruiting team does this manually, at least the biases are distributed across many reviewers who can disagree, calibrate, and learn. When an LLM does it, the biases of one prompt engineer harden into a repeatable function that runs at scale. The tool does not eliminate subjectivity. It archives it.

Use It as a Mirror, Not a Filter

HackerRank’s Hiring Agent is best understood as a prototype. It feels like a first draft, which is exactly what it is. It offers a fascinating early look at how AI hiring tools are constructed, but it lacks the calibration, testing, and diverse input of a real recruiting organization.

If you are building hiring tech, study it carefully. It shows how quickly arbitrary rules turn into automated gatekeeping. If you are a candidate,remember that these systems are not oracles. They are spreadsheets dressed up in natural language, and they carry the assumptions of whoever wrote the prompts.

Until these tools are tested for bias as rigorously as the engineers they judge, they should inform human conversation, not replace it.