If you have ever uploaded a resume to an applicant tracking system and wondered why a human never saw it, you already understand the black-box problem of AI hiring. Most resume-scoring tools hide their logic behind SaaS dashboards and polite rejection emails. HackerRank took a different route. Its Hiring Agent is open source, which means anyone can crack it open, trace the code, and see exactly how an LLM turns a PDF and a GitHub link into a number. One developer did exactly that. What they found is not a polished recruiting framework. It is a mirror showing us how easily automation codifies personal opinions.

Under the Hood

The pipeline is deceptively simple. A candidate’s PDF resume is converted to Markdown, then parsed into a rigid JSON structure with fields for work history, skills, education, and side projects. Python scripts shuttle the data from one stage to the next, but the actual thinking happens inside a chain of prompts. Each section gets its own prompt. The LLM reads the structured data, applies scoring rules written in plain English, and returns a grade.

This architecture matters. The heavy lifting is not happening in clever algorithms or training loops. It is happening in the wording of the prompts. Change a few adjectives in the instruction set, and the same engineer goes from a strong hire to a weak candidate. That makes the tool fragile. It also makes it honest. Most AI hiring vendors would never let you see the prompts. HackerRank’s prototype exposes the truth that resume scoring has always been about the rubric, not the code.

The Tyranny of the 35 Percent

The most striking bias hides in the scoring rubric. Open source contributions account for 35 percent of the total score. That is an enormous weight. To put it in perspective, a candidate’s entire work history, education, and skill set must compete with one slice of their extracurricular coding life for the other 65 percent.

The rules are even stricter than the weighting suggests. Personal GitHub repositories do not count. Maintaining your own library, no matter how useful, scores zero. The tool only rewards contributions to other people’s projects. The candidate must be a committer on someone else’s codebase to earn those points.

That preference carries real demographic weight. Engineers who maintain their own tools often do so because they solved a problem no one else was solving. They might also hold jobs that forbid external contribution, work in regions with fewer large open-source communities, or simply have family obligations that make unpaid coding after hours impossible. By writing the prompt this way, the tool does not measure raw engineering ability. It measures participation in a specific coding culture, then calls it objectivity.

When the Instructions Don't Land

The rubric also tries to reward startup experience. The prompt explicitly suggests giving extra points to founders and early-stage engineers. That sounds reasonable in theory. Startup veterans often wear many hats and ship under pressure. So the tester tried an experiment. They took a single resume and changed nothing except the most recent job title, running it through the agent three times with three different labels: Senior Java Engineer, Founding Engineer, and Co-founder / CTO.

The scores barely moved. The LLM essentially ignored the instruction.

This is one of the most important findings from the entire audit. It proves that a prompt rule is only a suggestion. Large language models are trained on vast corpora of text that contain their own stubborn biases about what signals quality. If the model’s training data associates prestige with certain titles, company names, or keywords rather than the phrase “founding engineer,” your carefully written instruction may simply bounce off. The prompt tells the model to care about startup titles, but the model has its own ideas, and it wins. That gap between human intent and machine behavior is dangerous when the output is a hiring score.

Points Without Purpose

Beyond the major weightings, the rubric is full of oddly specific micro-rules that feel less like data-driven decisions and more like someone’s late-night brainstorming session.

Bir LinkedIn profili tam olarak bir puan değerindedir. Profilin kalitesi değil. Tavsiye sayısı veya iş geçmişinin derinliği değil. Özgeçmişte sadece bir URL'nin bulunması, toplama tek bir puan ekler. Bu sırada, bir Google Summer of Code katılımcısı olmak beş puan değerindedir. Ve eğer bir aday GitHub'da fork'lanmış depolara sahipse, ajan kendi başına beşten az fork'u olan her fork'u görmezden gelir.

Bu kuralların her biri, sessiz bir katsayı kılığına girmiş yüksek sesli bir değer yargısı oluşturur. LinkedIn varlığı neden bir puan değerindedir ki? Bu, adayın bir sosyal ağı nasıl dolduracağını bildiğini gösterir, dağıtık bir sistemi nasıl tasarlayabileceğini değil. GSoC neden bir LinkedIn bağlantısından beş kat daha değerlidir? Belki de istemi yazan kişi programa saygı duyduğu içindir. Bu saygı artık bir işe alım politikasıdır. Peki neden sınırı beş fork'ta çiziyoruz? On kullanıcısı olan bir araç, kritik bir niş sorunu çözebilir. Bu sistem altında, sanki hiç yokmuş gibi sayılır.

Bu sayılar regresyon analizinden çıkmaz. Bireyler tarafından seçilmiştir. Bir kişi, açık kaynak katılımının bir mühendisin değerinin üçte birinden fazla olduğuna karar verdi. Bir diğeri ise bir LinkedIn profilinin 1 puan değerinde olduğuna karar verdi. Bu tahminleri otomatikleştirdiğinizde, onlara yazılımın otoritesini vermiş olursunuz.

Her İstem Bir Önyargıdır

Bir özgeçmiş puanlama ajanı oluşturmanın en zor kısmı PDF'leri ayrıştırmak veya bir API çağırmak değildir. Nelerin önemli olduğuna karar vermektir. Bir puanlama istemindeki her kelime, iyi bir mühendisi neyin oluşturduğuna dair bir değer yargısıdır. Yan projeler asıl işlerden daha mı önemli olmalı? Açık kaynak kodlar, özel kurumsal işlerden daha mı önemli olmalı? Bir sosyal medya profili hiç mi önemli olmamalı? Bu soruların matematiksel olarak doğru cevapları yoktur. Sadece kültürel tercihler vardır.

Bir işe alım ekibi bunu manuel olarak yaptığında, en azından önyargılar; anlaşmazlığa düşebilen, kalibre edebilen ve öğrenebilen birçok incelemeci arasında dağılır. Bir LLM bunu yaptığında, tek bir istem mühendisinin önyargıları, ölçeklenebilir şekilde çalışan tekrarlanabilir bir fonksiyona dönüşerek katılaşır. Araç öznelliği ortadan kaldırmaz. Onu arşivler.

Onu Bir Filtre Olarak Değil, Bir Ayna Olarak Kullanın

HackerRank’in Hiring Agent'ı en iyi bir prototip olarak anlaşılır. Bir ilk taslak gibi hissettiriyor ki tam olarak öyledir. Yapay zeka işe alım araçlarının nasıl inşa edildiğine dair büyüleyici bir erken bakış sunuyor ancak gerçek bir işe alım organizasyonunun kalibrasyonundan, testinden ve çeşitli girdilerinden yoksun.

Eğer işe alım teknolojisi geliştiriyorsanız, onu dikkatle inceleyin. Keyfi kuralların ne kadar çabuk otomatik bir engellemeye dönüştüğünü gösteriyor. Eğer bir adaysanız, bu sistemlerin birer kahin olmadığını unutmayın. Bunlar, doğal dille süslenmiş e-tablolardır ve istemleri yazan kişinin varsayımlarını taşırlar.

Bu araçlar, yargıladıkları mühendisler kadar titizlikle önyargı açısından test edilene kadar, insan konuşmasının yerini almamalı, aksine ona rehberlik etmelidir.