MARATTO

article · Informatics

Consolidating Access to Candidate Data for Recruitment Headhunting: Leveraging Explainable Machine Learning

Abstract

The recruitment headhunting process is time-intensive due to manual candidate searches across multiple job platforms, creating inefficiencies in identifying suitable candidates. Current AI-driven recruitment platforms frequently prioritize accuracy over explainability, limiting transparency for non-technical users such as recruiters. This study streamlines recruitment headhunting by (1) consolidating publicly available candidate data from multiple job portals using a professional data aggregation Application Programming Interface (API), and (2) implementing explainable machine learning for transparent candidate–job matching. We utilized the Coresignal API (v1) to aggregate and standardize candidate profiles (N = 587) sourced from LinkedIn and Indeed, including skills, experience, certifications, and education. Using Term Frequency–Inverse Document Frequency (TF-IDF) feature vectors and regression models (Ridge, Gradient Boosting, Random Forest), we matched and ranked candidates against a standardized Data Scientist job description. Shapash was incorporated to provide interpretable feature importance explanations accessible to non-technical users. Model performance was evaluated using stratified 5-fold cross-validation with statistical significance testing. Ridge Regression achieved superior performance (cross-validated R2 = 0.935, bootstrap R2 = 0.954, 95% confidence interval [0.939, 0.965], RMSE = 0.025) compared with Gradient Boosting (R2 = 0.840) and Random Forest (R2 = 0.733). Paired t-tests confirmed significant differences between all model pairs (all ps ≤ 0.001, Bonferroni corrected) with large effect sizes (Cohen’s d ≥ 1.992). Shapash analysis revealed that top-contributing features, such as “engineering”, “data science”, “machine learning”, and “python”, aligned precisely with job description requirements, validating the model’s feature-learning capability. This approach reduces repetitive manual searches across job portals while providing interpretable insights into candidate–job rankings. The methodology’s originality lies in combining professional data aggregation APIs that access publicly available profile data with interpretable models enhanced by user-friendly visualization tools, creating a practical, potentially transferable solution for transparent AI-driven recruitment.

Research topics

  • Employer Branding and e-HRM
  • Ethics and Social Impacts of AI
  • Expert finding and Q&A systems

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.3390/informatics13060094

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.