Resume-to-Score pipeline that extracts structured data from PDFs, enriches with GitHub signals, and outputs a fair, explainable evaluation.
- Context and intent
- Coverage
- Overview
- Architecture
- Installation and Setup
- Configuration
- How it works
- CLI usage
- Directory layout
- Provider details
- Contributing
- License
This project got a lot of attention recently, and some of the discussion surfaced misconceptions worth addressing directly.
What this is not:
- Not an ATS (Applicant Tracking System)
- Not used to screen HackerRank's open roles
- Not a product available to HackerRank customers
What it actually is:
Every year HackerRank receives 50,000–60,000 intern applications. No human can read that many resumes well. This tool was built to rank them — helping decide which resumes to read first. Resumes scoring below the cutoff are filtered out, but the cutoff is intentionally set very low so only candidates at the very bottom of the distribution are removed. The vast majority pass through to human review, where the real decisions are made.
Since this was built, HackerRank has also shipped AI Interviewer (Chakra) to automate the first round of interviews — so candidates are no longer assessed on their resume alone.
On the default model:
The repo ships with gemma3:4b as the default because it runs locally on most laptops without any cloud API key. Actual intern resumes at HackerRank are evaluated using a top-tier Gemini model. The repo ships with a demo config, not the production one.
Articles and discussions that have shaped how we think about improving this project:
| Article | Key takeaway |
|---|---|
| HackerRank open sourced its ATS. My resume scored 90/100. Oh wait 74/100. No — 88/100. Actually 83/100. — Dan Kinsky | Deep statistical analysis of score variance across 100 runs of the same resume. Isolates which categories are stable (technical skills) vs. noisy (project quality judgments). Points to LLM non-determinism as the root cause. |
| The Score Depends on the Roll of the Dice — Pinggy Blog | Reproduces the variance findings and surfaces a security issue: invisible text embedded in PDFs can inflate scores significantly. |
| The Hiring Rubric Inside — ByteIota | Breaks down the scoring weights and argues that a GitHub-centric rubric disadvantages engineers whose work is in private enterprise repos. Also notes the signal degradation risk as candidates optimize for the now-public rubric. |
| Analyzing resume scoring consistency — Mariano Gobea Alcoba, DEV Community | Proposes concrete fixes: standardized data formats, versioned evaluation models, ensemble scoring, and explainability layers to reduce variance and make the system more robust. |
| AI-Powered Pipeline for Explainable Resume Scoring — AIToolly | Covers the launch and highlights the transparency argument — making scoring logic public allows scrutiny that proprietary ATS systems never face. |
| Hacker News discussion | 200+ comment thread covering LLM determinism, GDPR Article 22 implications, and the broader ethics of automated resume filtering. |
Video coverage
- HackerRank Open-Sourced Their ATS? — YouTube Short
- HackerRank Open-Sourced ATS Tool for selecting Resume — YouTube Short
- HackerRank Custom ATS Released! Get Your Resume Score & Beat ATS Filters — full walkthrough video
Community tools built on this repo
- Resume Reality Check — hosted tool that lets candidates score their own resume against the same rubric
Hiring Agent parses a resume PDF to Markdown, extracts sectioned JSON using a local or hosted LLM, augments the data with GitHub profile and repository signals, then produces an objective evaluation with category scores, evidence, bonus points, and deductions. You can run fully local with Ollama or use Google Gemini.
|
Flow
|
Key modules
|
-
Python 3.11+
The repository pins
.python-versionto 3.11.13. -
One LLM backend (either of them)
- Ollama for local models
Install from the official site, then run
ollama serve. - Google Gemini if you have an API key, get it from here.
- Ollama for local models
Install from the official site, then run
$ git clone https://github.com/interviewstreet/hiring-agent
$ cd hiring-agent
$ python -m venv .venv
# Linux or macOS
$ source .venv/bin/activate
# Windows
# .venv\Scripts\activate
$ pip install -r requirements.txtPull the model you want to use. For example:
$ ollama pull gemma3:4bIf you want different results, you can pull other models such as:
# For higher system configuration
$ ollama pull gemma3:12b
# For lower system configuration
$ ollama pull gemma3:1bCopy the template and set your environment variables.
$ cp .env.example .envEnvironment variables
| Variable | Values | Description |
|---|---|---|
LLM_PROVIDER |
ollama or gemini |
Chooses provider. Defaults to Ollama. |
DEFAULT_MODEL |
for example gemma3:4b or gemini-2.5-pro |
Model name passed to the provider. |
GEMINI_API_KEY |
string | Required when LLM_PROVIDER=gemini. |
GITHUB_TOKEN |
optional | Inherits from your shell environment, improves GitHub API rate limits. |
Provider mapping lives in prompt.py and models.py. The config.py file has a single flag:
# config.py
DEVELOPMENT_MODE = True # enables caching and CSV exportYou can leave it on during iteration. See the next section for details.
1) PDF extraction
pymupdf_rag.pyandpdf.pyread the PDF using PyMuPDF and convert pages to Markdown-like text.- The
to_markdownroutine handles headings, links, tables, and basic formatting.
2) Section parsing with templates
prompts/templates/*.jinjadefine strict instructions for each section Basics, Work, Education, Skills, Projects, Awards.pdf.PDFHandlercalls the LLM per section and assembles aJSONResumeobject (seemodels.py).
3) GitHub enrichment
github.pyextracts a username from the resume profiles, fetches profile and repos, and classifies each project.- It asks the LLM to select exactly 7 unique projects with a minimum author commit threshold, favoring meaningful contributions.
4) Evaluation
evaluator.pyuses templates that encode fairness and scoring rules.- Scores include
open_source,self_projects,production, andtechnical_skills, plus bonus and deductions, then an explanation for evidence.
5) Output and CSV export
score.pyprints a readable summary to stdout.- When
DEVELOPMENT_MODE=Trueit creates or appends aresume_evaluations.csvwith key fields, and caches intermediate JSON undercache/.
Provide a path to a resume PDF.
$ python score.py /path/to/resume.pdfWhat happens:
- If development mode is on, the PDF extraction result is cached to
cache/resumecache_<basename>.json. - If a GitHub profile is found in the resume, repositories are fetched and cached to
cache/githubcache_<basename>.json. - The evaluator prints a report and, in development mode, appends a CSV row to
resume_evaluations.csv.
.
├── .env.example
├── .python-version
├── config.py
├── evaluator.py
├── github.py
├── llm_utils.py
├── models.py
├── pdf.py
├── prompt.py
├── prompts/
│ ├── template_manager.py
│ └── templates/
│ ├── awards.jinja
│ ├── basics.jinja
│ ├── education.jinja
│ ├── github_project_selection.jinja
│ ├── projects.jinja
│ ├── resume_evaluation_criteria.jinja
│ ├── resume_evaluation_system_message.jinja
│ ├── skills.jinja
│ ├── system_message.jinja
│ └── work.jinja
├── pymupdf_rag.py
├── requirements.txt
├── score.py
└── transform.py
- Set
LLM_PROVIDER=ollama - Set
DEFAULT_MODELto any pulled model, for examplegemma3:4b - The provider wrapper in
models.OllamaProvidercallsollama.chat
- Set
LLM_PROVIDER=gemini - Set
DEFAULT_MODELto a supported Gemini model, for examplegemini-2.0-flash - Provide
GEMINI_API_KEY - The wrapper in
models.GeminiProvideradapts responses to a unified format
Please read the CONTRIBUTING.md for detailed guidelines on filing issues, proposing changes, and submitting pull requests. Key principles include:
- Keep prompts declarative and provider-agnostic.
- Validate changes with a couple of real resumes under different providers.
- Add or adjust unit-free smoke tests that call each stage with minimal inputs.
MIT © HackerRank