Mohammad Ahmed
All work
Full-stack platformJul – Sep 2026Working application · not deployed

AI Recruitment Assistant

A hiring platform that turns CVs and job posts into structured data, then ranks applicants with an explainable hybrid score.

01

Overview

Two role-based experiences on one backend. HR creates and publishes jobs and triggers screening runs; candidates register, upload a PDF résumé and apply. Claude extracts structured requirements and résumé data, embeddings and rules produce a weighted score, and every run is stored with the weights it used.

02

Problem

First-pass CV screening is repetitive and inconsistent. A reviewer needs structured candidate data, a ranking they can explain, and a record of what each screening run actually did.

03

Architecture

  1. Next.js appApp Router + TypeScript. Auth context and role guards for HR and Candidate areas.
  2. FastAPI routersauth, jobs, profile, applications, screening, results, chat, health.
  3. Services layerLLM, résumé parsing, matching and screening logic kept out of the routers.
  4. PostgreSQL 16Async SQLAlchemy models with five Alembic migrations.
  5. Model providersClaude for extraction and analysis; OpenAI for embeddings only.
04

AI pipeline

  1. 01

    Input

    Job description text and a candidate's PDF résumé.

  2. 02

    Preprocessing

    PDF text extraction with pypdf.

  3. 03

    LLM extraction

    Claude is forced to call a schema'd tool; output is validated with Pydantic and retried if no tool call comes back.

  4. 04

    Retrieval / scoring

    Score = 0.5 semantic similarity + 0.3 skill match + 0.2 experience. Falls back to token overlap if embeddings fail.

  5. 05

    Analysis

    Per-candidate analysis, strengths, gaps and interview questions.

  6. 06

    Output

    A persisted screening run, with the weights it used, that HR can reopen and question.

  • Claude forced tool-use for schema-bound extraction
  • OpenAI text-embedding-3-small similarity
  • Hybrid scoring: semantic + skills + experience
  • Grounded Q&A over stored screening results
05

Engineering

Structured output, not parsed prose
Extraction goes through a forced tool call with a JSON schema, then Pydantic validation, with a retry when the model returns no tool block. Downstream code never parses free text.
Runs are auditable
Each screening run snapshots its scoring weights and stores a success or failed status per candidate, so one bad résumé cannot sink a run and old results stay interpretable if the weights change.
Graceful degradation
If the embedding call fails, similarity falls back to a bag-of-words cosine instead of failing the whole screening.
Tighter account model
Public registration only ever creates a Candidate. HR accounts are provisioned through a CLI script.
Chat explains, it doesn't re-rank
The chat endpoint answers from stored results, so a conversation can never silently change a ranking.
06

Challenges

  • Testing async SQLAlchemy with asyncpg: connections bind to an event loop, so the tests create a fresh engine per test.
  • Docker Compose environment variables were being shadowed by the host shell, so the stack uses an APP_ prefix.
  • The model occasionally returned no tool call, which led to the retry in the extraction path.
07

Results & limits

Measured

  • No quantitative evaluation has been run. Ranking quality, latency and extraction accuracy are not measured yet.
  • The backend has roughly 1,300 lines of pytest coverage across auth, jobs, applications, profile, screening and the database layer.

Known limits

  • Skill matching is a simple substring check, so short skill names can false-match.
  • The scoring weights are hand-chosen, not validated against labelled data.
  • Not deployed, and no public demo.
08

Stack

  • Next.js
  • TypeScript
  • Tailwind CSS
  • FastAPI
  • SQLAlchemy 2 (async)
  • PostgreSQL 16
  • Alembic
  • Pydantic v2
  • JWT
  • Docker Compose
  • pytest
10

Demo

No public demo for this one yet. The source repository has setup instructions.