Production-Deployed ยท Lead AI Developer Portfolio

I Built an Autonomous
AI Hiring Engine

Upload a CV. Press Search. Watch jobs appear from 15+ portals โ€” ranked, explained, and CV-optimised โ€” with applications auto-submitted and contacts found. End-to-end, zero manual steps.

โˆž
Jobs ingested
scales without limit
18
Pipeline stages
fully automated
15+
AI agents
orchestrated
Live
SSE streaming
to the browser
๐Ÿ Python ยท FastAPI ๐Ÿค– GPT-4o ยท pgvector โšก BullMQ ยท Redis SSE ๐Ÿ”ญ OpenTelemetry ๐Ÿ“ฆ Docker ยท Railway ๐ŸŽฏ Adaptive ML weights
โœ‰ Request Repository Access
The Problem

Job search is a data pipeline
nobody should run by hand

Hundreds of portals. Duplicate listings. Generic CVs. No ranking. No feedback loop. This is an engineering problem masquerading as a life problem โ€” so I automated it.

01
Fragmented signal
Relevant roles scattered across 15+ portals with no unified ranking or deduplication. Hours of manual triage every day.
02
Generic CVs get rejected
ATS systems discard non-tailored CVs before a human reads them. No scoring, no iteration โ€” just silence.
03
Zero observability
No data on which skills are in demand, which portals convert, or which applications lead to interviews. Flying blind.
The Product โ€” Live Search

15+ portals scraped in parallel.
Every job ranked against your CV.

One click triggers concurrent scraping across LinkedIn, Indeed, Jobber, Bundesagentur, Xing, Stepstone, Greenhouse, Lever, Ashby and more. Real-time progress streams to the browser via Redis SSE โ€” no polling, no refresh.

Live ยท Redis SSE

What happens behind
the scenes

Playwright launches headless browsers and navigates each portal. For every listing it extracts the job URL, title, company, and full description โ€” then inserts directly into Supabase, deduped by SHA-256 hash.

  • Playwright fetches job URL + full description from each portal page
  • Raw HTML normalised to Markdown before storage
  • SHA-256 content hash prevents duplicate inserts across portals
  • Jobs upserted into Supabase ยท embeddings computed via OpenAI
  • INGEST โ†’ RANK phases streamed live โ€” progress visible in real time
15+ Portals ยท Parallel Scrape
Live search progress dashboard
The Product โ€” Intelligent Ranking

Every job. Every dimension filterable.

Multi-factor scoring surfaces the right roles. Filter by match score tier, application status, role, location, and source portal โ€” all computed server-side via hybrid pgvector + FTS retrieval.

Hybrid pgvector + FTS

Role-aware ranking
across every portal

Not keyword search โ€” semantic vector similarity fused with full-text scoring and adaptive per-user weights. 194 jobs ranked by how well they match your actual CV.

  • Score tiers: โ‰ฅ85% ยท 70โ€“84% ยท <70% ยท Unscored
  • Status: Active ยท PDF Ready ยท Applied
  • Role breakdown: AI Engineer 68 ยท Data Eng 62 ยท Software Eng 57
  • Sources: LinkedIn 63 ยท Jobber 62 ยท LinkedIn-MCP 38 ยท Indeed 25
  • Sort: Best Match ยท Similarity % ยท Date Posted
Multi-portal ยท Hybrid Ranking
Smart filter panel with role and source breakdown
The Product โ€” AI Document Generation

One click. Tailored CV + Cover Letter.
Compiled to PDF.

Hit Generate on any job card. GPT-4o scores, rewrites, and iterates your LaTeX CV until it clears the ATS threshold โ€” then compiles both CV and cover letter to downloadable PDFs. Multiple jobs run in parallel.

4
Pipeline steps
Score โ†’ Optimize โ†’ Cover โ†’ PDF
94
ATS match score
achieved post-optimisation
โˆž
Parallel jobs
BullMQ corner bar
3
Optimisation loops
average to hit threshold
Parallel generation ยท Live progress
Generate corner bar with parallel CV generation
Corner bar โ€” multiple jobs generating simultaneously, real-time progress per task
Real CV output ยท 2 pages ยท PDF
Generated CV PDF output
ATS-optimised CV for Autom SRL โ€” compiled by tectonic LaTeX engine, stored in Supabase
The Product โ€” Market Intelligence

Real-time labour market data.
No third-party subscription.

Every scraping run feeds an analytics layer โ€” skill demand charts, role distribution, tooling trends, and a GPT-4o generated personalised learning roadmap. Data that normally costs thousands per year.

Aggregated ยท Live

Skill demand, role mix,
and tooling trends

The market page aggregates across all ingested jobs. No manual data entry โ€” it's a by-product of the pipeline running normally.

  • Average AI match score computed across all scored roles
  • Python #1 demanded skill โ€” 69% of all postings
  • AI Engineer top role โ€” 8 open positions tracked
  • Top tools: AWS ยท Airflow ยท Snowflake ยท TensorFlow ยท Kubernetes
  • GPT-4o personalised learning roadmap โ€” one click generate
AI-scored ยท GPT-4o Insights
Market intelligence dashboard with skill demand charts
The Product โ€” Network Discovery

Hiring contacts found.
Automatically.

The contact finder agent runs for every saved job โ€” discovering real LinkedIn profiles at the hiring company. Talent acquisition specialists, HR leads, recruiters โ€” surfaced with one click.

LinkedIn ยท Enrichment agent

Real people at
real target companies

Contacts at Schwarz Digits include SAP SuccessFactors consultants and talent specialists โ€” exactly the right people for a SAP-adjacent role. Found and stored automatically, no manual searching.

  • Contacts stored per company with LinkedIn URLs โ€” grows with every save
  • Schwarz Digits: SAP SuccessFactors consultants and TA specialists found
  • TAWO: Talent Acquisition Manager discovered and stored
  • Contact finder agent runs async via BullMQ
  • Direct "View Profile" links for immediate outreach
LinkedIn enriched ยท Auto-discovered
Connections page with real LinkedIn contacts
The Product โ€” Application CRM

Every application tracked automatically.
Every status. Every company.

Every auto-submitted application lands in the CRM โ€” company, portal, score, date applied, and CRM status. QuantumBlack (McKinsey), Axel Springer, Bending Spoons โ€” all tracked without a single manual entry.

Auto-tracked ยท Gmail sync

Application lifecycle
from submit to offer

The pipeline logs every application the moment it's submitted. Gmail OAuth syncs replies back as status transitions โ€” Awaiting Response โ†’ Interview โ†’ Offer โ€” with zero manual entry.

  • Every submitted application auto-logged โ€” grows with every run
  • Companies: QuantumBlack/McKinsey ยท Axel Springer ยท Bending Spoons ยท acemate.ai
  • Status stages: Applied ยท Interview ยท Offer ยท Rejected ยท Ghosted ยท Withdrawn
  • Sort by: Date Applied ยท Score ยท Company
  • Filter by: All ยท Applied ยท Interview ยท Offer ยท Rejected ยท Ghosted
Auto-tracked ยท Gmail sync
Applied Jobs CRM with 121 applications
Architecture

Three services. One coherent system.

Every boundary is intentional โ€” independently deployable, separately scalable, connected through typed contracts and a shared secret.

Browser Gateway + Dashboard
Next.js 16 App Router โ€” Port 3000
All user traffic enters here. Route Handlers enforce auth via httpOnly JWT cookies + Supabase RLS. Real-time search progress streamed via SSE. React 19 ยท TanStack Query ยท TailwindCSS 4 ยท Zustand.
โ†“ HTTP (X-Internal-Secret)    BullMQ queues
Python ML Compute โ€” Port 8000
FastAPI + uvicorn
Owns all ML ops: embedding (OpenAI text-embedding-3-small), hybrid retrieval (pgvector ANN + FTS), 18-stage pipeline, deterministic scoring, adaptive weights, Cohere reranking, GPT-4o explainability. 59KB pipeline.py.
Background Workers โ€” Node.js
BullMQ Consumers (cv-tasks ยท search-pipeline)
worker.ts handles score_job, generate_cv, generate_cover_letter, apply_job. search-pipeline-worker.ts delegates full scrape+rank runs to FastAPI. Both publish Redis pub/sub โ†’ SSE โ†’ browser in real time.
โ†“ psycopg3 pool    ioredis
Persistence
Supabase Postgres + pgvector
10 RLS-protected tables. HNSW vector index + GIN full-text index. 12 ordered migration files. Service-role key for workers (bypasses RLS); anon key for browser (enforces RLS). Zero privilege escalation possible.
Queue + Realtime Cache
Redis โ€” BullMQ ยท SSE ยท Cancellation
BullMQ queues with retry + exponential backoff. Per-search progress pub/sub. LLM explanation cache (24h TTL). Cancel flag propagation โ€” aborts in-flight pipeline runs without orphaned processes.
โ†“ OTLP exporter
Observability Stack
OpenTelemetry โ†’ Jaeger ยท Sentry ยท Langfuse
FastAPI + all httpx outbound calls auto-instrumented. OTLP span export to Jaeger. Sentry on Next.js for frontend error capture. Langfuse hook for LLM prompt versioning and call tracing. Per-stage timing in search_trace.py.
End-to-End Pipeline โ€” SA-00 โ†’ SA-18

18 stages. Zero manual steps.

From user clicking Search All to ranked, explained, CV-optimised matches with optional auto-apply โ€” fully automated, observable, and cancellable at any point.

01
Parallel Scraping
15+ portals fetched concurrently. Bundesagentur ยท LinkedIn ยท Jobber ยท Indeed ยท Xing ยท Stepstone ยท Greenhouse ยท Lever ยท Ashby ยท Crunchbase.
02โ€“03
Dedup + Embed
SHA-256 hash dedup across portals. HTML โ†’ Markdown. Batch embed via OpenAI text-embedding-3-small into pgvector HNSW.
04โ€“05
Hybrid Retrieval + Score
pgvector cosine ANN fused with Postgres FTS. Multi-factor composite score: semantic ยท role tags ยท location ยท recency ยท title.
06โ€“07
Rerank + Explain
Optional Cohere/cross-encoder reranking with wall-clock budget guard. GPT-4o match explanations cached 24h in Redis.
08โ€“10
Optimise โ†’ Apply
LaTeX CV rewritten by GPT-4o until ATS score โ‰ฅ threshold. Compiled to PDF via tectonic. Playwright auto-fills and submits.
SAP Lead AI Developer โ€” Role Alignment

Every requirement demonstrated
in production code

Not claimed on a CV โ€” implemented, deployed, and observable. Each capability maps directly to the SAP job description.

๐Ÿ
Advanced Python ยท FastAPI
Required ยท Backend frameworks
59KB async pipeline.py orchestrating 18 stages. Fully typed with Pydantic, psycopg3 connection pool, 50+ validated env vars, startup health checks, CORS and internal-secret middleware.
FastAPIasyncioPydanticpsycopg3
๐Ÿค–
Multi-Agent Architecture
Nice-to-have ยท LangChain, CrewAI
15 numbered agent modules (00 vector-store โ†’ 24 startup-funding-discovery), each with typed I/O contracts. Orchestrated via BullMQ โ€” decoupled, retryable, observable. GPT-4o feedback loops with safety guardrails.
GPT-4oBullMQ15 Agents
๐Ÿ”„
MLOps / LLMOps
Required ยท CI/CD, continuous training
Adaptive weight system derives per-user ranking deltas from behavioural signals โ€” zero-downtime retraining. CV scorer gates every application. Pipeline guardrails with wall-clock budget enforcement per stage.
Adaptive MLGuardrailsLLMOps
๐Ÿ“ฆ
Docker + CI/CD
Required ยท Docker, Kubernetes, GitHub Actions
Multi-stage Dockerfile bundles Node 20 + Python venv + tectonic LaTeX + Playwright Chromium into one reproducible image. docker-compose for full local stack. Railway CI/CD for production.
DockerComposeRailwayMulti-stage
๐Ÿ”ญ
Observability
Required ยท Reliability, compliance of AI solutions
OpenTelemetry SDK auto-instruments FastAPI and all httpx calls โ€” traces export to Jaeger. Sentry on Next.js. Langfuse for LLM prompt versioning. Structured pino logging with per-stage timing.
OpenTelemetryJaegerLangfuseSentry
๐Ÿ—„๏ธ
Data Engineering + Vector DB
Required ยท ETL, data modelling, distributed platforms
ETL: HTML โ†’ Markdown โ†’ embed โ†’ dedup โ†’ upsert. pgvector HNSW + FTS hybrid retrieval. 12 ordered migration files โ€” schema as code. RLS on every table. Service-role isolation for workers.
pgvectorHNSWRLSETL
Full Stack

Technology Inventory

Every tool chosen for a reason โ€” not for rรฉsumรฉ padding.

AI / ML
๐Ÿง  GPT-4o โ€” scoring ยท CV optimisation ยท cover letters ยท explainability ๐Ÿ“ text-embedding-3-small โ€” HNSW ANN retrieval ๐Ÿ” Cohere ยท Cross-Encoder โ€” optional second-stage reranking ๐Ÿ“Š Adaptive weights โ€” SA-18 behavioural feedback loop
Backend
๐Ÿ FastAPI + uvicorn ๐Ÿ”— psycopg3 connection pool ๐Ÿ“‹ Pydantic v2 schemas ๐ŸŒ Next.js 16 App Router Route Handlers โšก BullMQ + Redis queues ๐Ÿ“ก SSE pub/sub streaming
Data + Storage
๐Ÿ˜ Supabase Postgres ๐Ÿ” pgvector HNSW + GIN FTS hybrid ๐Ÿ” Row-Level Security ๐Ÿ“‚ 12 migration files โ€” schema as code โšก Redis โ€” cache ยท queues ยท SSE
Frontend
โš›๏ธ React 19 + Next.js 16 App Router ๐ŸŽจ TailwindCSS 4 ๐Ÿ”„ TanStack Query 5 ๐Ÿ“ฆ Zustand 5
Infrastructure + Observability
๐Ÿณ Docker multi-stage + Compose ๐Ÿš‚ Railway CI/CD ๐Ÿ”ญ OpenTelemetry โ†’ Jaeger ๐Ÿ› Sentry ๐Ÿ“Š Langfuse LLM tracing ๐ŸŽญ Playwright browser pool ๐Ÿ“„ tectonic + pdflatex
Private Repository

Want to see the code?

The repository is private โ€” but I'll add you as a read collaborator within the hour. Full access to commit history, open PRs, every file and migration. No tours, no filtered demos. Just the real codebase.

โœ‰ nihildev.nandakumar@gmail.com

Or reply directly โ€” response within the day.