The setting

Quillstack builds API mocking and testing tools used by about 30,000 developer teams. The company is 60 people and doubling its engineering team this year. You are the recruiter running technical hiring on your own, reporting to the head of talent.

Every mission on this path happens at the same company, so context carries over the way it does in a real job: the data you cleaned in mission two is the data the finance lead questions in mission four.

The missions

1. sourcing planstarter

Build the sourcing plan for Quillstack's senior backend opening

The eng lead just opened a senior backend engineer role and wants a full pipeline in three weeks. You have limited hours, so where you spend sourcing time is the whole game.

You deliver: A three-week sourcing plan that ranks channels by yield and sets concrete outreach targets.

Scored on: Channel prioritization, Funnel math, Concrete actions, Fit to the role.

Working from: role_brief.md, past_hires.csv.

2. screening rubricstarter

Turn the backend role into a phone-screen scorecard

Before screens start, you want every phone screen scored the same way so candidates are actually comparable at debrief. A consistent rubric is the fix.

You deliver: A phone-screen scorecard with weighted criteria, knockouts, and an anchored scoring scale.

Scored on: Reflects the priorities, Correct knockouts, Anchored scale, Built for consistency.

Working from: role_brief.md, hm_priorities.md.

3. interview kitcore

Assemble the onsite interview kit for the backend loop

The onsite loop is four interviewers. Right now it is not set up well, and a loop where two people test the same thing and nobody tests another wastes the candidate's day and produces a muddy debrief.

You deliver: An onsite interview kit that assigns competencies, questions, and slots with no gaps or overlaps.

Scored on: Full coverage, Fixes the conflict, Mapped questions, Shared anchors.

Working from: competency_model.md, panel_roster.csv.

4. candidate scoringcore

Compare the three finalists from the backend loop

The loop is done for three finalists and the scorecards are in. The eng lead wants your recommendation, and the honest read is not the one the raw averages give you.

You deliver: A hire recommendation comparing the three finalists with competency-level evidence.

Scored on: Reads beyond the average, Uses the debrief context, Calibration adjustment, Clear recommendation.

Working from: scorecards.csv, debrief_notes.md.

5. pipeline reportstretch

Report the Q3 hiring pipeline and unblock it

The head of talent wants the quarterly pipeline read: where candidates sit, how they convert by stage and source, how long they take, and where the pipeline is stuck. Three reqs are open and five hires are needed this half, and the current pace will miss it.

You deliver: A pipeline report with the funnel, the bottlenecks, and evidence-backed recommendations.

Scored on: Accurate funnel, Identifies the bottlenecks, Ties to targets, Actionable recommendations.

Working from: pipeline.csv, req_status.md.

How the scoring works

Each deliverable is graded against the rubric written for that mission. Separately, every mission on every path is graded on how you used AI, against the same four criteria:

  • Understood the task. The learner framed the goal for the assistant clearly instead of pasting the brief and hoping.
  • Grounded in the material. The learner directed the assistant into the provided files and based the work on them, not on invented facts.
  • Verified the output. The learner checked claims, numbers, or coverage against the source material before submitting.
  • Iterated with judgment. The learner refined weak parts of the draft with specific follow-ups rather than accepting the first answer.

Both scores, with the work behind them, go on your proof profile. That is what makes a claim like "I can use AI for recruiting and talent" something an employer can check.