Back to Guides
HR

How to Structure Interview Evaluations (and Cut Bias)


Most interview debriefs turn into a five-minute conversation where whoever speaks loudest decides the outcome. The notes each interviewer took rarely get compared side by side, so the same candidate can look strong or weak depending on who's in the room and how confidently they argue their impression. A hiring manager sitting through four back-to-back debriefs in one afternoon is especially prone to letting the last, most confidently argued opinion carry more weight than it deserves, simply because it's the freshest one in the room.

Warning

Unstructured debriefs are where bias creeps in easiest. If your evaluation criteria only exist in someone's head, they aren't being applied consistently, even when everyone in the room has good intentions.

A structured evaluation doesn't remove judgment from hiring. It just makes sure every candidate is judged against the same criteria.

Why side-by-side notes change the conversation

The core problem with an unstructured debrief isn't that people disagree. Disagreement is normal and often useful. The problem is that disagreement rarely gets surfaced explicitly. Two interviewers can walk away with opposite impressions of the same answer and neither one realizes it, because the conversation moves on before anyone compares notes closely.

Picture a debrief for a backend engineer candidate where one interviewer thought a system-design answer showed strong ownership, and another thought the same answer revealed a gap in scaling knowledge. In a five-minute verbal debrief, both impressions get mentioned in passing, the group nods, and the conversation moves to the next candidate without anyone realizing they were describing the same fifteen minutes of the interview in contradictory terms. Neither interviewer is wrong about what they noticed, but the disagreement itself, which is the actually useful signal, never gets examined.

Before: three sets of notes, each interviewer summarizing their own take out loud, no direct comparison.

After: one document, organized by criterion, where every interviewer's note on "technical depth" sits next to every other interviewer's note on the same thing.

That single change, comparing by criterion instead of by person, is what turns a debrief from a vibe check into an actual evaluation.


Turning notes into a decision

  1. 1

    Collect every interviewer's raw notes

    Don't summarize yet. Paste in the actual notes from each interviewer, even if they're messy or incomplete. Summarizing too early loses the specific details that make disagreement visible later.

  2. 2

    Ask Claude to organize by your criteria, not by interviewer

    Give Claude your actual evaluation criteria, the specific ones your team uses, not "communication skills" in the abstract, and ask it to sort every note underneath the right one.

  3. 3

    Ask for disagreement, not just consensus

    Explicitly ask Claude to flag anywhere interviewers rated the same thing differently. That's the part a five-minute debrief usually skips, and it's often the most useful part of the whole evaluation.

A structured evaluation should make it easy to see:

  • How the candidate scored on each specific criterion

  • Where interviewers agreed and where they didn't
  • Direct quotes backing up each rating, not just a number

Here's a prompt that puts this into practice:

Prompt

Here are notes from three interviewers on the same candidate for our backend engineer role. Organize them under our four evaluation criteria: technical depth, communication, ownership, and collaboration. Flag anywhere the interviewers rated the same thing differently, and include one direct quote per criterion.

What to do when interviewers actually disagree

Finding disagreement is only useful if you do something with it. The instinct is often to average it out, call it a 3 out of 5 and move on, but that treats a real signal like noise. A genuine disagreement between two experienced interviewers usually means the candidate's answer was actually ambiguous, and averaging the scores hides that ambiguity instead of surfacing it for the hiring team to weigh in on.

Tip

When interviewers disagree on the same criterion, ask Claude to lay out both interpretations side by side with the supporting quote from each. That turns a vague "we didn't agree" into something the hiring team can actually discuss and resolve.

Prompt

Two interviewers rated this candidate differently on "ownership." Show me both interpretations side by side, with the specific quote each interviewer is basing their rating on, so the hiring team can discuss which reading is more accurate.

Keeping the process fair across a whole hiring round

The same structure that helps one evaluation also makes it possible to compare candidates fairly against each other, not just against an abstract bar. Once every candidate's notes go through the same criteria and format, ranking them becomes a matter of reading four consistent documents side by side, instead of trying to remember which candidate said what three interviews ago. A hiring committee comparing four candidates for the same role benefits enormously from four documents that share an identical structure, since the comparison itself becomes mechanical rather than a memory exercise.

If your evaluation format changes candidate to candidate, you're not really comparing candidates. You're comparing whoever happened to write clearer notes.

Common mistake

Using a structured evaluation for some candidates in a hiring round and informal notes for others. The comparison only works if every candidate went through the same process.

Inside Claude Tutorial

Spotting when a request needs structure is transferable.

This kind of structuring shows up far beyond interviews. The app has a full lesson on it, with practice that applies to any request that needs more than one instruction to get right.

Coming Soon
Download on theApp Store
GET IT ONGoogle Play

Writing the final recommendation

Once the structured evaluation exists, the hiring decision still needs a clear written recommendation, something a hiring manager can act on without re-reading every interviewer's raw notes. This is a separate step from organizing the notes themselves.

Prompt

Based on the structured evaluation above, write a one-paragraph hiring recommendation. State the overall recommendation clearly in the first sentence, then reference the two or three pieces of evidence that most influenced it, including any unresolved disagreement between interviewers.

Tip

Ask for the recommendation to lead with the actual decision, not a summary of the process. "Recommend to hire, contingent on a stronger reference check around ownership" is more useful on a Friday afternoon than three paragraphs restating what everyone already said in the room.

Setting criteria before the interviews happen

Everything above assumes the evaluation criteria already exist. If they don't, the best time to define them is before the first interview, not while organizing notes afterward. Criteria written after the fact tend to bend toward whichever candidate the room already liked, which defeats the entire purpose of evaluating against a fixed standard.

Prompt

We're hiring a backend engineer. Help me turn this job description into four specific interview evaluation criteria, each with one or two concrete signals an interviewer could actually look for, not vague traits like "strong communicator."

Common mistake

Reusing the same generic criteria (communication, culture fit, technical skill) for every role, regardless of what the job actually requires. Criteria that don't reflect the specific role tend to produce ratings that don't predict anything useful about on-the-job performance.

Once criteria exist for a role, it's worth reusing and refining them across every hiring round for that same role, rather than redefining them from scratch each time a new requisition opens. A criteria set that's been used and adjusted across three hiring rounds is more reliable than one written fresh for each search.

Prompt

We used these four criteria for our last backend engineer search and they worked well, except "collaboration" was hard for interviewers to rate consistently. Help me sharpen that one criterion with more concrete signals, without changing the other three.

Continue reading