For hiring teams

Structured interview scorecard: a template and examples by role

Published Updated 13 min read By the Career1 team

A structured interview scorecard is a form that rates every candidate for a role on the same few competencies, from the same questions, on a scale whose levels are written down, with the evidence recorded next to each score. For hiring teams it turns an impression into ratings that can be compared and checked. Use four to six competencies, anchor every level, and score before anyone discusses.

What is a structured interview scorecard?

A structured interview scorecard is the written half of a structured interview: the list of what you are assessing, the questions that assess it, what each rating means, and a place to write the evidence. The US Office of Personnel Management's Structured Interviews guide draws the line between the two kinds of interview like this:

Swipe the table sideways to see every column.

Structured interviewUnstructured interview
All candidates are asked the same questions in the same orderCandidates may be asked different questions
All candidates are evaluated using a common rating scaleA standardised rating scale is not required
Interviewers agree on acceptable answersInterviewers do not need to agree on acceptable answers

A scorecard is how the right-hand column becomes the left. It has six parts:

  1. The role, in outcomes. What the person must deliver, which decides what is worth testing.
  2. Four to six competencies. OPM says a structured interview "is typically used to assess between four and six competencies". More than that and none of them is tested properly.
  3. One or two questions per competency, with the follow-up probes every interviewer is allowed to use.
  4. A rating scale with anchors: a sentence for each level saying what an answer at that level sounds like.
  5. An evidence box for what the candidate actually said.
  6. An overall recommendation, filled in last.

Why does a scorecard make interviews more consistent?

Because it removes the freedom that makes interviews unreliable. OPM's guide says unstructured interviews typically show "Low levels of reliability (rating consistency among interviewers)" and low to moderate validity, while structured interviews "have demonstrated a high degree of reliability, validity, and legal defensibility". The largest recent re-analysis of selection research, by Sackett, Zhang, Berry and Lievens, revised validity estimates for most methods downwards by .10 to .20 and still found that structured interviews emerged as the top-ranked selection procedure.

The mechanism is unglamorous. A scorecard makes each interviewer rate each competency separately, against written anchors, with the evidence in front of them, which blocks the rating errors and interviewing mistakes OPM's guide lists by name:

  • Halo: a strong answer on one competency lifting the scores on the others.
  • Central tendency: every candidate a 3 on a five-point scale.
  • Leniency and strictness: everyone high, or everyone low, whatever they said.
  • Similar to me: higher ratings for candidates who resemble the interviewer.
  • First impressions: deciding in the first few minutes and hearing the rest as confirmation.

A scorecard does not rescue bad questions. The consistency only helps if the competencies come from the job and the anchors describe real answers.

Where do the competencies and questions come from?

From the job, not from a list of generic virtues. OPM's guide starts every structured interview with a job analysis: the tasks, the competencies they need, and which of those are needed on day one. For a small team that can be one honest hour.

  1. Write the role as outcomes. What must be true ninety days after they start. Career1's free job description writer splits requirements into what is genuinely required and what is only preferred, which is most of this step.
  2. Pick the four to six competencies the required list depends on. Leave the preferences out of the scorecard.
  3. Write questions that ask for evidence. OPM's guide suggests past-behaviour questions with "superlative adjectives" such as most, last, worst or least, so the candidate describes one real incident rather than a habit. Situational questions, "what would you do if", suit people without the experience yet.
  4. Write the probes in advance. OPM says interviewers "should use very similar probes for all candidates", tailored to the answer but with the same meaning.

For questions by role with what a strong answer shows, the interview questions hub has seven role guides. They were written for candidates preparing, which makes each "what a strong answer shows" line a ready-made top anchor.

A structured interview scorecard template you can copy

Copy these three tables into a document or a spreadsheet, one card per candidate per interviewer.

The header

Swipe the table sideways to see every column.

FieldFill in
RoleTitle, and the three outcomes it must deliver
CandidateName
InterviewerName, and the stage: first round, panel, final
DateFilled in before the debrief, never after

The rating scale

Swipe the table sideways to see every column.

ScoreLabelWhat it means
4Strong evidenceA specific, first-hand example with their own decisions and a result, and it held up under follow-up
3Clear evidenceA specific example, with some detail missing or a result that was partly the team's
2Thin evidenceA general or hypothetical answer; follow-ups added little
1No evidenceNo example, or an example that shows the opposite
n/aNot assessedNot asked, or time ran out. Never scored as a 1

OPM's example uses five proficiency levels. Four removes the middle, which is where central tendency hides; either works if every level is written down.

One row per competency

Swipe the table sideways to see every column.

CompetencyQuestionProbes allowedWhat a 4 sounds likeEvidence, in their wordsScore
Name itThe question, word for wordTwo or three, agreed in advanceOne sentenceA quote or close paraphrase1 to 4

Then four rules that make the card worth having:

  • Evidence before score. Write what they said, then the number.
  • One competency at a time. Score each on its own row, which is the whole defence against halo.
  • Mark the must-haves. A 1 on a required competency is not averaged away by a 4 elsewhere.
  • Recommendation last, after every row is scored, and before any discussion.

What does a scorecard look like for different roles?

Four examples, four competencies each. They are starting points: replace any competency your job analysis does not support.

Backend developer

Swipe the table sideways to see every column.

CompetencyQuestionWhat a 4 sounds like
Technical depthWalk me through a service you built that is still running. What would you change now?Names the trade-offs they made, what failed, and how they knew
DebuggingTell me about the hardest production bug you fixed.A hypothesis, the evidence that tested it, the fix, and what stopped it recurring
CollaborationDescribe the last code review where you disagreed.The specific point, how it was settled, and whose mind changed
OwnershipTell me about something you shipped that went wrong after release.Their part without excuses, the rollback, and the follow-up

More questions, with what a strong answer shows, are in the Python developer guide.

Customer success manager

Swipe the table sideways to see every column.

CompetencyQuestionWhat a 4 sounds like
RetentionTell me about an account you kept that was about to leave.The early signal, what they did, and the renewal outcome
ExpansionDescribe the last upsell you found. Who closed it?How they spotted the need, and an honest account of their part
Difficult conversationsTell me about the angriest customer you handled.What they said, what they conceded, and what changed afterwards
PrioritisationHow did you decide which accounts got your time last month?A rule they actually use, with an example of it costing them something

The customer success manager guide has more, and Career1's sample report is for this role.

UX designer

Swipe the table sideways to see every column.

CompetencyQuestionWhat a 4 sounds like
ResearchTell me about research that changed a design decision.The finding, the decision it overturned, and who needed persuading
Craft and trade-offsWalk me through a shipped design and the constraint that shaped it most.The options they rejected and why
Working with engineeringWhen did an engineer say your design could not be built as drawn?What they changed, what they defended, and the result
OutcomesHow did you know a design you shipped worked?A measure chosen before launch, and what it showed

Pair it with a portfolio review, which a conversation does not replace. The UX designer guide has more questions.

Sales development representative

Swipe the table sideways to see every column.

CompetencyQuestionWhat a 4 sounds like
ResearchPick a company you would prospect. What would you say in the first line, and why?A specific reason to talk, from their own research
Objection handlingTell me about the last "not interested" you turned into a meeting.The objection word for word, and what they said back
ResilienceDescribe your worst week of outreach. What did you change?A real low point, and a change they measured
ClarityExplain what your last company sold, as you would to a prospect.Short, plain, and about the buyer's problem

How do you calibrate interviewers on a scorecard?

A scorecard only makes interviewers consistent if they read the anchors the same way. Calibration is how you find out, and it takes an hour.

  1. Agree the anchors before the first interview. For each question, write what a 4 and a 1 sound like, together.
  2. Rate the same interview separately. Everyone scores one recorded interview on their own, then compares. Talk through any competency where scores differ by more than one point, and fix the anchor that caused it.
  3. Score before the debrief, every time. OPM's guide says each panel member should rate "without discussion with other panel members" first, then compare notes, explore the discrepancies and reach a consensus. Whoever speaks first otherwise sets the answer.
  4. Look at each interviewer's scores after ten candidates. All 3s is central tendency, all 4s is leniency, the same score on every row is halo. Show people their own pattern; most correct it once they see it.
  5. Keep the cards. OPM's guide says notes should "Serve as documentation to support the employment decision". Months later, a card with quotes is the only record that explains a decision.
  6. Recalibrate when the role changes or a new interviewer joins.

Step 2 needs a recording. A recorded first round makes calibration material free: every interviewer can rate the same conversation without anyone booking a mock interview.

How do Career1's report dimensions map to a scorecard?

Career1's AI interviewer writes a report for every applicant after an interview of about eight minutes; how that interview runs is in what is an AI interviewer. Its parts line up with a scorecard like this:

Swipe the table sideways to see every column.

Scorecard partCareer1 report fieldHow to use it
Competency ratingsTechnical depth, communication, problem solving, experience fit and culture fit, each 0 to 100 with two or three sentences of reasoning citing answersA first rating to check against the transcript, not a final one
Evidence boxThe skill ledger: each skill, whether the resume claimed it, whether the interview demonstrated it, and the quote, or "not probed"The quote is your evidence; "not probed" is a question for your round
Must-havesMissing skills: required by the job and not demonstratedYour first check on the next round's scorecard
Overall recommendationMatch score, overall score, and strong hire, hire, maybe or no hireRead it last, as the scorecard says
NotesSummary, strengths, risks and detailed reasoningContext for the debrief
Next round's questionsFour to six suggested questions, based on the gapsAdd them to your human round's card

The sample interview report shows every field for a fictional customer success manager.

Where it is not a structured interview in OPM's sense, said plainly:

  • The questions differ by candidate. Career1 builds them from the job and each applicant's resume, within a fixed plan: an introduction, about three questions on the role, one on experience, one behavioural question and a close. What stays constant is the plan, the length, the job and the scoring dimensions, not the wording. OPM advises interviewers not to look at the resume during a structured interview; Career1 reads it deliberately, to test the claims on it.
  • The dimensions are fixed. Your competencies are role-specific. Map yours onto its five dimensions and the ledger, or keep your own card for the human round.
  • The scale has no written anchors. Do not convert 0 to 100 into 1 to 4 by division. Read the reasoning and the quote.
  • Culture fit is the riskiest row. It is where "similar to me" lives. Define it as job-relevant working behaviour, such as how someone handles feedback, or give it no weight.

Used that way, the report is the first round's evidence and your scorecard is the human round's structure. Your team decides; Career1 recommends and rejects nobody. Plans are on the pricing page.

Sources, checked 29 September 2026

Every page below was opened and read on 29 September 2026.

  • US Office of Personnel Management, "Structured Interviews: A Practical Guide", September 2008: opm.gov.
  • Sackett, Zhang, Berry and Lievens, "Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range", Journal of Applied Psychology 107(11), 2040 to 2068, doi 10.1037/apl0000994, abstract on the University of Minnesota's research portal: experts.umn.edu.
  • Career1's own report: the sample interview report.

Deliberately absent: the inter-rater reliability figures that circulate on scorecard template pages, because the study behind them could not be opened on the day.

Questions people ask

What is a structured interview scorecard?

It is a form that rates every candidate for a role on the same few competencies, from the same questions, on a rating scale whose levels are written down, with the candidate's own words recorded as evidence next to each score. It lets several interviewers' judgements be compared and checked, instead of debating impressions after the fact.

What should an interview scorecard include?

The role and its outcomes, four to six competencies, one or two questions per competency with the follow-up probes allowed, a rating scale with a written anchor for every level, an evidence box for what the candidate said, and an overall recommendation filled in last. Add the interviewer's name and the date, and mark which competencies are must-haves.

How many competencies should an interview scorecard have?

Four to six. The US Office of Personnel Management's structured interview guide says a structured interview is typically used to assess between four and six competencies. With more, each one gets too little time to test properly and the ratings drift back towards an overall impression. Put the rest of the role's wishes in the job description instead.

What rating scale should you use for interview scorecards?

A short one with every level written down. OPM's example uses five proficiency levels; a four-point scale removes the middle score where hesitant interviewers cluster. What matters more is the anchor: a sentence for each level describing what an answer at that level sounds like, plus a separate not-assessed option that is never counted as the lowest score.

What is the difference between a structured and an unstructured interview?

In a structured interview every candidate gets the same questions in the same order, is rated on a common scale, and the interviewers agree in advance what an acceptable answer is. An unstructured interview has none of that. OPM's guide says unstructured interviews show low rating consistency among interviewers, and structured ones a high degree of reliability and validity.

Are structured interviews better at predicting job performance?

The research says yes. A re-analysis of personnel selection studies by Sackett, Zhang, Berry and Lievens revised most validity estimates downwards and still found structured interviews the top-ranked selection procedure. The advantage comes from job-related questions and consistent scoring, so a structured interview built on vague competencies loses most of it.

How do you calibrate interviewers?

Agree what a top and a bottom answer sound like for each question before interviewing starts. Have every interviewer rate the same recorded interview on their own, compare, and fix any anchor where scores differ by more than a point. Always score before the debrief, and review each interviewer's pattern after ten candidates for leniency, strictness or identical scores on every row.

Should interviewers discuss candidates before scoring them?

No. Each interviewer should score on their own first and only then compare. OPM's guide tells panel members to rate without discussion with other panel members, then compare notes, explore the differences and reach a consensus. If people discuss first, whoever speaks first or most confidently tends to set the answer for everyone.

How do you write questions for an interview scorecard?

Start from the competency, not from a list of favourite questions. Ask for one real incident, using words such as most, last, worst or least, as OPM's guide suggests, or use a realistic scenario for people without the experience yet. Write two or three follow-up probes in advance and let every interviewer use the same ones.

Can you use the same scorecard for every role?

The structure, yes: the header, the rating scale, the evidence box and the rules. The competencies and questions, no. A backend developer and a customer success manager share little beyond communication, and a scorecard with generic competencies measures generic impressions. Write four to six role-specific competencies from the job description each time.

How do you score culture fit fairly in an interview?

Define it as job-relevant working behaviour, such as how someone handles feedback or works without close supervision, and ask for evidence like any other competency. Otherwise it becomes the similar-to-me rating error OPM's guide warns about, higher scores for people who resemble the interviewer. If you cannot define it, give it no weight.

How does an AI interview report compare to a scorecard?

It can supply the first round's evidence. A Career1 report scores technical depth, communication, problem solving, experience fit and culture fit with written reasoning, lists which resume skills were demonstrated with quotes, names missing skills and suggests questions for the next round. Its questions vary by candidate, so keep your own scorecard for the human rounds.

Try it on a real role, free

Post a role and every applicant gets an eight-minute spoken interview in the browser, with no scheduling. You get a score, strengths, risks, a recommendation, the transcript and the recording for each one.

The free plan vets two candidates a month and needs no card. Candidates never pay.

Keep reading

All articles

Career1 help

Answers in seconds, any time

Ask anything about Career1. Leave your email so we can reply if the answer needs a person.

Thinking…

Passed to a person. We will reply to .