WCapsuleM8

Interview Scorecard

$19

Score every interview against the same anchored criteria — evidence-based 1–5 scores per interviewer, red flags, calibration between interviewers and a side-by-side candidate comparison; the recruitment pipeline tracker follows candidates through stages, this scores the interviews themselves. Nothin

Version 1.0.0 · Updated Aug 7, 2026

Overview

Score every interview against the same anchored criteria — evidence-based 1–5 scores per interviewer, red flags, calibration between interviewers and a side-by-side candidate comparison; the recruitment pipeline tracker follows candidates through stages, this scores the interviews themselves. Nothing is uploaded.

Frequently asked questions

How does the Interview Scorecard licence work?

It is a one-time purchase for a downloadable tool — no subscription. You buy it once and the file is yours to keep and use.

Can I try the Interview Scorecard before buying?

Yes. Use the Try online button for a fully interactive demo with sample data already loaded — nothing to install and nothing is saved.

Does my data stay private?

Yes. The tool is a single HTML file that runs entirely on your computer and makes no network requests, so nothing you enter is ever uploaded or shared — which matters for people data.

Do I need Excel or any other software?

No. It replaces the spreadsheet template entirely: open the file in your browser (Chrome, Edge, Firefox or Safari) on Windows, Mac, Linux or a tablet, and start working.

How to use Interview Scorecard

The complete in-tool guidance, reproduced here so you can read it before you download.

What this tool does

CM8-267 is a structured interview scorecard. Each row is one score: a vacancy, a candidate, an interviewer, a criterion, an anchored 1–5 mark, and — the part that matters — the evidence actually heard in the room. From those rows it builds a ranking per vacancy, a per-type profile of each candidate, a calibration view of your interviewers, an audit trail of every red flag and extreme score, and a printable comparison for the decision meeting.

Everything runs inside this single file. There is no account, no upload and no network request of any kind, which matters here: an interview record contains candid judgements about identifiable people who may one day ask to read them.

Why structured interviews win

Decades of selection research point the same way: unstructured chats — different questions for each candidate, scored by overall impression afterwards — predict job performance poorly, and their verdicts are driven substantially by confidence, fluency and similarity to the interviewer. Interviews where every candidate faces the same criteria, defined before anyone is met, scored against written anchors, predict meaningfully better and are far easier to defend afterwards. Structure does not make interviewers objective; it makes their subjectivity visible, comparable and open to challenge, which is the best an interview can honestly do. That is the whole design of this tool: same criteria for everyone, a number only ever next to its evidence, and the disagreements between interviewers surfaced rather than averaged away.

Building the criteria

Write four to seven criteria from the job, before you see a single application, and use the same set for every candidate for that vacancy. Start from what the person must actually do in the first year — run a shift handover, plan labour against a schedule, challenge an unsafe act — and turn each into a criterion with a plain name, such as "Planning & organisation" or "Safety mindset".

Each criterion carries a type, and the type is a discipline in itself:

  • Skill — can they do it? Judged by watching or probing the doing.
  • Experience — have they done it? Judged on specifics: when, where, what happened.
  • Motivation — why this role, this organisation, now?
  • Values / team fit — judged on evidence of behaviour, never on whether you warmed to them. "Fit" scored on vibes is the polite name for bias.

Fewer than four criteria and one lucky answer dominates the average; more than seven and interviews become box-ticking. If a criterion cannot be scored from something a candidate could plausibly say or do in an interview, rewrite it or drop it.

Anchored scores and evidence discipline

Every score uses the same anchored scale: 1 clear gap, no evidence; 2 weak, thin evidence; 3 adequate, some solid evidence; 4 strong, repeated evidence; 5 exceptional — they could teach it to others. The anchors are about evidence heard, not impression formed: score what the candidate said and did, not what you inferred about the kind of person they probably are.

The evidence box is the tool's real product. A quote or a concrete description — "baseline 42 minutes, now 28" — lets the decision meeting weigh the judgement; "seemed strong" lets it weigh nothing. The tool enforces the discipline at the extremes: a 1 or a 5 will not save without evidence, and neither will a red flag. Extreme marks decide hires, so they carry the burden of proof.

Candidate average = sum of all criterion scores ÷ number of scores Interviewer spread (per candidate) = highest interviewer's average − lowest interviewer's average

A candidate's average only appears on the tiles once they have the minimum number of scores set in Settings — an average of two numbers is noise wearing a decimal point.

Behavioural questioning in one paragraph

The most reliable evidence comes from past behaviour, so ask for it directly: "Tell me about a time you had to re-plan a shift at short notice" — then follow up until you can see the event. What was the situation? What did you do, as opposed to the team? What happened next, and what would you do differently? Vague answers get one follow-up chance; hypotheticals ("I would…") are worth less than history ("I did…"); and the follow-ups are where the scoring evidence lives. Write down what you heard while it is fresh — the evidence box is designed to be filled in during or immediately after the interview, not reconstructed a week later.

Red flags vs dislikes

A red flag is a genuine disqualifier you observed: the candidate described bypassing a safety lockout and called it resourcefulness; they claimed a qualification the certificate contradicts; they spoke about former colleagues in a way you would not accept on your own floor. It is evidence-based, it would disqualify whoever showed it, and you must record what you observed — the tool will not save the flag without it, because a flag may be reviewed later, possibly by the candidate.

A dislike is different: an accent, a nervous manner, a career gap, a personality unlike yours. Dislikes are where interview bias lives. If something bothers you and you cannot write down observed evidence that connects it to a criterion, it is not a flag — leave it out of the record and out of the decision. The red-flags table keeps every flag and every extreme score with its evidence, interviewer and date: that is the audit trail for the strongest judgements made in the process.

Calibration between interviewers

Two honest interviewers can watch the same interview and score a point apart, because one carries a stricter internal anchor. The calibration chart shows each interviewer's average across everything they have scored — who marks hard, who marks soft — and the spread tile shows the widest gap between two interviewers' averages for the same candidate. When that gap exceeds your alarm threshold (one point by default), do not quietly average it away: get the two interviewers to compare evidence. Usually one heard something the other did not, and the conversation either surfaces real information or re-anchors the scale — both outcomes improve the decision. An outlying scorer is not a wrong scorer; an outlying scorer nobody talks to is a hidden thumb on the scale.

Comparing candidates fairly

Averages only compare candidates measured the same way: same criteria, same anchors, and scores on the full set. The coverage tile flags candidates missing scores on criteria that others have been scored on — a candidate scored only on their strong suits will beat a fully-scored rival without being better. Before the decision meeting, aim for full coverage of every finalist; where a gap cannot be filled, the comparison table shows a "—" for that type rather than pretending. The ranking chart, the shape chart and the comparison table all restrict themselves to one vacancy at a time — comparing averages across different vacancies with different criteria means nothing.

Keeping notes lawful and professional

In most countries candidates can lawfully request their interview notes, and notes can surface in a dispute. The working rule: write what you would be comfortable reading aloud to the candidate. Record behaviour and words, not diagnoses or protected characteristics; "did not give an example of leading through conflict" is a professional note, "not leadership material" is an opinion in search of trouble, and anything touching age, health, family plans or background does not belong in an interview record at all. If scorecards will circulate beyond the panel, consider anonymised candidate codes in the candidate field and keep the code key separately.

Scorecard vs recruitment pipeline

These are different documents doing different work. The recruitment pipeline tracker (CM8-24) follows candidates through stages — applied, screened, interviewed, offered — and measures the funnel: volumes, conversion, time to hire. This scorecard measures the interviews themselves: what was asked, what evidence was heard, and how it was scored, so the decision between finalists rests on more than the most confident voice in the debrief. Run both if you like; they answer different questions and neither replaces the other.

FAQ

Should every interviewer score every criterion? Not necessarily — panels often split criteria so each is probed properly. What matters is that every candidate ends up scored on every criterion by someone, and that any criterion used to compare two candidates was scored the same way for both.

What does the recommendation column mean if there is one per row? It is the interviewer's overall view of that interview, repeated on each of their rows — fill it identically across the interview, or leave it neutral until the interview ends. The comparison table counts it once per row, so keep it consistent.

Can we weight criteria? This tool deliberately does not. Weights add a second layer of judgement that is rarely evidenced and usually retro-fitted to justify a preference. If one criterion is truly decisive — safety in a supervisor role — treat a failure on it as a red flag, not as a number to be outvoted by charm elsewhere.

Does the highest average win? The average opens the conversation; the evidence closes it. A red flag is not offset by a high average, a thin coverage row is not comparable, and a large interviewer spread means the number is not yet settled. The decision is yours — the scorecard just makes it inspectable.

Saving your work

Scores, settings and the report header are written to this browser's local storage as you type, and the toolbar shows the time of the last save. That storage belongs to one browser on one computer: another browser, a private window, a second machine or a clean-up tool that clears site data will not have it.

Treat Export .json as the real save — one file containing everything, which Import .json restores anywhere. Export CSV gives you the register for spreadsheet work. Reset asks twice, then erases everything this tool has stored. There is no undo. Interview records are personal data about identifiable people — store and share exports accordingly, and delete them when your retention period ends.

Accuracy & disclaimer

The arithmetic here is deliberately simple — averages of the scores you enter — and the tool does it faithfully. Everything that matters sits underneath: whether the criteria reflect the job, whether the scores reflect evidence rather than affinity, and whether the panel discussed its disagreements instead of averaging them. Structure reduces the room bias has to work in; it does not eliminate it.

Employment law, discrimination law and data-protection rules on interviewing and interview records differ by country. This is a record-keeping and comparison aid, not legal advice and not a hiring decision — the decision, and the duty to make it fairly, stay with you.

Run a 9-box talent review — place people on performance and potential with evidence, calibrate the placements, and track flight risk, succession cover and development actions. Runs entirely in your browser — nothing is uploaded.

Download Runs in browserView

Track annual leave, sickness and every other absence in one register: entitlement and balance per person, Bradford Factor, cover clashes and a printable report. Runs entirely in your browser — nothing is uploaded.

DownloadView

Keep one reliable people register: contracts, hours, pay, probation ends and document expiries, with headcount, full-time equivalent and annualised cost worked out for you. Runs entirely in your browser — nothing is uploaded.

DownloadView

Run structured performance reviews — weighted objectives with evidence, anchored 1–5 ratings, development goals, the employee's own comments and a signed meeting record, printed as a report for the file. Nothing is uploaded.

Download Runs in browserView