Review quality doesn't live in the rating scale or the form — it lives in a handful of design choices that hold whatever tool you use. A high-quality process has seven: a shared, explicit standard of what "good" looks like at each level; multiple credible inputs instead of one lead's memory; evidence over impression; calibration so a score reflects the work, not the rater; a deliberate split between measuring and developing; an ending in a concrete forward plan, not just a number; and a cadence that actually runs and reaches everyone.
Get those right and almost any format works; get them wrong and no template saves you. The reason is blunt: when researchers decomposed the ratings of 4,492 managers, about 62% of the variance came from the individual rater's idiosyncrasies and only ~21% from the person's actual performance (Scullen, Mount & Goff, Journal of Applied Psychology, 2000). A review process is high-quality when it engineers that noise back down.
Why isn't a good form enough?
Because the form is the cheapest part of the system and the least predictive of the outcome. Teams spend their energy choosing a 5-point scale versus a no-ratings model versus a narrative template, then run the same thin process behind it: one manager, one memory, assembled the night before. The result is a number that mostly measures the rater. A single supervisor's rating of overall performance has an interrater reliability of just .52 — internally consistent, but two raters agree only moderately (Viswesvaran, Ones & Schmidt, Journal of Applied Psychology, 1996).
The cost of that isn't only inaccuracy. Feedback delivered badly is not neutral: across 607 studies it raised performance on average, but over a third of feedback interventions actually made performance worse (Kluger & DeNisi, Psychological Bulletin, 1996). And after a century of research there is surprisingly little evidence that appraisal on its own improves performance at all (DeNisi & Murphy, Journal of Applied Psychology, 2017). So "high-quality" can't mean "a slicker form." It means the design choices that decide whether the process produces a fair, defensible number — and whether anything changes afterwards.
What actually makes a review process high-quality?
Seven design choices. They sit underneath whatever scale you pick, and they are where the quality is won or lost.
- A shared, explicit standard. A written career framework that says, in concrete behaviours, what "good" looks like at each role and level — the practitioner standard for defining roles (CIPD, Competence and competency frameworks). Without a common bar, every reviewer invents their own and a "Grade 3" means something different to each lead.
- Multiple credible inputs. Evidence from the leads, peers and (where relevant) clients the person actually worked with — not one manager's recency-weighted recollection. More independent vantage points cancel noise a single rater can't, the direct answer to that .52 reliability problem.
- Evidence over impression. Every judgment anchored to a concrete example tied to the framework, not a global gut feel. This is what blunts recency, halo and leniency bias and turns "I think she's a 3" into "here's the behaviour at this level."
- Calibration across reviewers. A deliberate step where the leads who rated different people reconcile their scores against the shared bar. This attacks the 62% rater-idiosyncrasy problem head-on (Scullen, Mount & Goff, 2000) — it keeps a score about the work, not the rater's generosity.
- Separate "measure" from "develop." A defensible rating and an honest growth conversation pull in opposite directions; when both ride on one number, candour gets punished. High-quality processes design the two so a person can hear hard feedback without it instantly threatening their grade.
- End in action, not a score. A review changes behaviour only with a concrete, forward-looking plan and follow-up. Improvement after feedback is generally small and conditional — it shows up when coaching and a plan follow, not from the act of rating (Smither, London & Reilly, Personnel Psychology, 2005).
- It runs and it reaches everyone. On cadence, with high completion, even when the team is billable. A beautifully designed process that only half the firm completes is not high-quality — it's a pilot.
What actually drives a rating?
Mostly the rater, unless you design against it. When Scullen, Mount and Goff decomposed the ratings of 4,492 managers, the largest single share of the variance traced not to how people performed but to the quirks of whoever was holding the pen (2000). That is the empirical case for choices 2, 3 and 4 above: multiple inputs, evidence and calibration exist precisely to shrink the violet bar and grow the teal one.
Does the rating scale matter at all?
Less than the design choices, but it isn't nothing. The well-known revolt against ratings came from cost and futility, not from the scale itself: 58% of executives said their performance-management approach drove neither engagement nor performance, and Deloitte alone was spending close to two million hours a year on forms, meetings and ratings (Buckingham & Goodall, Harvard Business Review, 2015). The dissatisfaction hasn't gone away — as of 2023, only 29% of HR leaders were confident their current process effectively helps people perform (Gartner, 2023).
The lesson isn't "abolish ratings" or "keep them." It's that the scale is a downstream choice. Ratings, no-ratings, or both can each be high-quality if the seven design choices are in place, and each is low-quality without them. Pick the scale that fits how you make staffing and promotion decisions — then spend your real effort on the standard, the inputs, the calibration and the follow-up.
Why does this matter for agencies and consulting boutiques?
Because the generic answer assumes one manager who watched the person all year — and in a project-based firm, that manager rarely exists. People are staffed across several engagements under different leads, rated per engagement by whoever ran it, and the score feeds grades, raises and the partner track. The 62% rater-idiosyncrasy finding (Scullen, Mount & Goff, 2000) isn't an abstraction here; it's the difference between a fair promotion case and a political one.
So "high-quality" translates directly into your reality. A shared bar lets different leads mean the same thing by a grade. Multiple credible inputs reconstruct a year no single lead saw. Calibration across leads is what makes a rating defensible when it decides the partner track. And because billable pressure eats the time, the process has to be light enough to actually run at high completion without burning the hours that pay for it. Get this right and a review is fair and trusted; get it wrong and the cost lands on the senior talent a boutique can least afford to lose — an opaque verdict reads as an unfair one.
A quick self-check: is your review process high-quality?
Score one point per "yes". Six or seven and your process is genuinely high-quality; four or fewer and you're relying on the form to do work it can't.
- There is a written, shared standard (a career framework) for what "good" looks like at each level.
- Ratings draw on multiple credible inputs, not one lead's memory.
- Every judgment is anchored to concrete evidence tied to the framework.
- Leads calibrate scores against the shared bar before they're final.
- "Measure" and "develop" are designed so candour isn't punished.
- Every review ends in a concrete forward plan with a follow-up date.
- The cycle runs on cadence and reaches everyone, even when the team is billable.
Scored four or fewer? Book a call and we'll walk through where your process is leaking and what to fix first.
FAQ
Is a performance review process about the rating scale?
No. The scale (5-point, no-ratings, narrative) matters far less than the design choices around it: a shared standard, multiple inputs, evidence, calibration, and a forward plan. Most of the variance in a rating traces to the rater rather than the performance (Scullen, Mount & Goff, 2000), so quality is won by engineering that noise down — not by re-drawing the form.
What's the single biggest driver of a fair rating?
Calibration on top of multiple inputs. A single supervisor's rating of overall performance has an interrater reliability of only .52 (Viswesvaran, Ones & Schmidt, 1996); adding independent vantage points and then reconciling them against a shared bar is what makes a score about the work rather than the rater.
Should we drop ratings entirely?
Not automatically. The backlash was driven by cost and low usefulness — 58% of executives saw no engagement or performance benefit (Buckingham & Goodall, 2015), and only 29% of HR leaders were confident in their process in 2023 (Gartner, 2023). But ratings, no-ratings or a hybrid can each be high-quality with the seven design choices in place, and low-quality without them.
Why doesn't a good review automatically improve performance?
Because appraisal on its own rarely does (DeNisi & Murphy, 2017), and feedback can even backfire — over a third of interventions reduced performance (Kluger & DeNisi, 1996). Improvement shows up when a plan and follow-up come after the conversation (Smither, London & Reilly, 2005), which is why "end in action" is a design choice, not a nicety.
How do we keep quality high when everyone is billable?
Make the heavy parts one-time and the recurring parts light. The framework and templates are built once; each cycle is then mostly assembling inputs you can collect asynchronously around client work, plus a short calibration. High completion comes from a process that fits the billable week, not one that fights it.

