Feature
9-Box Grid: How to Run Evidence-Based Talent Reviews with Bias Safeguards
By Priya Ellison ·

Overview
The 9-box grid is a 3×3 matrix that places each employee at the intersection of two ratings: current performance and future potential. HR and talent teams use it to structure succession planning, development conversations, and talent reviews. Its value depends on three things: job-related criteria defined before rating, calibration across managers, and a documented follow-up action for every placement.
AIHR describes the grid as a well-known talent management tool in which employees are segmented into nine groups based on performance and potential. Leapsome frames it as a succession planning tool, sometimes called a performance-potential matrix or talent matrix, that assesses employees on current performance and future growth potential. Beyond succession, the grid supports career-path discussions, retention planning, and broader workforce conversations, because it puts an entire team or population on one visual canvas that leaders can debate together.
The grid does not decide anything on its own. As Worknice puts it, the 9-box is a calibration tool, not a decision tool: its job is to make leaders defend their ratings to each other, not to assign rewards or exits. This guide walks through how to define the axes, run and document a review, translate each box into action, and decide when the model needs safeguards or supplementation.
Define job-related standards before comparing employees
The single biggest design decision in a 9-box process is what “performance” and “potential” actually mean, and that decision has to happen before anyone is rated. Folks HR states this directly: before placing anyone on the grid, get alignment on what performance and potential mean in your organization.
Performance is what an employee has achieved against agreed expectations in their current role. Worknice describes it as backward-looking and anchored to objective outputs such as goals met, formal review ratings, and customer or commercial outcomes, not general impression. AIHR notes the advantage of using the objective job requirements defined in the organization’s job structure as performance criteria, which keeps the rating tied to the role rather than to the rater.
Potential is harder because it is a forecast, not a record. The workable answer is observable behaviors. Leapsome recommends objective indicators such as learning agility, feedback responsiveness, and leadership behaviors; Folks HR adds adaptability and openness to new responsibilities. Worknice suggests defining three to five observable behaviors, for example “actively seeks feedback and adjusts” or “leads cross-functional initiatives without formal authority,” and circulating them to every manager before placements begin.
That leaves the question of role standards versus peer comparison. The sources genuinely differ here. Leapsome describes the grid as ranking team members against peers in similar roles rather than producing absolute labels, while AIHR and Personio anchor ratings to job-specific criteria. A defensible way to combine them:
- Rate first against defined, job-related standards for the role.
- Use peer comparison during calibration to test whether managers applied those standards consistently, as Worknice’s calibration question implies: “is this person strong relative to the rest of the population in this box?”
- Do not impose a forced distribution. None of the supplied sources supports requiring a fixed percentage of employees per box, and a distribution quota is a property of the population, not evidence about any individual’s performance.
Peer comparison is a consistency check on the raters. It should never replace the job-related criteria that make the rating explainable to the employee and defensible afterward.
How to run a 9-box talent review
A complete 9-box review moves from purpose to documented action in a defined sequence. Skipping steps, especially calibration, is where the process typically breaks. Worknice frames the review as a process that starts with criteria, ends with development plans, and never skips calibration.
- Set purpose and population. Decide what decision the review informs (development planning, succession pipeline, talent discussion) and which employees are in scope. The purpose shapes what evidence matters and who needs to participate.
- Define and circulate criteria. Personio advises reviewing the 9-box structure and methodology with the team and confirming how performance and potential will be measured, using each employee’s specific job description and responsibilities to create a customized assessment framework.
- Gather evidence before rating. AIHR is explicit: collect objective data to support the evaluation before placing anyone, because it reduces subjectivity and leads to more consistent decisions across teams.
- Draft placements independently. Worknice recommends each manager place their direct reports on a draft grid before the calibration session, with two or three sentences of evidence per person, done asynchronously so managers are not anchored by what they hear in the room.
- Calibrate across leaders. AIHR calls calibration a critical step; without it, different managers apply inconsistent standards. Worknice frames the 9-box as a calibration tool whose job is to make leaders defend their ratings to each other, so placements should be calibrated across managers and leaders, with HR involvement where appropriate. Personio frames positioning as a collaborative exercise between management, HR, and leadership.
- Document the outcome. Record the final placement, the evidence behind it, and any placements changed during calibration, with the rationale for the change.
- Assign a follow-up action with an owner. AIHR states that placing employees in the grid is only useful if it leads to clear follow-up actions. Worknice puts it more bluntly: a 9-box review without a development action for each placement is a labelling exercise.
Two operational details make the difference between a discussion and a process. First, collect draft placements centrally, in an HRIS, talent platform, or a structured shared sheet, so the facilitator can see the whole population at once (Worknice). Second, capture each agreed action where it will be tracked and reviewed at the next cycle rather than in meeting notes that disappear.
A practical 9-box record template
Each employee’s assessment should exist as a structured record, not just a dot on a slide. A minimal record ties the rating to its evidence and to the action it produced, which is what makes the placement explainable later. Based on the workflow documented by Worknice, Personio, and AIHR, each record needs the following fields:
- Employee and current role (with the job description used as the assessment baseline)
- Assessment period covered by the evidence
- Performance rating (low / moderate / high) against defined role expectations
- Potential rating (low / moderate / high) against defined observable behaviors
- Supporting evidence: two to three sentences per axis citing goals met, review-cycle outcomes, or observed behaviors
- Grid placement, plus any change made during calibration and why
- Follow-up action: development plan, succession note, performance conversation, or a deliberate “stay the course” entry
- Action owner and target timeline
- Review trigger: the next scheduled cycle or a material event such as a role change
Store these records in whatever system the organization already uses for talent data, so actions are tracked and revisited at the next review rather than recreated from scratch.
Calibration-session checklist
Calibration exists because unmanaged rating styles distort placements. Consider a generalized failure mode: two managers rate equally capable employees, but one manager rates generously and the other conservatively, and neither works from written criteria. The lenient manager’s reports land in high-potential boxes and receive development investment; the strict manager’s reports do not. The difference in access to opportunity reflects rating style, not employees. SIGMA Assessment Systems notes that standardized definitions enhance objectivity, and AIHR warns that without calibration, managers apply inconsistent standards.
A facilitator should verify the following in every session:
- The right participants are present: the managers who rated, plus HR facilitation (Personio, Worknice)
- Every draft placement arrives with written evidence, not impressions (AIHR, Worknice)
- The same defined criteria are applied to every team; Folks HR states the same standard should apply whether assessing one department or ten
- Outlier placements are challenged; Leapsome advises challenging any assumptions that feel unfounded and calibrating across teams to maximize fairness
- Placements changed in the room are recorded with their rationale
- Every final placement leaves with an action, an owner, and a follow-up date
Close the session by reading back changed placements and assigned actions. If a placement cannot be defended with evidence against the written criteria, it should change, and the record should say so.
The nine boxes and the actions they should trigger
Each of the nine intersections describes a combination of demonstrated performance and assessed potential, and each should trigger a conversation, not an automatic outcome. AIHR notes that combining the two dimensions creates nine possible employee profiles, and Leapsome emphasizes that each block should feed into development plans and career-path conversations rather than fixed labels.
| Potential ↓ / Performance → | Low performance | Moderate performance | High performance |
|---|---|---|---|
| High potential | Possible role misfit or unmet needs. Ask whether expectations were clear and whether the role fits. Personio notes some employees in this position are simply poorly placed and improve in a better-suited role. | Growing talent. AIHR suggests mentoring and targeted development to convert potential into results. | Top talent. AIHR suggests stretch roles or leadership programs; a succession plan named for a specific role and horizon fits here (Worknice). |
| Moderate potential | Inconsistent contributor. Diagnose the performance gap first: skills, workload, management, or fit. Set clear expectations with a timeline before drawing conclusions. | Core of the workforce. Steady development, skill broadening, and a deliberate “stay the course” note where appropriate (Worknice). | Strong performer with room to grow. Offer targeted stretch work or deepening in the current track; discuss which direction the employee actually wants. |
| Low potential | Underperformer. AIHR notes this placement may lead to performance improvement plans or role changes; verify evidence and context before any consequential step. | Solid contributor at the right level. Support and retain; do not treat the absence of upward trajectory as a problem to fix. | High-performing specialist or expert. Recognize, retain, and consider mentoring roles; do not force a management path. |
Two readings of this table matter more than the labels. First, similar placements can lead to different paths depending on role and context. A high-performance, low-potential engineer may become a technical mentor; a high-performance, low-potential team lead in the wrong function may be a candidate for role realignment. Personio documents that employees with potential but low performance can improve in another, better-suited role or department, which means “low” placements are diagnostic questions, not verdicts.
Second, no box outcome is automatic. Paylocity notes that analyzing the grid can inform career-path strategies such as promoting, retraining, or ending employment, but Personio’s caution is the controlling principle: the grid should not make the final decision on an employee, their potential, or their future; it establishes a position and begins a productive conversation. Development actions belong to the review. Consequential employment decisions require a separate, evidence-based process, covered later in this guide.
Keep placements provisional, reviewable, and useful to employees
A placement is a dated judgment, not a durable attribute of a person. Folks HR makes both halves of this point: the grid captures a moment in time, and the grid is a living document, not a static label. AIHR similarly advises tracking progress over time instead of treating the grid as a one-off exercise. The supplied evidence does not establish one correct review cadence, so the practical rule is trigger-based: revisit a placement at each regular review cycle and after material changes such as a new role, new manager, completed development plan, or new evidence. Retain prior placements and their supporting evidence so movement is visible; an employee who moved from a middle box to a high-performance box after clearer expectations or coaching is exactly the signal the process exists to surface.
Communication is a separate decision from recordkeeping. Folks HR describes a common practice: the grid itself stays internal, and what gets shared is a development conversation covering what is expected, where the gaps are, and what the plan is going forward. AIHR notes that some companies decide not to communicate the potential score to employees. The supplied evidence does not resolve whether disclosing an exact box placement helps or harms outcomes, so treat full disclosure as a policy choice, not a best practice with proof behind it. What is supported is giving employees a route into the process: SIGMA describes employees participating through self-assessment, and EEOC guidance recommends explaining the reasons for employment decisions to affected persons and allowing employees, without negative consequences, to have appraisals reviewed and corrected when appropriate. A placement built on an inaccurate appraisal should be correctable through the same route.
Handle edge cases without turning context into a label
Snapshot ratings misfire when the snapshot itself is unrepresentative. Folks HR names the problem: the grid does not account for an employee who is mid-transition, navigating a difficult manager relationship, or going through a major life change. The supplied sources do not prescribe scoring rules for these cases, so the sound approach is to qualify the rating, gather more evidence, or defer placement, and to write that qualification into the record rather than forcing a number.
For new hires and recent promotions, the evidence base for a performance rating may simply not exist yet; note the short observation window instead of guessing. For employees returning from leave, be alert that the period covered by the evidence may not reflect capability; Worknice reports that potential ratings skew against employees on parental leave when criteria are not anchored. For employees in weak role fit, remember Personio’s point that low performance can reflect placement, not capability.
A contrast exposes the hidden assumption. An illustrative high-performing specialist who does not want a management role is not a problem case; a low potential-for-bigger-roles rating alongside high performance describes a valuable expert to retain, possibly as a mentor. An illustrative low performer three months into a role transition is not a confirmed underperformer; the honest record says “insufficient evidence, reassess next cycle.” Treating both as fixed labels would misallocate development in one case and damage a career in the other.
Benefits, limitations, and bias safeguards
The grid’s core benefit is that it makes a complex talent conversation visible and shared. Folks HR calls visual simplicity its greatest strength, and Personio notes it is especially useful as a visual tool for HR and talent management. Paylocity summarizes it as simple, user-friendly, holistic, and cost-efficient. Done well, Worknice argues, it surfaces bias by forcing leaders to defend their ratings to each other.
The limitations are equally well documented, and they cluster around what the two axes cannot capture:
- Subjectivity, especially on potential. Paylocity notes the model lacks an objective ranking method and is prone to bias. BambooHR states that because it is a subjective appraisal, it is vulnerable to conscious or unconscious bias. Worknice reports that potential is the axis most vulnerable to bias, skewing against women, employees on parental leave, and underrepresented groups when criteria are not anchored.
- Omitted context. AIHR notes that reducing employees to two dimensions can leave out specific skills, motivation, role complexity, learning agility, and external circumstances.
- Rigidity and snapshot effects. Paylocity flags rigidity; Folks HR flags the moment-in-time problem, where a placement made under temporary circumstances hardens into a durable label if never revisited.
The honest framing is that safeguards reduce these risks; they do not eliminate them. Anchored criteria, evidence requirements, and calibration narrow the room for individual rating style to determine outcomes. StaffCircle adds that gathering input from a variety of team members can limit bias risk, since a single rater’s blind spots weigh less. None of this converts a subjective judgment into a measurement, which is why the grid should inform conversations rather than settle consequential decisions.
Governance controls for a more defensible process
Because 9-box outputs can influence pay, promotion, and other employment decisions downstream, the process needs governance controls, not just facilitation habits. The controls below draw on U.S. EEOC guidance on evaluating employment decisions; readers outside the United States should map the same principles to their own legal context rather than treating this as a universal legal checklist.
- Use job-related criteria. EEOC guidance recommends analyzing the duties, functions, and competencies relevant to jobs, then creating objective, job-related qualification standards tied to them.
- Apply criteria consistently. The same guidance calls for consistency: comparable job performances should receive comparable ratings regardless of the evaluator, and appraisals should be neither artificially low nor artificially high.
- Document rationale. EEOC guidance advises making decisions transparent to the extent feasible and documented, with reasons well explained to affected persons and records retained.
- Monitor for patterns. The guidance recommends monitoring performance appraisal systems for patterns of potential discrimination and conducting self-analyses to determine whether practices disadvantage protected groups.
- Provide a correction route. Allow employees, without negative consequences, to have appraisals reviewed and corrected when appropriate.
- Limit access. Define who can view full grid placements versus individual development plans. The supplied evidence supports keeping the grid internal (Folks HR) but does not specify jurisdiction-level privacy rules; set access policy deliberately and check local requirements.
- Separate the grid from consequential decisions. When outputs feed decisions with adverse-impact potential, EEOC discussion of selection procedures points to validation: evidence that a tool measures job-related skills or attributes rather than relying on arbitrary or implicitly biased features.
None of these controls is exotic. Most are documentation and consistency disciplines that a well-run calibration process already produces, provided the outputs are actually retained.
Use the grid for calibration, not as a final employment decision
The decision boundary is simple to state: a 9-box placement is an input to diagnosis and discussion, not a sufficient basis for promotion, compensation, succession appointment, or exit. Personio states that the grid should not be used to make the final decision on an employee, their potential, or their future within the company; it is a starting point for a productive conversation. SIGMA describes it the same way, as a starting point for conversations about employee development. Worknice draws the line most sharply: done badly, the grid labels people, entrenches bias, and gets misused for pay or termination decisions it was never designed to make.
The mechanism behind this boundary is worth spelling out. A placement compresses a year or more of context into two three-level ratings, at least one of which (potential) is a subjective forecast. That compression is acceptable for prioritizing development conversations, because a conversation can recover the lost context. It is not acceptable as the direct trigger for a consequential decision, because the decision inherits the compression and the rater’s subjectivity along with it. Testimony to the EEOC on automated and selection systems describes the standard that applies when a tool does drive selection outcomes: a job analysis to understand what the job requires, examination of the relationship between the tool and the job requirements, and a validity study supporting the interpretation made from the score. A typical 9-box process has none of that.
The operational rule: let the grid decide what conversation happens next. Let a separate, documented, job-related process, with its own evidence standard, decide who is promoted, paid differently, appointed as a successor, or exited.
When to supplement or replace the 9-box grid
Whether to keep the grid depends on the decision being made, not on whether the model is fashionable. BambooHR captures the mainstream position: the 9-box provides an accessible starting point but is not always the most comprehensive, and in most cases it should be combined with a more robust model. The supplied evidence does not establish that any named alternative performs better; what it supports is matching the tool to the question.
- Development dialogue and talent-review calibration: the standard grid fits, provided criteria are anchored and follow-up actions are tracked (SIGMA, AIHR).
- Role-competency evidence: the grid omits specific skills and role complexity (AIHR), so pair it with competency-based assessment against the job’s defined requirements.
- Near-term succession readiness: the potential axis signals direction but not timing; add readiness definitions naming the target role and horizon (Worknice’s succession-plan output points this way).
- Skills visibility across a workforce: a two-axis matrix cannot represent a skills inventory; use a skills-based framework alongside it.
- Consequential employment decisions: replace the grid as the decision mechanism entirely and use a documented, job-related process, per the EEOC material above.
Modifying the grid, for example by publishing behavioral anchors for potential or adding structured self-assessment (SIGMA), addresses many objections at lower cost than replacement. Replace it only when the decision at hand needs information the two axes cannot carry.
How accurate are 9-box potential ratings?
The supplied evidence does not establish a predictive-accuracy or inter-rater-reliability estimate for 9-box potential ratings. No source in this article’s evidence base reports empirical data on how well a potential rating forecasts later performance, promotion success, or leadership readiness, so no honest number can be given.
That gap has a practical consequence rather than a rhetorical one. EEOC meeting testimony on selection procedures describes what accuracy evidence looks like when a tool influences employment outcomes: empirical data showing the procedure is significantly correlated with important elements of job performance, plus documented evidence of validity and reliability, because “I think this test looks good” is not sufficient. A typical 9-box process carries no such evidence. The reasonable response is not to abandon the grid but to keep its role proportionate to what is known: use it to structure development conversations and calibrate manager judgments, document the evidence behind each rating, and require a separately validated, job-related process before any placement influences a consequential decision.