Inclusive Hiring: Shared Criteria and Reviewable Decisions
Unstructured interviews quietly penalize underrepresented candidates. Shared criteria help teams compare evidence and examine differences in outcomes.
Ernest Bursa
Shared questions, behavioral rating anchors, and independent reviews make hiring decisions easier to inspect. They can reduce some sources of bias, but a scorecard alone does not make a process fair or establish where differences in outcomes arise.
Preparing questions, calibrating reviewers, and auditing outcomes take time. The evidence supports examining structured interviews; it does not establish a cost-free intervention or a guaranteed result from adding software.
Where the advancement gap actually gets built
Differences may arise in sourcing, requirements, screening, interviews, or final decisions. Examine each stage rather than assuming that the review step is almost always responsible.
Think about the everyday face of it: the “weak no, didn’t quite click” with no criterion attached. That is affinity bias wearing a casual outfit. It feels like judgment, but it is a reaction to similarity, communication style, or shared background, dressed up as a hiring signal. Anchored review forces the only question that matters: didn’t click on what job-relevant dimension? More often than not, the answer evaporates, and a qualified candidate who was about to be filtered out stays in.
You cannot fix this by hiring more people into the top of the funnel while the review step keeps leaking. You fix it by changing what the review step is allowed to measure.
Why unstructured interviews disadvantage underrepresented candidates
Unstructured interviews disadvantage underrepresented candidates because they maximize discretion, and discretion is exactly where bias operates. Improvised questions, holistic “gut” scoring, and ratings formed mid-conversation are the points where affinity bias, halo effects, and confirmation bias quietly drive outcomes.
Huffcutt and Roth (1998), Table 2, report Black-White mean rating differences of d = 0.32 for low-structure and d = 0.23 for high-structure interviews, not the often-repeated 0.56 versus 0.23. Bobko and Roth (2013) report a structured-interview difference near d = 0.25. Original Huffcutt and Roth paper.
The fix is to remove the points of discretion one by one: ask everyone the same job-related questions, define what each score looks like in observable behavior, have interviewers score independently before they talk, and combine ratings mechanically instead of debating to a vibe.
Do structured interviews reduce bias?
The cited research associates greater interview structure with smaller average rating differences between the groups studied. That is not a guarantee that adding anchors to a particular hiring process reduces bias by a specified amount.
The statistic d measures a standardized difference between group means. It is not an advancement rate or the proportion of discriminatory decisions. A value near zero alone does not establish fairness, and a difference requires investigation of its causes.
Structure includes several components. These comparisons do not establish that each added component produces a fixed additional improvement in validity or subgroup differences, or that rating anchors alone cause the observed effects.
What a behaviorally anchored rating scale actually is
A behaviorally anchored rating scale (BARS) replaces abstract labels with described behavior, so a “3” means the same thing to every reviewer. Instead of scoring “communication” from 1 to 5 in the abstract, the scale spells out each level: a 5 might be “structured the answer, surfaced tradeoffs unprompted, checked my understanding”; a 2 might be “answered the question but needed prompting to go deeper.” The ETS research on building BARS for structured interviews (Kell et al., 2017) ties their use to higher reliability and lower bias. Anchors are what stop a scale from drifting back into a personality contest. They are the difference between a rubric that improves fairness and one that just adds paperwork.
The rare double win: fairer and more predictive
Predictive validity and subgroup mean differences answer different questions. Consider both, without treating either as proof that a process is fair.
Here is why the contrast is so stark. Compare the methods on both axes at once:
| Method | Predictive validity | Black-White subgroup gap (d) |
|---|---|---|
| Low-structure interview | r ≈ .20 | ≈ 0.32 |
| Structured / anchored interview | r ≈ .57 | ≈ 0.23 |
| Cognitive-ability test | r ≈ .51 | ≈ 1.0 |
The table combines historical estimates from different analyses. It is not a head-to-head experiment, a current product benchmark, or evidence that one method guarantees a diverse and high-performing team. The .20–.57 validity estimates concern entry-level jobs and different levels of interview structure.
One precision note, because this is where most articles get caught overstating. The .20-to-.57 range comes specifically from Huffcutt and Arthur’s (1994) four-level structure taxonomy, not from the famous Schmidt and Hunter (1998) figures (which report .51 structured versus .38 unstructured). Both support the thesis. Conflating them is the most common error in the secondary literature, and citing the merged version marks work that copied a competitor’s blog rather than reading the research. We covered the validity side in depth in structured interview scorecards and predictive validity; this article is about the equity side of the same change.
The opaque-AI shortcut makes it worse, not better
Using the same vendor’s screening algorithms across employers can create correlated outcomes. Whether a particular tool adds bias requires evidence about that tool, its use, and the relevant applicant population.
The 2026 study “Algorithmic Monocultures in Hiring” (Bommasani et al., FAccT ’26) analyzed 4,197,168 applications from 3,372,132 applicants across 156 employers using one vendor. 25.87% of applications submitted by Black applicants went to positions meeting the study’s adverse-impact criterion for that group. This does not mean that 25.87% of people were rejected because of race or that every rejected applicant was rejected everywhere. Authors’ paper.
Anchored human review is the inverse architecture. The criteria are explicit, the evidence is shared, a person makes the call on the record, and the decision is auditable and correctable. The goal is not to remove humans from hiring; it is to give the human a structure that caps how much bias can enter and a paper trail that lets you check whether it did. We unpacked the broader failure mode in how AI hiring tools produce industry-wide exclusion.
How to make interviews more inclusive
You make interviews more inclusive by removing discretion at every point where bias enters, then auditing the outcome. Four moves, in order:
- Ask everyone the same job-related questions. Fix the question set before you see a single candidate. Improvised questions are where confirmation bias steers the conversation toward people who already impressed you in the first two minutes.
- Score against anchored criteria, not impressions. Use a BARS so a “4” means the same observable behavior to everyone. Practice applying the anchors and investigate disagreements.
- Record independent scores before the debrief. Independent ratings submitted before discussion reduce exposure to the anchor where the first or most senior voice sets the reference point. Combine the scores mechanically; do not debate to a feeling.
- Audit advancement rates by group. Where lawful and appropriate, analyze who advances at each stage by group. Establish the legal basis, safeguards, and adequate sample size before collecting or interpreting demographic data. This is how you catch a leak in real time instead of discovering the gap a year later in a headcount report.
An audit can identify differences that need investigation; it does not certify fairness. Keep the process manageable for candidates with limited scheduling flexibility. We discuss why too many interview rounds lose candidates.
How Kit builds anchored, auditable review in
Kit supports shared criteria and review records. It does not turn published effect sizes into a product guarantee.
- Team review criteria. Set criteria and weights on a team review stage, then describe relevant answers in the instructions. Reviewers supply the evidence.
- Blind review. When enabled, reviewers cannot see others’ cards until submitting their own. Cards remain editable.
- Configured decisions. Votes can automatically advance an application, and a lead veto can trigger rejection. Unresolved cases enter the decision queue. Authorized MCP tools can also change application state.
- Process records. Ratings and application history help the team inspect decisions. They do not constitute a built-in demographic-disparity report or automatic legal compliance assessment.
Prepare the questions and anchors, record evidence before discussion, and revisit unclear criteria after the hire. Start a free trial to configure a team review in Kit.
Related articles
Ready to hire smarter?
Start free for 30 days. Cancel before it ends and you pay nothing. Set up your first hiring pipeline in minutes.
Start hiring free