Sign UpLogin With Facebook
Sign UpLogin With Google

Personality Tests for Hiring: What the Evidence Actually Says

What these tests predict, which ones hold up, what the law requires in 2026, and how to run one without wrecking your funnel

Author: Michael Hodge
Published: 19th August 2026

The short answer

A well built personality test is a useful tie-breaker, not a filter. The best current meta-analysis puts work-framed conscientiousness at an operational validity of .25 for predicting job performance, against .42 for a structured interview. Personality explains real variance, but roughly a third as much as the interview you should already be running.

Use it if you want a cheap, low-bias signal that tells you what to ask about in the interview. Do not use it to auto-reject, do not use MBTI for selection (the publisher says so itself), and do not buy a test that has never published a validity study on a job like yours.

Somewhere between two thirds and four fifths of employers now put candidates through some kind of pre-hire assessment. One widely cited 2018 survey of HR professionals found 79% used testing for external hires and 72% for internal moves, though that figure covers every test type, from typing speed to cognitive ability, not personality alone. Clean prevalence data specific to personality testing is genuinely scarce, and most numbers circulating online trace back to vendor marketing pages rather than published research.

What is not scarce is evidence about whether these tests work. The field was substantially rebuilt in 2022, and much of what you will read elsewhere still quotes figures that industrial and organisational psychologists have since abandoned. This guide uses the current estimates, says plainly where the evidence is thin, and gives you a runbook you can execute this quarter.

What personality tests actually predict

For twenty-five years the reference table in personnel selection was Schmidt and Hunter (1998). In 2022, Paul Sackett and colleagues published a re-analysis in the Journal of Applied Psychology showing that the standard corrections for range restriction had been applied in a way that systematically inflated almost every number in that table. Their corrected estimates are lower nearly across the board, and they reorder the field: cognitive ability testing is no longer the top predictor.

Here is the current table. Operational validity is the correlation between the predictor and job performance, corrected for measurement error in the performance rating and for range restriction. Higher is better. Around .10 is weak, .30 is respectable, and .40 is strong for a single predictor.

Operational validity for predicting job performance. Source: Sackett, Zhang, Berry and Lievens (2022), as tabulated in Sackett et al. (2023). Blue rows are personality measures.
Selection methodSchmidt & Hunter 1998Sackett et al. 2022Black / White d
Structured interview.51.42.23
Job knowledge test.48.40.54
Biodata, empirically keyed.35.38.33
Work sample test.54.33.67
Cognitive ability (GMA).51.31.79
Integrity test.41.31.10
Personality-based emotional intelligencen/a.30.22
Assessment centre.37.29.52
Situational judgement testn/a.26.34 to .39
Conscientiousness, work-framedn/a.25-.07
Vocational interests.10.24.33
Emotional stability, work-framedn/a.23.09
Extraversion, work-framedn/a.21.16
Conscientiousness, generic.31.21-.07
Unstructured interview.38.19.32
Agreeableness, work-framedn/a.19.03
Openness, work-framedn/a.12.01
Extraversion, genericn/a.11.16
Agreeableness, genericn/a.10.03
Emotional stability, genericn/a.09.09
Years of job experience.18.07.49
Openness, genericn/a.06.10

Three things jump out of that table, and they should drive every decision you make about assessment.

  • Personality is mid-table, not top. The strongest single personality predictor lands at .25, well behind a structured interview at .42. If you are choosing between spending budget on an assessment vendor or on training your panel to run structured interviews, train the panel first.
  • Framing changes the answer more than the trait does. Conscientiousness measured with work-specific wording scores .25. The same trait measured generically scores .21, and emotional stability jumps from .09 to .23 on that change alone. See the next section, because this is the cheapest improvement available to you.
  • Personality is the cleanest predictor on subgroup differences. The last column is Cohen's d between Black and White applicants. Cognitive ability sits at .79 and work samples at .67, both large. Conscientiousness sits at -.07, effectively zero and marginally favouring Black applicants. That is the strongest practical argument for personality testing, covered in full further down.
.10 .20 .30 .40 Structured interview.42 Job knowledge test.40 Biodata, empirically keyed.38 Work sample.33 Cognitive ability.31 Integrity test.31 Assessment centre.29 Situational judgement test.26 Conscientiousness, at work.25 Vocational interests.24 Emotional stability, at work.23 Conscientiousness, generic.21 Unstructured interview.19 Openness, at work.12 Emotional stability, generic.09 Openness, generic.06 Blue bars are personality measures. Grey bars are other selection methods.
Operational validity for predicting job performance, Sackett et al. (2022). Personality sits in the middle of the pack, and how you word the items matters as much as which trait you measure.
Be careful with the raw numbers

Operational validity is a corrected figure, and the uncorrected correlation is lower. Watrin and colleagues (2023) ran a deliberately conservative meta-analysis of conscientiousness and found a raw r of .17 across 102 samples and 23,305 people, dropping to .14 in real applicant samples. They also found that only about 12% of studies in this literature used actual applicants in a predictive design, which is the only design that matches how you would really use the test. Treat .25 as an optimistic ceiling and .14 as a pessimistic floor.

What a validity of .25 buys you in practice

Correlations are hard to feel, so translate them. Run the standard Taylor-Russell calculation: if half your current hires work out, and you start selecting the top 25% of applicants on a predictor with a validity of .25, roughly 63% of the new cohort works out. That is a 13 point improvement, which is worth having, especially at volume. Run the same calculation on a structured interview at .42 and you get about 72%. The interview is not marginally better than the test, it is close to twice the gain.

Those figures also assume a clean process. Both numbers get swamped by a broken interview or a job advert that attracts the wrong people, so fix those first.

The honest framing is this. A personality test is worth adding when it is cheap, when it does not lengthen your funnel much, when it feeds your interview rather than replacing it, and when your alternative is guessing. It is not worth adding if it costs you good candidates through drop-off, or if it hands a hiring manager a scientific-looking reason to act on a hunch they already had.

The biggest lever: ask about work, not about life

If you take one practical thing from this page, take this. A personality item that says "I am organised" is measuring something different from an item that says "at work, I keep my commitments organised." The second one predicts job performance roughly twice as well.

Shaffer and Postlethwaite (2012) meta-analysed the published and unpublished studies comparing generic and work-contextualised personality measures. Generic measures averaged a validity of .11. Work-framed measures averaged .24, roughly double. A decade later the Sackett table above reproduces the same pattern trait by trait, from an entirely separate analysis.

TraitGeneric wordingWork-framed wordingChange
Emotional stability.09.23+.14
Extraversion.11.21+.10
Agreeableness.10.19+.09
Openness.06.12+.06
Conscientiousness.21.25+.04

There is a second, quieter finding buried in that table. The residual standard deviation for work-framed conscientiousness is .00, meaning its validity barely varies from job to job. Generic conscientiousness has an SD of .15, so in some jobs it predicts well and in others it predicts nothing at all. Work framing does not just raise the average, it makes the measure dependable.

Why does this work? Most likely because people genuinely behave differently in different contexts. Someone can be chaotic at home and meticulous at work, and only one of those is relevant to you. Asking about the wrong context adds noise, and noise is exactly what a validity coefficient punishes.

  • What to do with this: when you evaluate a vendor, ask to see actual item wording. If the items read like a magazine quiz about your personality in general, you are buying the .11 version. If they reference work situations, colleagues, deadlines and customers, you are buying the .24 version.
  • If you are writing items yourself: embed the context in every single item, not in one instruction at the top of the form saying "answer these thinking about work." Per-item framing is what the research actually tested.
  • Do not confuse this with a situational judgement test. An SJT presents a scenario and asks what you would do. A contextualised personality item still asks about a typical tendency, just anchored to the workplace. They are different instruments with different validities, .26 for SJTs.

Which personality test should you use for hiring?

The question people usually ask is which test is best. The better question is which tests are defensible for selection at all, because the popular ones and the defensible ones are almost entirely different lists.

Common instruments and whether they belong in a selection decision.
InstrumentWhat it measuresUse for hiring?Notes
Big Five / Five FactorConscientiousness, emotional stability, extraversion, agreeableness and openness on continuous scales Yes The research standard. Every validity number on this page comes from Big Five style measurement. Use a work-contextualised version.
HEXACOThe Big Five plus honesty-humility Yes The sixth factor overlaps with integrity testing and adds signal for roles with theft, fraud or safety exposure.
Integrity testsCounterproductive work behaviour, rule-following, reliability Yes Validity of .31 and a Black/White d of just .10. One of the best value-for-risk instruments available, though overt versions are transparent enough to fake.
Hogan (HPI, HDS, MVPI)Bright-side traits, derailers, values With care Built for selection and well validated, but expensive and usually needs certified interpretation. Ask for the technical manual, not the brochure.
Predictive Index, Caliper and similar vendor toolsProprietary trait models, often Big Five adjacent With care Quality varies enormously. Insist on a published technical manual with validity evidence for jobs like yours and adverse impact data before you sign.
DISCDominance, influence, steadiness and conscientiousness as a behavioural style No Designed as a communication and team-development tool. It sorts people into styles rather than measuring degrees, and there is little peer-reviewed evidence it predicts job performance.
MBTISixteen types across four preference pairs No The Myers-Briggs Company states that "it is unethical to require job applicants to take the assessment if the results will be used to screen out applicants." Its own reliability page notes about half of people get a different four-letter type on retest.
EnneagramNine motivational types No No credible criterion validity evidence for selection. Fine as a coaching conversation starter, nothing more.
CliftonStrengthsRanked talent themes No Ipsative by design, so scores are comparable within a person but not between candidates. That makes it structurally unsuited to ranking applicants.
Colour-based tools and free online quizzesVaries No No technical manual, no norms, no adverse impact data. Using one in a hiring decision is an unforced legal error.
The most common mistake on this list

MBTI is the most widely used personality instrument in corporate life and one of the least appropriate for hiring, and the publisher agrees. Their position is not a technicality. The instrument was built for self-discovery, it uses transparent items a candidate can easily game, it has no validity scale to catch that, and it forces continuous traits into binary categories, which throws away information precisely at the boundary where your borderline candidates sit. If your organisation already runs MBTI for team development, keep it there and keep it out of the funnel.

How to choose between the defensible options

  1. Ask for the technical manual. Not a case study, not an ROI calculator. A document with reliability coefficients, norm groups and criterion validity studies. If the vendor cannot produce one, that is your answer.
  2. Check the validity evidence covers a job like yours. A validity study on call-centre agents tells you very little about hiring a field engineer.
  3. Ask for adverse impact data by race, sex and age. A vendor selling into the United States in 2026 who has never run this analysis is selling you their legal risk along with their product.
  4. Check the item wording is work-contextualised, for the reasons set out above.
  5. Check how long it takes. Anything over 20 minutes early in the funnel will cost you candidates, and the people who drop out are disproportionately the ones with other offers.
  6. Check what the report tells a hiring manager. A percentile with no interpretation gets misused. A report that says "this candidate scores low on structure, so probe how they managed the last project that had no process" gets used well.

If you would rather build a work-framed Big Five assessment with role-specific trait targets than buy an off-the-shelf one, SeeMyPersonality's hiring assessment builder lets you set target ranges per role before candidates respond, then reports each candidate against those targets alongside interview questions to probe the gaps. Building an assessment and sending invitations is free, which makes it a low-commitment way to trial the approach before signing a per-seat vendor contract.

Can candidates fake a personality test?

Yes, obviously, and they do. The interesting question is whether it matters.

Birkeland and colleagues (2006) meta-analysed 33 studies comparing real job applicants with people taking the same tests with nothing at stake. Applicants scored higher on the traits that look job-relevant.

TraitApplicants vs non-applicants (d)Reading
Conscientiousness.45Nearly half a standard deviation of inflation
Emotional stability.44Same magnitude
Openness.13Small
Extraversion.11Small on average, considerably larger for sales roles

Two nuances change what you should do about this. First, the inflation is largest on exactly the traits you most want to measure. Second, and more usefully, candidates fake toward what they believe the job wants. Birkeland found the rank order of trait inflation shifted substantially for sales roles, where extraversion inflation rose sharply. Faking is not random noise, it is targeted noise, which means the more obvious your target profile is, the more it gets gamed.

The counter-intuitive part is that faking does less damage to validity than you would expect. Because most candidates inflate, and inflate by similar amounts, the rank ordering is disturbed less than the absolute scores are. That is cold comfort at the top of the distribution, where the people who fake hardest are the ones who reach your shortlist.

What actually reduces faking

  • Forced-choice formats. Instead of rating "I stay calm under pressure" on a scale, the candidate picks which of several equally desirable statements is most like them. Martinez and Salgado's 2021 meta-analysis found quasi-ipsative forced-choice formats cut the conscientiousness faking effect to d = .49 against d = 1.27 for fully ipsative formats, and found faking is consistently smaller in real selection settings than in lab studies where people are instructed to fake.
  • A warning that responses are checked. Telling candidates that inconsistent or implausible response patterns are flagged reliably reduces inflation, and it costs nothing.
  • Response-consistency and social desirability indices. Use them to flag a profile for a closer look, never to auto-reject. False positives here are people with genuinely consistent personalities.
  • Verification in the interview. If someone scores at the 95th percentile on structure and discipline, ask them to walk you through how they planned their last major project. Faking survives a questionnaire far better than it survives a follow-up question demanding specifics.
  • Not publishing your target profile. If the job advert says you are looking for highly detail-oriented self-starters, you have handed out the answer key.
The practical position

Assume every score you see is inflated by roughly half a standard deviation on conscientiousness and emotional stability, and treat scores as relative rather than absolute. Comparing candidates against each other within the same requisition is defensible. Comparing a candidate's raw score against a general population norm is not, because your applicants are all inflating and the norm group was not.

Bias, adverse impact, and why personality often looks good here

Adverse impact is conventionally measured with the four-fifths rule: if the selection rate for any protected group is below 80% of the rate for the highest-scoring group, that is prima facie evidence of impact and you need a job-relatedness justification. The relevant question is not whether your test is biased in the abstract, but whether it produces different pass rates in your funnel.

On this dimension, personality testing is the strongest tool available. Return to that last column of the main table.

MethodValidityBlack / White dWhat that combination means
Cognitive ability.31.79Good prediction, large subgroup gap. The classic validity and diversity tradeoff.
Work sample.33.67Widely assumed to be fair. The data says otherwise.
Job knowledge test.40.54Strong prediction, substantial gap.
Structured interview.42.23Best prediction, small gap. The reason it is the first thing to invest in.
Integrity test.31.10Respectable prediction, negligible gap.
Conscientiousness.25-.07Modest prediction, no gap. Adds signal without adding impact.

That is the real case for personality assessment, and it is a better case than the one most vendors make. Personality is not the strongest predictor. What it is, is the predictor that adds signal without adding adverse impact, which makes it genuinely useful as a supplement to methods that do carry impact.

There are caveats worth holding onto.

  • Small average gaps do not guarantee small gaps in your funnel. Meta-analytic d values describe applicant populations in aggregate. Run your own four-fifths analysis on your own data, by race, sex, age and disability status where you lawfully hold it.
  • Disability is the blind spot. Race and sex gaps on personality measures are small, but items about sociability, energy or emotional steadiness can systematically disadvantage candidates with anxiety, autism or ADHD. That is both an ADA exposure and a fairness problem, and it is far less studied than race and sex.
  • Cut scores create impact that the raw scale does not. A trait with a d of -.07 can still produce a failing four-fifths ratio once you impose a hard threshold, particularly a high one on a small applicant pool. Test the rule you actually apply, not the instrument in isolation.
  • Combining predictors changes the arithmetic. Adding a cognitive test to a personality test does not average the two impact figures. Model the composite you actually use.

How to run a personality test in hiring, properly

Most failures are process failures rather than instrument failures. This is the sequence that survives both a validity check and a legal one.

  1. Start with a job analysis, not a test

    Write down what the person will actually do, then which behaviours separate strong performers from weak ones in that specific role. Interview two or three of your current strong performers and their managers. If you cannot articulate why conscientiousness matters for this job, you have no business measuring it. This step is also the documentation that establishes job-relatedness if anyone ever asks.

  2. Define target ranges, not maximums

    More is not always better. Very high agreeableness can be a liability in procurement or collections. Very high conscientiousness can slow down a role that needs fast, imperfect iteration. Set a target band per trait and treat both ends as worth a conversation. Bands also make the output far more useful to a hiring manager than a raw percentile.

  3. Put it after the screen, before the panel

    Too early and you lose candidates to drop-off, and you burn assessment budget on people who will not clear a phone screen. Too late and it cannot inform the interview, which is its main job. The sweet spot is normally straight after the recruiter screen, so results are in the panel's hands before they write their questions.

  4. Never use it as a hard filter on its own

    A validity of .25 does not support auto-rejection. Use the result as an input to a human decision and as a source of interview probes. If your applicant tracking system can auto-reject on a score, turn that off. This is also the single behaviour that most reliably turns an assessment into a legal problem.

  5. Turn every notable score into an interview question

    This is where the value actually is. A candidate scoring low on structure is not disqualified, they are a prompt: "tell me about a project with no process in place, what did you put in place and what did you deliberately leave alone." Write the probes in advance, ask them of every candidate whose profile triggers them, and score the answers against a defined scale. That converts a .25 predictor into an input to a .42 predictor.

  6. Keep the interview structured

    Same questions, same order, defined rating anchors, independent scoring before discussion. Every number on this page says the structured interview is your best tool. The personality result should sharpen it, never replace it. If your panel debriefs by talking first and scoring afterwards, the loudest voice sets the outcome. Collect independent scores privately first, for example with an anonymous team voting poll, and only then open the discussion.

  7. Monitor adverse impact from day one

    Log every score, every decision, and the demographics you lawfully hold. Run the four-fifths analysis quarterly rather than when a complaint arrives. If a group falls below 80%, you need either a documented job-relatedness case or a change to the process.

  8. Validate locally once you have the data

    After 12 to 18 months you should have enough hires to correlate pre-hire scores against performance ratings or objective output. Do it. Local validity evidence is worth more than any vendor's meta-analysis, both for improving the process and for defending it. If the correlation is near zero for your roles, drop the test. That is a legitimate and useful outcome.

Trait targets by role

Conscientiousness predicts performance across essentially every job, which is why it dominates the literature. The other four traits are role-dependent and the effects are modest, so treat this as a starting hypothesis to test against your own data rather than as a rulebook.

Role typeTraits that tend to matterWhat to probe in the interview
Sales, new businessConscientiousness (achievement facet), extraversion, emotional stabilityRejection recovery, pipeline discipline across a full quarter, how they handle a lost deal. Expect the heaviest faking here, so verify hard.
Customer supportAgreeableness, emotional stability, conscientiousnessHandling a hostile customer, staying accurate at volume, when they escalate rather than persist.
Operations, admin, complianceConscientiousness (order and dutifulness facets)Error-catching habits, how they handle a process that is clearly wrong but mandated.
Engineering and technicalConscientiousness, openness. Extraversion largely irrelevant.Debugging persistence, how they choose between the clean fix and the quick one. Weight the work sample far above the personality result.
Creative and researchOpenness, and lower agreeableness can helpWhere they have pushed back on a consensus and what happened. Note that openness is the weakest predictor in the whole table, so hold it loosely.
People managementEmotional stability, conscientiousness, moderate agreeablenessHandling a poor performer, delivering unwelcome news, a decision they made that their team disliked.
Safety-critical and cash-handlingConscientiousness, integrity, honesty-humilityRule-following under time pressure, a time they reported a problem that made them look bad.
A warning about facets

Broad traits hide useful detail. Conscientiousness contains achievement striving, orderliness, dutifulness and self-discipline, and these can pull in different directions for a given job. A salesperson usually needs high achievement striving and can survive middling orderliness. A compliance officer is the reverse. If your instrument only reports five numbers, you are losing information a facet-level report would give you. That is one of the better reasons to pay for a proper instrument rather than use a free one.

What the questions actually look like

People searching for sample pre-employment personality test questions usually want one of two things: to know what they will face as a candidate, or to sanity-check a vendor's item quality. Here is what each format looks like, using work-contextualised wording.

Likert, contextualised

The standard format

A statement plus an agreement scale, usually five or seven points. Simple, fast, and the easiest to fake.

  • At work, I finish tasks well before the deadline
  • I stay calm when a project changes direction late
  • I prefer to check my work twice before submitting it
  • I find it easy to raise a concern with a senior colleague

Rated from strongly disagree to strongly agree.

Forced choice

The faking-resistant format

Statements matched for social desirability. The candidate picks most and least like them, so there is no obviously correct answer.

  • I plan my week in detail in advance
  • I build relationships quickly with new colleagues
  • I stay level-headed when priorities change
  • I look for better ways to do routine tasks

Pick one as most like you and one as least like you.

Frequency based

The behavioural format

Asks how often something actually happened rather than what the candidate is like. Harder to inflate because it invites specifics.

  • In the last month, how often did you miss an internal deadline?
  • In the last three months, how often did you volunteer for work outside your remit?
  • How often do you re-read an email before sending it?

Rated on a frequency scale rather than an agreement scale.

Avoid entirely

Items that create legal risk

These appear in older instruments and in cheap online quizzes. Any one of them can turn your assessment into a medical inquiry or a privacy claim.

  • Anything about mood disorders, treatment or medication
  • Anything about religious belief or practice
  • Anything about sexual orientation or sexual history
  • Anything about alcohol or drug use history
  • Anything about family, pregnancy or caring responsibilities

Scale design matters more than people expect. A five-point agreement scale and a five-point frequency scale are not interchangeable, and neither is a scale you can safely average without thinking about it. If you are building your own instrument, our guide to nominal, ordinal, interval and ratio scales covers which arithmetic each scale type actually supports.

Where personality testing in hiring goes wrong

This is the section people are looking for when they search for the problem with using personality tests for hiring. The problems are real, and every one of them is avoidable.

Failure 1

Using it as a filter

A predictor with a validity of .25 rejecting candidates on its own will discard good people at scale. It is also the fastest route to a disparate-impact claim, because a hard cut-off manufactures adverse impact that the underlying scale does not have.

Failure 2

Hiring for culture fit

Fit in practice usually means similarity to the incumbents, which is how a team converges on one personality type and loses the friction that catches mistakes. Hire against the role's requirements, not against the team's average profile.

Failure 3

Using a development tool for selection

MBTI, DISC, Enneagram and CliftonStrengths were all built to help people understand themselves. None was designed to rank strangers, and using them that way is indefensible if challenged.

Failure 4

Believing the number over the evidence

A precise-looking percentile carries unearned authority. Hiring managers routinely weight a 73rd percentile score above two hours of interview evidence. Report bands and interview prompts instead of raw numbers and the problem largely disappears.

Failure 5

Never checking it works

Most employers never correlate pre-hire scores with actual performance. Without that, you cannot tell a useful instrument from an expensive ritual, and you have no evidence to defend the process.

Failure 6

Ignoring the candidate experience

A 45-minute unexplained assessment before a human conversation reads as disrespect. Candidates with options simply leave, which biases your funnel toward the least in-demand applicants.

Failure 7

Testing traits the job does not need

Measuring extraversion for a solo research role adds noise and creates impact for no return. If a trait is not in your job analysis, do not score it.

Failure 8

Letting the vendor set the cut score

Vendor defaults are built on their norm group, not your applicant pool or your role. A default threshold is a guess about your business made by someone who has never seen it.

What to use instead, or alongside

If you have limited budget and attention, spend it in this order.

  1. A structured interview, validity .42. The highest-validity method available, with a small subgroup gap, and it costs process discipline rather than money. Same questions, same order, anchored rating scales, independent scoring. If you do nothing else on this page, do this.
  2. A job knowledge test at .40 or a work sample at .33, where the role has concrete, testable content. Both carry substantial adverse impact, so pair them with monitoring.
  3. Empirically keyed biodata at .38. Underused and strong, though it needs enough historical data to key against your own outcomes, which puts it out of reach for small employers.
  4. An integrity test at .31, with a subgroup d of .10. Excellent value wherever reliability, safety or cash handling matter.
  5. A work-contextualised personality measure at .25. Cheap, fast, near-zero adverse impact, and it makes your interview better. A good fifth thing to add, and a poor first thing.

One note on stacking. Berry, Lievens, Zhang and Sackett rebuilt the selection-method correlation matrix on the 2022 estimates and ran a dominance analysis with six methods in a single model. Conscientiousness accounted for only about 3% of the predictable variance once structured interviews, biodata, integrity tests, cognitive ability and situational judgement tests were in play, and its regression weight actually turned slightly negative. Personality overlaps heavily with the others, particularly biodata. So the incremental value of adding a personality test is largest when your existing process is thin, and smallest when you already run a rigorous multi-method battery.

If you only change one thing

Write down five behavioural questions per role, define what a 1, a 3 and a 5 answer sounds like for each, and make every interviewer score independently before the debrief. That is free, it takes an afternoon, and on the current evidence it outperforms every assessment product on the market.

Measuring whether it actually worked

An assessment you never evaluate is a cost centre with good branding. Four feedback loops will tell you whether yours is earning its place, and none of them needs a data science team.

1. The local validity study

Once you have 12 to 18 months of hires, correlate pre-hire trait scores against a performance measure. Manager ratings work, objective output works better where it exists. You want at least 50 hires in comparable roles before the number means much. A correlation near zero is a real finding: it means the test is not earning its place in that role.

2. Adverse impact monitoring

A quarterly four-fifths analysis on the assessment stage specifically, not just on the overall funnel. Impact often hides at a single stage while the end-to-end numbers look acceptable.

3. Candidate experience

Ask people who completed the assessment, and people who abandoned it, what the experience was like: length, clarity, relevance, and whether the questions felt fair. Two or three questions sent immediately afterwards will get you a usable response rate, and you can build that with a free form builder in a few minutes. Watch for nonresponse bias here, because the people most annoyed by your process are also the least likely to answer a survey about it.

4. Hiring manager confidence and 90-day outcomes

Ask hiring managers at 90 days whether the assessment told them anything they did not learn in the interview. If the honest answer is consistently no, the test is decoration. Pair that with a new-hire check-in on whether the role matched what was described. Our employee engagement pulse questions cover the onboarding period, and the survey question library has ready-made wording for both sides of that loop.

  • Keep the loop anonymous where you can. Candidates who want the job and new hires still in probation both have obvious incentives to be positive. Anonymity buys you honesty.
  • Keep it short. Three questions answered by 60% of people beats fifteen questions answered by 8%.
  • Close the loop out loud. Tell hiring managers what you changed as a result. It is the only thing that keeps response rates up next quarter.

If you are the candidate taking one

A large share of searches around this topic come from candidates rather than employers, usually asking how to pass. Here is the honest version.

  • There is no answer key, but there is a direction. Almost every employer scores conscientiousness and emotional stability positively. Answering as your best professional self rather than your most candid private self is normal and expected, and the research shows nearly everyone does it.
  • Do not invent a different person. Wildly inconsistent answers get flagged by response-consistency checks, and even if they do not, you will have to sustain the story through an interview and then a job. Optimising your way into a role that suits someone else is a bad outcome for you.
  • Read the job advert for the target profile. If it emphasises attention to detail and process, the assessment will weight orderliness. If it emphasises pace and autonomy, it will weight achievement striving.
  • Answer as you are at work, not at home. Well-built tests ask about the workplace explicitly. If yours does not, answer as your working self anyway, because that is the behaviour being predicted.
  • Do not overthink the timed items. Timed personality questions measure consistency, not knowledge. First instinct is usually fine.
  • You can ask for an accommodation. If a disability affects your ability to take the test in its standard form, you are entitled to request an adjustment, and requesting one is not a mark against you.
  • Practice tests are of limited use. There is no content to learn. Taking one once to remove the surprise is worthwhile, and beyond that it is wasted effort.

If you are curious what a properly built work-focused profile actually says about you, SeeMyPersonality shows the same Big Five report structure employers see, including the fit bands and the interview questions a hiring manager would be handed.

Frequently asked questions

The questions people most often ask about using personality tests in hiring, answered against the evidence set out above.

Do personality tests actually work for hiring?
Partly. Work-framed conscientiousness predicts job performance at an operational validity of .25, which is real but modest, and below a structured interview at .42. Personality tests work best as a supplement that sharpens your interview, not as a standalone screening decision. A conservative 2023 meta-analysis put the uncorrected correlation as low as .14 in real applicant samples, so treat vendor claims of dramatic accuracy with suspicion.
Are personality tests legal for hiring?
Yes in the United States, with conditions. They must not function as a medical examination before a conditional offer under the ADA, they must be job related if they screen out a protected group at a materially lower rate, and in New York City, California and Illinois there are additional notice, audit or record-keeping obligations. Clinical instruments such as the MMPI have been held to be medical examinations, so keep those out of your pre-offer process.
Should personality tests be used for hiring at all?
Use one if it is cheap, short, work-contextualised, feeds your interview rather than filtering candidates, and you monitor its impact. Skip it if it would be your first investment in the process, if you would auto-reject on the score, or if you cannot articulate which traits the job needs and why. The strongest argument in favour is not accuracy, it is that personality adds signal with almost no adverse impact.
Which personality test is most frequently used for hiring decisions?
MBTI and DISC are the most widely used personality instruments in corporate life, and neither is appropriate for selection. Among defensible tools, Big Five based instruments dominate, including Hogan, while Predictive Index and Caliper are common in mid-market recruiting. The most commonly used is not the most appropriate, and that gap is the central problem in this market.
What is the best personality test for hiring?
A work-contextualised Big Five measure with a published technical manual, norms relevant to your applicant pool, and adverse impact data. Brand matters less than those four things. Ask any vendor for reliability coefficients, criterion validity studies on jobs like yours, subgroup difference data and sample item wording. A vendor who cannot supply all four has answered your question.
Can you use MBTI for hiring?
No. The Myers-Briggs Company states that "it is unethical to require job applicants to take the assessment if the results will be used to screen out applicants." Their reasoning is that the items are transparent and easy to fake, there is no validity scale to catch it, and the instrument measures preferences rather than competence. Their own reliability data shows roughly half of people receive a different four-letter type on retest.
Can candidates fake or cheat a personality test?
Yes. Real applicants score about 0.45 standard deviations higher on conscientiousness and 0.44 higher on emotional stability than people with nothing at stake. Because nearly everyone inflates, rank ordering suffers less than absolute scores do. Forced-choice formats reduce the effect substantially, a warning that responses are checked helps, and verifying high scores with specific interview questions is the most reliable defence.
Are personality tests discriminatory?
Less so than most alternatives on race and sex. Conscientiousness shows a Black to White standardised difference of -.07, against .79 for cognitive ability and .67 for work samples. The genuine risk is disability: items about sociability, energy or emotional steadiness can disadvantage candidates with anxiety, autism or ADHD, which is both an ADA exposure and a fairness problem. Hard cut scores can also create adverse impact that the underlying scale does not have.
What is the four-fifths rule?
If the selection rate for any protected group is below 80% of the rate for the highest-scoring group, that is treated as prima facie evidence of adverse impact and you need to show the tool is job related and consistent with business necessity. Run it on the assessment stage specifically, because impact often hides at one stage while the end-to-end funnel looks acceptable.
Did the 2025 executive order make disparate impact go away?
No. Federal enforcement retreated: an April 2025 executive order directed agencies to deprioritise disparate-impact liability, the DOJ concluded the EEOC guidelines including the Uniform Guidelines are unconstitutional, and in June 2026 the EEOC adopted an enforcement plan dropping disparate-impact theories. But the theory remains codified in Title VII, private plaintiffs can still sue in federal court, no court has adopted the DOJ reading, and numerous states provide their own route. Keep validating.
When in the hiring process should the test happen?
Straight after the recruiter screen and before the interview panel. Earlier costs you candidates through drop-off and wastes assessment budget on people who will not clear a screen. Later means results arrive too late to shape the interview, which is where most of the value is.
How long should a hiring personality assessment be?
Ten to twenty minutes. Long enough for adequate reliability, short enough that people finish. Beyond about 20 minutes early in the funnel, drop-off rises sharply and it falls hardest on candidates who have other offers, which biases your pool in exactly the wrong direction.
Should I set a cut score?
Preferably not for personality. A validity of .25 does not support a hard threshold, and cut scores manufacture adverse impact. Use bands that trigger interview probes instead. If your process genuinely requires a threshold, set it from your own applicant data, document the job-relatedness rationale, and test the four-fifths ratio at that exact threshold before you apply it.
How many companies use personality tests for hiring?
Nobody has a reliable figure. A widely cited 2018 HR survey found 79% used some form of testing for external hires, but that covers every test type from typing speed to cognitive ability. Most percentages circulating online trace back to vendor marketing rather than published research. Adoption is clearly common among larger employers and lower among small ones, and beyond that the honest answer is that the data is poor.
Are there free personality tests suitable for hiring?
Free consumer quizzes are not suitable, because they lack technical manuals, norms and adverse impact data. Some legitimate platforms let you build and send a work-focused assessment at no cost and charge for detailed reporting at volume, which is a reasonable way to trial the approach. The test to apply is not price, it is whether the provider can show you reliability, validity and subgroup data.
What is a contextualised or frame-of-reference personality measure?
One where every item is anchored to the workplace, for example "at work, I finish tasks before the deadline" rather than "I finish tasks before the deadline." In meta-analysis, generic measures averaged a validity of .11 and work-framed ones .24. The context must sit in the items themselves, not in a single instruction at the top of the form.
Is DISC valid for hiring?
No. DISC was built as a communication and team-development framework. It sorts people into behavioural styles rather than measuring traits by degree, and there is little peer-reviewed evidence that DISC profiles predict job performance. Keep it for team workshops.
What about cognitive ability tests instead?
Cognitive ability predicts performance at .31, roughly on par with integrity tests and below structured interviews and job knowledge tests. It also carries the largest subgroup difference of any common method, at .79. It remains a legitimate tool, but it is no longer the automatic first choice it was under the older evidence base, and the adverse impact cost is high.
Do we need a bias audit?
In New York City, yes, if the tool substantially assists a hiring decision: an annual independent bias audit, published publicly, plus candidate notice. California sets anti-bias testing expectations and four-year retention of automated-decision data. Illinois requires notice but no formal audit. If you hire across jurisdictions, build to the strictest standard that applies to you.
What should the report give a hiring manager?
Bands rather than raw percentiles, a plain-language summary of working style, the two or three areas worth probing, and specific interview questions for each. A bare percentile invites over-interpretation. A report that ends in questions turns a .25 predictor into an input to a .42 one.
Can we use personality tests for promotions and internal moves?
Yes, and the same rules apply. Karraker v. Rent-A-Center was a promotion case, not a hiring case. Internal use carries the added risk that a low score follows someone around their career, so keep results out of the general HR record and delete them on a documented retention schedule.
How do we validate the test on our own data?
Wait 12 to 18 months, then correlate pre-hire trait scores against performance ratings or objective output for at least 50 hires in comparable roles. Manager ratings work, objective output works better. A near-zero correlation means the test is not earning its place in that role, and dropping it is the correct response.
What are the main downsides?
Modest predictive power, meaningful faking, disability-related risk, over-interpretation by hiring managers who trust a number over two hours of evidence, candidate drop-off from long assessments, and a tendency to encourage hiring for similarity to the existing team. All are manageable, and all get worse if the test is used as a filter.
What is the single highest-return change we could make?
Structure the interview. Same questions in the same order, defined rating anchors, and independent scoring before any group discussion. It is the highest-validity method in the table at .42, it has a small subgroup gap at .23, and it costs process discipline rather than budget.

Methods and sources

Validity figures on this page are operational validities, meaning correlations corrected for measurement error in the performance criterion and for range restriction. They are not raw correlations, which are lower. Subgroup differences are Cohen's d between Black and White applicants, where a positive value favours White applicants. Where the evidence is weak or contested, this page says so rather than picking a convenient number.

  1. Sackett, P. R., Zhang, C., Berry, C. M., and Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040 to 2068. PubMed record. The source of the main validity table.
  2. Sackett, P. R., Zhang, C., Berry, C. M., and Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology, 16(3). Cambridge Core. Contains the side-by-side comparison with Schmidt and Hunter (1998) and the subgroup d values reproduced above.
  3. Shaffer, J. A., and Postlethwaite, B. E. (2012). A matter of context: a meta-analytic investigation of the relative validity of contextualized and noncontextualized personality measures. Personnel Psychology, 65, 445 to 494. Wiley. The source of the .11 versus .24 comparison between generic and work-framed measures.
  4. Berry, C. M., Lievens, F., Zhang, C., and Sackett, P. R. Insights from an updated personnel selection meta-analytic matrix: revisiting general mental ability tests' role in the validity-diversity tradeoff. Working paper. The source of the dominance analysis showing conscientiousness contributing about 3% of predictable variance alongside five other methods.
  5. Birkeland, S. A., Manson, T. M., Kisamore, J. L., Brannick, M. T., and Smith, M. A. (2006). A meta-analytic investigation of job applicant faking on personality measures. International Journal of Selection and Assessment, 14(4). Open access record. The source of the applicant faking effect sizes.
  6. Martinez, A., and Salgado, J. F. (2021). A meta-analysis of the faking resistance of forced-choice personality inventories. Frontiers in Psychology, 12. Frontiers in Psychology. The source of the forced-choice comparison.
  7. Watrin, L., Weihrauch, L., and Wilhelm, O. (2023). The criterion-related validity of conscientiousness in personnel selection: a meta-analytic reality check. International Journal of Selection and Assessment. Wiley. The conservative uncorrected estimate and the point about applicant predictive designs.
  8. The Myers-Briggs Company. MBTI facts, including the official position on selection and the retest consistency figure. Publisher statement.
  9. Karraker v. Rent-A-Center, Inc., 411 F.3d 831 (7th Cir. 2005). Holding that MMPI-based testing constituted a medical examination under the ADA.
  10. California Civil Rights Council, automated-decision systems employment regulations, effective 1 October 2025. Final text of the regulations.
  11. New York City Local Law 144 of 2021, automated employment decision tools, enforced from 5 July 2023. NYC rules.
  12. United States Department of Justice, Office of Legal Counsel opinion on the EEOC disparate-impact guidelines, together with the June 2026 EEOC National Enforcement Plan. DOJ announcement.
  13. American Psychological Association, Speaking of Psychology, with Fred Oswald: can a personality test determine if you are a good fit for a job. APA podcast page. The source of the 2018 testing prevalence figures.

Nothing on this page is legal advice. Employment testing law varies by state and country and has changed materially since 2025. Have counsel review any assessment before it goes into a live hiring process.