Headquarters: South Jordan, United States
Assessed: Structured video interview and game-based assessment.
Rubric: version 1. Search protocol applied and all sources accessed: 2026-09-22.
Documented 5, partial 3, absent 0, of eight criteria. This records what is published about the validity and fairness of the scores. It is not an assessment of the product, the model or the underlying science.
| # | Criterion | Level |
|---|---|---|
| C1 | Technical manual or validity report publicly available without a sales conversation | D |
| C2 | Validity coefficients reported with sample size and uncertainty | D |
| C3 | Adverse impact reported by named subgroup | D |
| C4 | Scored construct named and defined, not described only by outcome | D |
| C5 | Evidence of comparability across languages, versions or delivery modes | P |
| C6 | Human oversight documented, with a candidate route to contest a score | P |
| C7 | Independent review named and the reviewer identified | D |
| C8 | Reliability reported, and reported by group | P |
C1. Technical manual or validity report publicly available without a sales conversation
Documented.
"HIREVUE’S ASSESSMENT SCIENCE", "WHITE PAPER OCTOBER / 2021", authored and attributed: "Dr. Kiki Leutner: Director, Assessment Innovation; Dr. Josh Liff: Director, Assessment Psychometrics and Applied Research; Dr. Lindsey Zuloaga: Chief Data Scientist; Dr. Nathan Mondragon: Chief I/O Psychologist".
Source: https://www.hirevue.com/wp-content/uploads/2023/01/2021_10_HireVue_Assessment_Science_white_paper-FINAL.pdf (accessed 2026-09-22).
A 21-page dated document, downloadable without a form or a login, with named authors and a reference list.
C2. Validity coefficients reported with sample size and uncertainty
Documented.
"Interview assessments have predictive validities of r = .25 to r = .49 with various job-related outcomes"; a table gives per-study values with "INITIAL SAMPLE SIZE" (710, 404, 696, 380, 53,194, 6,345), "AUC VALUE" and "CORRELATION COEFFICIENT" with significance flags ("**Multiple R is based on out-of-sample model performance using stratified k-fold cross-validation").
Source: https://www.hirevue.com/wp-content/uploads/2023/01/2021_10_HireVue_Assessment_Science_white_paper-FINAL.pdf (accessed 2026-09-22).
Coefficients are reported with their criteria, per-study sample sizes and an out-of-sample estimation method. Confidence intervals are not reported.
C3. Adverse impact reported by named subgroup
Documented.
A worked table by named group: "Service Orientation: Black (n=412) 52.2%, White (n=582) 47.9%, Asian (n=108) 47.2%, Hispanic (n=961) 50.7%", with columns "PROTECTED GROUP", "PASSING RATE", "ADVERSE IMPACT RATIO" (0.92, 0.90, 0.97), "COHEN’S H", "FISHER’S EXACT TEST", "CHI-SQUARED TEST" and "STATISTICAL EVIDENCE OF ADVERSE IMPACT?".
Source: https://www.hirevue.com/wp-content/uploads/2023/01/2021_10_HireVue_Assessment_Science_white_paper-FINAL.pdf (accessed 2026-09-22).
Named groups, per-group sample size, impact ratios, an effect size and two significance tests.
C4. Scored construct named and defined, not described only by outcome
Documented.
"HireVue assessments are psychometric assessments that measure traits and competencies associated with performance at work"; competencies are named and illustrated at item level, for example "a question designed to measure Dependability is: ‘Tell me about a time you had a challenge keeping…’".
Source: https://www.hirevue.com/wp-content/uploads/2023/01/2021_10_HireVue_Assessment_Science_white_paper-FINAL.pdf (accessed 2026-09-22).
Constructs are named, mapped to the reported score and illustrated at item level. Convergent validity against a published reference instrument is also reported.
C5. Evidence of comparability across languages, versions or delivery modes
Partial.
"To ensure that candidates are only compared to others in their country or region, local norming groups are implemented. We localize our assessment content (instructions, test items, and interview questions) using experts in psychometric test translation. All models that rely on verbal behavior are language-specific."
Source: https://www.hirevue.com/wp-content/uploads/2023/01/2021_10_HireVue_Assessment_Science_white_paper-FINAL.pdf (accessed 2026-09-22).
The statement describes translation process, local norming and language-specific models. No invariance, differential item functioning or equating study is published.
C6. Human oversight documented, with a candidate route to contest a score
Partial.
"The scoring models are monitored and updated to ensure that there is no adverse impact or scoring anomalies present at potential cut scores"; assessments are "developed and monitored by an interdisciplinary team of Data Scientists and Industrial and Organizational (I/O) Psychologists".
Source: https://www.hirevue.com/wp-content/uploads/2023/01/2021_10_HireVue_Assessment_Science_white_paper-FINAL.pdf (accessed 2026-09-22).
Oversight of the scoring models is documented and the reviewing professions are named. No candidate-facing route to contest a score was found on the pages examined.
C7. Independent review named and the reviewer identified
Documented.
"The audit was conducted by distinguished professor and CEO of Landers Workforce Science LLC, Dr. Richard Landers." Scope: "Our science team met with Dr. Landers for 18 hours of meetings, he reviewed nearly 1,000 pages of documentation". Published "April 7th, 2021".
Source: https://www.hirevue.com/blog/hiring/independent-audit-affirms-the-scientific-foundation-of-hirevue-assessments (accessed 2026-09-22).
The full report is reached through a resource page rather than downloaded directly. The review is dated 2021.
C8. Reliability reported, and reported by group
Partial.
Convergent validity coefficients are tabulated per competency ("TABLE 2: Convergent Validity of Game and Interview Assessments") and the summary states the assessments "meet strict standards of reliability and validity".
Source: https://www.hirevue.com/wp-content/uploads/2023/01/2021_10_HireVue_Assessment_Science_white_paper-FINAL.pdf (accessed 2026-09-22).
No internal-consistency or test-retest coefficient with a sample size per scale was found, and no reliability breakout by group.
Absent records the outcome of the published search protocol on 2026-09-22. Where a source is supplied that the search did not reach, the cell is re-scored under the corrections policy, with the date of the change recorded and the previous level retained.
Back to the index, the rubric and the methods note