Headquarters: Amsterdam, Netherlands
Assessed: Multi-test screening library covering cognitive ability, personality, skills and role-specific tests.
Rubric: version 1. Search protocol applied and all sources accessed: 2026-09-22.
Documented 3, partial 5, absent 0, of eight criteria. This records what is published about the validity and fairness of the scores. It is not an assessment of the product, the model or the underlying science.
| # | Criterion | Level |
|---|---|---|
| C1 | Technical manual or validity report publicly available without a sales conversation | D |
| C2 | Validity coefficients reported with sample size and uncertainty | P |
| C3 | Adverse impact reported by named subgroup | P |
| C4 | Scored construct named and defined, not described only by outcome | P |
| C5 | Evidence of comparability across languages, versions or delivery modes | P |
| C6 | Human oversight documented, with a candidate route to contest a score | D |
| C7 | Independent review named and the reviewer identified | P |
| C8 | Reliability reported, and reported by group | D |
C1. Technical manual or validity report publicly available without a sales conversation
Documented.
A per-test fact sheet is published on the open web with no account: "The results are published in each test’s fact sheet, so you can read the numbers rather than take our word for them", and the fact sheets themselves are viewable without login.
Source: https://www.testgorilla.com/test-library/cognitive-ability-tests/numerical-reasoning/facts/ (accessed 2026-09-22).
Fact sheets are published per test and are viewable without an account.
C2. Validity coefficients reported with sample size and uncertainty
Partial.
"Candidates with higher scores on this test received higher average ratings from the hiring team during the selection process (r = .16, N = 885)."
Source: https://www.testgorilla.com/test-library/cognitive-ability-tests/numerical-reasoning/facts/ (accessed 2026-09-22).
The coefficient, the criterion and the sample size are reported per test. No confidence interval or standard error is reported.
C3. Adverse impact reported by named subgroup
Partial.
Gender differences are reported with an outcome status of "Acceptable"; age and ethnicity differences are marked "Pending".
Source: https://www.testgorilla.com/test-library/cognitive-ability-tests/numerical-reasoning/facts/ (accessed 2026-09-22).
The result is reported as a category rather than as impact ratios or selection rates. Age and ethnicity are marked as pending.
C4. Scored construct named and defined, not described only by outcome
Partial.
"We ground that definition in established frameworks like the US Department of Labor’s O*NET skills database or the European Commission’s ESCO framework."
Source: https://www.testgorilla.com/assessments-technology/ (accessed 2026-09-22).
The definitional source is named. Individual constructs are not defined on the public pages.
C5. Evidence of comparability across languages, versions or delivery modes
Partial.
"Available languages: English, Dutch, French, German, Spanish, Portuguese (Brazil), Japanese, Italian"
Source: https://www.testgorilla.com/test-library/cognitive-ability-tests/numerical-reasoning/facts/ (accessed 2026-09-22).
Language versions are named per test. No invariance, differential item functioning or equating evidence is published across them.
C6. Human oversight documented, with a candidate route to contest a score
Documented.
"You see what was evaluated, how the response measured against each criterion, and why it got the score it did. Agree with it, adjust it, or override it."
Source: https://www.testgorilla.com/assessments-technology/ (accessed 2026-09-22).
The route described is the reviewer route. No candidate-facing route is described.
C7. Independent review named and the reviewer identified
Partial.
"the content is checked for accuracy, technical correctness, and alignment by experts independent of whoever built it"
Source: https://www.testgorilla.com/assessments-technology/ (accessed 2026-09-22).
Independence is described as internal to the review process and no reviewer is named.
C8. Reliability reported, and reported by group
Documented.
"Cronbach’s alpha coefficient = .66", published per test with the sample size band stated on the fact sheet.
Source: https://www.testgorilla.com/test-library/cognitive-ability-tests/numerical-reasoning/facts/ (accessed 2026-09-22).
The coefficient is published per test with its sample size. It is not broken out by group or by language.
Absent records the outcome of the published search protocol on 2026-09-22. Where a source is supplied that the search did not reach, the cell is re-scored under the corrections policy, with the date of the change recorded and the previous level retained.
Back to the index, the rubric and the methods note