Every instrument makes a claim: its score means what it says, and means the same thing for everyone it is used on. Most instruments were never tested against that claim, and a translation, a new site or a new population quietly breaks it. The audit tests the claim and tells you, in writing, what your scores can and cannot support. It takes instruments written in English; the same audit for Montenegrin and regional-language instruments runs from psihologija.me.
The pre-screen reads a draft before the pilot. The audit diagnoses an instrument you already run. The development builds the one you do not have. Start wherever you are.
Two tiers, chosen by what you have
The desk review needs only the questionnaire. The full audit adds your pilot or field data. Both are fixed in scope, price and turnaround, and both end in a verdict page you can hand to a board, an ethics committee or a funder.
Tier 1. Instrument Review
A structured expert review of the instrument as it stands. No data needed.
You send
- The instrument as administered: items, instructions, response options, scoring key
- Your construct definition, or the document that stands in for one
- Any translated versions, with whatever translation record exists
You receive
- A scored item table with the reason for every call: keep, revise or cut
- A construct-coverage, response-scale and scoring verdict
- A translation equivalence-risk table for each language version
- A prioritised fix-list in three severity bands
- A signed verdict page and a thirty-minute findings call
1,800 EUR, ten working days from receipt of the instrument. One instrument, up to 60 items or eight subscales, up to three language versions.
Tier 2. Full Psychometric Audit
Tier 1 plus a quantitative pass on your own pilot or field data.
You send
- Everything in Tier 1
- One data file with one row per respondent, and a codebook
- A grouping variable, where you want groups or languages compared
You receive
- Item analysis, reliability (alpha and omega) and factor structure against the claimed model
- The measurement-invariance cascade across groups or languages, with partial invariance sought where a level fails
- A technical report to APA seventh edition, with every table ready to paste into your own reporting
- The scripts that reproduce every number
- The fix-list, a signed verdict page and a sixty-minute findings call
5,900 EUR, four weeks from receipt of clean data. Up to 60 items, 5,000 respondents and four groups on one grouping variable, one measurement occasion.
Prices exclude VAT; Montenegrin VAT of 21 percent is added where it applies.
Where an audit stops and a study begins
The ceiling is the product. Beyond it, the work is a validation study and is quoted as one. Outside both tiers, and quoted separately: new data collection or a pilot; norms, reference distributions, cut scores or standard setting; criterion or predictive validity against an external outcome; longitudinal or multi-occasion invariance; rewriting the instrument or authoring new items. The fix-list in either tier is written so that you can commission exactly the follow-on you need, or none. Where the verdict is that the instrument should be rewritten rather than repaired, the rewrite is a priced product: Instrument Development, which starts from the audit’s fix-list at 7,500 EUR ex VAT. Where what needs reviewing is a manuscript rather than an instrument, the entry product is the Pre-Submission Methods Review, a signed methods referee report before you submit, from 550 EUR ex VAT.
More than one expert’s opinion
The best current evidence on structured quality review is a 2026 preprint, not yet peer reviewed, in which minimally trained raters assessed the methodological rigour of 52 and then 110 psychology papers against written criteria (Etzel et al., 2026, PsyArXiv 4w7rb). Overall rigour scores agreed well across raters, and typical recent papers met under a tenth of the criteria. Two limits, stated plainly: the paper rates papers, not instruments, and agreement was strong for checkable criteria and weak for judgement calls. The audit is built to that finding. Every call in the report is tied to a written criterion and a quoted item so a second psychometrician can check it, and a two-auditor agreement study is running on the first audits.
A public worked example: a measurement audit of a multilingual language-model benchmark, asking how much of its language ordering is actually estimable, with the data, code and every results table open. Open the deposit on Zenodo.
Sample reports
- Tier 1 sample: desk review of the Rosenberg Self-Esteem Scale (PDF, 5 pages)
- Tier 2 sample: technical appendix, three-language invariance run (PDF, 5 pages)
Both samples are demonstrations on public-domain or synthetic material. A client receives the same documents on their own instrument and data.
Start an audit
Tell us which tier and describe or paste the instrument. We reply within two working days with a start date and the written scope; a data file is only ever sent once that scope is agreed.
Prefer to attach the file, or the form gives you trouble? Send it to milos@centerforpsychology.me.