EV-0002
The strongest challenge — integrity test validity is substantially lower under stricter criteria
Van Iddekinge, C. H., Roth, P. L., Raymark, P. H., & Odle-Dusseau, H. N. (2012). Journal of Applied Psychology, 97(3) — updated meta-analysis, with the authors' reply to · Source
What this entitles us to say
A meta-analysis applying stricter inclusion criteria to the same literature found integrity tests predicted job performance far more weakly than previously reported — around ρ = .15 in applicant samples and ρ = .20 in incumbent samples.
What it found
104 studies representing 134 independent samples. Applying stricter inclusion criteria than the 1993 analysis, validity for job performance fell to roughly ρ = .15 among applicants and ρ = .20 among incumbents — substantially below the earlier estimates in EV-0001.
The methodological objection is as important as the number: on review, only about 30% of the primary studies used in the earlier meta-analysis met the newer inclusion criteria, and a central concern was the weight previously given to unpublished studies authored by the firms selling the tests.
The paper produced a published exchange — comments from the original authors and from test publishers, and a reply from Van Iddekinge and colleagues on research questions, inclusion criteria and transparency.
Why it matters here
This entry exists because an institute that only cites supportive findings is not an institute. A sophisticated buyer's advisor will already know this literature is contested, and being the ones who raised it is a considerably stronger position than being the ones who omitted it.
It also disciplines our own claims. If the best-funded, most-studied integrity instruments in the world are still arguing about their validity, we have no business implying our free 37-question self-assessment has settled the matter.
The limits
This is itself contested. Test publishers and the original meta-analysts both published comments defending the earlier findings, and Van Iddekinge et al. replied in turn. The debate is genuinely unresolved, and an honest reading is that the true value sits somewhere between the two camps — not that either side has been refuted. Note also that a weaker validity for hiring screens says nothing directly about whether character development works.