Science backed methodology
Every claim, with its receipt.
Prismona exists because the most popular personality tests have weak retest reliability and poor predictive validity, while the scientifically defensible ones are enterprise-priced and consultant-gated. This page is the entire method, in the open.
Instruments
- Big Five — quick
- Mini-IPIP (Donnellan et al., 2006): 20 items, four per domain. A validated screening tier — we say so on every quick-tier result.
- Big Five — full
- IPIP-NEO-120 (Johnson, 2014): 120 items, 30 facets, developed on 619,150 protocols. Facet resolution is where equal domain scores stop hiding different people — Orderliness and Industriousness are both “Conscientiousness,” and decisive for different things.
- Big Five — standard
- A facet-balanced selection from the IPIP-NEO-120: one item from each of the six facets per domain (36 scored items, α ≈ .80), round-robin interleaved so no two consecutive items probe the same domain — content breadth instead of repetition, by design, against respondent fatigue. Two instructed attention checks (Meade & Craig, 2012) are embedded and excluded from scoring.
- Honesty-Humility
- IPIP HEXACO markers (Ashton, Lee & Goldberg, 2007): 6 items. The H factor is the strongest known trait predictor of workplace deviance (ρ ≈ −.48; Pletzer et al., 2019) — the trust layer most instruments omit.
- Interests
- O*NET Mini Interest Profiler (Rounds et al.): 30 items, five per RIASEC scale (Holland, 1997), scale α ≈ .70–.75, r = .95–.96 with the 60-item Short Form. Scored ipsatively to a Holland code; interests supply career direction, traits the performance estimate.
- Licensing
- All personality items are public domain via the International Personality Item Pool (Goldberg et al., 2006); interest items are public domain via the U.S. Department of Labor O*NET program. No proprietary instrument is imitated or licensed.
Scoring, exactly
Responses are five-point Likert. Reverse-keyed items are reflected (6 − x). Each scale is the mean of its answered items; unanswered items are never imputed. Scale means are standardized against provisional adult norms — z = (m − μ)/σ — and expressed as percentiles via the normal CDF. Emotional Stability is reversed Neuroticism throughout.
Uncertainty bands. Every score carries ±1 standard error of measurement, SEM = √(1 − α) in z units, using published internal consistencies: quick domains α ≈ .70, full domains α ≈ .88, facets α ≈ .72, Honesty-Humility α ≈ .76. A point score without error is a small lie; no consumer competitor draws the band.
Norms are provisional — approximate values from IPIP community samples — and labeled so on every screen. They will be re-estimated from our own user base at scale, and we will publish the revision.
Stability facets. The six facets under Emotional Stability are the IPIP Neuroticism facets, reversed and renamed for one-direction reading: Composure = Anxiety (reversed); Even Temper = Anger (reversed); Buoyancy = Depression (reversed); Self-Assurance = Self-Consciousness (reversed); Moderation = Immoderation (reversed); Resilience = Vulnerability (reversed).
The twenty-second clock
Three reasons, in honesty order. First-instinct responses reduce impression management. Latency profiles help flag careless or faked protocols (Fine & Pirak, 2016; Meade & Craig, 2012) — though the literature also shows limits to latency-based detection (Röhner & Thoss, 2022), so timing informs a confidence indicator, never an accusation. And a visible clock keeps pace, which protects completion without rushing anyone: twenty seconds is generous. Timeouts simply record the item as unanswered.
Archetypes, disciplined
Large datasets show density clusters in trait space (Gerlach et al., 2018) — and a sharp critique shows such types may be neither robust nor exhaustive (Freudenstein et al., 2019). We honor both findings: eight narrative archetypes are matched by distance in six-trait z-space and always reported as a gradient blend over your dimensional blueprint. You are the percentages, not the label. This is the explicit antithesis of type-first instruments, whose bimodality the evidence does not support.
The dyad engine
Compatibility scoring is purpose-specific because the evidence is. For romance, actor and partner effects dominate: a partner's emotional stability, agreeableness and conscientiousness predict the other's satisfaction (Malouff et al., 2010; Dyrenforth et al., 2010), while raw similarity adds little — we say so rather than selling resemblance. For founding teams, complementary trait spreads predict venture success (McCarthy et al., 2023; Bell, 2007); we flag dual-low conscientiousness and any low Honesty-Humility pairing as gates, and forecast conflict style from the agreeableness-by-stability interaction.
The output is always a gauge plus the top frictions, each with a structured conversation — never a binary verdict. The deepest reason is Joel et al. (2020): across 43 longitudinal couples studies, relationship-specific perceptions outpredict individual traits. An instrument that promised more would be lying; one that structures the right conversation is useful.
Share codes carry six quantized trait scores, a date, and a checksum — twenty-one characters, no identity, no answers. Comparisons run entirely in the browser.
Privacy & ethics
Personality data is sensitive, so the architecture is the policy: no account, no server-side scoring, no analytics on answers, nothing transmitted. Blueprints live in your browser's local storage and die with it. The one exception is explicit and opt-in: the results page offers an anonymous norms contribution — the share-code payload, an optional coarse age band, and a network-level country code, no IP stored — so percentiles can eventually rest on our own published norms and trait distributions can be compared across countries in aggregate (details in the Privacy Policy, §V). We will never sell blueprints or assess anyone without their consent — inferred blueprints of third parties are an ethical boundary, not a roadmap item. Any future hiring module ships only after criterion validation and adverse-impact analysis under the AERA/APA/NCME Standards.
Citations
- Donnellan, M. B., Oswald, F. L., Baird, B. M., & Lucas, R. E. (2006). The Mini-IPIP scales: Tiny-yet-effective measures of the Big Five. Psychological Assessment, 18(2), 192–203. link
- Johnson, J. A. (2014). Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: Development of the IPIP-NEO-120. Journal of Research in Personality, 51, 78–89. link
- Ashton, M. C., Lee, K., & Goldberg, L. R. (2007). The IPIP–HEXACO scales: An alternative, public-domain measure of the personality constructs in the HEXACO model. Personality and Individual Differences, 42(8), 1515–1526. link
- Goldberg, L. R., et al. (2006). The International Personality Item Pool and the future of public-domain personality measures. Journal of Research in Personality, 40(1), 84–96. link
- Roberts, B. W., Kuncel, N. R., Shiner, R., Caspi, A., & Goldberg, L. R. (2007). The power of personality. Perspectives on Psychological Science, 2(4), 313–345. link
- Soto, C. J. (2019). How replicable are links between personality traits and consequential life outcomes? Psychological Science, 30(5), 711–727. link
- Gerlach, M., Farb, B., Revelle, W., & Amaral, L. A. N. (2018). A robust data-driven approach identifies four personality types across four large data sets. Nature Human Behaviour, 2, 735–742. link
- Freudenstein, J.-P., et al. (2019). Four personality types may be neither robust nor exhaustive. Nature Human Behaviour, 3, 1045–1046. link
- Malouff, J. M., Thorsteinsson, E. B., Schutte, N. S., Bhullar, N., & Rooke, S. E. (2010). The Five-Factor Model of personality and relationship satisfaction of intimate partners: A meta-analysis. Journal of Research in Personality, 44(1), 124–127. link
- Dyrenforth, P. S., Kashy, D. A., Donnellan, M. B., & Lucas, R. E. (2010). Predicting relationship and life satisfaction from personality in nationally representative samples. Journal of Personality and Social Psychology, 99(4), 690–702. link
- Joel, S., et al. (2020). Machine learning uncovers the most robust self-report predictors of relationship quality across 43 longitudinal couples studies. PNAS, 117(32), 19061–19071. link
- McCarthy, P. X., et al. (2023). The impact of founder personalities on startup success. Scientific Reports, 13, 17200. link
- Bell, S. T. (2007). Deep-level composition variables as predictors of team performance: A meta-analysis. Journal of Applied Psychology, 92(3), 595–615. link
- Pletzer, J. L., et al. (2019). Comparing domain- and facet-level relations of the HEXACO personality model with workplace deviance: A meta-analysis. Personality and Individual Differences, 152, 109539. link
- Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1–26. link
- Judge, T. A., Bono, J. E., Ilies, R., & Gerhardt, M. W. (2002). Personality and leadership: A qualitative and quantitative review. Journal of Applied Psychology, 87(4), 765–780. link
- Higgins, E. T. (1987). Self-discrepancy: A theory relating self and affect. Psychological Review, 94(3), 319–340. link
- Markus, H., & Nurius, P. (1986). Possible selves. American Psychologist, 41(9), 954–969. link
- Hudson, N. W., & Fraley, R. C. (2015). Volitional personality trait change: Can people choose to change their personality traits? Journal of Personality and Social Psychology, 109(3), 490–507. link
- Holland, J. L. (1997). Making Vocational Choices: A Theory of Vocational Personalities and Work Environments (3rd ed.). Psychological Assessment Resources. link
- Rounds, J., Wee, C. J. M., Cao, M., Song, C., & Lewis, P. Development of an O*NET Mini Interest Profiler (Mini-IP) for Mobile Devices: Psychometric Characteristics. National Center for O*NET Development. link
- Fine, S., & Pirak, M. (2016). Faking fast and slow: Within-person response time latencies for measuring faking in personnel testing. Journal of Business and Psychology, 31, 51–64. link
- Meade, A. W., & Craig, S. B. (2012). Identifying careless responses in survey data. Psychological Methods, 17(3), 437–455. link
- Röhner, J., & Thoss, P. (2022). Challenging response latencies in faking detection. link
- Wasserman, N. (2012). The Founder's Dilemmas. Princeton University Press. link
The full 65-paper annotated bibliography lives in the project repository. Ready to be measured? Begin the Test.