Research · no sales talk
We make a living teaching the Enneagram. That does not make this page neutral, but it makes it necessary: if we do not write down what is supported and what is not, someone else will – and they are right about more than most providers admit.
That is what it looks like when you read the literature rather than the marketing. The rest of this page is the documentation.
There is one systematic review of Enneagram research: Hook, Hall, Davis, Van Tongeren and Conner in the Journal of Clinical Psychology, 2021. They went through 104 independent samples. The key sentence of the abstract reads: «After reviewing 104 independent samples, we found mixed evidence of reliability and validity.»
One fact says almost everything about the maturity of the field: «Of the 104 independent samples in the present review, about half (i.e., 49) were published.» The rest were doctoral dissertations, master’s theses, conference presentations and unpublished manuscripts. The authors themselves write that the research «is still mostly relegated to unpublished dissertations or journals not indexed in PsycINFO, which may explain its relatively poor reputation among some psychologists».
Their advice to practitioners is equally plain: «Given that research is still in a preliminary stage, we suggest that clinicians proceed with caution. Enneagram theory remains largely untested.»
Reliability is where the Enneagram does best – and where the figures swing most.
| Instrument | Internal consistency | Test-retest |
|---|---|---|
| RHETI, original version | 0.35–0.84 depending on the study | 0.72–0.94 |
| RHETI, Likert version | Above 0.70 on all scales | – |
| WEPSS, later version | 0.78–0.88 in Wagner (1999), 0.85–0.93 in Thrasher (1994) | 0.62–0.91 after 6 weeks, 0.55–0.86 after 8 months |
| ETASI (Turkish, n = 3,531) | 0.67–0.87 | 0.29–0.51 after 4 weeks |
| OEPS (honours project, n = 1,039) | Mean 0.46 | 0.68 overall |
Note two things. The kindest WEPSS figures come from the instrument’s own manual, that is from a party with a commercial interest. And the ETASI figures of 0.29 to 0.51 after four weeks mean that a large share of participants did not get the same result again shortly afterwards.
The hardest single figure in the whole literature appears in the 2021 review: «42% of participants were classified into the same type by the RHETI and the Wagner Enneagram Personality Style Scales.» Two of the most widely used Enneagram tests agree on the type in fewer than half of cases.
That is also why we – and any serious Enneagram teacher – treat a test result as a qualified suggestion to be checked in conversation. The test is a way in, not a verdict.
«The RHETI is 72 per cent accurate.» The figure comes from a much-cited instrument directory which writes: «The internal consistency reliability scores show that the RHETI ranges from 56% to 82% accurate on various types; with an overall accuracy of 72%.» The values 0.56 to 0.82 are Cronbach’s alpha from the study by Newgent and colleagues in 2004 – a measure of whether the items in a scale hang together. They are not hit rates, and they cannot be converted into a percentage correct.
«The test is 95 per cent accurate.» The claim appears twice in marketing material for a commercial Enneagram instrument, alongside a testimonial about 100 per cent accuracy. There is no description of method, no sample size and no reference to a published study anywhere in the material.
We use neither, and you should be sceptical of any provider that does.
This is where the distance between marketing and documentation is greatest. The short version:
| Claim | Status |
|---|---|
| The symbol is 2,500 years old and Babylonian | Not documented. The origins are claimed, not evidenced |
| The Enneagram is an ancient Sufi teaching | Not documented from Sufi sources |
| G. I. Gurdjieff brought the symbol forward | Documented. He is the only known source for the figure |
| Gurdjieff attached the nine types to the symbol | No. He used the figure for other purposes, among them sacred dances |
| Oscar Ichazo formulated the nine types | Documented, among other places in court records from 1992 |
| Claudio Naranjo translated them into Western psychology from 1970 | Documented |
One detail is worth knowing, because it shows how easily the ancient-origin claim spreads: even the 2021 systematic review repeats it in its introduction, citing an article that is not a historical primary source. Appearing in academic literature does not make a claim documented. It merely makes it repeated.
The model is therefore about as old as the DISC instruments and younger than the MBTI. That does not make it worse. But we say it as it is.
Personality is probably not categorical. In 2012 Haslam, Holland and Kuppens reviewed all published taxometric research – 177 articles, 311 findings, 533,377 participants – and concluded that «most latent variables of interest to psychiatrists and personality and clinical psychologists are dimensional». Once methodological weaknesses were controlled for, the share of findings pointing to genuine categories fell to 14 per cent. Normal personality was not among them. That hits every typology equally hard, the MBTI and DISC included.
An expert panel has placed the Enneagram among discredited methods. In 2006 Norcross, Koocher and Garofalo asked a panel of 101 psychologists how discredited a range of methods were, on a scale from 1 to 5. The Enneagram for personality assessment scored 3.37 in the first round and 4.14 in the second. An important caveat is rarely mentioned: the same table shows that 65.9 per cent of the panel in the second round said they were not familiar enough with the Enneagram to judge it. The 4.14 therefore rests on about a third of the panel. It is a survey of professional opinion, not a psychometric analysis – but it is the source a sceptical occupational psychologist is most likely to reach for.
Recognition is not measurement. A frequently cited figure is that 87 per cent recognised themselves in the description of their test result. It is used as an argument for the Enneagram. It is just as good an argument against: high self-recognition in kindly worded type descriptions is exactly what the Barnum effect predicts. Set it beside the fact that two instruments agree in only 42 per cent of cases, and it becomes hard to claim the recognition comes from measurement.
Wings and lines are largely untested. The 2021 review says it directly: «there is little research supporting secondary aspects of Enneagram theory, such as wings and intertype movement». Two studies tested movement between types, and neither supported the hypothesis. We still teach it, because it makes sense to participants, but we do not call it documented.
Three things can be said with reasonable confidence.
1. The scales measure something real. Across nine studies, Enneagram scales correlate with the five-factor model in the direction the theory predicts. Type 1 was positively related to conscientiousness in 9 of 9 studies, type 5 negatively to extraversion in 9 of 9, type 8 negatively to agreeableness in 9 of 9. That is not random noise.
2. The types differ on work-relevant measures. In 2013 Sutton, Allinson and Williams studied 416 working adults and found that all nine types differed significantly on five-factor traits, on eight of ten personal values and on all three implicit motives. The types predicted job involvement and job self-efficacy on a par with values and motives.
The same study is also a warning. When the researchers tried to predict people’s Enneagram type from the rest of their profile, they were right in 32 to 36 per cent of cases. Chance would give about 11 per cent, so that is roughly three times better – but almost two thirds were still placed incorrectly.
3. The programmes work. The 2021 review summarises: «Enneagram interventions resulted in positive changes in some variables, such as leadership versatility, self-consciousness, communication, interpersonal relationships, team effectiveness, productivity, and staff turnover.»
Note what it says and what it does not. The programmes worked. That is not the same as the typology being true. A structured conversation about yourself with a competent teacher does good almost regardless of which model it hangs on. That is an honest objection to our own field, and it belongs here.
We have worked with the Enneagram in business since 2003, and we are behind the Enneagram test 360indicator. So here is what we think the model can and cannot be used for.
Use it to: give people a language for why they react as they do. Understand what happens to a colleague under pressure. Open a conversation in a team that otherwise does not get had. Work with leadership, where knowing people’s working style is not enough.
Do not use it to: choose between candidates. Put someone in a box they do not recognise. Explain a conflict away. Or tell anyone that their type is a fact – it is a model they can use if it fits.
Our own data says something about what people experience: among 6,330 completed full 360indicator profiles, between 84 and 92 per cent of participants confirmed the type the test pointed to when they saw the result. That is not a measure of validity, and we do not present it as one. It is a measure of recognition – with exactly the caveat the Barnum effect calls for.
As far as we know there is no Danish or Nordic peer-reviewed research on the reliability or validity of the Enneagram. That is a real gap, and we would like to help close it.
Read on: Enneagram or MBTI · Enneagram or DISC · Enneagram or the five-factor model · The nine types
Sources: Hook, Hall, Davis, Van Tongeren & Conner (2021), Journal of Clinical Psychology 77(4), 865–883. · Haslam, Holland & Kuppens (2012), Psychological Medicine 42, 903–920. · Norcross, Koocher & Garofalo (2006), «Discredited Psychological Treatments and Tests: A Delphi Poll», Professional Psychology: Research and Practice 37(5), 515–522. · Newgent, Parr, Newman & Higgins (2004), Measurement and Evaluation in Counseling and Development 36(4), n = 287. · Sutton, Allinson & Williams (2013), European Management Journal 31(3), n = 416. · Yanartaş, Malakcıoğlu, Acarkan & Akça (2022), Anatolian Clinic 27(3), n = 3,531. · Sutton (2012), The Enneagram Journal. · Cusack (2020), Aries 20(1). · Arica Institute, Inc. v. Palmer, 970 F.2d 1067 (2d Cir. 1992). · Carroll, The Skeptic’s Dictionary.
This page is written by Balder and Mia Vendt-Striim, who make a living teaching the Enneagram and are behind 360indicator. We have taken care to reproduce the sources as they stand, including where they speak against us. If you find an error, we would like to hear from you at info@enneagram.dk.
Tell us whether you want an open course, an in-house course or a talk, and we'll find the right solution and send a no-obligation quote.