Studies
The Natasha Demkina Test
Also Known As The 2004 Discovery Channel test
Citation Formats
General Reference
APA Style
BibTeX
The value of this episode lies in how badly a pre-agreed threshold survived contact with an ambiguous result. The investigators set the bar at five of seven precisely so that the outcome could not be argued about afterwards, and it was argued about afterwards anyway, at length and in public, by people who had agreed to it. Demkina's supporters pointed out that four of seven is unlikely by chance and that the sample was tiny. The investigators pointed out that a threshold accepted in advance is not renegotiable after the fact and that a small sample was the reason the bar had been set high. Both are right about their own point, which is why the test settled nothing.
Facts
Disputed
Test ResultFour correct matches out of seven, against a pre-agreed pass threshold of five 1 Ray Hyman, Richard Wiseman and Andrew Skolnick, for the Committee for Skeptical Inquiry and the Commission for Scientific Medicine and Mental Health, tested Natasha Demkina's claimed ability to diagnose medical conditions by sight in a protocol both the investigators and Demkina's representatives agreed to in writing beforehand. Demkina correctly matched four of seven test subjects to their diagnoses, a result Hyman's report Testing Natasha (2005) records as above chance expectation for the task but below the numerical threshold both sides had accepted beforehand as a pass. The method was a pre-registered blind matching protocol, fixing the passing score in advance specifically to prevent post hoc argument over what would count as success. Because the agreed threshold was not met even though the raw score exceeded chance, each side read the same documented outcome differently, Demkina's supporters pointing to the above-chance score and the investigators holding to the pre-agreed threshold, which is why the test result is recorded as debated rather than a clean pass or failure. Evidence
Replication StatusA single test by design. Both sides accepted the criteria beforehand and then disagreed about what the outcome meant, since four of seven is above chance expectation for the task and below the threshold everyone had signed up to. No further formal test followed. 1 The study's replication record as reported in the literature. Where replication failed or is disputed, that is stated rather than softened. Method
InvestigatorsRay Hyman, Richard Wiseman and Andrew Skolnick, for the Committee for Skeptical Inquiry and the Commission for Scientific Medicine and Mental Health 1 Taken from the study's own record of who conducted it. Learn More
Designing a Test Both Sides Sign First
Natasha Demkina, a teenager from Saransk, had become known in Russia and then in the British press for saying she could see inside people's bodies and identify what was wrong with them. The investigators designed a test intended to be decisive in a single sitting, which is difficult, and the design shows the effort. Seven volunteers were recruited, each with a documented condition that would be dramatic to anyone who could genuinely see inside a body: one had had a large organ removed, one had a metal plate in the skull, one had a surgically reshaped chest. Cards naming the seven conditions were laid out, the volunteers sat in a row, and she was asked to assign each condition to a person. Chance alone would give roughly one correct.
The threshold for passing to a fuller study was fixed in advance at five, a number chosen so that a pass would be hard to attribute to luck. All parties agreed to it before she began. She was given as long as she wanted and took several hours. Two further features of the design are worth recording. The conditions were chosen to be structurally distinct rather than merely different diagnoses, so that a genuine ability to see inside a body would not require fine discrimination: a missing organ and an intact one are not a subtle contrast. And the volunteers were seated together rather than examined one at a time, so that the task was assignment across a fixed set rather than a series of independent judgements, which is what allowed a chance expectation to be calculated at all.
Four of Seven, and an Argument Nobody Won
She matched four. That is above what chance predicts and below the threshold everyone had agreed, and the argument that followed was long, public and bad-tempered on both sides. Her supporters said four of seven is unlikely enough to warrant the fuller study, that the sample was too small for a five-of-seven bar to be fair, and that the televised setting was hostile. The investigators said a threshold accepted in writing beforehand cannot be renegotiated once the number is known, and that the bar had been set high precisely because the sample was small. Both arguments are sound within their own terms, which is what makes the episode instructive.
The methodological lesson is that a threshold protects against post-hoc reinterpretation only if the parties genuinely accept its authority, and acceptance is easy to give in advance and hard to honour afterwards. It is worth adding, since this concerns a living person who was seventeen at the time, that the test established nothing about her honesty or her beliefs, and that nobody involved alleged she was cheating. No further formal test was arranged. A test with a small sample and a high bar can only ever produce a pass that means something and a failure that means very little, and the investigators accepted that asymmetry when they designed it.
Cross-Tradition Connections
Associated With
2004, Years The Natasha Demkina Test, a controlled test, is dated here from 2004. Whether a young Russian woman who said she could see inside bodies could match seven specific medical conditions to the seven people who had them.
The sole subject of the test, which she agreed to with the pass threshold fixed in advance.
Conducted By
One of the three investigators who designed the protocol and the pass threshold. Documented in his own published account.
One of the three investigators. Documented in the published accounts of the test.
Product Of
Designed and run by investigators associated with the committee and filmed for television. The televised format was itself criticised by both sides.
Sources
Reader Challenges (0 open reader challenges)
No disputes yet. Spotted an error or a better source? Open the first one.
Sign in to dispute this or suggest a correction.
View At A Past Year
The atlas records no dated fact of its own for this entry, so there is no other year to choose.