What is the scientific value of administering the standard Rorschach test to LLMs when the training data is almost certainly contaminated? (R) + [D]
Our take
The recent study published in *JMIR Mental Health* by Csigó and Cserey (2026) raises intriguing questions about the intersection of psychology and artificial intelligence, particularly in the context of how we evaluate the cognitive capabilities of large language models (LLMs). By administering the standard Rorschach test to multimodal LLMs such as GPT-4o, Grok 3, and Gemini 2.0, the researchers sought to analyze these models' responses through a framework designed primarily for human subjects. However, the methodological concerns surrounding this approach warrant a closer examination, especially given the implications for both psychological assessment and AI development.
One of the most significant issues highlighted in the article is the potential for data contamination. The Rorschach inkblots and the extensive literature surrounding them are widely available online, making it highly probable that these models have been exposed to similar stimuli during their training. This raises the question: are we genuinely assessing the models' perceptual abilities, or merely their capacity to retrieve pre-existing knowledge? When we consider the nature of LLMs as sophisticated pattern matchers, it becomes clear that the study may be more indicative of the models' ability to regurgitate learned associations rather than their understanding of visual ambiguity. This concern is compounded by the study's use of a small sample size and lack of control conditions, which further diminishes the validity of the findings.
Moreover, the authors acknowledge the limitations of their study, admitting that the models likely encountered both the Rorschach images and the scoring concepts during training. This acknowledgment raises a critical point about the scientific value of applying traditional psychological assessments to AI systems. If the objective is to explore how AI processes visual stimuli, why not use novel or controlled images that could minimize the influence of training data? By doing so, researchers could better isolate the models' perceptual capabilities from their retrieval functions, ultimately leading to more meaningful insights into how AI interprets complex information.
The implications of this study extend beyond mere academic curiosity. As AI continues to evolve and integrate into various sectors, understanding the cognitive processes behind these models becomes increasingly vital. For instance, the insights gleaned from refining psychological assessments for LLMs could inform advancements in AI applications across fields such as mental health, education, and creative industries. This mirrors discussions in our recent article, "I Let CodeSpeak Take Over My Repository," where we explore how AI-native workflows can transform project management, highlighting the need for rigorous evaluation of AI capabilities.
As we look to the future, questions remain about the methodology employed in studies like this one. How can we ensure that research in AI and psychology is conducted with the rigor it demands? Moreover, as we continue to explore the depths of AI capabilities, how can we utilize innovative experimental designs to better understand the intricacies of model interpretation and perception? The answers to these questions will be critical in shaping the future of AI research and its applications, guiding us toward more robust and insightful evaluations of these powerful technologies.
A recent paper published in JMIR Mental Health (Csigó & Cserey, 2026) caught my attention. The researchers administered the 10 standard Rorschach inkblot cards to three multimodal LLMs (GPT-4o, Grok 3, Gemini 2.0) and coded their responses using the Exner Comprehensive System. They analyzed the models' "perceptual styles," determinants (like human movement vs. color), and human-related content themes.
However, I am seriously struggling to understand the methodological validity of this setup, and I’m curious what the scientific community thinks. My main concerns are:
Massive Data Contamination: The 10 standard Rorschach cards, along with decades of psychological literature, scoring manuals (like the Exner system), and typical human responses, are widely available on the internet. It is highly probable that this data is already embedded in the models' training weights.
Testing Retrieval, Not Perception: Because they used the standard, century-old inkblots instead of novel, AI-generated, or strictly controlled ambiguous images, aren't they just testing the models' ability to retrieve the most statistically probable lexical associations for those specific images from their training data?
Lack of Controls: As I understand according to the paper, the researchers used the public web interfaces with default settings (no API, no temperature control) and seemingly only ran the test once per model, generating a tiny sample size.
Ironically, the authors explicitly admit in their "Limitations" section that the models likely encountered the stimuli and scoring concepts during training, which could influence outputs independently of any image understanding. So, methodologically what is the actual scientific value of conducting projective psychological tests on LLMs without using novel stimuli to - at least try - rule out data contamination? What do you think, based of mechanisms of LLMs, does a study like this tell us anything meaningful about how AI processes visual ambiguity, or is it merely demonstrating advanced pattern matching and text completion based on widely known psychometric data? And - how do studies with such glaring methodological loopholes regarding LLM training data contamination make it through peer review in decent journals? Maybe I'm a little bit critical here, I just wanted to be a little provocative. Here is the study: https://mental.jmir.org/2026/1/e88186?fbclid=IwY2xjawRd27dleHRuA2FlbQIxMQBzcnRjBmFwcF9pZBAyMjIwMzkxNzg4MjAwODkyAAEe-wkKP6fKZRmAAuNvtN6BjknolIGcfTGu0-cLFs6CC49kZ1gcR6ccdcaRiWA_aem_7hHg5G96xjDZ-04YlSs1Ew
[link] [comments]
Read on the original site
Open the publisher's page for the full experience