College of Social and Behavioral Science
104 3d Models of the Pelvis as Offloading Working Memory and Increasing Visual Search Efficency: An Experiment in Virtual Reality and Interdisciplinary Approach to Psychology
Kyle Hudson Guttman and Alex Detrich
Abstract
Viewing and diagnosing abnormalities in volumetric scans, like MRIs and CTs, is an increasingly large part of a radiologist’s workload. Volumetric scans involve a radiologist scrolling through a “stack” of 2D images, individual “slices,” which put together represent a 3D anatomical structure. Viewers of a volumetric image never engage with a 3D model entirely but must contextualize where individual slices sit in space using what slices come before and after. Viewers are therefore presented with “pseudo-spatial” information: they are given enough information to understand slices spatially but must do so utilizing their own cognitive resources. We called this “mental translation:” the process of converting 2D slices to a 3D mental representation. In the current experiment, we tested whether enough cognitive load was being used for mental translation to impact visual search efficiency. We asked twenty-seven novice participants to identify fractures in three conditions and collected data on speed, accuracy, and confidence. In one condition, participants saw a 3D model of a pelvis, in the second condition participants saw a volumetric image, and in the third condition participants saw both the volumetric image and the 3D model. We found no effect between conditions on accuracy or confidence. However, we did find a significant difference in reaction time across conditions. Our findings do not reflect the current literature and thus our hypotheses were inconsistent with our data. We expect that this might be due to limitations of the study, specifically the low visual acuity of 3D models causing ambiguity for if certain features were fractures or not. We also did not collect trial-by-trial data which means we were unable to discuss possible tradeoffs between speed and accuracy in some conditions. Nonetheless, the current study presents a novel paradigm and was important in establishing what kinds of stimuli ought to be presented for optimal data collection.
3d Models Of The Pelvis Offloading Visual Working Memory
Introduction
Radiologists are tasked with locating abnormalities in images from hundreds of patients a day where the result of a miss or false alarm could have significant impacts on patient health and lifestyle or result in a malpractice lawsuit (Alexander et al., 2022). Radiology is difficult and puts a significant cognitive load on the viewer, with abnormalities being missed up to 30% of the time (Austin et al., 1992; Bird et al., 1992; Birkelo et al., 1947; Guiss & Kuenstler, 1960). However, radiologists’ expertise and deployment of visual search behaviors that lead to the best outcomes have been studied, usually because those behaviors in this expert group require less cognitive load (Morra et al., 2021). Specifically, global processing, with one model being holistic visual processing, has been found in expert radiologists (Kundel et al., 1978; Kundel et al., 1984; Kundel et al., 1991; Kundel et al., 2007; Swensson 1980).
Most of the research that has been done to understand radiologists’ visual search and expert strategies has involved 2-dimensional stimuli, like a radiograph. But as technology develops and more complicated image types are used in the medical field, the task of the radiologist becomes more complicated. Volumetric images, like MRIs, are one such example. Volumetric images consist of a “stack” of 2D images that can be moved through at will (see Figure 1, Alexander et al., 2022; Drew et al., 2021; Williams & Drew, 2019; Williams et al., 2021). Some volumetric scans can include upwards of 1000 images, providing much more information to the radiologist and placing a large computational burden on them for finding all salient targets (Williams & Drew, 2019). It is unclear if the paradigms used to test global processing in 2D images applies to volumetric images, nor is it clear which visual search strategies viewers use when performing global processing. (Williams et al., 2021) Our theoretical position is that global processing of volumetric images might be a result of a “mental translation.” Viewers create a 3D mental representation of the anatomical object using the many individual slices of the MRI. The current thesis tested fully virtual 3D models as a possible alternative to using volumetric images for visual search, given a 3D model should offload the spatial transformation that is currently needed for radiologists to understand sets of 2D images as 3D. With less bandwidth occupied by the mental translation, more cognitive resources can be spent on the task of visual search.

What is Global Processing? Global processing is a phenomenon where a viewer can detect targets even when shown an image for a fraction of a second (Williams et al., 2021). Global processing is found in experts because experts have developed mental representations of the typical image such that abnormalities “pop out” without requiring a systematic search of the entire image (Ivy et al., 2023; Morra et al., 2021; Williams et al., 2021). Studies have showcased this skill with radiologists detecting abnormalities within the first second of exposure (Donovan & Litchfield, 2013; Kundel et al., 2007; Kundel et al., 2008). Global processing has been modeled as a two-pathway approach: a first global pathway and a second, selective, local pathway. The viewer grasps a global understanding of the image after which why can focus on specific targets with a local search (Drew et al., 2013; Ivy et al., 2023; Williams et al., 2021).
Holistic visual processing (HVP) is a form of global processing and is the dominant model in the medical image perception literature (Reingold & Sheriden, 2011; Sheriden & Reingold 2017) but is not always supported by empirical research (Litchfield & Donovan, 2016). In the holistic visual processing model, experts hold schemas about the kinds of images they are looking at and thus are able to quickly detect abnormalities in the image. Their behaviors are typically measured by saccadic amplitude, the distance the pupil travels from one point of fixation on an image to another. Expert radiologists have been found to have a longer saccadic amplitude in the beginning of a search, followed by smaller saccadic amplitude later (Nodine & Kundle, 1987; Sheridan & Reingold, 2017). A higher saccadic amplitude allows for more time to collect visual data, especially through peripheral vision, which has been testing using a gaze-contingent viewing paradigm that locks to eye tracking and only allows users to see 5 degrees of visual information away from the pupil (Ivy, 2021; Ivy, 2023). The first long saccades of the search are indicative of the global pathway, where the expert relies on their mental representation of the kind of image they are looking at and where abnormalities are often found overall, whereas the small saccades are indicative of the local pathway that are needed to focus it on a potential abnormality. A more accurate model of visual search that understands both the global and local pathways that radiologists employ could then be used to optimize training and reduce errors in detection.
Although HVP is a dominant model for understanding search, research on HVP in volumetric images is not as comprehensive as in 2D radiographs (Sheridan & Reingold, 2017; Venjakob and Mello-Thoms, 2015; Williams & Drew, 2019). Some research has found evidence that the holistic visual processing model might not apply to volumetric images. Volumetric images are unique because of their extra spatial information. Although they are comprised of 2D images, the ability to scroll quickly through the stack of images provides the ability to contextualize individual slices spatially. Drew et al. (2013) characterized the difference in viewing behavior by distinguishing between “drillers” and “scanners” where drillers scroll faster through the stack and scanners spend more time on each individual slice. William et al (2021) worked with the unique qualities of volumetric images and developed measures of eye movement and scrolling behavior that in sum might capture HVP. However, despite the use of these new measure, like “depth passes,” there was no statistically significant evidence that radiologists deployed HVP when viewing volumetric scans. Typically, and in line with the HVP model, novices present with shorter saccadic amplitude and are slower to locate a target with diagnostic accuracy (Sheridan & Reingold, 2017). However, eye tracking of experts viewing volumetric images sometimes shows experts as having smaller saccadic amplitude than novices, but this could be because of a rapid scrolling through the stack of slices rather than holistically searching each 2D slice (Bertram et al., 2013; Drew et al., 2013, 2021; Sheridan & Reingold, 2017; Williams & Drew, 2019; Williams et al., 2021).
Despite the evidence that HVP does not apply to volumetric images, it still may be the case that the two-pathway approach of global and local search applies, even if it is not realized through larger saccadic amplitudes. Volumetric images are merely sets of 2D images placed in order, so spatial information does not appear in the volumetric image itself. Rather, spatial information must be gleaned by the viewer using the position of a single slice relative to other images in the stack. By scrolling through the images, the viewer can gather the 3-dimensional qualities of the structure the volumetric image represents. It is up to the cognitive resources of the viewer to combine the 2D images to understand how they represent a three-dimensional object. The viewer performs a ‘mental translation:’ they use cognitive resources to translate a set of 2D images into a 3D representation. This 3D mental representation might be considered the global processing of the model and locating individual abnormalities in 2D slices is the local search. We hypothesize that the viewer could hold a 3D mental representation in working memory at the same time as performing the local search, in line with the two-pathway model (Drew et al., 2013), but that they would also show a high load on working memory, given that working memory is a limited resource (Oberauer, 2019). If a viewer is performing the mental translation when looking at volumetric images, they would have limited bandwidth for other tasks like visual search. Visual search efficiency has been shown to decrease with increased load on working memory (Oh & Kim, 2004; Recarte & Nunes, 2003). Given that visual search is the main task of the radiologist, the strain on working memory presented by the volumetric image due to mental translation is contrary to the radiologists’ and patients’ interests.
Real three-dimensional models of anatomy or body parts might be an alternative to volumetric images. Recent developments in computer engineering have allowed for the generation of 3D models from volumetric images through new software (Talanki et al., 2021). 3D models include the spatial information the radiologist must piece together themselves, leaving more bandwidth for the radiologist to perform visual search. The global pathway – contextualizing slices relative to each other and three-dimensional space – is offloaded onto the model. 3D models have been used throughout medicine, specifically for surgical preparation and education. Benefits have been seen in oncology (Fitzgerald et al., 2023; Lindemann et al., 2023; Oderda et al., 2023), neurosurgery (Klein et al., 2013), and throughout medicine generally (Aimar et al., 2019; Diment et al., 2017; Gross et al., 2014; Jones et al., 2016; Silberstein et al., 2014; Ventola, 2014). Radiologists who are tasked with performing visual search on a 3D model might perform better than doing visual search on volumetric images given the decreased load on the radiologists working memory, at least for certain tasks where 3D information may be hard to represent in 2D images.
In the context of 3D models, creating a virtual or augmented reality version of the 3D model is much more cost efficient than printing a real, physical model. Printing a 3D model can be expensive and often depends on the complexity of the anatomical part which is being studied. (Chen, Dang & Dang, 2021) A new high quality VR headset costs around a thousand dollars at consumer price (Wirecutter, n.d.). This cost is not negligible, but VR headsets can be reused, whereas a 3D printed model is only applicable to one patient and one scan of that patient. The current study tests a cost-effective technology that can be performed in-house, thus positively impacting patients with less cost.
The aim of the current study is twofold. First, we will compare visual search performance for volumetric images to 3D models in VR, adding to the literature as to which method leads to more accurate and efficient searching. Second, we will compare performance between volumetric images and 3D models to determine whether performance does or does not support the theoretical claim that viewers of volumetric images have added processing needed in order to perform a mental translation. A statistically significant number of viewers performing better on visual search with the 3D model than the volumetric image might be evidence that the 3D model offloads work typically done by working memory. Overall, the work also serves to situate 3D virtual representation in the global processing literature to assess possibilities for future use as well as testing 3D mental representations as they relate to the global processing literature. This thesis also includes a second section. The aim of the second section is to situate our observations from the empirical study amidst current philosophical literature on mental imagery. Studies of mental imagery from the discipline of philosophy is an important inclusion because philosophy, and particularly philosophy of mind, is interdisciplinary. The problem of mental imagery is understood as a neurological, behavioral, and phenomenological process. Specifically, we argue that volumetric images can be understood as cases of amodal completion, a case of mental imagery where partially occluded objects still activate perceptual representation in the brain. By including philosophical work on mental imagery, we have expanded our current study beyond understanding how 3D models impact visual search performance and make empirically informed philosophical claims regarding the mental imagery processes at play during visual search of 3D models and volumetric images.
Our study tests novices’ visual search performance with volumetric images and 3D models of a pelvis with or without fractures. We decided to use the pelvis because it is an important test case where global processing of 3D objects in volumetric images is difficult. The pelvis is a rather complex bone, and it is often the case the pelvic fractures are occluded when viewed as volumetric images. Visual search performance is measured by the speed with which an abnormality was correctly identified, how accurate the participant’s selection of the abnormality was when an abnormality was present, and the participant’s confidence in their selection. Each participant was tested on a volumetric image, a 3D model, and a volumetric image and a 3D model presented at the same time, such that they could reference both images in order to make their selection. We had several hypotheses: 1) Participants would be most confident in the volumetric 3D (vol-3D) condition, followed by the 3D only condition, with the volumetric (vol.) only condition being the least confident, 2) participants would be most accurate in the vol-3D condition, followed by the 3D condition and least accurate in the vol. only condition, 3) participants would be faster in identifying abnormalities (fractures) in the vol. condition, followed by the 3D condition, and would be slowest in the vol-3D condition. Our rationale for the first two hypothesis is that with less information, participants are less accurate and less confident. Our rationale for the third hypothesis is that the volumetric condition will require less cognitive load and increase the speed of visual search. However, whereas in other cases (accuracy and confidence) more information always caused a benefit, we think that the vol-3D condition provides more information to the point of a slower visual search.
Methods
Participants
Data were collected from twenty-seven undergraduate psychology students (n = 27, 15 male, 12 female). Undergraduates had no experience in radiology viewing and received course credit for their participation.
Materials
Participants wore a Vive Pro Eye virtual reality headset, paired with two HTC VIVE Pro Wireless Controllers (model 2PR7200). The headset collected eye tracking data and displayed the virtual reality environment where the study took place. Two Steam VR Base Station 2 (model 1004) towers were used to track the headset as part of the virtual reality system. The headset, towers, controllers, and software used for the experiment ran through a Lenovo ThinkPad E14 Gen 6 computer running Windows 11 Pro for Workstations. A Cleanbox CX1 was used to disinfect the headset after each experimental session. Accuracy, confidence, and speed data was collected through the software designed for this experiment and downloaded directly to an encrypted hard-drive as the experimental session took place.
Participants were placed in a virtual environment that contained radiology images dependent on the condition and trial number. The virtual environment consisted of a pastel-colored room that, when in the vol. or vol-3D condition, had the same MRI image displayed on all four walls, with the ability to navigate the stack of MRI images using a plane in the center of the room the participant could “grab”. When in the 3D or vol-3D condition, the room had a 3D modeled pelvis skeleton in the center of the room. The model was noticeably larger than the average person’s pelvis. This was done for viewers to more easily see abnormalities. In all three conditions participants had access to four semi-transparent colored boxes. The participant could move the boxes by placing the controller inside the box and holding the trigger button. When the control was inside the box it would change color, and when selected using the trigger button it would change color again to signify the box could be moved around. The boxes were how participants would select abnormalities, as they would move the box around until it contained the abnormality. If abnormalities were larger than a single box, participants were asked to use multiple boxes to capture the abnormality. Boxing was used for both vol. and 3D conditions. In the vol-3D condition, a box placed in the render would simultaneously appear in the volumetric image. A light blue 2D plane perpendicular to and intersecting the 3D model was present. Users could select the plane by placing the controller in it and holding the trigger button, upon which the plane would change color. Once selected, users could move the plane along the y axis. The slice where the plane was placed on the 3D model would determine which slice of the MRI was displayed on the walls. (see Figure 2)
Nineteen anonymous MRI images were collected through collaboration with the University of Utah Radiology department. In the 3D condition, participants saw a 3D modeled pelvis made from the anonymous MRIs. Of the nineteen images, four were not included in data collection. Two were used as the two training images participants saw after providing informed consent, and two were used in the introduction scenes.

Procedure
Before participants entered the lab room the headset and controllers were calibrated to the towers and connected to the computer. Participants were greeted and confirmed as being in the correct study. Participants were then handed a sheet with three questions: “is there a possibility that you are pregnant? Do you have a heart condition? Have you ever had a seizure?” If participants answered in the affirmative to any of the three questions, they were ineligible for the study and thanked for their time before leaving. Credit was provided to ineligible participants. If participants did not answer in the affirmative for any of the eligibility questions, they would sign a consent form and complete demographics information through a Qualtrics survey on the lab computer.

Before donning the headset and entering the virtual environment, participants were briefly trained on what fractures looked like in volumetric images and 3D models of a pelvis. First, participants were given a paper with two images of a normal pelvis, one in MRI and one in 3D. Participants were guided to look at the margins, or cortex, of the bone, marked with blue arrows, as well as the joints, marked with yellow arrows. The participant was asked to notice how in the normal volumetric image the cortex was smooth without interruption and the joints were spaces between two normal bone cortexes. The experimenter then showed the abnormal pelvis images (Figure 3) and asked the participant to notice how the cortex was interrupted by the location of the fracture.
Next, participants were asked to sit in a safety chair calibrated to the position and height of the headset and asked to put on the headset. Participants were handed the controllers and entered an introduction scene to familiarize them with the VR environment, as well as what a normal pelvis case looks like. When the participant indicated they were comfortable with the virtual environment the experimenter explained the controls necessary. This included explaining the trigger button to select boxing and blue plane, as well as the touchpad which moved the viewer to a different position in the virtual environment, allowing them to gain a different perspective of the model without leaving the safety chair. The experimenter explained how to use the boxes to select fractures, and how to move the MRI plane to view different slices. The transparent annotation boxes also had a red dot in the center of the box, and participants were asked to put the red dot in the center of the fracture. The participant was then given three minutes or until they were ready to continue looking at the sample render of a normal pelvis case in the vol-3D condition. When participants said they felt familiar or time was up, another vol.-3D case was shown, this time with four fractures already boxed using the annotation method. This was done to show what participants are looking for in VR, and what to do with the annotation boxes. Participants were then given a maximum of three minutes to familiarize themselves with the fractured pelvis case. After this, the experimental trials began.
All participants underwent all three conditions, but the order in which conditions appeared was randomized. At the beginning of each condition, the experimenter explained that there would be zero to four fractures present, and that they should move to the next render if they think they have found all the fractures or there are no fractures. Participants had unlimited time and were given five cases in each condition. After each case the participants saw a screen that asked how confident they were in their identification of a fracture or that there was no fractures. After the experiment, participants completed a short Qualtrics survey on which image type they favored and if they had any input as to what changed they would love to see in these viewing conditions. Participants were then thanked for their time and awarded credit.
Results
Our primary interest was to see how speed, accuracy, and confidence compared across conditions. Our hypothesis was that participants would perform most accurately and confidently on the 3D condition, next best on the vol-3D render condition, and worst on the vol. only condition. However, we hypothesized the volumetric only condition would be the fastest, with the 3D condition next fastest, and the vol-3D render condition the slowest. The 3D model and vol-3D condition provide more information that allow for accuracy, but are a somewhat novel representation of the pelvis, along with a novel interaction method (the use of controllers). Our hypothesis that participants would be more accurate in the 3D condition and vol-3D conditions follows our theory that the mental translation is offloaded onto the 3D model decreasing cognitive load, whereas our hypothesis for lower speed was due to the novelty of the search paradigm participants were exposed to, although in theory participants ought to have performed faster in the 3D condition.
In order to test these hypotheses, we performed three repeated-measures one-way ANOVA tests, one for each hypothesis. For each dependent variable (speed, accuracy, and confidence) we ran a repeated-measures ANOVA with the viewing condition as the independent variable (vol., vol-3D., and 3D) and the respective measure as the dependent variable. We expected a main effect of condition for each test and ran post-hoc tests to compare across viewing conditions if an effect was found.
Confidence Ratings. For confidence, a repeated-measures ANOVA with a Greenhouse-Geisser correction determined that there was no statistically significant difference between confidence and any of the three viewing conditions, (F(1.66, 39.89) = 1.74, p = 0.192, η2 = 0.07). Post hoc analysis with a Bonferroni adjustment (see Figure 4) confirmed this, with no statistically significant difference in confidence between. The volumetric only condition (M = 2.96, SD = 0.17) was not different from the 3D only condition (M = 2.93, SD = 0.13), p = 0.841 or the 3D/volumetric condition (M = 3.16, SD = 0.13), p = 0.158.
Accuracy. Our accuracy value consisted of a ratio of the total number of boxes placed divided by the distance between the center of a fracture and the center of a box when placed on a fracture. This calculation reveals how close the correct guesses were to the ground truth of the center of a fracture (as designated by a trained radiologist). For reference, a value closer to zero is the most accurate, because the values are the distance of error between the target and the center of the placed box. Five participant values were removed because of data corruption. A repeated-measures ANOVA with a Greenhouse-Geisser correction determined that there was no statistical significance between accuracy and any of the three viewing conditions, (F(1.67, 38.38) = 1.33, p = 0.273, η2 = 0.05). The volumetric only condition (M = 69.21, SD = 5.61) was not different from the 3D only condition (M = 80.21, SD = 7.12), p = 0.789 or the volumetric/3D condition (M = 66.11, SD = 5.85), p = 1.00.
Reaction Time. Our speed variable consisted of the average time per render per participant for a particular condition. For example, participant 1 spent an average time of 120.962 seconds on each render presented in the volumetric only condition. A Greenhouse-Geisser correction determined that there was a significant main effect of average time per trial per participant differed statistically significantly between viewing conditions, (F(1.88,46.92) = 7.30, p = 0.002, η2 = 0.23). Post-hoc analysis with a Bonferroni adjustment (see Table 1) revealed that the difference between the volumetric condition (M = 127.06, SD = 11.43) and the 3D condition (M = 124.78, SD = 10.95) was not statistically significant, p = 1.000. However, the difference between the volumetric condition (M = 127.06, SD = 11.43) and the volumetric/3D condition (M = 87.94, SD = 9.30) was statistically significant, p = 0.002, with participants being significantly faster to respond when both 2D and 3D information was available to them for search. Further, the difference between the 3D condition (M = 124.78, SD = 10.95) and the volumetric/3D condition (M = 87.94, SD = 9.30) was statistically significant, p = 0.013, suggesting that 3D information alone was not enough to account for the speeded response.
Discussion
In this experiment we sought to understand how visual search compared between volumetric images and 3D models of pelvis scans. We tested participants with no medical image viewing experience on locating fractures in pelvis scans across three conditions. We decided to only test novice participants because with novices we would not need to overcome how long-term training experience with 2D scans might bias interpretation of the 3D model, or confidence ratings on the viewing conditions with a volumetric image. In one condition participants only saw the volumetric image, in another condition participants only saw a 3D model of the pelvis, and in the last participants saw both the volumetric image and 3D model. Our hypotheses were not fully supported. Participants did not show confidence in the volumetric/3D condition or the 3D condition compared to the volumetric only condition. There were also no differences in accuracy of identifying fractures across conditions. However, we did find that participants were fastest in identifying abnormalities (fractures) in the volumetric/3D condition compared to the other two conditions.
Hypothesis 1: Confidence. Our rationale behind the first hypothesis was participants would be less confident in their selections when they have less visual information to specify the location of a fracture. Previous literature shows that as images become more complicated, participants are less confident in their selection of targets (Deza, Xiao & Eckstein, 2016; Lee, Yeung & Summerfield, 2021; Mamassian, 2016).
The volumetric image, then, would have the least visual information of the three conditions because it does not directly portray spatial information. We predicted participants must use cognitive resources to transform the 2D slices into a cohesive representation of a three-dimensional object. The 3D model would have more information than the volumetric image because it includes the spatial information the volumetric image does not, and the volumetric/3D condition would have the most information because it combines the two image types. We found no statistical difference between the volumetric and 3D condition as well as between the volumetric and volumetric/3D condition. Contrary to our predictions, we only found a statistically significant difference between the volumetric/3D condition and the 3D only condition.
We expected the 3D model to provide more information and make the participant more confident in their reading of the medical image, but one interpretation of these results is that the 3D model did the opposite. Given that there was no difference in confidence between the conditions, it seems like the volumetric condition provided the necessary information for participants to be confident in their selection. Previous literature shows a clear increase in diagnostic confidence with an increase in experience (Crowe et al., 2018; Crowley et al., 2003; Donovan & Litchfield, 2013), and this literature might help explain the surprising lack of results observed in our study. Our study was the first to compare volumetric images with 3D models, which makes extrapolating to other research difficult. Nonetheless, other studies have indirectly tested spatial reasoning and confidence. Crowe et al. (2018) tested novices and experts on locating brain tumors throughout slices of an MRI scan. Because the participants had to select all slices with abnormalities somewhat calls forth three-dimensional spatial thinking. Unfortunately, the authors of this study did not compare confidence when locating tumors in the entire MRI with a single slice. Still, given that previous literature shows an increase in confidence with expertise, the results of our study might be explained by a lack of expertise. Since confidence is dependent on experience, and the participants in our study had no experience, it makes sense that there is no data consistent with literature that has historically compared novices and experts (Drew et al., 2019). Although we originally theorized that the 3D model would offload spatial reasoning, it could be the case that this only matters when participants have an understanding of what they are looking at, rather than being exposed to a pelvis model for the first time. Future studies ought to continue the tradition in medical visual search literature in comparing experts and novices.
A second limitation which may have impacted the confidence results in this confusing way is the low acuity of the 3D modeled scan. As pictured in Figure 7 some of the 3D models had a fuzzy texture to them which could have resulted in misinterpretation of the image or lowered the confidence in what was pictured. All participants in this experiment were novices with no medical image viewing experience and only a short training on what fractures looked like in either image type, so oddities like this would not have been easily dismissed as they may have been from a more experienced participant. The limitation of 3D models with low acuity might explain why we did not see 3D models increase confidence although they do contain more information than volumetric images: the 3D model was potentially ambiguous for the spatial information to provide the boost in confidence we hypothesized. Future studies could use models with higher visual acuity and less ambiguity, give participants a more comprehensive training that makes it easier to differentiate between low acuity and fractures, or test expert radiologists. There is some literature that suggests that perfect virtual trainings can lead to impairment in future tasks (Lefor et al., 2020; Massoth et al., 2019).

These possible interpretations of the confidence data allow for speculation about the use of 3D models to provide increased visual search efficiency. We might gather that 3D models for medical image viewing need to have high acuity to provide benefits, but we ought to do another experiment and see if 3D models provide benefits at all when there is high acuity to the model. This could be a significant contribution because it calls into question the cost of resources to develop a 3D model from volumetric scans compared to the possible benefit. In our own experiment, converting volumetric images to 3D models took a lot of time and we still had the issue of low acuity. Having a model with higher acuity may take longer to create, and then the cost of the process from patient scanning to radiologist reading to patient treatment is still a long process. Even if a 3D model with high acuity provides benefits to visual search efficiency, time and energy resources are spent on the development of the model itself. As discussed in §I.1, 3D modeling is used more in surgical preparation and education than in radiology image reading, and with 3D printed models instead of virtual reality (Aimar et al., 2019; Diment et al., 2017; Fitzgerald et al., 2023; Gross et al., 2014; Jones et al., 2016; Klein et al., 2013; Lindemann et al., 2023; Oderda et al., 2023; Silberstein et al., 2014; Ventola, 2014). It might be the case that the literature and development of this technology moving more towards practical work with a model, rather than quick image reading, is indicative of the resource cost associated with the development of 3D radiology scans, at least for now. More research should be done to clarify the many questions raised by our confidence data, including extending the testing paradigm to expert radiologists.
Hypothesis 2: Accuracy. Our second hypothesis was that participants would be most accurate in the vol.-3D condition, followed by the 3D condition and least accurate in the vol. only condition. However, our results showed no significant differences in accuracy among the three conditions. This finding is consistent with a previous preliminary experiment. In a preliminary experiment where we tested accuracy of fracture identification on medical students before and after an intervention, we also found no statistically significant differences in the experimental group before and after an intervention (Stefanucci et al., 2025). In the preliminary study, the experimental group was given a training to identify fractures in 3D models as an intervention. In the second case set, which was received after the intervention, participants were exposed to a volumetric and 3D condition (see Figure 8 for experimental design). That we saw no difference in accuracy between conditions in either the preliminary experiment or the main experiment suggests that either our paradigm does not allow the user to engage with the model in a way that allows for accurate visual search (or maybe features of the model or user interaction hinder accurate visual search), or there is no change in accuracy when looking at a 3D model compared to a volumetric image. We air more on the side of the former because of the limitation of low acuity of the 3D model in both the preliminary and main experiments.

One explanation for our findings might be the lack of expertise in our sample size. Much of the literature supports this explanation but there are caveats. It seems that the piecing together of 2D images to create a 3D mental model of the body part when viewing a volumetric image is not dependent on expertise. Expertise is not required to comprehend volumetric images. Even without medical expertise, viewers can still look at a volumetric image and contextualize the individual slices to grasp the spatial qualities of the body part being captured as evidenced by the data from the current study as well. If volumetric images are still comprehensible for novices, then the argument that benefits from viewing 3D models pertain only to experts relies on a mechanism independent from the spatial features of volumetric images. An example of a non-spatial skill gained through expertise is perceptual expertise (Waite, 2019). Reliance on a non-spatial expertise skill to explain differences in performance across conditions presenting different spatial information is difficult because there is nothing to link the expertise to the conditions being tested. In other words, perceptual expertise functions the same across all three conditions, despite their unique spatial features. More research is needed to see how expertise influences the performance on accuracy of abnormality selection between spatially unique conditions. One might also question our hypothesis given our preliminary experiment and the accuracy results we got. Our thinking was that the significant difference in paradigm, specifically being immersed in VR and having more control of the volumetric image and 3D model might significantly change the results.
Hypothesis 3: Speed. Our third hypothesis was that participants would be faster in identifying abnormalities (fractures) in the vol. condition, followed by the 3D condition, and would be slowest in the vol-3D condition. We found two statistically significant differences between conditions with the volumetric/3D condition being significantly faster compared to the volumetric condition and the 3D condition. These results in contrast to our hypothesis and theorizing about the role of visual data influencing cognitive load. Our thinking was that, while we thought 3D models provided better data for fracture identification because it more accurately represented the 3D human body, the 3D model also provided more visual data for the viewer to process. The volumetric and 3D condition had even more data, given the many layers of 2D scans provided by the volumetric image. Our hypothesis was inconsistent with the collected data. When participants were presented with the volumetric and 3D condition, they were able to perform significantly faster. Previous literature does not support these findings. In one study (Alvarez & Cavanagh, 2004) that compared speed of visual search between simple and complex visual objects, more complex visual objects led to slower visual search rates, which is what we would expect. Other studies have also found that visual search speed was slowed with more or less visual information (Barbosa et al., 2024; Wolfe, 2012). One possibility for why the literature does not reflect our data is that we did not have trial-by-trial data for accuracy. So, although we found participants to be faster when viewing the volumetric and 3D condition, we do not know if they sacrificed accuracy for that boost in speed. Future studies using this paradigm ought to test accuracy as well as speed on a trial-by-trial basis so we can know if faster visual search is only possible with a loss in accuracy.
One possible interpretation of this data is that the volumetric and 3D conditions were too ambiguous for participants to move quickly, and the increased cognitive processing of making decisions about ambiguous abnormalities may have caused the decrease in speed. There is, however, one flaw with this interpretation, and that is that the volumetric images are not necessarily “ambiguous.” Nonetheless, it is still necessary to scroll through the stack of images, which takes more time. The design of the 3D models had limitations, seen in Figure 7, where conversion from volumetric image to 3D model led to fuzziness or low acuity. However, the quality of the volumetric images was the same as when medical professionals perform analysis on patient scans. One way out of this hole is to point towards the naivety of the participant in this study. It may be the case that the volumetric image was of a better quality than the 3D model, and thus both image types did not have the same kind of ambiguity, but that does not necessarily mean that the volumetric image was not confusing to the point of increasing speed. This is most likely a feature of having novice participants, as the ambiguity of the volumetric image could come simply from a lack of experience with medical images. A plethora of research shows increase in speed, accuracy, and confidence with an increase in viewing experience, which shows that as viewers have experience with a certain stimuli they can process it better (Drew et al., 2019; Williams & Drew, 2021). However, there is other data that complicates this interpretation. Our preliminary experiment (Stefanucci et al., 2025) found a decrease in reaction time after the intervention when participants were exposed to a similar volumetric and 3D condition. Future work is needed to flush out alternative hypotheses.
A major limit of our experiment was that we did not collect trial-by-trial data and only analyzed averages across participants and conditions. For speed this is particularly deleterious because we do not know if participants were being faster in the volumetric and 3D condition while being less accurate. It could be the case that participants were exposed to so much information that they simply did not want to deal with it and lost out on accuracy as a result. The first interpretation of this data, that the volumetric and 3D conditions were too ambiguous for participants to move quickly, may still hold. However, trial-by-trial data would provide important information for making stronger claims and better understanding the implications of those claims.
Another limit of this study was that we did not include a measure of working memory load. A major theory that motivated the experiment was about working memory load when viewing a 3D model compared to a volumetric image. Future research utilizing this experiment paradigm ought to include a measure of working memory across all three conditions as well, in order to compare WM load between image types.
Conclusion. In the present study we tested speed, accuracy, and confidence of novice participants viewing medical images of a pelvis and locating fractures. Medical images were of two types across three conditions. In one condition, participants only saw only a 3D modeled pelvis. In the second condition, participants saw a volumetric image, and in the final condition participants saw both the volumetric image and 3D model. We had three hypotheses going into this experiment, one for each main dependent variable captured in the experiment (speed, accuracy, and confidence). Our data was mostly inconsistent with our hypotheses, which provides information for how to develop an experiment that aligns more with the literature on visual search and expertise. The major limitations of this study were the places where the most information can be gleaned, specifically about testing expert visual search ability in 3D compared to volumetric image viewing paradigm and collecting trial by trial data in order to understand speed accuracy tradeoffs. Future directions in this research could continue to use the same paradigm to better understand how 3D models change visual search performance compared to other types of medical images.
MENTAL TRANSLATION AS A MODAL COMPLETION
Introduction: Review of Experiment
In the experimental portion of this paper, we described an experiment in a virtual reality environment with the aim of testing participants on accuracy, speed, and confidence locating abnormalities (fractures) when viewing virtual 3D models compared to viewing volumetric images compared to viewing both 3D models and volumetric images at the same time. Our empirical study had two goals: 1) to gain data that could be used to develop radiology workstations that increase radiology accuracy and efficiency, and 2) to use the data to make a claim about “mental translation,” the process we believed to be at play when viewers look at volumetric images. In this philosophically oriented section of the paper, I will expand on the second aim of the study, fleshing out the philosophical work behind what mental translation is and how our data relates to this previous research.
Briefly summarizing, volumetric images are stacks of 2D images through which the viewer can scroll. The viewer never encounters a 3-dimensional model of the structure captured by the volumetric image but must gather spatial information themselves using the sequential series of 2D images (see Figure 1). Further, abnormalities are rarely contained to a single 2D slice of a volumetric image, so the radiologist has stake in understanding the information as a 3D object and not merely a series of 2D images. The process described – taking a set of 2D images and gaining a 3-dimensional mental representation – is what we termed “mental translation.” Our thinking was that such a process would require significant cognitive resources. Therefore, if we saw an increase in speed, accuracy, and confidence when participants were viewing a 3D model (a case where they didn’t need to perform a mental translation), then they were holding a 3D mental representation in their head when viewing the volumetric images.
Our data did not support our theory, as we observed no statistically significant effects that went in the same direction as we predicted. We thought the 3D model would provide more information, and thus speed, accuracy, and confidence would increase, we found no effects showing this. In fact, the data seems to show that the volumetric image allowed for more accurate and more confident visual search. Despite this data, the following philosophical analysis is still applicable and provides important information to understanding the research question of how viewers engage with volumetric images. As I describe in §I.4, there were significant limitations to the research paradigm and the kind of data we were able to collect. For example, the 3D model was often fuzzy making it hard to interpret problems with rendering from legitimate fractures. Another issue was that we could not collect trial-by-trial data but could only gather averages for participants across the different conditions. So, while we observed participants being faster in the volumetric&3D condition, we do not know if they were more, less, or just as accurate as in other conditions when going faster. The data we collected does not point in the direction of any “mental translation” process, but these limitations mean that we did not perform an experiment which could accurately collect data on this process at all. I will provide a philosophical analysis for mental translation as we expected it would function given our theorizing, our data collection process and experience with the research paradigm, as well as the most recent philosophical literature on mental imagery. One goal of this research project was to have an experiment that could be used to create some empirically informed philosophy. Unfortunately, that is not possible but there is still useful critical philosophical work that can be done.
The language used to describe the mental translation phenomenon we want to explore already points towards a possible candidate to contextualize this process. Mental imagery and mental representation are well studied phenomenon and the theories explaining these phenomenon graft well onto mental translation. Further, the approach to studying mental imagery steers away from introspection and more towards empirically informed philosophy, the method which is deployed in this section. My plan for the rest of the paper is as follows. First, I will explain a dominant theory of mental imagery and the evidence in support of the theory. Next, I will fit mental translation into that theory and demonstrate the explanatory power gained by calling mental translation a case of mental imagery. Next, I will make an argument for exactly what kind of mental imagery the mental translation processes might be, settling on an unconscious “hybrid between perception and mental imagery” (Nanay, 2023) I will rest my case there, but speculate that perhaps mental translation is a complex kind of amodal completion. Next, by discussing issues with my analysis, specifically the lack of neurological data, data which much of the mental imagery claims depend on. I will finish by discussing a possible future direction for research on this topic.
Nanay’s Definition of Mental Imagery
Bence Nanay’s seminal work Mental Imagery: Philosophy, Psychology, and Neuroscience published in 2023 paints a thorough picture of mental imagery and the role it plays throughout our cognition. Nanay defines mental imagery as “perceptual processing (or perceptual representation) that is not directly triggered by sensory input.” (Nanay, 2021, p. 4) By perceptual processing,[1] Nanay targets the specific neurological structures which are activated during typical perception, what Nanay calls “online” perception. (Nany, 2021) Online perception is the use of sensory mediums, like touch or smell or sight, to activate perceptual processes in the brain. For example, our vision, specifically vision as online perception, works through light first activating photoreceptors, then moving through the brain until it reaches the occipital lobe and sends signals through two different streams depending on the input. (Zachariou et al., 2014) Mental imagery, specifically visual, would be if these same neurological structures, those normally used during online perception, were also activated but without the transduction of light. This is where the next part of his definition comes into play, where he says “not directly triggered by sensory input.” Mental imagery is only when the sense organs for the corresponding perceptual process are not used in the activation of the perceptual processing. If they are directly used, then it is online perception.
Directness is the vital defining line in what counts as mental imagery. Importantly, Nanay draws that line such that many more things count as mental imagery than we might think at first. Any kind of perceptual processing that does not directly involve the sensory organs is considered “indirect” activation, counting as mental imagery. Indirect includes any top-down, lateral or cross modal activation of neurological perceptual processes. (Nanay, 2021, pg. 7) An example of mental imagery would be visualizing an apple with your eyes closed, as this is a top-down activation of perceptual visual representation. Less obviously would be lateral perceptual activation, which can be happen within a single sense modality or “cross modally.” In a single sense modality, bottom-up perceptual processes can trigger other perceptual representation, making those latter processes mental imagery as they are no longer directly triggered through the sense organs. As Nanay says
If the visual representation of the center of the visual field is triggered by input in the periphery of the visual field (say, because the center of the visual field is occluded by an empty white piece of paper) … the visual input in the periphery leads to the visual representation of the contours in the periphery. These visual representations trigger, laterally, the visual representation of the contours at the middle of the visual field. This process is mediated by the visual representation in the periphery. Hence, it is not direct. It counts as mental imagery.
Besides explaining his definition of mental imagery, Nanay demonstrates that mental imagery is much more prolific throughout our daily life than we might think at first. The “filling in” of our blind-spot, the point on our retina that does not contain any photoreceptors because it is where the optic nerve converges and moves towards the brain, is a lateral activation of perceptual processing. The blind-spot itself literally cannot be seen, yet we do not walk around with a black hole in the center of our vision. Instead, the perceptual processing from the rest of our visual field laterally activates perceptual processing later in the visual stream and fills in the blind spot. (Nanay, 2021, Pg. 33)
Competing views of mental imagery
Nanay’s definition of mental imagery is not the first, but it is unique and disagrees with previously proposed definitions. Nanay is aiming to have a cohesive definition of mental imagery that can capture the large breadth of mental imagery capabilities. Some people cannot perform mental imagery, while others have extremely vivid mental imagery. (Nanay, 2021) Further, Nanay wants a technically useful definition that hold some critical explanatory power for cases where mental imagery is at play. (Nanay, 2021, pg. 3) Nanay critiques three previously provided definitions of mental imagery and describes why his is the best, and in doing so builds out the limits of mental imagery in our cognition.
The first possible view Nanay explores is the claim that mental imagery “[refers] to representations and the accompanying experience of sensory information without a direct external stimulus” (Pearson et al., 2015). The Pearson definition is almost exactly like Nanay’s own definition. Representations of sensory information refers to perceptual processing without the use of sensory organs to activate processing. However, this definition adds the “accompanying experience” qualifier. Doing so requires that mental imagery is a conscious experience, which might not be something we would want to accept. Nanay makes three arguments as to why. (Nanay, Chapter 4) Firstly, both Nanay and Pearson conceptualize mental imagery as a perceptual process, but one that is not directly activated by sensory input. Perception can be unconscious, like when a stimulus is presented for a very short periods of time and perceptual processing takes place, but the subject has no conscious experience of it. (Nanay pg. 23; Kentridge et al., 1999; Goodale & Milner 2004; Kouider and Dehaene 2007). While ‘perception’ is most often thought of as phenomenological, in the context of an empirically informed philosophy it is neurologically defined. Perception is the activation of perceptual processing in the brain, and Nanay’s definition of mental imagery is similarly defined. It is the case that perceptual processing can happen unconsciously, and mental imagery is the same neurological system as perception, only mental imagery is not directly triggered by sensory input. There is no reason why the perceptual system could function unconsciously in regard to direct triggers (through sensory organs) but not through indirect triggers like top down or cross modal processing.
An example of unconscious mental imagery could be the filling in of the blind-spot, and another would be amodal completion. As mentioned, the blind spot is the area on the retina without photoreceptors, where light cannot be translated to neurological activation. Instead of experiencing a blind spot, our perceptual representation of the blind-spot is activated laterally, by our perceptual representation directly triggers by the light we do capture. We do not consciously “perform” mental imagery to fill in the blind spot. Instead, it seems to happen unconsciously. One might argue that the Pearson (2015) definition does account for this case, given that we do have an “accompanying experience” of the filled-in blind spot. On the one hand, there is conscious mental imagery because we are conscious of the end result of a process of mental imagery: the filled in blind-spot. One the other hand, there is unconscious mental imagery because the mental imagery processing is activated unintentionally and unconsciously. Perhaps the Pearson definition can account of the filling in of the blind spot, although the “accompanying experience” of a filled in blind spot is distinct from the conscious generation of a mental imagery.
Another case in daily life of mental imagery is totally unconscious, that of amodal completion. Amodal completion is when we hold complete representations of objects that are partially occluded (Nanay, pg. 56). For example, if a book is sitting on my desk behind my computer, such that I can only see half of the book, I will still hold a representation of the book in it’s entirety. Another example is the completion of the sides of three-dimensional solid objects that we cannot see. For example, holding a representation of the back of a wooden cube. When we perceive occluded objects, our perceptual systems complete representations of the parts of the object we cannot see. (Ekroll et al., 2016; Nanay et al., 2010) Amodal completion might make a stronger case for truly unconscious mental imagery: mental imagery where we have no experience of a mental image. One might say that there is still a kind of ‘experience’ of amodal completion, at the very least the experience of knowing that the objects we’re looking at don’t simply end when we stop being able to see them, but this is no doubt dependent on the person. Just as the vividness of mental imagery is different for different people, the intensity of the experience of amodal completion is probably different depending on the person. Some people may not understand ambiguous cases of occluded objects, like in Figure 6. There are then, most likely, at least some cases where mental imagery, through amodal completion, is not consciously recognized and is an objection to the Pearson definition of mental imagery.

If the blind-spot and amodal completion are not convincing enough to demonstrate unconscious mental imagery, Nanay provides an example from aphantasia: people who are unable to perform visual mental imagery, or can only create a very mental image. Nanay makes a point to say that this example of unconscious mental imagery does not rely on the existence of unconscious perceptual processes, since some people are wedded to a belief that unconscious perceptual representation, of any kind, is impossible. (Nanay, 2021, pp. 23) In a study (Jacobs et al., 2018) with someone who qualified as having aphantasia through the Vividness of Mental Imagery Questionnaire, both controls and the participant with aphantasia were presented with a series of stimuli. (Figure 7) First, participants would see the name of the shape which they are supposed to imagine. Then, they either saw the shape or the corners of a box in which the shape would appear. Participants then saw a brief noise stimulus to remove the benefit of any afterimage from the shape, and then they saw a dot on the screen. Participants were asked how confident they were that the dot was within the boundaries of the shape. Interestingly, the performance of AI, a subject with aphantasia, was not significantly different from controls on either of these tasks…The straightforward explanation is that the subject does use mental imagery and uses it in a very similar way to the control subjects when performing the mental imagery task. But while the controls use conscious mental imagery, the subject uses unconscious mental imagery. (Nanay, 2021, pp. 29)
The Pearson definition of mental imagery requires that mental imagery have an “accompanying experience,” but there are several cases where there is unconscious indirect activation of perceptual processes.

Another definition of mental imagery that Nanay disagrees with comes from Kosslyn, Behrmann, & Jeannerod (1995). They write that “Visual mental imagery is “seeing” in the absence of the appropriate immediate sensory input, auditory mental imagery is “hearing” in the absence of the immediate sensory input, and so on” (Kosslyn et al. 1995). The Kosslyn definition of mental imagery is attractive, but has room to be more specific. Exactly what “seeing” or “hearing” means is questionable. One could input a whole manner of claims that we might not want to accept, or have already raised objection to. For example, “seeing” might mean a conscious mental image experience, and we have already restated Nanay’s arugment for unconscious mental imagery. “Seeing” might be making claims about the phenomenology of mental imagery, such that mental imagery and bottom-up visual experience are phenomenologically similar in important ways. This claim is problematic. For some it is the case that their mental imagery is hyper-realistic (called hyperphantasia), but as we know some do not have the ability for mental imagery. In between, there is a spectrum of capability (Nanay, 2021, pp. 19). Kosslyn’s claim about “seeing” cuts out the aphantasia, or people who have an experience of mental imagery distinct from seeing. Even for people who are somewhere in between aphantasia and hyperhantasia, visual mental imagery is not the same as seeing, calling for the augmentation of Kosslyn’s definition.
Kosslyn’s definition also fails to explain the full breadth of sense modalities to which mental imagery can apply. Kosslyn only lists seeing and hearing, but, as Nanay writes, “mental imagery, in spite of the connotations of the word “image,” is not necessarily visual: mental imagery can be auditory, olfactory, and tactile as well: it can happen in all sense modalities, not just in vision” (Nanay, pg. 4). One might argue that the Kosslyn definition covers itself by saying “and so on,” perhaps signaling to the other sense modalities that are there. Even if one extrapolated the Kosslyn definition as making this move, it would be insufficient. Simply putting “seeing” and calling forth the entire visual process is not very specific, at least not in a way that provides significant explanatory power for understanding mental imagery and applying mental imagery to research.
Mental Translation as Amodal Completion
Nanay’s definition of mental imagery is perceptual processing without direct sensory input. Although there are competing definitions of mental imagery, Nanay’s definition is stronger given the defense described in §II.2 In the remainder of the paper, I will describe how Nanay’s definition of mental imagery as perceptual processing without direct sensory input can be used to explain the mental translation act, or at least the requirement of people viewing volumetric images to understand 3D spatial information gleaned from stacks of images.
Nanay spends a significant portion of his book describing mixed cases of mental imagery and how they fit into his view (Nanay, Chapter 9). Mixed cases refers to times when our online perceptual processing is simultaneously mixed with offline perception (mental imagery) in order for us to have the kind of experience in the world that we do. To reuse an example, filling in the blind spot is a type of lateral processing where the direct sensory activation of the retina laterally activates perceptual processing down the visual stream to fill in the blind spot. The visual experience of the filled in blind spot does not happen from sense organ activation because there are no photoreceptors to be activated where the blind spot is. The online perception comes from the transduction of light through the photoreceptors on the retina, and the offline perception comes from perceptual processing later in the visual stream. Filling in the blind spot, then, is a “mixed case” between online perception and mental imagery.
Amodal completion, another example already discussed, is an example of mixed perception and I will argue that when viewers are looking at volumetric images, they are performing amodal completion, or something like amodal completion. Very briefly, amodal completion is when viewers perceive (have perceptual representation of) the occluded parts of a partially occluded object. Examples are something like realizing there is an entire chair in a scene even when the bottom half is occluded by the table. We are often presented with parts of object, rather than a complete view of those object, and amodal completion is how we perceptually represent the world in these cases.
There are a few premises we can layout first, for which a strong argument is not necessary. Firstly, I am assuming it is the case that people viewing a volumetric image are, at the very least, using some sort of spatial reasoning. Although we conducted a psychology experiment, described in §I, to test this hypothesis and hopefully provide empirical evidence, we did not find any significant effects that could be used to support this premise through that study. This doesn’t mean that it’s not the case spatial reasoning is used when looking at volumetric images. Several studies explore the spatial component of the volumetric image (Williams & Drew, 2021; Drew et al., 2019) and intuitively it seems like spatial reasoning is required; an abnormality rarely sits within one “slice” of the volumetric image, so scrolling between slices and contextualizing the images with those that come before and after seems like it would be important. As I discuss in §I. 4, there were several significant limitations to our study that may have biased out results not making it necessarily the case that people viewing volumetric images do not utilize spatial reasoning. Therefore, I will continue with the premise that spatial reasoning is used when viewing a volumetric image.
The second premise follows from the first, and it’s that if it’s the case that spatial reasoning is used when viewing volumetric images, then viewing volumetric images uses mental imagery. Again, this seems intuitively right. Mental imagery is perceptual representation without sensory input, and there is no sensory input of a 3D object when viewing a volumetric image. Yet, understanding the multitude of volumetric slices requires some sort of spatial reasoning, or some sort of representation of the anatomical object pictured in the volumetric image as three dimensional. When a viewer uses spatial reasoning to combine the 2D slices of the volumetric image into a cohesive three dimensional model, they are generating a perceptual representation without direct sensory input of a 3D model. In other words, they are performing mental imagery. Here is where the biggest limitation of my argument lies. Nanay’s definition of mental is easy to test using fMRI or other neurological studies. If we want to see if perceptual processing is happening when viewing a volumetric scan, we ought to collect data on what kind of perceptual processing is taking place when viewers are presented in a volumetric scan compared to a 3D model of the same anatomical object and see if any differences in activation are present. Unfortunately, the current study does not collect data like this, so I can’t support my argument with any neurological empirical data. Future directions of this research ought to collect data like this. Barring this lack of empirical data, and the neat notch in Nanay’s definition where empirical data would fit, we use modus ponens on the previous two premises to say that mental imagery is involved when viewing volumetric images.
The next premise is that the kind of mental imagery involved when viewing a volumetric image is a mixed case of mental imagery, like amodal completion or filling in the blind spot. People viewing the volumetric image are collecting online perceptual information; they are looking at and scrolling through the slices of the scan, using the visual sense organ to perform perceptual processing. However, as has been established, there is simultaneous mental imagery going on with the perceptual representation of the anatomical object represented through a series of 2D images as three-dimensional. The mental imagery happening, then, is mixed. There is both online perception through perceptual processing directly activated through the visual sense organ and offline perception of perceptual processing, of the 3D anatomical object, without direct activation.
Of the kinds of mixed cases of mental imagery Nanay describes, amodal completion provides the most explanatory power for understanding the viewing of volumetric images. Amodal completion has already been used to explain how we perform spatial reasoning in the world. For example, how we understand that object have a backside to them although it is occluded by the object being opaque (Ekroll et al., 2016; Nanay et al., 2010). One might say that there is some sort of “reaching” or “striving” going on when looking at partially occluded objects. When we see a partially occluded object, we do not accept the components of the object that we see as the only components of the object. Rather, we reach to understand what we see as part of a cohesive object, one that exists in a three-dimensional world. How we approach volumetric images can be characterized in the same way. We are presented with incomplete information: cross-sections of a three-dimensional anatomical object which we can only engage with sequentially and one at a time. We understand that the information we have is incomplete and we strive to grasp a cohesive perceptual representation of the object being pictured in the volumetric scan. We have already established that mental imagery and mixed mental imagery are at play when viewing volumetric images. Given these criteria, amodal completion explains how mixed cases of mental imagery apply.
Despite the success of Nanay’s definition of mental imagery in explaining mental translation as amodal completion, some may argue that other definitions of mental imagery can perform the work done by Nanay’s definition. However, this is not the case, as other definitions fail to explain the phenomenon we tried to observe as described in the experimental portion of this paper. Firstly, the Pearson definition which defines mental imagery as “[refers] to representations and the accompanying experience of sensory information without a direct external stimulus” (Pearson et al., 2015). The major problem with the Pearson definition is that is adds the “accompanying experience” qualifier to mental imagery. In order for perceptual representation without direct sensory input to qualify as mental imagery there needs to be an experience of it. As seen in §II.3, there are many cases in perception generally where perceptual representation without direct sensory input but without an associated experience. In the case of viewing volumetric images and performing mental representation, this the same limitation applies. It is conceivable of someone who is performing amodal completion, and hence mental imagery, when viewing a volumetric image but does not have an experience of that happening. One scenario that demonstrates the plausibility of amodal completion when viewing volumetric images without an associated experience is in cases where a viewer performs amodal completion to contextualize a slice of the volumetric image but does not create a whole 3D mental image of the anatomical object. One might argue that even contextualizing an individual slice is a kind of experience, but that would require more arguments than what Pearson provides. It seems that an experience would be a phenomenological act beyond the interpretation of information. However, using amodal completion to contextualize slices (amodal completion) without generating an entire mental image doesn’t seem like it would have a strong phenomenology, and perhaps none at all.
In §I we discussed the act of “mental translation,” which was defined as the act of understanding a series of 2D images (the volumetric image) as a three-dimensional object. Amodal completion matches the loose definition of mental translation. There is a series of 2D images in the volumetric image and some cognitive process is used to grasp a cohesive perceptual representation of the object the volumetric image depicts. Using amodal completion and what is said about mental imagery explains what was called “mental translation” significantly more than was discussed in §I. Rather than a vague definition to describe a complex cognitive process, mental translation becomes amodal completion: a robust neurological and philosophical process.
Bibliography
Aimar, A., Palermo, A., & Innocenti, B. (2019). The Role of 3D Printing in Medical Applications: A State of the Art. Journal of healthcare engineering, 2019, 5340616. https://doi.org/10.1155/2019/5340616
Alexander, R., Waite, S., Bruno, M. A., Krupinski, E. A., Berlin, L., Macknik, S., & Martinez-Conde, S. (2022). Mandating Limits on Workload, Duty, and Speed in Radiology. Radiology, 304(2), 274–282. https://doi.org/10.1148/radiol.212631
Albers, F., Trypke, M., Stebner, F., Wirth, J., & Plass, J. L. (2023). Different types of redundancy and their effect on learning and cognitive load. British Journal of Educational Psychology, 93(3), 12592. https://doi.org/10.1111/bjep.12592
Alvarez, G. A., & Cavanagh, P. (2004). Psychological Science, 15(2), 106-111. DOI:10.1111/j.0963-7214.2004.01502006.x
Andriole, K. P., Wolfe, J. M., Khorasani, R., Treves, S. T., Getty, D. J., Jacobson, F. L., … Seltzer, S. E. (2011). Optimizing analysis, visualization, and navigation of large image data sets: one 5000-section CT scan can ruin your whole day. Radiology, 259(2), 346–362.
Austin, J. H., Romney, B. M., and Goldsmith, L. S. (1992). Missed bronchogenic carcinoma: radiographic findings in 27 patients with a potentially resectable lesion evident in retrospect. Radiology 182, 115–122. doi: 10.1148/radiology.182.1.1727272
Barbosa, A., Ruarte, G., Ries, A. J., Kamienkowski, J. E., & Ison, M. (2024). Frontiers in Human Neuroscience.DOI:10.3389/fnhum.2024.1436564.
Bertram, R., Helle, L., Kaakinen, J. K., & Svedström, E. (2013). The effect of expertise on eye movement behaviour in medical image perception. PLoS ONE, 8(6), e66169. https://doi.org/10.1371/journal.pone.0066169
Bird, R. E., Wallace, T. W., and Yankaskas, B. C. (1992). Analysis of cancers missed at screening mammography. Radiology 184, 613–617. doi: 10.1148/radiology.184.3.1509041
Birkelo, C. C., Chamberlain, W. E., Phelps, P. S., Schools, P. E., Zacks, D., and Yerushalmy, J. (1947). Tuberculosis case finding. A comparison of the effectiveness of various roentgenographic and photofluorographic methods. J. Am. Med. Assoc. 133, 359–366. doi: 10.1001/jama.1947.02880060001001
Chen, J. V., Dang, A. B. C., & Dang, A. (2021). Comparing cost and print time estimates for six commercially-available 3D printers obtained through slicing software for clinically relevant anatomical models. 3D printing in medicine, 7(1), 1. https://doi.org/10.1186/s41205-020-00091-4
Chua, K. W., Richler, J. J., & Gauthier, I. (2015). Holistic processing from learned attention to parts. Journal of Experimental Psychology: General, 144(4), 723.
Crowe, E. M., Gilchrist, I. D., & Kent, C. (2018). New approaches to the analysis of eye movement behaviour across expertise while viewing brain MRIs. Cognitive research: principles and implications, 3(1), 12. https://doi.org/10.1186/s41235-018-0097-4
Crowley, R. S., Naus, G. J., Stewart, J., & Friedman, C. P. (2003). Development of visual diagnostic expertise in pathology-an information-processing study. Journal of the American Medical Informatics Association, 10, 39–51. https://doi. org/10.1197/jamia.M1123.
Cymek D. H. (2024). Effects of blinded and nonblinded sequential human redundancy on inspection effort and inspection outcome in low prevalence visual search. Scientific reports, 14(1), 23003. https://doi.org/10.1038/s41598-024-72210-8
Deza, A., Xiao, J., & Eckstein, M. P. (2016). The influence of visual clutter on search guidance with complex scenes. Journal of Vision, 16(9), 20. https://doi.org/10.1167/16.9.20.
Diment L. E., Thompson M. S., Bergmann J. H. M. Clinical efficacy and effectiveness of 3D printing: a systematic review. BMJ Open. 2017;7(12) doi: 10.1136/bmjopen-2017-016891.e016891
Donovan, T., & Litchfield, D. (2013). Looking for cancer: Expertise-related differences in searching and decision making. Applied Cognitive Psychology, 27(1), 43–49. https://doi.org/10.1002/acp.2869
Drew, T., et al. (2013). Informatics in radiology: What can you see in a single glance and how might this guide visual search in medical images? Radiographics, 33(1), 263–274. https://doi.org/10.1148/rg.331125023
Drew, T., Lavelle, M., Kerr, K. F., Shucard, H., Brunyé, T. T., Weaver, D. L., & Elmore, J. G. (2021). More scanning, but not zooming, is associated with diagnostic accuracy in evaluating digital breast pathology slides. Journal of Vision, 21(11), 7-7.
Drew, T., Vo, M. L., Olwal, A., Jacobson, F., Seltzer, S. E., & Wolfe, J. M. (2013). Scanners and drillers: characterizing expert visual search through volumetric images. Journal of vision, 13(10), 3. https://doi.org/10.1167/13.10.3
Ekroll, V., Sayim, B., Van der Hallen, R., & Wagemans, J. (2016). Illusory visual completion of an object’s invisible backside can make your finger feel shorter. Current Biology, 26(8), 1029–1033. https://doi.org/10.1016/j.cub.2016.02.001
Fitzgerald, C. W., Hararah, M., McLean, T., Woods, R., Dogan, S., Tabar, V., Ganly, I., Matros, E., & Cohen, M. A. (2023). Virtual surgical planning and three-dimensional models for precision sinonasal and skull base surgery. Cancers, 15(20), 4989. https://doi.org/10.3390/cancers15204989
Goodale, M. A., & Milner, A. D. (2004). Sight unseen. Oxford University Press.
Gross BC, Erkal JL, Lockwood SY, et al. Evaluation of 3D printing and its potential impact on biotechnology and the chemical sciences. Anal Chem. 2014;86(7):3240–3253
Guiss, L. W., and Kuenstler, P. (1960). A retrospective view of survey photofluorograms of persons with lung cancer. Cancer 13, 91–95. doi: 10.1002/1097-0142(196001/02)13:1<91::AID-CNCR2820130117>3.0.CO;2-K
Ivy, S., Rohovit, T., Stefanucci, J., Stokes, D., Mills, M., & Drew, T. (2023). Visual expertise is more than meets the eye: an examination of holistic visual processing in radiologists and architects. Journal of medical imaging (Bellingham, Wash.), 10(1), 015501. https://doi.org/10.1117/1.JMI.10.1.015501
Ivy, S., Rohovit, T., Lavelle, M., Padilla, L., Stefanucci, J., Stokes, D., & Drew, T. (2021). Through the eyes of the expert: Evaluating holistic processing in architects through gaze-contingent viewing. Psychonomic Bulletin & Review, 28(3), 870-878.
Jacobs, C., Schwarzkopf, D. S., & Silvanto, J. (2018). Visual working memory performance in aphantasia. Cortex, 105, 61-73.
Johnson, J. S., & Olshausen, B. A. (2005). The recognition of partially visible natural objects in the presence and absence of their occluders. Vision research, 45(25-26), 3262-3276.
Jones D. B., Sung R., Weinberg C., Korelitz T., Andrews R. Three-dimensional modeling may improve surgical education and clinical practice. Surgical Innovation. 2016;23(2):189–195. doi: 10.1177/1553350615607641.
Kundel, H. L., Nodine, C. F., & Carmody, D. (1978). Visual scanning, pattern recognition and decision-making in pulmonary nodule detection. Investigative radiology, 13(3), 175–181. https://doi.org/10.1097/00004424-197805000-00001
Kentridge, R. W., Heywood, C. A., & Weiskrantz, L. (1999). Attention without awareness in blindsight. Proceedings of the Royal Society of London. Series B: Biological Sciences, 266(1430), 1805–1811. https://doi.org/10.1098/rspb.1999.0850
Keogh, R., & Pearson, J. (2011). Mental imagery and visual working memory. PloS one, 6(12), e29221. https://doi.org/10.1371/journal.pone.0029221
Klein GT, Lu Y, Wang MY. 3D printing and neurosurgery—ready for prime time? World Neurosurg. 2013;80(3–4):233–235
Kosslyn, S. M., Behrmann, M., & Jeannerod, M. (1995). The cognitive neuroscience of mental imagery. Neuropsychologia, 33(11), 1335–1344. https://doi.org/10.1016/0028-3932(95)00067-4
Kouider, S., & Dehaene, S. (2007). Levels of processing during non-conscious perception: A critical review of visual masking. Philosophical Transactions of the Royal Society B: Biological Sciences, 362(1481), 857–875. https://doi.org/10.1098/rstb.2007.2093
Kundel H. L., Nodine C. F., Conant E. F., Weinstein S. P. (2007). Holistic component of image perception in mammogram interpretation: gaze-tracking study. Radiology 242 396–402. 10.1148/radiol.2422051997
Kundel, H. L., Nodine, C. F., Krupinski, E. A., & Mello-Thoms, C. (2008). Using gaze-tracking data and mixture distribution analysis to support a holistic model for the detection of cancers on mammograms. Academic Radiology, 15(7), 881–886. https://doi.org/10.1016/j.acra.2008.01.023
Kundel H. L., Nodine C. F., Toto L. (1984). “Eye movements and the detection of lung tumors in chest images,” in Theoretical and Applied Aspects of Eye Movement Research eds Gale A. G., Johnson F. (Amsterdam: Elsevier; ) 297–304.
Kundel H. L., Nodine C. F., Toto L. (1991). Searching for lung nodules: the guidance of visual scanning. Invest. Radiol. 26 777–781. 10.1097/00004424-199109000-00001
Lee, A. L. F., Yeung, N., & Summerfield, C. (2021). Global visual confidence. Journal of Experimental Psychology: Human Perception and Performance, 47(6), 960–976. https://doi.org/10.1037/xhp0000885
Lefor, A. K., Iqbal, S., & Ota, K. (2020). The effect of simulator fidelity on procedure skill training: a systematic review. Journal of Surgical Education, 77(5), e160–e172. https://doi.org/10.1016/j.jsurg.2020.06.014
Lindemann, M. C., Glänzer, L., Roeth, A. A., Schmitz-Rode, T., & Slabu, I. (2023). Towards Realistic 3D Models of Tumor Vascular Networks. Cancers, 15(22), 5352. https://doi.org/10.3390/cancers15225352
Litchfield, D., and Donovan, T. (2016). Worth a quick look? Initial scene previews can guide eye movements as a function of domain-specific expertise but can also have unforeseen costs. J. Exp. Psychol. Hum. Percept. Perform. 42, 982–994. doi: 10.1037/xhp0000202
Mamassian, P. (2016). Visual confidence. Annual Review of Vision Science, 2, 467–491. https://doi.org/10.1146/annurev-vision-111815-114630
Massoth, C., Steigerwald, S., Sauter, P., Winkel, M., Lipprandt, M., & Huber, G. (2019). High-fidelity is not superior to low-fidelity simulation but leads to overconfidence — a randomized controlled trial. BMC Medical Education, 19, 130. https://doi.org/10.1186/s12909-019-1464-7
Morra, M., Braund, H., Hall, A. K., & Szulewski, A. (2021). Cognitive load and processes during chest radiograph interpretation in the emergency department across the spectrum of expertise. AEM Education and Training, 5(4), e10693.
Nanay, B. (2023). Mental imagery: Philosophy, psychology, neuroscience. Oxford University Press.
Nanay, B. (2010). Perception and imagination: Amodal perception as mental imagery. Philosophical Studies, 150(2), 239–254. https://doi.org/10.1007/s11098-009-9407-5
Nodine, C. F., and Kundel, H. L. (1987). “The cognitive side of visual search in radiology,” in Eye Movements: From Physiology to Cognition, eds J. K. O’Regan and A. Levy-Schoen (Amsterdam: Elsevier), 573–582.
Oberauer, K. (2019). Working Memory and Attention – A Conceptual Analysis and Review. Journal of Cognition, 2(1), 36. https://doi.org/10.5334/joc.58
Oderda, M., Calleris, G., D’Agate, D., Falcone, M., Faletti, R., Gatti, M., Marra, G., Marquis, A., & Gontero, P. (2023). Intraoperative 3D-US-mpMRI Elastic Fusion Imaging-Guided Robotic Radical Prostatectomy: A Pilot Study. Current Oncology, 30(1), 110–117. https://doi.org/10.3390/curroncol30010009
Oh, S. H., & Kim, M. S. (2004). The role of spatial working memory in visual search efficiency. Psychonomic bulletin & review, 11(2), 275–281. https://doi.org/10.3758/bf03196570
Pearson, J., Naselaris, T., Holmes, E. A., & Kosslyn, S. M. (2015). Mental imagery: Functional mechanisms and clinical applications. Trends in Cognitive Sciences, 19(10), 590–602. https://doi.org/10.1016/j.tics.2015.08.003
Recarte, M. A., & Nunes, L. M. (2003). Mental workload while driving: effects on visual search, discrimination, and decision making. Journal of experimental psychology. Applied, 9(2), 119–137. https://doi.org/10.1037/1076-898x.9.2.119
Reddan, M. C., & Wager, T. D. (2018). Modeling Pain Using fMRI: From Regions to Biomarkers. Neuroscience bulletin, 34(1), 208–215. https://doi.org/10.1007/s12264-017-0150-1
Reingold, E. M., and Sheridan, H. (2011). “Eye movements and visual expertise in chess and medicine,” in The Oxford Handbook of Eye Movements, eds S. P. Liversedge, I. D. Gilchrist, and S. Everling (Oxford: Oxford University Press), doi: 10.1093/oxfordhb/9780199539789.013.0029
Scarry, E. (1985). The body in pain: The making and unmaking of the world. Oxford University Press.
Sheridan, H., & Reingold, E. M. (2017). The Holistic Processing Account of Visual Expertise in Medical Image Perception: A Review. Frontiers in psychology, 8, 1620. https://doi.org/10.3389/fpsyg.2017.01620
Silberstein, J. L., Maddox, M. M., Dorsey, P., Feibus, A., Thomas, R., & Lee, B. R. (2014). Physical Models of Renal Malignancies Using Standard Cross-sectional Imaging and 3-Dimensional Printers: A Pilot Study. Urology, 84(2), 268–273. https://doi.org/10.1016/j.urology.2014.03.042
Stefanucci, J. K., Detrich, A. N., Guttman, K. H., Mills, M., Auffermann, W., Barber, N., Visintainer, K., & Orlosky, J. (2025, February). Training perceptual expertise in radiologists [Poster presentation]. 1U4U Conference, Salt Lake City, UT.
Swensson, R. G. (1980). A two-stage detection model applied to skilled visual search by radiologists. Percept. Psychophys. 27, 11–16. doi: 10.3758/BF03199899
Talanki, V. R., Peng, Q., Shamir, S. B., Baete, S. H., Duong, T. Q., & Wake, N. (2021). Three-dimensional printed anatomic models derived from magnetic resonance imaging data: Current state and image acquisition recommendations for appropriate clinical scenarios. Journal of Magnetic Resonance Imaging, 54(4), 1071–1084. https://doi.org/10.1002/jmri.27744
Thelen, J., Sant Fruchtman, C., Bilal, M., Gabaake, K., Iqbal, S., Keakabetse, T., Kwamie, A., Mokalake, E., Mupara, L. M., Seitio-Kgokgwe, O., Zafar, S., & Cobos Muñoz, D. (2023). Development of the Systems Thinking for Health Actions framework: a literature review and a case study. BMJ global health, 8(3), e010191. https://doi.org/10.1136/bmjgh-2022-010191
Venjakob, A. C., & Mello-Thoms, C. R. (2016). Review of prospects and challenges of eye tracking in volumetric imaging. Journal of medical imaging (Bellingham, Wash.), 3(1), 011002. https://doi.org/10.1117/1.JMI.3.1.011002
Ventola C. L. Medical applications for 3D printing: current and projected uses. Pharmacy and Therapeutics. 2014;39(10):704–711.
Waite, S., Grigorian, A., Alexander, R. G., Macknik, S. L., Carrasco, M., Heeger, D. J., & Martinez-Conde, S. (2019). Analysis of Perceptual Expertise in Radiology – Current Knowledge and a New Perspective. Frontiers in human neuroscience, 13, 213. https://doi.org/10.3389/fnhum.2019.00213
Williams, L. H., Carrigan, A. J., Mills, M., Auffermann, W. F., Rich, A. N., & Drew, T. (2021). Characteristics of expert search behavior in volumetric medical image interpretation. Journal of Medical Imaging, 8(4), 041208-041208.
Williams, L. H., & Drew, T. (2019). What do we know about volumetric medical image interpretation?: a review of the basic science and medical image perception literatures. Cognitive research: principles and implications, 4(1), 21. https://doi.org/10.1186/s41235-019-0171-6
Wirecutter. (n.d.). The best standalone VR headset. The New York Times. Retrieved October 27, 2025, from https://www.nytimes.com/wirecutter/reviews/best-standalone-vr-headset/
Wolfe, J. M. (2012). Journal of Experimental Psychology: General. DOI:10.1037/a0027406. (See related PMC: 2013)
Zachariou, V., Klatzky, R., & Behrmann, M. (2014). Ventral and dorsal visual stream contributions to the perception of object shape and object location. Journal of cognitive neuroscience, 26(1), 189–209. https://doi.org/10.1162/jocn_a_00475
Zhang, J., & Wang, H. (2009). An Exploration of the Relations between External Representations and Working Memory. PLoS One, 4(8)https://doi.org/10.1371/journal.pone.0006513
- 1. An important colloquial point here is that Nanay uses perceptual processing and perceptual representation interchangeably. ↵