Authors: Sarah Carneiro Pereira, Gilles Vannuscorps
Categories: Article, Facial expression, Motor simulation, Action recognition, Audio-visual integration, Human behaviour, Emotion, Visual system
Source: Scientific Reports
Authors: Sarah Carneiro Pereira, Gilles Vannuscorps
Efficient recognition of facial expressions is often assumed to require covert and unconscious imitation of the observed facial movements – a “motor simulation”. At odds with this hypothesis, some individuals with congenital facial paralysis can recognize facial expressions with typical speed and accuracy. However, efficiency also implies minimal cognitive effort. Motor simulation may reduce the need for effortful inferential processing. To explore this, we asked 16 individuals born with congenital facial paralysis and typically developed participants to categorize emotional sounds (e.g. anger vs. fear) while viewing congruent or incongruent facial expressions. The facial expressions were irrelevant to the task, but facial identity was important for a subsequent memory test. If motor simulation is necessary for effortless facial expression recognition, individuals born with congenital facial paralysis should not be influenced by the facial expressions. If motor simulation contributes to effortless facial expression recognition without being necessary, they could be less influenced than the typically developed participant. However, the results did not support either of these predictions. Effortless facial expression recognition can be achieved without motor simulation.
Much of human communication is nonverbal. Body movements and facial expressions convey emotions, intentions, and reactions allowing individuals to navigate social interactions, foster emotional connections, and understand one another quickly and efficiently. The ability to interpret these cues accurately, rapidly, and with minimal cognitive effort, is essential for effective social functioning. To date, the cognitive mechanisms underlying the efficient recognition of body movements and facial expressions remain incompletely understood. A key question concerns the role of motor simulation in interpreting these nonverbal signals.
Motor simulation refers to the activation of imitative motor plans in the observer’s brain when they observe facial and body postures/movements^1–5^. Motor simulation has been proposed to contribute to action^6–10^ and facial expression recognition^11–16^. However, whether, and if so when and how, motor simulation contributes to the recognition of body and facial movements remains unclear.
The interpretation of neuroimaging, behavioral, and transcranial magnetic stimulation (TMS) studies that have been conducted to address these questions remains unsettled (see for discussion^17,18^). At first sight, one might expect clearer evidence to come from studies examining how damage to the brain areas involved in motor simulation affects individuals’ ability to recognize others’ actions and expressions. However, findings from brain-damaged patients are inconsistent and often contradictory. Some studies report that motor impairments are associated with difficulties in recognizing actions and facial expressions^19–22^, but other find no such impairments^23–25^. The reasons for these discrepancies remain unclear. They may stem from differences in the location and extent of motor system damage (e.g., specific regions within the parieto-frontal circuit or primary motor cortex; see e.g^26^), the involvement of non-motor brain regions (e.g., areas responsible for visuo-spatial or conceptual processing), or the nature of the resulting functional impairments (e.g., various forms of apraxia, executive dysfunction, or visuo-spatial deficits). Variability in the action recognition tasks themselves—ranging from basic motion detection to distinguishing correct from incorrect actions or identifying specific actions—may also play a role, as could differences in task sensitivity and the type of visual stimuli used (e.g., videos, photographs, pantomimes, or point-light displays). A key challenge in interpreting this evidence is that theories proposing a role for motor simulation in action recognition do not specify in detail what brain areas support motor simulation, nor do they predict the nature and severity of deficits that should result from damage to these areas.
In the last ten years, we and others have sought to overcome these ambiguities by systematically testing a series of specific hypotheses about possible contributions of motor simulation through the study of individuals who, due to congenital conditions such as facial paralysis or absent limbs, cannot covertly imitate observed postures and movements involving these effectors^12,18,25,27–42^. The results of these studies have provided insights into the boundaries of motor simulation’s involvement, suggesting that while it is not universally required across all tasks, it may support certain aspects of action and facial expression recognition.
Individuals born without upper limbs exhibited specific difficulties in tasks requiring them to maintain hand postures in memory for short periods, suggesting a role for motor simulation in working memory^39^. Moreover, they struggled to recognize manual actions presented as point-light displays, suggesting that motor simulation may support action recognition when the stimuli are unfamiliar^36,40^. However, these individuals were able to perceive, anticipate, and recognize upper limb actions depicted in more familiar formats, such as pictures and video-clips, as accurately and quickly as typically developed individuals^25,36^.
Interpreting the collective results of previous studies on individuals with congenital facial paralysis is somewhat more complicated. This condition often arises in the context of Moebius Syndrome (MS), which typically affects not only the facial nerves involved in facial expression execution, but also the individuals’ visual, perceptual, cognitive, and social abilities to various extents^18,27,43^. This makes it difficult to interpret the findings of previous studies reporting that individuals with Moebius Syndrome (IMS) react to facial expressions differently than normotypical participants when no control task is included^12,34^.
Indeed, studies that include control tasks often show that IMS struggle not only with facial expression recognition but also with these control tasks. For example, in Bate et al., five out of six IMS who were impaired on at least one of three facial expression recognition tasks also showed deficits in facial identity recognition and/or tests assessing low-level vision and object recognition^27^. Similarly, Japee et al. (2023) found that IMS required slightly longer stimulus exposure not only to categorize facial expressions but also to recognize bodily expressions^32^.
These findings suggest that facial expression recognition difficulties in IMS stem, at least in part, from general visuo-perceptual deficits rather than from facial paralysis alone. Supporting this, in a previous study, we found a strong and significant correlation between IMS’ facial expression recognition performance and mid-level perceptual abilities (three rs > 0.5, ps < 0.05), but no correlation with the paralysis severity^18^. Furthermore, two IMS with very severe facial (1) showed no sign of difficulty in five sensitive facial expression recognition experiments; (2) reached a performance in these tasks that equaled or exceeded their results in control tasks (face identity recognition and vocal emotion recognition); and (3) performed these tasks in a way that was qualitatively similar to that of neurotypical controls. Together with findings from previous studies on lip reading^37^ and facial expression recognition^28,29^, these results indicate that fast and accurate facial expression recognition is possible without motor simulation. Of course, this does not rule out a potential role for motor simulation in other contexts. Indeed, recent studies suggest that motor simulation may help detect subtle differences between facial expressions^33,35^.
In sum, these results motivate a nuanced view according to which motor simulation may contribute to some specific aspects of action and facial expression recognition without being necessary across all tasks and stimuli. This highlights the complexity of this issue and the need for precise, targeted investigations into the specific roles that motor simulation may—or may not—play in each case. The aim of this study was to continue to progress in this direction by testing the specific hypothesis according to which motor simulation, although not the only mechanism available to recognize actions, may be the only mechanism which “allows the observer to understand the actions of others without needing inferential processing”^10^ (p. 268); see also^44,45^. Hence, according to this view, there are two potential routes to action one is a direct, non-inferential route grounded in motor simulation, which allows for fast and automatic recognition by mapping observed actions onto one’s own motor representations and the other is an inferential route, which relies on deliberate, effortful reasoning and the conscious interpretation of perceptual cues.
To test this hypothesis, we asked individuals born with congenital facial paralysis and typically developed participants to complete three tasks. In the first task, they were asked to categorize emotional sounds (e.g. anger vs. fear) while concurrently viewing congruent or incongruent facial expressions. Importantly, the facial expressions were irrelevant to the participants were told to focus on the voices, and the only instruction regarding the faces was that they might be asked to recall the identity of the actor later, as part of a memory test.
This allowed us to examine whether facial expressions are recognized automatically, that is, quickly, effortlessly, and without deliberate effortful inferential any influence of facial expressions on participants’ categorization of the sounds—manifested as a congruency effect (i.e., better performance when the emotion in the face matches that in the voice)—would suggest that facial expression recognition occurred incidentally, without requiring inferential processing. This logic parallels that used in Stroop-like paradigms, where the automatic processing of irrelevant stimulus features (e.g., reading the word “red” even when naming the ink color) is inferred from interference or facilitation effects. In our case, if participants are affected by the facial expressions they were not instructed to recognize, this indicates that these expressions were processed automatically, through a route that does not rely on inferential processing.
Thus, this first task allowed to test three distinct theoretical hypotheses regarding the role of motor simulation in effortless facial expression recognition, each associated with a specific analytical strategy. First, if motor simulation is necessary for effortless recognition of facial expressions, then individuals with Moebius Syndrome (IMS) should not show a congruency effect. That is, their performance should not differ between trials in which the facial expression is congruent versus incongruent with the emotion conveyed by the voice. This prediction will be tested by conducting individual-level analyses for each IMS participant comparing their performance in congruent and incongruent trials.
Second, if motor simulation contributes to effortless facial expression recognition without being strictly necessary, IMS participants might still show a congruency effect, but it should be smaller than that observed in typically developed control participants. This prediction will be tested by comparing the magnitude of the congruency effect between groups using Bayesian analyses in order to assess the relative likelihood of a greater effect in controls (in line with the hypothesis that motor simulation contributes to effortless facial expression recognition) versus the null hypothesis assuming no difference between groups.
Third, if motor simulation is neither necessary nor contributing to effortless facial expression recognition, then congruency effects should not be smaller in the IMS than in the controls. In this case, automatic facial expression recognition would appear to occur independently of motor simulation mechanisms.
The second and third tasks were designed to clarify the results of the first. A facial identity task assessed participants’ memory of the identities of the actors and actresses from the dual task, ensuring they had paid attention to the images—a prerequisite for interpreting a putative failure to detect an influence of facial expressions in the dual task. Finally, an emotional sound categorization task evaluated participants’ ability to recognize emotional sounds in isolation, ensuring that the influence of facial expressions in the dual task, if any, could be not attributed to a strategic use of the visual stimulus caused by the impossibility of performing the task based solely on auditory input.
To assess whether motor simulation is necessary for effortless facial expression recognition, we planned to conduct individual analyses testing for differences in performance between congruent and incongruent trials in each individual with Moebius Syndrome (IMS). Given the extremely low prevalence of this condition (estimated at 1 in 50,000 to 1 in 500,000 newborns^46^), our goal was to recruit as many IMS as possible over a 24-month period. To maximize recruitment, we adopted a flexible testing approach that allowed for either in-person or remote participation, depending on each individual’s availability and ability to travel.
Sixteen individuals with congenitally severely reduced or completely absent facial movements in the context of Moebius Syndrome (MS, see Table 1) eventually participated in the study. Among them, three (IMS 1–3: 3 woman, 2 right-handed, mean age = 41, SD = 23, age 19–65 years) were tested in person and thirteen (IMS 4 to IMS 8 women; 13 right-handed; mean age = 34.8; SD = 15.1; age 17–62 years) were tested online.
The size of the sample of IMS that we succeeded in recruiting (N = 16) provided a unique opportunity to further investigate whether motor simulation may contribute to facial expression recognition without being strictly necessary, by comparing the magnitude of the congruency effect in this group to that observed in typically developed controls. We aimed to recruit approximately 50 control participants, with at least 25 tested in person and 25 remotely, to ensure that the testing modality did not influence significantly the size of the congruency effect. Although this control sample size was not determined by a formal power analysis, it was considered a suitable compromise between practical constraints and the study’s objectives to obtain a stable estimate of baseline performance and to yield informative Bayes factors for the directional hypothesis (H₁: controls exhibit a larger congruency effect).
Fifty-six typically developed control participants eventually participated in the study. Among these control participants, one was excluded prior to data analysis because of a clear lack of focus during testing, and three were excluded after their data was analyzed because their response profile in one of the tasks suggested a misunderstanding of the instructions (see data analysis section). This resulted in a final sample of 52 control participants (34 women; 3 left-handed and 2 ambidextrous; mean 29.1; SD: 12.4; age 18–64 years), 27 tested in person (19 woman; 2 left-handed; mean 22.5; SD: 3.1; age 18–29 years) and 25 tested online (16 woman; 2 ambidextrous and 1 left-handed; mean age = 36.28.6; SD = 14.6; age 18–64 years).
Information on participants’ sex assigned at birth, native and schooling language, hand-laterality, educational level, motor, auditory, psychiatric, and neurological history was obtained through an interview. Self-report about the presence/absence of hearing loss has been shown to have a predictive value of 82% for mild hearing loss, 98% for moderate and 100% for marked hearing loss^47^. For participants with Moebius Syndrome (MS), additional questions were included to assess the presence of associated malformations and surgical history related to Moebius Syndrome.
Participants were all either native French speakers or had completed their entire schooling in French. None of the control participants reported a history of motor, psychiatric, or neurological disorders, and all had normal or corrected-to-normal vision with no hearing loss. In contrast, most IMS participants exhibited one or more clinical features typically associated with their condition. As shown in Table 1, most IMS presented with oculomotor paralysis and three of them showed signs of mild hearing impairment. In addition, six IMS reported some neurological history including head trauma during childhood (n = 2), narcolepsy (n = 1), peripheral neuropathy with hypersensitivity to touch (n = 1), seizure episodes during childhood (n = 1) or during adulthood (n = 1). Three IMS reported past history of depression (not present at the time of testing). Two IMS were born without a left hand. Moreover, all but two IMS reported a history of surgeries linked to MS, including hand syndactyly (n = 2), clubfoot (n = 6) and strabismus (n = 4) correction, stump (n = 1) and hip (n = 1) surgeries, arthrodesis (n = 1), implant for eyelid closure (n = 2), lachrymal duct surgery (n = 4), tongue frenectomy (n = 1) or frenuloplasty (n = 1), bilateral myoplasty (n = 4), nerve transfer in the mouth area (n = 1) and facial reconstruction surgery (n = 1).
To protect participants’ anonymity, we chose not to specify which IMS had which neurological or psychiatric condition, malformation, or surgical history. This level of detail was not essential for interpreting the study’s findings or assessing the validity of our conclusions, and given the rarity of Moebius Syndrome, disclosing such information to the general public could potentially make individual participants identifiable and unnecessarily expose private health details. However, this information is available upon request to fellow researchers.
Table 1Demographic and clinical information about the IMS.Sex assigned at birthAge (years)Hand lateralityEducation (years)VisionHearing lossIMS01F39Left14Corrected-to-normalNone reportedIMS02F19Right13NormalNone reportedIMS03F65Right17Corrected-to-normalNone reportedIMS04M19Right13Corrected-to-normal; No eye abductionNone reportedIMS05F17Right11Corrected-to-normal; No eye abductionNone reportedIMS06F33Right15Corrected-to-normal; No eye abductionNone reportedIMS07M62Right14Corrected-to-normal; No eye abductionSelf-reported hearing difficulties only in noisy environments (left ear)IMS08M19Right8Corrected-to-normal; No eye abductionNone reportedIMS09F43Right15Corrected-to-normal; No eye abductionNone reportedIMS10F45Right16Corrected-to-normal; No eye abductionHearing impairment corrected with a hearing aidTinnitusIMS11F50Right12Corrected-to-normal; No eye abductionMild hearing lossTinnitusIMS12F26Right12Corrected-to-normalNone reportedIMS13F19Right13Corrected-to-normal; No eye abductionNone reportedIMS14M54Right20Corrected-to-normal; No eye abductionNone reportedIMS15M39Right15Corrected-to-normalNone reportedIMS16F27Right14Corrected-to-normal; No eye abductionNone reported
Information about the IMS participants’ ability to perform facial expressions was obtained by assessing their ability to execute all the Action Units (AUs) governed by facial nerve function involved in producing the six primary facial expressions—joy, sadness, anger, fear, surprise, and disgust^48,49^. The IMS were shown animated depictions of each AU^49^, and asked to imitate the movements. Their imitations were recorded, and the recordings were reviewed by the first author. Each AU performed was then categorized into one of four (1) Absent, when no relevant movement was detected; (2) Very limited, when only minimal movement was observed (i.e., a few millimeters); (3) Limited, when the movement was not complete but more marked than the very limited ones; (4) Complete, when the AU was marked bilaterally. The results are shown in Fig. 1. A detailed description of the IMS’s ability to produce the Action Units categorized as “very limited” and “limited” is available on OSF (https://osf.io/vkt42/).
Fig. 1IMS’s ability to execute action units.
The study was approved by the biomedical ethics committee of the Cliniques Universitaires Saint-Luc, Brussels, Belgium (#B403201835404). All participants, as well as IMS05’s legal guardians, provided informed consent prior to the study. All methods were performed in accordance with the relevant guidelines and regulations, and in accordance with the Declaration of Helsinki.
Stimuli were 136 pictures, 48 sounds and 96 bimodal stimuli. This set of stimuli allowed us to collect experimental data from 40 audio stimuli presented in both congruent and incongruent conditions, which provided statistical power slightly above 90% to detect either (a) a minimum asymmetry of 10% in error rates between congruent and incongruent conditions, estimated using a one-tailed exact McNemar test at α = 0.05 with the powerPaired2×2() function from the exact2×2package (version 1.6.9^50^; in R (version 4.4.3^51^;, or (b) an effect size of at least Cohen’s d = 0.5 in response latencies, estimated using a one-tailed paired t-test at α = 0.05 with G*Power (version 3.1.9.7).
The pictures were selected from the FACES database^52^. The original pictures, which included the full head and neck, were transformed to remove any potential information from neck muscle contractions. Only the facial contour remained visible in the final stimuli (800 × 800 pixels), ensuring that facial recognition relied exclusively on facial features.
The pictures were divided in three sets. The “familiarization” and “test” sets included respectively 16 (2 genders x 2 stimuli x 4 emotions) and 80 (2 genders x 10 stimuli x 4 emotions) pictures of different actors and actresses expressing a clear and prototypical facial expression of either anger, fear, disgust or sadness. The “memory” set included 48 pictures of different actors (2 genders x 20 stimuli) with a neutral facial expression. Twenty-four of these pictures (the “old” pictures) depicted actors and actresses also included in the familiarization (2 genders x 2 stimuli) and test set (2 genders x 10 stimuli). The twenty-four other pictures (the “new” pictures) were of actors and actresses that were not included in the other sets.
The auditory stimuli were non-verbal affective bursts of the French vowel /a/ of variable duration. Forty “test” stimuli (mean duration = 1.18 s; SD = 0.8) were sourced from the Montreal Affective Voices database^53^ and featured five actors and five actresses, each recorded to convey anger, fear, disgust, and sadness (2 genders x 5 stimuli x 4 emotions). Eight similar “familiarization” stimuli (mean duration = 1 s; SD = 0.3) were recorded locally by one actor and one actress, following the same recording procedure as those from the MAV database (2 genders x 4 emotions).
The bimodal stimuli were created by pairing the audio files with pictures matched in gender. Sixteen “familiarization” bimodal stimuli were generated by combining each of the eight “familiarization” audio files with both a congruent and an incongruent picture from the “familiarization” set. Similarly, 80 “test” bimodal stimuli were created by pairing each of the 40 “test” audio files with one congruent and one incongruent image from the “test” picture set. Congruent bimodal stimuli paired an audio file with an image depicting the same emotion. In contrast, incongruent bimodal stimuli paired an audio file with an image representing a different fear-related audio files were associated with pictures of anger, sadness-related audio files were paired with pictures of disgust, and vice versa.
Stimuli from databases may be obtained directly from the references cited in the text. The stimuli created for this study, as well as the filenames of the stimuli selected from the databases, are available on OSF (https://osf.io/vkt42/).
Participants performed a dual task, involving concurrent emotional sound categorization and facial identity memorization, and an emotional sound categorization task, in that order.
The dual task was conducted both in person and online. For the participants tested in person, the experiment was controlled by Psychopy v.2023.2.3^54^ and run on 15.6inch DELL Latitude laptop operated by Windows 10. Online testing was controlled by the testable.org interface (http://www.testable.org), which allows for precise spatiotemporal control of both online and offline experiments. Participants were tested on their own computer under supervision of the experimenter through a video conference system. The procedure was otherwise the same for all participants. The task was divided into two the first contrasted fear versus anger, and the second contrasted disgust versus sadness.
Each block began with a brief introduction to the type of audio files. Participants were successively presented with the four “familiarization” audio stimuli corresponding to the emotions contrasted in that block.
Participants were then briefed on the dual-task and completed a brief training session. In each of eight trials, participants were presented with one “familiarization” bimodal stimulus corresponding to the emotions contrasted in that block (e.g., anger and fear in block 1). Pictures subtended approximately 20° size of visual angle and were presented on the center of the screen until response. They were asked to perform two tasks simultaneously. The first was to decide as accurately and as fast as possible which of the two emotions contrasted in the block the sound corresponded to (e.g., anger or fear) and to press the corresponding keyboard key (left arrow for anger and sadness; right arrow for fear and disgust). Participants responded with their left and right index fingers, except two IMS who have no left hand. They used either the thumb and middle finger of the right hand, or the middle and little finger of the right hand. In addition, participants were instructed to pay attention to the identity of the actor depicted on the screen, which they were told they would need for a later memory test. Each response provided by participants was followed by a central cross lasting either 500ms (online) or 1000ms (on-site), and then the next trial.
After the dual task training phase, participants were involved in a short facial identity recognition task. In each of the four trials, participants saw one picture from the “memory” set (two “old” and two “new”) and were asked to decide whether they had already seen the actor depicted on the picture by pressing the corresponding keyboard key (left and right arrow for “yes” and “no”, respectively). The focus was on accuracy, not speed. After each response, a fixation cross appeared for 500ms (online) or 1000ms (on-site) before the next trial was initiated.
Then, participants had the opportunity to ask any questions before placing their fingers on top of the arrow’s keys and pressing the spacebar to begin the dual-task experiment. Pressing the space bar initiated a three second countdown followed by the beginning of the task. The task and procedure were identical to the training phase, except that participants were presented with the 40 test bimodal stimuli corresponding to the emotions contrasted in the block (e.g., anger and fear).
Finally, participants took part in a longer facial identity recognition task involving 10 “old” and 10 “new” actors and actresses from the “memory” set. The procedure was the same as that of the shorter task. The end of this task at the end of the first block initiated the second block, which followed the same procedure.
Across all these phases, blocks and tasks, stimuli were presented in pseudo-random order, identical across participants. In addition, in the first block of the dual-task an equal number of audio files appeared first in congruent and incongruent pairings. In the second block, however, an issue occurred during the pseudo-randomization process and 11/20 audio files appeared first in an incongruent pairing.
The emotional sound categorization task was controlled by the testable.org interface (http://www.testable.org). For the participants tested in person, the experiment was run on 15.6inch DELL Latitude laptop operated by Windows 10. Online participants were tested on their own computer under supervision of the experimenter through a video conference system. The procedure was otherwise the same for all participants.
The task was divided in two the first contrasted fear versus anger, and the second contrasted disgust versus sadness. Each block began with a familiarization phase, during which participants listened to the four “familiarization” audio files corresponding to the emotions being contrasted in that block (e.g., fear and anger), with an inter-stimulus interval (ISI) of 500 ms. Participants were then briefed on the task and completed a brief training session, which included the same four “familiarization” audio files. During this training, each sound was presented twice. Participants were instructed to determine which of two emotions the sound conveyed (anger or fear in the first block, disgust or sadness in the second) and to select the corresponding option displayed on the computer screen. The focus was on accuracy, not speed. Following each response, a central cross was displayed on the screen for 500 ms, followed by the next sound. After the training phase, participants had the opportunity to ask any questions before pressing the spacebar to begin the experiment. The task and procedure were identical to the training phase, but this time participants heard the 20 “test stimuli” corresponding to the emotions contrasted in the block. All stimuli were presented in a pseudo-random order that was kept constant across participants.
The first step in data analysis was to examine individual participant data for each task to ensure accurate data encoding and identify potential anomalies. As a result of this quality check, three control participants were two for providing no correct responses in the second block of the dual task, and one for scoring below 15% accuracy in the incongruent condition.
Following this, we conducted four sets of planned analyses. The first three focused on examining the congruency effect from the results of the dual task, while the final set addressed performance in the facial identity and emotional sound categorization tasks included to help the interpretation of the main results.
The first set of analyses, conducted in R (version 4.4.3^51^), addressed our main research question by testing the congruency effect on each participant’s correct responses (CRs) and response latencies (RLs) in the dual task. For CRs, a series of exact McNemar’s tests were run to compare the proportion of correct and incorrect responses across paired (congruent and incongruent) trials (analysis scripts are available on https://osf.io/vkt42/).
RL analyses were conducted on correct trials, which represented 86.5% of trials for the control group and 81.9% for the IMS group. In addition, bimodal stimuli for which a participant’s RL difference between the two conditions (congruent and incongruent) deviated by more or less than three median absolute deviations (MAD-robust) from the median of the RL difference of that participant were considered outliers and discarded (5.54% and 5.07% of trials for controls and IMS, respectively). Finally, depending on the distribution of the data—assessed via the Shapiro-Wilk test and visual inspection of Q-Q plots—paired t-tests or Wilcoxon signed-rank tests were applied.
The large sample of IMS that we succeeded in recruiting provided a unique opportunity to investigate whether motor simulation contributes to facial expression recognition without being strictly necessary, by comparing the magnitude of the congruency effect in this group to that observed in typically developed controls. As a preliminary step, we examined whether the mode of participation (online or in person) had an influence on the size of the congruency effect. This analysis was crucial to determine whether data from both conditions (and groups) can be pooled or must be treated separately.
To do so, we first computed each participants’ congruency effect over both correct responses (CR) and response latency (RL). The congruency effect on CR was calculated by subtracting the number of CR in the incongruent condition from the number CR in the congruent condition for each participant. Therefore, a higher positive difference indicated a stronger congruency effect on accuracy. RL analyses were conducted on correct trials only (87% of trials for the onsite control group and 86% for the online control group). In addition, trials for which a participant’s RL difference between the two conditions (congruent and incongruent) deviated by more or less than three median absolute deviations (MAD-robust) from the median of the RL difference of that participant were excluded (6% of trials for the onsite control group and 5% for the online control group). Then, the congruency effect on RL was calculated by subtracting the mean RL in the congruent condition from the mean RL in the incongruent condition for each participant. Therefore, a positive difference between conditions indicates a stronger congruency effect on response latencies (i.e., longer response latencies in the incongruent condition).
Two ANCOVAs were conducted in IMB SPSS Statistics (version 27) to assess whether the mean congruency effect on accuracy (CR) and response latency (RL) differed between control participants tested online versus in person, while controlling for age, education, and gender. All model assumptions were checked and met. As no significant differences were found between testing conditions (see Results for details), data from both groups were pooled for the group-level analyses.
The third set of analyses aimed to compare the magnitude of the congruency effect in the IMS group to the baseline established by the control participants. Congruency effects on accuracy (CR) and response latency (RL) were calculated as described in the “Individual Analyses” section.
Before turning to these analyses, it was important to identify and remove potential outliers’ participants that could disproportionately influence the results of the controls and IMS. Based on group-level outlier detection, two control participants were excluded due to congruency effect values on CR that exceeded ± 3 median absolute deviations (MADs) from the group median. Additionally, two other control participants and IMS 15 were excluded for the same reason based on their congruency effect values on RL.
These statistical analyses were performed using the BayesFactor package (version 0.9.12–4.7^55^) from R (version 4.4.3^51^). Planned Bayesian independent samples t-tests were conducted on both CR and RL to compare the strength of evidence for the hypothesis that the Control group presents a larger congruency effect than the IMS group (H₁: Control > IMS) to the strength of the evidence for the alternative (H₀: Control = IMS).
Finally, we counted the number of correct responses of each participant in the facial identity and emotional sound categorization tasks.
All the data and analysis scripts reported in this article are available on OSF (https://osf.io/vkt42/). A series of control analyses was performed by rerunning all analyses after excluding the two last audio stimuli that were presented first in the incongruent condition from Block 2 to achieve perfect counterbalancing of the presentation order and condition (see Procedure section). The results of these additional analyses did not differ meaningfully from those reported in the Results section (see Appendix A and B).
Figure 2 displays participants’ mean response latency and number of correct responses for the congruent and incongruent conditions. To provide an indication of the task’s ability to reliably elicit congruency effects in typically developing individuals, we first examined the performance of control participants. The results indicated that 15/52 control participants (28%) showed significantly more correct responses, and 24/52 control participants (46%) showed significantly shorter response latencies in the congruent condition (all ps < 0.05). Overall, 34 control participants (65%) showed a congruency effect on either CR or RLs – a proportion that significantly exceeded what would be expected by chance (5% false positives under the null; both binomial tests ps < 0.001).
Turning to the question of interest, analyses of the IMS revealed that 9/16 IMS participants (56%) had a significantly larger number of correct responses, and 6/16 IMS participants (37.5%) had significantly shorter response latencies in the congruent condition (all ps < 0.05). Overall, 13 IMS (81%) showed a congruency effect on either CR or RLs – a proportion that significantly exceeded what would be expected by chance (5% false positives under the null; both binomial tests ps < 0.001).
To ensure the robustness of these individual-level findings in the context of multiple comparisons, we also applied a False Discovery Rate (FDR) correction^56^ to the results of the McNemar and paired-sample tests within each group. This more conservative approach yielded a consistent pattern of results, with a substantial proportion of participants in both groups still showing significant congruency effects (56% of the IMS and 40% of the controls; see Table 2 for detail). This indicates that motor simulation is not necessary for effortless, automatic facial expression recognition. The results (including test statistics and p-value) are available for each participant in Table 2.
Fig. 2Participants’ accuracy and response latency in the two conditions of the dual task. Note. Each asterisk indicates a unilateral p-value < 0.05. (A) For accuracy, participants are ordered from the smallest to the largest p-value. (B) For response latency, participants are displayed in the same order as in (A) to ensure participants’ alignment across both measures.
Table 2Individual results.ParticipantCongruentIncongruent t dfW p-value (uni.)Cohen’s dCRRLCRRLCRRLIMS14401.34 ± 0.63181.35 ± 0.580.1816 < 0.001* 0.430.02IMS12391.2 ± 0.52221.36 ± 0.541.0219 < 0.001* 0.160.29Control391.48 ± 0.48231.59 ± 0.530.820 < 0.001* 0.220.2Control380.92 ± 0.35251.13 ± 0.362.9822 < 0.001*
0.003* 0.59IMS01391.19 ± 0.26261.17 ± 0.35− 0.2224 < 0.001* 0.58− 0.05Control381.45 ± 0.4271.48 ± 0.330.3425 < 0.001* 0.370.07IMS09401.55 ± 0.54291.83 ± 0.492.5727 < 0.001*
0.008 0.55IMS16392.26 ± 1.14282.63 ± 1.281.5426 < 0.001* 0.0680.31IMS07352.6 ± 1.07242.78 ± 1.110.5822 0.002* 0.290.16Control380.95 ± 0.33291.02 ± 0.321.2128 0.002* 0.120.24Control390.99 ± 0.2291.04 ± 0.21.2224 0.003* 0.120.26Control390.91 ± 0.21291.04 ± 0.252.127 0.003*
0.023 0.53IMS08330.78 ± 0.2230.9 ± 0.211.8520 0.006*
0.040 0.62Control370.94 ± 0.25300.96 ± 0.270.2426 0.008 0.410.05IMS02391.52 ± 0.49321.59 ± 0.46263 0.008* 0.39Control381.52 ± 0.54301.6 ± 0.590.6328 0.011 0.270.14Control391.06 ± 0.4331.17 ± 0.41.5629 0.016 0.0650.27Control381.15 ± 0.37321.5 ± 0.443.6931 0.016
< 0.001* 0.87Control370.98 ± 0.19311.09 ± 0.32.8229 0.016
0.004* 0.42Control370.93 ± 0.27300.98 ± 0.270.8727 0.019 0.200.17Control381.35 ± 0.46331.51 ± 0.61.6230 0.031 0.0580.28Control371.07 ± 0.25321.19 ± 0.451.830 0.031
0.041 0.3Control390.97 ± 0.29331.06 ± 0.371.2431 0.035 0.110.27IMS13361.72 ± 0.49301.74 ± 0.580.1127 0.035 0.460.03Control380.86 ± 0.28331.02 ± 0.262.93300.063 0.003* 0.57Control371.25 ± 0.46331.32 ± 0.430.61320.0630.270.16Control391.23 ± 0.41351.41 ± 0.362.16330.063 0.019 0.45Control371.22 ± 0.36321.16 ± 0.331550.0630.58Control370.95 ± 0.34320.97 ± 0.340.33270.0900.370.07Control391.02 ± 0.33351.05 ± 0.340.52330.110.3040.09Control340.87 ± 0.32301.02 ± 0.292.54270.11 0.009* 0.5Control390.95 ± 0.19361 ± 0.171.6350.130.0590.25Control381.21 ± 0.38351.27 ± 0.370.92340.130.180.15Control401.56 ± 0.41371.73 ± 0.631.87360.13 0.035 0.3Control371.11 ± 0.36341.33 ± 0.432.7330.13 0.005* 0.54Control371.25 ± 0.3341.38 ± 0.421.53330.130.0680.34Control381.43 ± 0.43341.53 ± 0.471.33300.150.0960.23Control381.08 ± 0.36341.21 ± 0.461.72310.15 0.048 0.33IMS03370.98 ± 0.28331.19 ± 0.492.16270.15 0.02 0.5Control361.13 ± 0.3321.19 ± 0.310.8250.170.220.2Control351.26 ± 0.4311.38 ± 0.41.42220.170.0850.32Control310.99 ± 0.22271.03 ± 0.190.67190.170.260.21Control360.89 ± 0.25320.94 ± 0.240.93230.170.180.21IMS05361.38 ± 0.67321.56 ± 0.741.12260.170.140.26IMS06371.48 ± 0.51331.69 ± 0.541.94290.17 0.031 0.4Control381.11 ± 0.39351.35 ± 0.373880.19 < 0.001* IMS10371.05 ± 0.39341.48 ± 0.53.76290.19 < 0.001* 0.96Control340.96 ± 0.3300.81 ± 0.25−2.15220.210.98− 0.57Control361.48 ± 0.44331.81 ± 0.643.34300.23 0.001* 0.58Control361.09 ± 0.33331.1 ± 0.350.12290.230.450.03Control360.75 ± 0.17330.91 ± 0.24.66270.23 < 0.001* 0.86Control371.03 ± 0.25351.13 ± 0.361.84340.25 0.037 0.33Control370.84 ± 0.28350.94 ± 0.263.19310.25 0.002* 0.35Control381.18 ± 0.34361.36 ± 0.412.45290.25 0.010* 0.48Control381.29 ± 0.26361.28 ± 0.23− 0.34310.250.63− 0.05IMS04340.73 ± 0.22310.84 ± 0.21.83240.27 0.040 0.57Control391.09 ± 0.44371.34 ± 0.612.43340.31 0.010* 0.45Control351.24 ± 0.45331.35 ± 0.371.11300.310.140.25Control351.31 ± 0.47331.45 ± 0.532.18240.31 0.019 0.28IMS11361.5 ± 0.43341.57 ± 0.441.07290.340.150.17Control360.91 ± 0.35341.17 ± 0.592.75270.36 0.005* 0.49Control370.81 ± 0.23350.79 ± 0.22− 0.41300.360.66− 0.07IMS15321.55 ± 0.93301.35 ± 0.6− 0.84210.390.80− 0.25Control272.04 ± 0.83252.26 ± 1.040.86180.400.200.24Control321.19 ± 0.37311.38 ± 0.461.9270.50 0.034 0.45Control370.97 ± 0.21361.05 ± 0.281.79320.50 0.041 0.34Control361.02 ± 0.24361.16 ± 0.342.76290.50 0.005* 0.45Control352.26 ± 1.05352.57 ± 1.162.7320.50 0.005* 0.37Note. For each participant, the table reports the number of correct responses (CR) in the congruent and incongruent conditions, along with the mean and standard deviation (SD) of response latencies (RL). Test statistics are reported as t and df for the paired t-test, and W for the Wilcoxon test. The “p-value (uni.)” column corresponds to the unilateral p-value obtained from the CR analyses (exact McNemar test) and RL analyses (paired t-test or Wilcoxon test). Significant p-values are shown in bold. Asterisks indicate p-values that remain significant after applying the FDR correction^56^ within each group. Effect sizes are reported for the paired t-test.
The ANCOVAs conducted to test whether the congruency effect on accuracy (CR) and response latency (RL) differed between testing conditions (online vs. onsite) did not reach statistical significance (CR: F(4, 47) = 1.21, p =.320, η~p²= 0.09; RL: F(4, 47) = 2.27, p =.076, ηp² = 0.16). Moreover, there was no significant effect of the testing condition in either model (CR: F(1, 47) = 1.92, p =.173, ηp² = 0.04; RL: F(1, 47) < 0.001, p =.995, ηp~² < 0.001), indicating that the congruency effect did not differ significantly between testing conditions.
Figure 3 displays the two groups’ congruency effects on accuracy (CR) and response latency (RL). The control group exhibited shorter mean response latencies (RLs) and higher accuracy in the congruent (Mean RL = 1.13 s, SD = 0.24; Mean CR = 92%, SD = 5.7) than in the incongruent condition (Mean RL = 1.24 s, SD = 0.27; Mean CR = 81.8%, SD = 6.7). The IMS group showed a similar pattern, though the difference between conditions was somewhat larger for both RLs (Congruent: Mean RL = 1.42 s, SD = 0.50; Incongruent: Mean RL = 1.58 s, SD = 0.54) and accuracy (Congruent: Mean CR = 92%, SD = 6.1; Incongruent: Mean CR = 71.7%, SD = 12).
Planned Bayesian independent samples t-test (one-sided and using a Cauchy prior width of 0.707) revealed that the data were 7.86 times more likely under the null hypothesis than under the alternative hypothesis predicting a larger congruency effect in the control group compared to the IMS group, for RLs (BF₀₁ = 7.86), and 14.93 times more likely under H₀ for CRs (BF₀₁ = 14.93).
Additional Bayesian independent samples t-tests (one-sided and using a Cauchy prior width of 0.707) conducted to gain additional insight into the observed group differences revealed that the data were in fact slightly for RLs (BF₁₀ = 1.48) to largely for CRs (BF₁₀ = 248.09) more likely under the hypothesis of a larger congruency effect in the IMS groups than under the null hypothesis assuming no difference between groups. Similar results emerged from exploratory generalized linear mixed model (GLMM) analyses on accuracy and linear mixed model (LMM) analyses on response latencies, including Group (Control, IMS) and Condition (Congruent, Incongruent) as fixed effects, and random intercepts and slopes for Condition across both Participants and Audio stimuli (see Appendix C). The results of these additional analyses are therefore incompatible with the hypothesis that motor simulation contributes to effortless facial expression recognition. If this hypothesis were correct, IMS participants should have shown a smaller congruency effect than control participants, rather than the (trend towards) larger effect that was observed.
Fig. 3Group-level distribution of the congruency effect on accuracy and response latency. Note. Box-plots depicting the distribution of the congruency effect on accuracy (left) and response latency (right) in the IMS and Control groups.
The recognition of emotional sounds presented in isolation was high for both control individuals (mean rate = 93.6%, SD = 4.6; all 82.5> %) and IMS (mean rate = 92.3%, SD = 3.7; all 82.5> %). Despite the clear influence of facial expressions in the dual task, both the controls (Mean CR = 52.2%, SD = 8) and the IMS (Mean CR = 52.34%, SD = 7) performed only slightly above chance level (Control: t(51) = 1.95, p =.028; IMS: t(15) = 1.34, p =.100), suggesting that the task was highly challenging.
The aim of this study was to test a specific hypothesis according to which motor simulation, although not the only mechanism available to recognize actions, may be the only mechanism which “allows the observer to understand the actions of others without needing inferential processing”^10^ (p. 268). At odds with that hypothesis, individuals born with congenital facial paralysis categorized emotional sounds significantly less accurately and more slowly when exposed concurrently to an incongruent than to a congruent facial expression, even though these facial expressions were unnecessary and irrelevant to the task at hand. This indicates that facial expression recognition can occur automatically – quickly, effortlessly, and involuntarily - without the involvement of motor simulation.
In addition, we failed to find any support for the possibility that motor simulation, although not necessary, may contribute to automatic facial expression recognition. There was no indication that the IMS could be less influenced by the facial expressions than control participants. On the contrary, exploratory Bayesian analyses and hierarchical modelling provided evidence for a larger influence of facial expressions on sound categorization accuracy in the IMS group. This suggests that the observed effects cannot be dismissed as mere false negatives due to limited task sensitivity or insufficient statistical power. This demonstrates that it is possible to account for effortless facial expression recognition without motor simulation.
These findings align with previous research suggesting that motor simulation is not universally required for the recognition of facial expressions^18,27–29^ and actions more generally^25,40^. In particular, they corroborate the result of a previous study in which we found that individuals born with facial paralysis showed a typical McGurk effect - a perceptual phenomenon resulting from an automatic influence of lip reading on auditory speech perception^37,57^.
However, the type of visual information required for the speech visual signal to influence auditory perception in the McGurk effect is relatively rudimentary (e.g., whether the lips touch or not). The results reported here further demonstrate that motor simulation is not a prerequisite for the automatic recognition of facial gestures, even when more in-depth perceptual analysis is needed to discriminate between facial expressions that share many characteristics. For instance, the facial expressions of fear and anger, contrasted in the first block, share similar facial movements, such as brow lowering, upper eyelid raising, and eyelid tightening.
It remains an open question whether these results and conclusions generalize to other stimuli. While this is ultimately an empirical question that requires further investigation, we believe it is unlikely that our conclusions apply to all types of stimuli. The facial expressions used in this study were familiar and prototypical. Previous research has shown that individuals born without upper limbs struggle to recognize hand actions when presented in particularly unfamiliar and ambiguous formats—such as point-light displays depicting movements as 12 small white dots placed on an actor’s main joints^40^. Similarly, individuals born with facial paralysis may have difficulty detecting subtle differences between facial expressions^33,35^. Therefore, it is possibly, in our view likely, that motor simulation may play a role in the ability to recognize actions and facial expressions when stimuli are more unfamiliar or ambiguous (see^40^ for detail).
Similarly, although our results show that motor simulation is not necessary for automatic facial expression recognition in theory, this does not imply that it is not involved in practice. In theory, there is no reason to believe that individuals born with facial paralysis possess unique brain mechanisms that would allow them to recognize facial expressions without motor simulation. Therefore, we argue that our conclusion can be generalized to the population at large in it must be possible for anyone to develop the ability to recognize facial expressions without motor simulation.
However, in practice, it could be that individuals with facial paralysis, due to their lifelong experience of being exposed to facial expressions they cannot mimic or simulate, have developed an extraordinary ability in this regard—an ability that, while potentially available to everyone, is more developed in those who do not rely on motor simulation. Thus, while these results provide evidence that the ability to recognize facial expressions without motor simulation can develop at a normal level of efficiency, they also leave open the possibility that neurotypical individuals may still rely on motor simulation to some degree.
In conclusion, the findings from this study provide strong evidence that motor simulation is not necessary for the automatic and efficient recognition of facial expressions, thereby providing new information about the boundaries of motor simulation’s involvement in facial expression recognition. These results offer new insights into the cognitive mechanisms underlying this essential aspect of social functioning. Future research should further investigate the (other) conditions and abilities to which motor simulation may contribute (or not) and explore fresh hypotheses about how individuals with motor impairments may compensate for the lack of motor cues.