Authors: Hardik Kothare (aModality.AI, Inc., San Francisco, CA), Vikram Ramanarayanan (aModality.AI, Inc., San Francisco, CA; bUniversity of California, San Francisco), Michael Neumann (aModality.AI, Inc., San Francisco, CA), Jackson Liscombe (aModality.AI, Inc., San Francisco, CA), Vanessa Richter (aModality.AI, Inc., San Francisco, CA), Linnea Lampinen (bUniversity of California, San Francisco), Alison Bai (bUniversity of California, San Francisco), Cristian Preciado (bUniversity of California, San Francisco), Katherine Brogan (bUniversity of California, San Francisco), Carly Demopoulos (bUniversity of California, San Francisco)
Categories: Speech
Source: Journal of Speech, Language, and Hearing Research : JSLHR
Authors: Hardik Kothare, Vikram Ramanarayanan, Michael Neumann, Jackson Liscombe, Vanessa Richter, Linnea Lampinen, Alison Bai, Cristian Preciado, Katherine Brogan, Carly Demopoulos
We investigate the extent to which automated audiovisual metrics extracted during an affect production task show statistically significant differences between a cohort of children diagnosed with autism spectrum disorder (ASD) and typically developing controls.
Forty children with ASD and 21 neurotypical controls interacted with a multimodal conversational platform with a virtual agent, Tina, who guided them through tasks prompting facial and vocal communication of four emotions—happy, angry, sad, and afraid—under conditions of high and low verbal and social cognitive task demands.
Individuals with ASD exhibited greater standard deviation of the fundamental frequency of the voice with the minima and maxima of the pitch contour occurring at an earlier time point as compared to controls. The intensity and voice quality of emotional speech were also different between the two cohorts in certain conditions. Additionally, facial metrics capturing the acceleration of the lower lip, lip width, eye opening, and vertical displacement of the eyebrows were also important markers to distinguish between children with ASD and neurotypical controls. Both facial and speech metrics performed well above chance in group classification accuracy.
Speech acoustic and facial metrics associated with affect production were effective in distinguishing between children with ASD and neurotypical controls.
https://doi.org/10.23641/asha.28027796
Autism spectrum disorder (ASD) is a neurodevelopmental disorder (American Psychiatric Association, 2013) with an estimated overall prevalence of one in 36 children aged 8 years in the United States (Maenner, 2023). A defining feature of ASD, according to the Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition (DSM-5) criteria, is impairment in nonverbal communicative behaviors used for social interaction (American Psychiatric Association, 2013) such as facial expression, which can manifest as absent/minimal, more intense/exaggerated, poorly integrated, or inappropriate to context. Vocal features such as fundamental frequency (F0) and prosody, which carry emotion-specific information (Nussbaum et al., 2022), have been described as atypical when produced by children with ASD (Nadig & Shaw, 2012). Facial expressions of emotion have also been characterized as less natural and more intense in individuals with ASD (Faso et al., 2015). Prior studies have reported atypical production of vocal and facial affect during emotional speech (Hubbard et al., 2017; Loveland et al., 1994) and poor cross-modal coordination between facial expression and emotional speech production in ASD (Sorensen et al., 2019).
The feasibility and potential clinical utility of speech biomarkers for automated assessment of atypical vocal and facial expression in ASD and other neurodevelopmental disorders have been established by prior work (Ramanarayanan et al., 2022). Acoustic–phonetic and lexical features extracted from short, unstructured conversations accurately identified the diagnostic status of children with ASD 66% of the time and that of typically developing children 86% of the time (Cho et al., 2019). Automated assessment of prelinguistic vocalizations has also been shown to be helpful in predicting future diagnosis of ASD (Pokorny et al., 2017). Atypical facial expressions, defined by facial action unit (FAU) intensities, in ASD can be automatically quantified through the use of computer vision and FAU intensities (Leo et al., 2018). FAUs are the components of facial expressions defined by the movement of a muscle or set of muscles (Ekman & Friesen, 1978). FAU intensities are a way to quantify emotional expressions through temporal and geometric analysis of FAUs. When individuals with ASD were cued to mimic facial expressions that carried either positive or negative valence (expressions that were inferred to carry positive or negative values), the degree of facial movements did not depend on the emotional valence and the movements were fleeting, exaggerated, and jerky (Zane et al., 2019) as compared to a control group. Prior work has demonstrated the feasibility and utility of computer vision in a standalone or multimodal framework in ASD research (Bangerter et al., 2020; Samad et al., 2017; Sorensen et al., 2019). Given that there are no standardized measures of facial or vocal affect production ability currently available, these automated, objective measurements have potential clinical utility in quantifying domain-specific nonverbal communicative behavior.
Our prior work has demonstrated the utility of a cloud-based multimodal conversational platform (Ramanarayanan et al., 2023, 2024; Suendermann-Oeft et al., 2019) that uses a virtual human guide, Tina, to conduct self-driven assessments that elicit speech and facial expressions through a variety of tasks for detection and progress monitoring of various neurological and mental health disorders like amyotrophic lateral sclerosis (Neumann et al., 2021, 2024), depression (Neumann et al., 2020), Parkinson's disease (Kothare et al., 2022), and schizophrenia (Richter et al., 2022). During an interactive session with Tina, analytic modules extract objective metrics in real time that can be accessed by researchers or clinicians through a user-friendly dashboard. In prior work in ASD (Kothare et al., 2021), we showed that atypical affect production, measured using a novel affect production task (APT), correlates with accuracy in recognition of vocal and facial affect in children with ASD. Furthermore, we identified a positive correlation between jaw kinematic measures and the motor speed of the dominant hand, which supports the hypothesis that there is a coupling between speech motor coordination and fine motor skills in ASD (Talkar et al., 2020).
Building on this foundation, the current work aims to identify facial and vocal markers that show significant differences between children with ASD and neurotypical controls (NTCs). Objective audiovisual metrics of affect production in ASD may be used to quantify expressive aspects of nonverbal communication. Impairment in nonverbal communication is one of the diagnostic criteria for ASD (American Psychiatric Association, 2013). As such, quantification of this symptom domain also has potential clinical utility in tracking clinical presentation over time or in response to interventions. This is particularly salient, as objective measures of symptom presentation in ASD are lacking. As such, clinical trials currently must rely on subjective observation and informant report measures.
We leverage the aforementioned objective multimodal metrics to answer the following research questions in this
The study was approved by the institutional review board of the University of California, San Francisco (UCSF IRB Approval 11-05249 and 21-33613). Informed consent from the participants' guardians and written assent from the participants were obtained prior to enrollment. The study was conducted onsite at the University of California, San Francisco. Data from 40 participants with ASD (14 female, mean age ± standard deviation = 12.50 ± 2.68 years) and 21 NTC participants (11 female, mean age ± standard deviation = 12.52 ± 2.88 years) who completed an interactive session on the cloud-based multimodal dialogue platform (see Table 1) between December 2019 and December 2022 were included in the analysis. Inclusion criteria for the NTC group no neurological or psychiatric diagnosis and a Social Communication Questionnaire score in the nonclinical range (Rutter et al., 2003). To minimize differences across participants and cohorts, these sessions were conducted in the same controlled environment on the same device (a MacBook Pro with an Intel Core i7 processor) in the presence of a clinical psychology doctoral student who accompanied the participant in the testing room to help with any technical difficulties and provide behavioral support during data collection (e.g., redirecting attention during breaks). Diagnoses in the ASD cohort were confirmed according to DSM-5 criteria by a licensed clinical psychologist (author C.D.) who established research reliability on the Autism Diagnostic Observation Schedule–Second Edition (ADOS-2; Lord et al., 2000) and the Autism Diagnostic Interview–Revised (ADI-R; Lord et al., 1994). Information obtained from the ADI-R and the ADOS-2 (used as an observational tool only, as scoring was not possible due to deviation from standard administration because of COVID-19 masking mandates) was used to inform diagnostic determinations along with parent report measures of social, emotional, behavioral, and adaptive functioning via the Behavior Assessment System for Children–Third Edition (Reynolds & Kamphaus, 2015), performance-based measures of language skills via the Clinical Evaluation of Language Fundamentals–Fifth Edition (CELF-5; Wiig et al., 2013), and general intellectual abilities on the Wechsler Intelligence Scale for Children (WISC; Wechsler, 2014) and the Test of Nonverbal Intelligence–Fourth Edition (TONI-4; Brown et al., 2010). Standardized test scores are included in Table 1 and Figure 1.

The ASD cohort had a lower average score (see Table 1 and Figure 1) on the WISC Full-Scale IQ (t test; t = −3.25, p = .0019), the CELF-5 Expressive Language Index (t test; t = −2.30, p = .0249), and the CELF-5 Receptive Language Index (t test; t = −2.83, p = .0064) but not on the TONI-4 Nonverbal IQ (t test; t = −1.78, p = .0805).
The APTs presented in the interactive session asked the participants to produce one of four emotions—happy, sad, angry, and afraid—through the subtasks listed below. All tasks are performed under directed conditions in which the emotion the participant is expected to communicate is explicitly stated, with the exception of the imitation task in which the emotion is not specified and the participant is simply asked to mimic each stimulus. The session begins with a speaker test, background noise check, and a microphone test. The speaker test determines if the participant is able to hear sounds. The virtual guide, Tina, says a number from zero to nine, and the participant is asked to enter the number in a text field. The background noise measures ambient background noise in decibels while the participant remains silent. During the microphone test, the participant is asked to speak, and it is determined whether the intensity of the participant's speech is at least 40 dB. The participant is asked to ensure that their face is fully visible with no face coverings or shadows obscuring their face. Participants are also asked, “How do you identify?” with four options to choose a boy, a girl, nonbinary, or other than a boy or a girl. Tina then welcomes the participant to the session.

For all the tasks described above, there is an option to repeat each turn, which the examiner would select if the child was not responding to the prompt, not attending to the task, or otherwise had an unusable turn. See Figure 3 for a schematic of the interactive session.

Speech audio data were collected by the platform at a sampling rate of 48 kHz. All speech acoustic metrics were extracted using Praat (Boersma & van Heuven, 2001). These metrics were spectral domain metrics (F0 [Hz]; jitter [difference of difference of periods, %]; F1, F2, and F3 formant frequencies [Hz]; F2 slope [Hz/s]; cepstral peak prominence [CPP; dB]; harmonics-to-noise ratio [dB]), signal energy metrics (shimmer [%], signal-to-noise ratio [SNR; dB], intensity [dB]), and duration metrics (speaking duration [s], articulation duration [s], time point of maximum and minimum F0).
To extract facial metrics, MediaPipe face detection based on BlazeFace (Bazarevsky et al., 2019) was used to determine framewise x and y coordinates of the face. Facial landmarks were then generated by the MediaPipe face mesh algorithm (Kartynnik et al., 2019), 14 of which are key landmarks in the computation of jaw kinematics, lip aperture, mouth surface area, eyebrow height, and so forth. All facial metrics were normalized by dividing them by the intercaruncular distance (see Figure 4) to account for cross-participant positional variability relative to the camera (Roesler et al., 2022).

See Table 2 for an overview of the metrics and Supplemental Material S1 for a glossary. All metrics went through a two-step automatic outlier detection. Metrics are not excluded on a participant level but on a turn level in this method. First, all metric values beyond 5 SDs, that is, extreme outliers, from the mean metric value were removed. These extreme outliers likely arise from incorrect task performance or noncompliance. Second, the mean of the distribution was recomputed, and any values beyond 3 SDs were flagged as outliers and removed from the analysis in accordance with the three sigma rule (Upton & Cook, 2008).
Since each subtask had multiple turns per emotion, metrics were averaged across turns within every emotion. All metrics were z-scored by sex at birth across the entire sample to account for sex-specific differences. To identify which metrics showed statistically significant differences, nonparametric Kruskal–Wallis tests (Kruskal & Wallis, 1952) were run for all metrics. Effect sizes, as measured by Glass's Δ (Glass et al., 1981), were calculated for all metrics, and only those metrics with an absolute Glass's Δ of greater than 0.6 or large effect sizes (Panzarella et al., 2021) are reported in this article.
To assess if group differences in facial and vocal metrics were associated with differences in affective communication as opposed to general differences in facial and vocal expression, two human raters (with at least average scores on the DANVA-2 Facial and Vocal Affect Recognition subtests; Nowicki & Duke, 1994) classified the emotion (happy, sad, angry, afraid, or neutral) produced by participants in response to each prompt. Raters were blinded to the prompted emotion and group. Facial affect was classified from video responses of each participant in the absence of vocal audio. Likewise, vocal affect was classified in the absence of video. Thus, raters made affective judgments based solely on the facial and speech behavior of the participant, respectively. Percent accuracy values of emotion judgment using video and audio were calculated for each rater and were then averaged across raters. Nonparametric independent-samples Mann–Whitney U tests were then run to identify differences between the two cohorts in percent accuracy of the rater's perception of facial and vocal affective expression.
To evaluate whether group differences varied according to task demands (i.e., production length, comprehension of contextual narrative), the above-mentioned Kruskal–Wallis tests were performed on task–metric combinations.
To test whether an interaction effect between emotion and cohort was present when it came to differences in metric values, we ran a two-step analysis. For this analysis, we aggregated all 30 metrics across the two comparable tasks that involved noncontextual and contextual monosyllabic production (Tasks 1 and 3) by averaging them. In Step 1, we ran a one-way repeated-measures analysis of variance (ANOVA; Vallat, 2018) to test for an effect of emotion without controlling for cohort. In Step 2, to test for an interaction effect between cohort and emotion, we ran a mixed-design ANOVA (Murrar & Brauer, 2018) for the metrics that showed a significant effect of emotion in Step 1. The between-subjects factor was cohort and the within-subject factor was emotion to account for repeated measurements. For metrics that showed a significant interaction effect between emotion and cohort, post hoc Wilcoxon signed-rank tests were run for pairwise comparison.
To evaluate whether the cohorts differed in vocal affect imitation, which assesses the ability to produce vocal sounds that convey emotion without requiring knowledge of how to use vocalization for the purpose of communication emotion, we looked at metrics showing differences in Task 2 (vocal imitation of a noncontextual monosyllable).
For all metrics with Glass's Δ greater than 0.6, a leave-one-out logistic regression classifier model using the Scikit-learn Python module (Pedregosa et al., 2011) was run using speech acoustic metrics alone, facial metrics alone, and the combination of both modalities. Receiver operating characteristic (ROC) curves were plotted for these three models. Area under the curve (AUC; the higher the better) and the Brier score (Brier, 1950) measuring the accuracy of probabilistic predictions (the lower the better) were generated for these models. This process was repeated for each individual emotion to test classification of groups while controlling for emotion.
The effect sizes of metrics, as measured by Glass's Δ, that showed a statistically significant difference between the ASD cohort and controls and had an absolute value greater than 0.6 are shown in Figure 5. A positive effect size denotes a greater median value for the ASD cohort, and a negative effect denotes a smaller median value for the ASD cohort. Note that nonsignificant effects are represented by blank spaces, irrespective of the size or direction of the actual effect.

With regard to vocal metrics, participants with ASD exhibited a larger standard deviation of the F0 of their voice when conveying sadness during noncontextual monosyllabic imitation and production tasks. Moreover, the time point of the minimum and maximum values of F0 during the imitation task occurred earlier in the ASD cohort when the emotion to be conveyed was sad or happy. The maximum value of F0 also occurred earlier in the ASD cohort while repeating an afraid “oh.” The SNR during the noncontextual monosyllabic imitation and production tasks conveying fear was lower in the ASD cohort. The ASD cohort also had a greater jitter value during the noncontextual monosyllabic production task and a lower CPP value while imitating an afraid “oh.”
In the case of facial metrics, when participants with ASD produced an angry “oh” for the noncontextual monosyllabic production task, an afraid “oh” for the contextualised monosyllabic production task, and a noncontextual sentential production conveying fear, they had a smaller eye opening than the control group. Acceleration of the lower lip during a sad noncontextual monosyllabic production of “oh” was higher in the ASD group and lower during an angry sentential production. During a happy noncontextual monosyllabic production, children with ASD had a smaller lip width than controls. Also, during a happy sentential production, the ASD group had smaller average eyebrow vertical displacement.
Average human rater accuracy for facial affect recognition was 54% (47% for the ASD cohort and 65% for NTC). Average human rater accuracy for vocal affect recognition was 56% (53% for ASD and 63% for NTC). For all tasks and emotions, human rater accuracy in emotion perception for facial data, defined by the agreement between the rater's emotion classification and the prompted emotion, was significantly different between the two cohorts (U = 487.00, p = .002), indicating that the ASD group was less effective in communicating the prompted emotions than the control group. Lower human rater accuracy for the ASD cohort was seen across Task 1 (U = 529.5, p = .007), Task 2 (U = 535.00, p = .018), Task 3 (U = 436.00, p = .002), and Task 4 (U = 420.00, p < .001). When rater accuracy of video data was split by emotions, lower accuracy for the ASD cohort was observed for three of the four happy (U = 562.50, p = .016), sad (U = 576.00, p = .023), afraid (U = 393.00, p < .001), and angry (U = 703.50, p = .334).
Overall human rater accuracy of emotion perception based on speech data was not significantly different between the two cohorts (U = 463.50, p = .142). Such differences were also not seen when data were split by Task 1 (U = 474.00, p = .277), Task 2 (U = 469.00, p = .201), Task 3 (U = 404.00, p = .052), Task 4 (U = 511.50, p = .531). While overall rater accuracy and accuracy for combined emotions for each tasks did not show significant group differences for speech data, when perceptual accuracy of speech data was compared for each emotion across tasks, the ASD group demonstrated significantly poorer vocal communication of happiness (U = 343.50, p = .005) and sadness (U = 411.00, p = .048). Significant group differences were not identified for vocal communication of fear (U = 523.50, p = .535) and anger (U = 649.00, p = .373).
From Figure 5, it can be observed that differences in facial and vocal metrics between the two cohorts can be captured even with monosyllabic or shorter utterances. In fact, only three metrics showed differences when the participants produced sentence-length utterances (Task 4). Interestingly, all three metrics were facial metrics (eye opening, lower lip acceleration, and eyebrow displacement).
The prompt during Task 3 included an illustrated narrative providing additional emotional context. Notably, only one facial metric (eye opening during the expression of fear) showed differences between the two cohorts in Task 3. All other differences in objective metrics were captured when additional narrative context was not provided.
Twenty-seven of the 30 metrics aggregated across Tasks 1 and 3 showed a significant effect of emotion (see Supplemental Material S1). The three metrics that did not show an effect of emotion F2 slope, time point of minimum F0, and mean symmetry ratio of the mouth surface area. Four of the 27 metrics showing a significant effect of emotion also showed a significant interaction effect between cohort and emotion (see Supplemental Material S1). These metrics were average lip width, average velocity of the lower lip, average acceleration of the lower lip, and average jerk of the lower lip. Post hoc pairwise Wilcoxon signed-ranks tests were run to test for which emotion the metrics were significantly different between cohorts. Only one pairwise test was statistically significant (see Table 3); mean velocity of lower lip was significantly higher in the NTC cohort as compared to the ASD cohort when the emotion was happy (p = .03, Hedges's g = 0.60).
There were differences only in speech metrics and not facial metrics when participants were asked to repeat a monosyllable after listening to an audio stimulus, which is not surprising given that participants were only asked to imitate the vocal expression. The differences in vocal imitation were related to the standard deviation of F0, time points of maximum and minimum values of F0, SNR, and CPP.
ROC curves for the classification experiment between cohorts can be seen in Figure 6. When all metrics, across emotions, with an absolute effect size greater than 0.6 were used as features in the classifier, both facial (AUC = .79, Brier score = 0.17) and speech metrics (AUC = .74, Brier score = 0.20) performed well above chance with the facial metrics outperforming the speech metrics. The performance, as measured by the AUC, of the classifier model was slightly better than individual modalities when metrics from both modalities were considered (AUC = .8, Brier score = 0.18). When metrics related to happy utterances were considered, facial metrics (AUC = .67, Brier score = 0.21) again outperformed the speech metrics (AUC = .62, Brier score = 0.22) in classification of the two cohorts. A multimodal model, with both speech and facial metrics, was again slightly better than the individual modalities alone (AUC = .69, Brier score = 0.20). For sad utterances, the speech metrics performed much better (AUC = .73, Brier score = 0.20) than the facial metrics (AUC = .59, Brier score = 0.22), and the performance did not improve drastically when a combination of both modalities was used (AUC = .72, Brier score = 0.20). Since there were no differences in speech metrics for angry utterances between the two cohorts, a classifier model with the two facial metrics showing a difference between the two cohorts was run (AUC =.64, Brier score = 0.21), and its performance, while not being superlative, was well above chance. For afraid utterances, speech metrics (AUC = 0.71, Brier score = 0.20) were slightly better at classifying the cohorts than the facial metrics (AUC = .67, Brier score = 0.22). A combination of both modalities had a much better performance for afraid utterances (AUC = .78, Brier score = 0.18).

In this study, we investigated which audiovisual metrics associated with affect production and imitation showed significant differences between children with ASD and NTCs. We examined effects of task demands and specific emotions on group differences and used a leave-one-out logistic regression classifier model to evaluate the efficacy of facial and vocal metrics in classifying groups.
With regard to Research Question 1, we identified group differences in objective facial and vocal metrics within each emotional category and across-task conditions. We also identified relatively lower human rater accuracy in identifying affect conveyed by the ASD cohort, which is to be expected given that this group is partially defined by deficits in nonverbal communication. These differences in human accuracy, in conjunction with prior work demonstrating prediction of human rater performance from objective metrics (Demopoulos et al., 2024), suggest that the metrics are capturing differences in ability to communicate affect as opposed to nonspecific differences in facial movement and vocal expression. Specifically, in this prior study, we found that the linear combination of objective facial metrics predicted 32%–60% of the variance in human rater accuracy for facial APTs and the linear combination of objective vocal metrics predicted 41%–58% of the variance in human rater accuracy for vocal APTs. This suggests that the automatically extracted metrics are measuring information that human raters are using in making affective judgments.
With regard to Research Question 2, effects of task demands, while objective metrics extracted from all tasks were useful in distinguishing between the two cohorts, there were more metrics that significantly distinguished groups in Tasks 1 and 2 (noncontextual monosyllabic production and imitation) than those from Tasks 3 and 4 (contextualized monosyllabic and noncontextual sentence-length production). This suggests that the task with most minimal expressive and receptive language demands (i.e., brief verbal instructions and requiring production of only a monosyllabic utterance) was equally, if not more, effective in identifying differences in affect production associated with autism. This also suggests that assessment via the APT can be accessible to individuals who require minimal language demands for valid assessment. Notably, for Tasks 3 and 4, only facial and no vocal metrics showed differences between children with ASD and controls. The nature of the tasks, prompted monosyllabic speech after a narrative/picture stimulus and noncontextual sentence-length productions, may have a role to play in this observation. Specifically, both of these tasks have greater verbal demands in different ways. For example, greater receptive language skills are necessary to understand the narrative, even though only a monosyllabic utterance is required for vocal response to the contextual monosyllabic condition. In contrast, greater speech/expressive language skills are necessary to produce the longer sentence-length utterance, while the semantic content of the sentence is not meaningful to affective vocal production in and of itself.
Regarding Research Question 3, interaction effects between group and prompted emotion, a main effect of emotion was identified across most objective metrics, as expected given that these metrics were selected based on their relevance to communicating affect. An interaction between group and emotion was also identified for several facial metrics related to mouth movements and positions. Taken together with human rater data indicating overall less effective communication of facial affect in the ASD group, these interaction effects suggest that poor affect production in the ASD group may be associated with ineffective use of mouth movements and position during facial expression. Indeed, we observed smaller lip width during monosyllabic production of speech conveying happiness in the ASD cohort. A happy vocalization is often accompanied by smiling where the mouth orifice is widened (Shor, 1978; Tartter, 1980). Smaller lip width in the ASD cohort during happy emotional speech production may indicate the absence of an accompanying smile, therefore indicating an inability to express the emotion successfully.
Several other group differences were identified under specific emotion conditions and specific tasks. For example, when conveying a sad emotion, participants with ASD had a greater standard deviation of the F0 of their voice as compared to controls during both noncontextual monosyllabic production and imitation tasks conveying a sad emotion. Relatedly, according to the human rater accuracy data, the ASD cohort demonstrated poorer vocal communication of sadness. Indeed, increased pitch variation during speech production in general and emotional speech production in particular has been observed quite consistently in studies of individuals with autism described as “high functioning” (Diehl et al., 2009; Edelson et al., 2007; Fosnot & Jun, 1999; Nadig & Shaw, 2012). This increased pitch variability has been shown to be language agnostic and is not associated with the language ability of the ASD participants (Bonneh et al., 2011; Green & Tobin, 2009; Sharda et al., 2010).
We also observed that during emotional speech production, maximum F0 and minimum F0 time points occurred earlier in the ASD cohort. Atypical prosody has been documented in both receptive and expressive speech in high-functioning autism (McCann & Peppé, 2003; McCann et al., 2007; Peppé et al., 2007). Individuals with ASD are said to experience difficulties with social acceptance due to the atypical prosody of their speech (Paul et al., 2005; Shriberg et al., 2001), underscoring that prosodic differences may functionally impact paralinguistic aspects of vocal communication. Specifically, it has been postulated that extreme pitch variation in ASD could be placed arbitrarily in the utterance, thus rendering the acoustic cues of the speech nonmeaningful to listeners (Nadig & Shaw, 2012). Understanding the root of these prosodic differences could direct novel approaches to improving communication skills via targeting the barriers to effective use of ancillary acoustic cues of vocalization (not only what is said but how it is said).
The SNR was lower in the ASD cohort during the noncontextual monosyllabic imitation and production of “afraid.” Furthermore, metrics indicative of voice quality (i.e., jitter and CPP) were higher and lower, respectively, in the ASD cohort for the afraid utterances. Group differences in human rater accuracy of fear were not identified, however, suggesting these may be vocal differences not associated with affective communication.
Additionally, we also observed lower average vertical displacement of the eyebrows during happy noncontextual sentence-length production and a smaller average eye opening during noncontextual monosyllabic productions of fear and anger in the ASD cohort. These results, combined with the lower human rater accuracy for the ASD cohort, suggest that individuals with ASD exhibit reduced expressivity of nonverbal cues of emotion during affect production. The eyes and eyebrows have a significant role to play in the expression of distinct human emotions (Perveen et al., 2012). Furthermore, reduced eye contact is also common in individuals with ASD. The findings of reduced facial expressivity in the eye region of the face may be associated with lack of experience in reading those cues in others or, alternatively, a lack of salience for those movements resulting in a tendency not to produce them or watch for them in others.
With regard to Research Question 4, group differences in affect imitation (Task 2), only speech acoustic metrics showed large differences. No emotional cue was provided in this task apart from the paralinguistic information in the audio stimulus (a noncontextual monosyllabic vocalization). Furthermore, there was no facial stimulus to imitate and no instructions regarding production of facial expression. Differences in speech metrics but not in facial metrics for this task may suggest that both cohorts had a similar range of facial motoric productions during these vocal imitations, but the ASD cohort differed in the reproduction of the paralinguistic information in the actor's speech sounds. Interestingly, the average rater accuracy was still relatively lower for facial imitation in the ASD group despite failure to identify differences in objective facial metrics. This may suggest that both groups performed a similar range of facial movements under nondirected conditions, but the typically developing controls group produced more facial expressions that were congruent with the vocal affect of the stimulus prompt and the ASD group produced facial expressions that were less communicative of the emotion being conveyed in the vocal affect of the stimulus prompt. These results are consistent with prior studies reporting reduced multimodal affective communication (Hubbard et al., 2017; Loveland et al., 1994; Sorensen et al., 2019).
Finally, for Research Question 5, we evaluated classifier models to distinguish between ASD and controls. In these analyses, we observed that classifiers performed best when metrics across all emotions were considered. However, there were slight differences in model performance when individual emotions were considered. For utterances that were supposed to convey sadness and fear, speech acoustic metrics performed better than facial metrics. When it came to happy utterances, there was an equal number of speech and facial metrics showing differences between the two cohorts, but the facial metrics had better classification performance. Interestingly, for anger, only facial metrics showed large differences between the two cohorts, and their classification performance was well above chance. For most of the classification models though, the combination of both modalities (speech and facial) was either more effective than individual modalities or equivalent to the performance of the better performing modality (speech metrics for sadness). This observation underscores the importance of a multimodal framework approach in studying a complex disorder like ASD with significant cross-domain atypicalities (C.-P. Chen et al., 2017; J. Chen et al., 2020; Kothare et al., 2021; Samad et al., 2017). The observed AUCs are comparable to prior classification studies with multimodal features (Cho et al., 2019).
The current study comes with a set of limitations. First, while this study focused on measurement of objective features of facial and vocal expression of emotion under cued conditions in order to standardize the emotion intended to be communicated across participants, it is not the same as measuring emotional expression under spontaneous environmentally provoked conditions. The objective metrics extracted essentially measure the participants' ability to consciously and intentionally produce or act out the emotions. This distinction is important as natural expression of emotion may not communicate the experienced emotion directly but may be influenced by what one intends to communicate (i.e., one may not wear one's heart on one's sleeve for certain reasons and in certain contexts). There may be differences in communicative intentions that would distinguish ASD from other groups in a more naturalistic assessment of expressive emotional communication. Future studies should investigate the differences between objective measures of emotional expression upon cue and emotional expression in the natural environment. Second, the control cohort was smaller than the ASD cohort, and future investigations should consider larger and more equally matched cohorts. Third, it cannot be ascertained that group differences between the cohorts arise solely due to functional deficits. Differences observed could just be a reflection of how the two groups responded to task prompts. It must be noted that the results of this study may only be generalizable under certain contexts and prompts. Lastly, although the doctoral student did not interact with the participants during data collection, their presence in the room may have affected performance differentially between cohorts. Future studies involving remote data collection in natural environments may help answer some of these questions. That being said, it is worth underscoring that the platform described in this study can also be accessed remotely using any device equipped with a webcam and a microphone (Ramanarayanan et al., 2020). This is especially important considering that there is a growing need to remove barriers and expand telehealth services in children with neurodevelopmental disorders (Masi et al., 2021).
In conclusion, we found that audiovisual metrics extracted through a multimodal conversation-based dialogue platform show significant differences between children with ASD and NTCs and may have potential for monitoring behavior in ASD. Both monosyllabic and sentence-length prompts may have their own advantages in evoking emotional expressions. Monosyllabic utterances would make the task more accessible to people who are minimally verbal. Sentence-length utterances may capture nuances of speech acoustics and facial movement due to the greater length of data available per utterance; however, our current findings suggest that a monosyllabic utterance may be sufficient to effectively measure affect production. We also observed that emotional context in the form of a narrative, which requires more skill in receptive language than other tasks, was not necessary to evoke group differences in emotional expressions. Future research examining psychometric performance of these different task conditions is needed to determine minimum required task demands for sensitive measurement while maximizing inclusivity and accessibility of the task. Emotion-specific and task-specific differences in metrics and model performance were also observed. Furthermore, we found that a multimodal approach is important to classify children with ASD from controls. This is even more important because of the emotion-specific differences in classification performance of the individual modalities.
The de-identified data (speech and facial metrics) generated and analyzed during the current study is available from the corresponding author on reasonable request. Note that the raw audio and video are not shareable as they contain personally identifiable information regarding the participants.