Authors: Yoonji Kim, Diana Sidtis, John J. Sidtis
Categories: Speech
Source: Journal of Speech, Language, and Hearing Research : JSLHR
This study examined spontaneous, spoken-to-a-model, and two sung modes in speakers with Parkinson's disease (PD), speakers with cerebellar disease (CD), and healthy controls. Vocal performance was measured by intelligibility scores and listeners' perceptual ratings.
Participants included speakers with hypokinetic dysarthria secondary to PD, those with ataxic dysarthria secondary to CD, and healthy speakers. Participants produced utterances in four vocal spontaneous speech, spoken-to-a-model, sung-to-a-model, and spontaneous singing. For spoken-to-a-model and sung-to-a-model modes, written material was provided the model. For spontaneous singing, participants sang songs that they endorsed as familiar. Dependent In Experiment I, listeners orthographically transcribed the audio samples of the first three vocal modes. In Experiment IIa, raters evaluated the accuracy of the pitch and rhythm of the spontaneous singing of familiar songs. Finally, familiar songs and sung-to-a-model utterances were rated on a competency scale by a second group of raters (Experiment IIb).
Results showed increases in intelligibility during the spoken-to-a-model mode compared with the spontaneous mode in both PD and CD groups. Singing enhanced the vocal output of speakers with PD more than in speakers with CD, as measured by percent intelligibility. PD participants' pitch and rhythm accuracy and competency in singing familiar songs was rated more favorably than those produced by CD participants.
The findings reveal a vocal task effect for spoken utterances in both groups. Sung exemplars, more impaired in CD, suggest a significant involvement of the cerebellum in singing.
https://doi.org/10.23641/asha.21809544
The basal ganglia and cerebellum are neurological structures that are important to vocal motor control. These brain structures enable integration, coordination, and refinement of motor movements during vocal production. These two brain structures form relatively separate control circuits and influence the planning and programming of speech movements (Duffy, 2013). Lesions to these structures result in speech impairment identified as different forms of hypokinetic dysarthria secondary to Parkinson's disease (PD) associated with basal ganglia damage, and ataxic dysarthria following cerebellar damage (Kent et al., 2000). In this study, the two dysarthric groups were a focus of investigation to explore how differing vocal task demands affect the ability of persons with basal ganglia or cerebellum dysfunction to speak and sing.
Decades of clinical observations have documented that speech and singing differ in clinical conditions. Dysfluencies are reduced or eliminated during singing in persons who stutter (Andrews et al., 1982; Colcord & Adams, 1979; Davidow et al., 2009; Healey et al., 1976). The fluency-inducing effect of singing has also been demonstrated in other neurologic disorders associated with basal ganglia dysfunction associated with PD (Kempler & Van Lancker, 2002; Van Lancker Sidtis et al., 2012). A recent study by Harris et al. (2016) found that singing minimized the dysarthria in PD such that, when singing, PD patients were perceptually indistinguishable from healthy controls (HCs). Dissociations of the elements of singing and talking were documented in two dysprosodic persons (Van Lancker Sidtis et al., 2021). Persons with nonfluent aphasia are able to sing familiar lyrics, but can only effortfully speak the same words (Brust, 2003; Gerstman, 1964; Geschwind, 1971; Geschwind et al., 1968; Hébert et al., 2003; Mills, 1904; Racette et al., 2006; Sparks et al., 1974; Warren et al., 2003; Yamadori et al., 1977). A complementary pattern, spared ability to talk but not sing, occurred following right hemisphere stroke (Schön et al., 2004).
The cerebellum has been studied for its possible involvement in singing (Callan et al., 2007). Building on the proposed link between singing and cerebellar functions, this study proposes that motor components essential for singing may be mediated at least in part by the cerebellum (Abernethy et al., 2004). The prediction is that articulated speech in cerebellar disease (CD) will not benefit from the sung mode, unlike those with the above-mentioned other clinical conditions, where singing yielded enhanced production efficiency in pathological speakers. To date, the effect of cerebellar dysfunction on singing has been only sparsely examined (Callan & Manto, 2013). This study addresses this gap in our understanding of cerebellar contributions to sung production, as well as contributing to broader debates regarding the neural bases of speech and singing.
A body of research has centered on how the use of external cues in the visual (Bagley et al., 1991), auditory (Thaut et al., 1996), or combined (de Oliveira Souza et al., 2015) modality alleviates gait abnormalities present in PD. For instance, rhythmic auditory stimulation (RAS) is a gait intervention that provides external rhythmic cues to which participants synchronize with their footfalls. Hausdorff et al. (2007) reported that in PD, gait parameters, such as speed, stride length, and swing time, were significantly improved during walking with RAS compared with the noncued walking. Similar observations have been reported for arm reach, where a lighted trajectory facilitated smooth and targeted extended movement of the arm (Georgiou et al., 1993). One potential explanation for such improvement in complex gestures proposes compensatory mechanisms aiding the dysfunctional basal ganglia (Amirnovin et al., 2004; Debaere et al., 2003; Glickstein & Stein, 1991).
In the speech domain, external cueing has been provided in the form of repetition (auditory cue) and reading (visual cue), and represented as a contrasting condition to spontaneous monologue or conversation (Kempler & Van Lancker, 2002; Van Lancker Sidtis et al., 2012; Weir-Mayta et al., 2017). Accumulating evidence suggests that externally cued speech tasks (reading, repetition) elicit significant improvements in PD speech compared with spontaneous speech, as revealed by acoustic and intelligibility measures (Kempler & Van Lancker, 2002; Van Lancker Sidtis et al., 2010, 2012). Various explanations for improved speech measures in reading and repetition have been offered.
While it is known that individuals with PD are vulnerable to the effects of vocal tasks, relatively little is known about the influence of different task conditions on speech motor behavior in individuals with CD. It has been suggested that the provision of external cues does not facilitate speech production in individuals with CD (Cannito & Marquardt, 1997; Kent et al., 2000). One study that addressed this issue compared listeners' perceptual ratings of understandability (focusing on articulation) and naturalness (focusing on prosody) in the reading and spontaneous conversation of individuals with PD and CD (Weir-Mayta et al., 2017). As expected, for individuals with PD, significant improvement in both measures (understandability and naturalness) was found during reading versus conversation; for individuals with CD, only understandability, not naturalness, was rated higher during reading relative to conversation. The use of external cues in speech tasks enhanced listeners' perception of PD speakers, whereas with CD speakers, external cues induced these effects in a more limited way. Yet, evidence directly comparing speech abnormalities arising from basal ganglia disease and cerebellar disease is scarce. This study tested the hypothesis that an externally cued speech task would result in different performance outcomes than those measured in the spontaneous mode in one or both clinical groups. The study will provide insight into the broader issue of whether speech motor control is task specific (Case & Grigos, 2020; Kleinow & Smith, 2006; Reuterskiöld & Grigos, 2015; Sadagopan & Smith, 2013).
Intelligibility has been widely used in the pediatric and adult motor speech literature as a measure that quantifies the adequacy of speech function and reflects severity of speech impairment (Bunton & Keintz, 2008; Hustad et al., 2019; Levy et al., 2017; Stipancic & Tjaden, 2022). According to Yorkston et al. (1996), intelligibility is “the degree to which the acoustic signal (utterance produced by the speaker) is understood by a listener” (Yorkston et al., 1996, p. 55). Clinically, reduced intelligibility has been used to index the impaired functioning of the speech production system. Decreased intelligibility has been found in PD (Miller et al., 2007; Tjaden, 2008) as well as cerebellar ataxia (Hilger et al., 2022). In studies examining factors that influence speech motor performance in persons with dysarthria (e.g., speech task, instruction), intelligibility has been frequently employed to gauge the compromised integrity of the speech motor systems across conditions (Kempler & Van Lancker, 2002; Lam & Tjaden, 2013; Tjaden & Wilding, 2011).
Measures of intelligibility are typically obtained using two orthographic transcription and scaled subjective ratings by listeners. In transcription-based approaches, listeners transcribe what they believe the speaker to have said; the transcripts are scored based on the number of words identified correctly. In scaled rating approaches, listeners provide subjective, perceptual ratings of the speech sample using scales (e.g., Likert scales, visual analogue scales, direct magnitude estimation; Stipancic et al., 2016; Weismer & Laures, 2002). In this study, a transcription-based percent score was used as a reliable measure of intelligibility.
In order to explore the auditory-acoustic characteristics of PD and ataxic speech in relation to different vocal modes, four acoustic measures were articulation rate, coefficient of variation of fundamental frequency (f
o), intensity standard deviation (SD), and smoothed cepstral peak prominence (CPPS) for voice quality. These measures were selected based on known characteristics of the respective Increased or variable rate of speech is commonly associated with hypokinetic dysarthria, whereas slow speech rate is often reported in ataxic dysarthria; decreased f
o variation and intensity variability were linked to hypokinetic dysarthria as opposed to increased f
o variation and intensity variability in ataxic dysarthria (Duffy, 2013). Although no differences in CPPS were found between PD and CD speakers (Jannetts & Lowit, 2014), the CPPS was included as a potential indicator of voice quality changes across vocal tasks.
Articulation rate has been widely used to differentiate dysarthric from healthy speech (Nishio & Niimi, 2006) and to assess the effect of speech instructions on dysarthric speakers (Goberman & Elmer, 2005; Lam & Tjaden, 2016). Studies investigating PD's speech rate have produced variable results. Skodda and Schlegel (2008) noted that there was no significant difference in speech rate between PD and control speakers. In contrast, Martínez-Sánchez et al. (2015) observed a significant decrease in speech rate in Spanish PD speakers, and short rushes are commonly observed (Duffy, 2013). Reduced speech rate for CD speakers in comparison with HCs has been observed (Ackermann & Hertrich, 1994; Schalling & Hartelius, 2004). These differences in dysarthric profiles might be expected to bear on intelligibility data and the study independent variables, speech versus singing.
Reduced variation in f
o is a characteristic of PD speech, leading to the perception of monopitch (Duffy, 2000; Gamboa et al., 1997; Skodda et al., 2011). In contrast, excessive variations of f
o are characteristic features of ataxic speech (Boutsen et al., 2011).
Reduced vocal intensity (hypophonia, or soft speech) is a distinguishing characteristic of PD speech. Monoloudness, a lack of variation in vocal intensity, is also commonly present (Duffy, 2013). By contrast, individuals with CD have been reported to have excess variations in loudness, referred to as explosive speech (Duffy, 2013; Swanberg et al., 2007).
Abnormal voice qualities have been reported in PD and CD speakers. Increased breathiness is associated with PD speakers, and the vocal quality of CD speakers is often described as “harsh.” These perceptual impressions agree with findings in acoustic studies, which documented elevated levels of jitter and shimmer in PD (Gamboa et al., 1997; Holmes et al., 2000; Larson et al., 1994) and CD (Ackermann & Ziegler, 1994) speakers. Research suggests that measures in smoothed CPPS correlate with perceptual ratings of dysphonia severity (Awan & Roy, 2005; Heman-Ackah et al., 2002) and reliably discriminate dysphonic from normal voices (Eadie & Baylor, 2006; Watts & Awan, 2011). Jannetts and Lowit (2014) revealed that the cepstral measures were better suited to reflect the perceptual impression of dysphonic voices than traditional perturbation measures, such as harmonics-to-noise ratio, jitter, and shimmer, in hypokinetic and ataxic dysarthria.
The goal of this study was to examine the effects of basal ganglia and cerebellar dysfunction on vocal motor performance. The specific aims of the study were (a) to investigate whether the sung mode elicits different intelligibility outcomes in speakers with PD versus CD (Experiment I); (b) to examine whether the two clinical groups perform differently in the spontaneous versus spoken-to-a-model (external cue) mode, as captured by intelligibility measures (Experiment I); and (c) to assess the accuracy of pitch and rhythm (Experiment IIa) and singing competency (Experiment IIb) as observed in the two groups. Neurologically healthy speakers served as controls (HC) to provide normal values for the measures. Intelligibility measures used in the study refer to transcription accuracy.
Four types of vocal modes were recorded from PD, CD, and HC spontaneous, spoken-to-a-model, sung-to-a-model, and spontaneous singing. The first type of vocal production, spontaneous mode, refers to utterances that were initiated by speakers in the absence of any explicit modeling of the speech. It constitutes narratives from topic prompts that speakers could choose during interviews. The second type, spoken-to-a-model mode, is based on a written model, where speakers were visually presented with a written model of an utterance (on a computer screen) and prompted to say it as naturally as possible after its withdrawal from their visual field. In the third vocal production, sung-to-a-model mode, participants were presented with a model in the same way as in the second type, spoken-to-a-model mode, but instead of saying the utterance, they were prompted to sing the utterance after the written model disappeared. Finally, in the fourth mode, the role of singing was assessed by obtaining spontaneously sung familiar songs. The four types of vocal modes are also referred to as vocal “tasks” in the paper.
In Experiment I, neurologically healthy listeners provided an indirect intelligibility measure by transcribing audio samples obtained from the three groups. In Experiment IIa, musically trained listeners rated the pitch and rhythm fidelity of familiar songs produced by participants in the three study groups. In Experiment IIb, healthy listeners rated familiar songs and sung-to-a-model utterances on a competency scale from very bad to very good.
Two neurologically impaired groups with different lesion loci were the focus of individuals with PD, having dysfunctional basal ganglia with intact cerebella, and individuals with CD, having dysfunctional cerebella with intact basal ganglia. The broader rationale of this study includes furthering our understanding of the relative contributions of the basal ganglia and cerebellum to vocal motor control. Specifically, the following research questions were
First, would the sung-to-a-model mode improve vocal performance in individuals with one disease more than individuals with the other, as measured by intelligibility? With respect to the effect of singing on production efficiency, it was predicted that the sung-to-a-model mode would result in improved intelligibility for PD speakers, but not for CD speakers, compared with HCs. This hypothesis was derived from the proposal that the cerebellum plays a major role in singing and that damage to this structure was expected to disrupt singing.
Second, does efficiency of the sung mode differ in the two neurological disorders, as measured by the accuracy ratings of pitch and rhythm and competency ratings in sung productions? It was expected that healthy listeners' ratings of pitch and rhythm accuracy and singing competency in familiar songs and sung-to-a-model productions would reveal better performance by persons with PD than those with CD.
Thirdly, would the spoken-to-a-model (external cue) and spontaneous speaking (without external cue) conditions yield different levels of intelligibility in the speech of individuals with PD versus those with CD? With regard to the impact of external cueing on speech motor performance, it was anticipated that the two speech tasks, the spontaneous and spoken-to-a-model, would yield different intelligibility outcomes for speakers with PD compared with those with CD. External cueing (i.e., providing a model) was hypothesized to be more effective at reducing speech disturbance in PD than in CD. By contrast, differences in CD speakers' performance under spontaneous and spoken-to-a-model mode were expected to be minimal or limited only to certain measures.
Vocal samples in four tasks, spontaneous, spoken-to-a-model, sung-to-a-model, and singing of familiar songs, were obtained from three groups of native speakers of English: 12 adults with PD (average age = 72.8 years, age 62–84 years), 12 adults with cerebellar ataxia (average age = 53.9 years, age 29–63 years), and 12 typical adults (average age = 64.4 years, age 52–79 years). Detailed demographic information of the speakers can be found in Tables 1 –3. The 36 speakers included in the study were native speakers of American English, born and educated in the United States, except for one woman, a British English speaker who had been living in the United States for 36 years. All participants had adequate hearing and vision with correction by self-report. The majority were right-handed except for one PD and one HC who reported being left-handed.
Study participants were not required to have musical experience or training to be included in the study. Professional singers who received significant singing training were excluded. Seven participants (three PD, three CD, and one HC) had previously sung in an amateur choir (e.g., church, community, or college choir) as adults. Other musical experience (e.g., playing an instrument) was noted as part of the musical background.
PD and CD study participants were recruited from local support groups or outpatient clinics of local medical facilities. Inclusion criteria for the PD participants were (a) a diagnosis of idiopathic PD, (b) no DBS implant, and (c) presence of mild to moderate hypokinetic dysarthria confirmed by a trained listener. Inclusion criteria for the CD participants were (a) a diagnosis of spinocerebellar ataxia confirmed from medical or neurology records, and (b) presence of mild to moderate ataxic dysarthria as verified by a trained listener. Exclusion criteria for both PD and CD participants included (a) premorbid speech and language impairments caused by neurological illnesses other than PD or cerebellar ataxia, and (b) a severe impairment in hearing.
The presence and severity of dysarthria were determined by a trained listener, who has extensive experience in dysarthria, based on telephone conversations with the participants and face-to-face conversations during visits. The ratings were made on a four-point scale with labels “0 = none,” “1 = mild,” “2 = moderate,” and “3 = severe.” Using the rubric of Hartelius et al. (2008) and Yorkston et al. (2004), mild dysarthria was characterized by an identifiable speech disturbance with minimal effect on intelligibility. Moderate dysarthria was characterized by intelligibility that was noticeably reduced. Participants' dysarthria was considered severe when their speech was mostly unintelligible.
To minimize the impact of medication, study participants with PD were asked to meet with the experimenter for the study approximately 2 hr after a single oral dose of carbidopa-levodopa (Sinemet). All study participants provided written informed consent. They completed language background and musical history questionnaires.
Three types of excerpts were prepared for this spontaneous utterances, spoken-to-a-model, and sung-to-a-model utterances, all of which served as test stimuli. Spoken-to-a-model and sung-to-a-model utterances were excerpted from a spontaneous speech corpus produced by healthy speakers for previously published studies (Sidtis et al., 2012; Sidtis & Sidtis, 2017; Van Lancker Sidtis et al., 2010). These healthy speakers from the earlier study were 11 right-handed healthy native speakers of American English (age 41–73 years) who served as age-matched controls for individuals with PD. The stimuli consisted of 85 different utterances (72 for experimental trials, three for video demonstration, and 10 for practice trials), each consisting of six to seven words, excerpted from transcribed samples of conversational speech obtained from older adult healthy speakers.
Spontaneous speech excerpts in this study were selected from discourse spoken by the study participants as part of the intake procedure. In all cases, speakers spoke about a topic of their choice, such as family, hobbies daily activities, or something else. The flowchart of the study's experimental design is illustrated in Table 4.
The three types of excerpts, spontaneous speech, spoken-to-a-model, sung-to-a-model shared the following They (a) ranged from six to seven words and eight to nine syllables in length (contractions, e.g., “they're,” were counted as one word; syllable count was based on standard pronunciation), (b) were singable in one breath, (c) were made up of a clause or phrase, and (d) were exclusive of proper nouns, fixed expressions, jargon, and infrequent words. Examples include “They travel to all different places,” and “He used to be a baseball umpire.” The full set of utterances is listed in Appendix A.
Familiar songs were utilized to allow for measures of pitch and rhythm and for competence ratings. A listing of songs offered to the participants consisted of seven well-known childhood and traditional songs (e.g., Happy Birthday, Deck the Halls, Silent Night, Amazing Grace, My Country 'Tis of Thee, Jingle bells, and Eensy Weensy Spider). Each participant selected two or three familiar songs from this list.
The protocols used in this study were approved by the New York University Institutional Review Board for studies involving human subjects. Each participant attended a single session, seated comfortably in a private testing booth at the university or a quiet room in the library or their home. All vocal samples were recorded using the same recording setting onto a Marantz digital recorder PMD661 via a microphone (Shure SM10A) placed at a distance of approximately 5–10 cm from the mouth. They were recorded at a 44.1-kHz sampling rate with a 16-bit precision.
The speech samples obtained from the participants for the experiment consisted of a set of three vocal modes (spontaneous, spoken-to-a-model, and sung-to-a-model utterances) and a set of familiar songs. For the spontaneous mode in this study, participants talked for 3 min about any topic; suggested topics were family, wedding, birthday, hobby, or job. To elicit spoken-to-a-model utterances, speakers were presented with a target utterance (test stimuli described above) in the center of the screen for 3 s and instructed to read the utterance silently; the sentence was then removed from the screen, and participants said the utterance as naturally as possible. In the sung-to-a-model condition, a different target utterance appeared on the screen, in the same manner. After the target sentence disappeared, study participants sang the utterance to any melody. They were told that it was not important what melody they took. For the familiar song production task, participants chose two or three familiar songs from a list.
1) Elicitation of spoken-to-a-model and sung-to-a-model utterances.
To record the spoken-to-a-model and sung-to-a-model utterances, each participant was presented with a 4-min training video, which features a demonstration of the research protocol. This procedure was followed by a practice session and an actual recording session. The purpose of the training video was to assure that uniform information was conveyed consistently by presenting the verbal instructions and demonstrating the elicited spoken-to-a-model and sung-to-a-model formats (see Supplemental Material S1 for the training video). In the video, the experimenter, acting as a trainer, provided a description of the behaviors required to perform the task and demonstrated how to perform the task using sample utterances (see Appendix B). The trainee in the video then rehearsed performing the task modeling the trainer's demonstration. The trainer provided corrective or supportive feedback to the trainee as they practiced performing the task. Three utterances were displayed on PowerPoint slides one at a time in the video, presented in the spoken mode first and then in the sung mode.
Immediately following the training session, practice sessions were conducted until the participants exhibited full understanding and appeared comfortable with the procedure (see Appendix C). The aim of the practice sessions was to ensure that although singing can be a daunting and possibly embarrassing experience for some people, participants would be at ease for this portion of the protocol. Once they were demonstrably able to perform as instructed (read a written utterance silently then speak it naturally; read a written utterance silently then sing out loud), they underwent 20 practice trials using 10 utterances (10 sentences × 2 modes [spoken-to-a-model/sung-to-a-model] = 20 trials) to demonstrate their ability to comply with the protocol. The order in which the practice items were presented was the same as in the experimental trials so that half of the speakers performed the spoken-to-a-model condition first followed by the sung-to-a-model condition; the other half of the speakers performed the condition in the order sung-to-a-model followed by spoken-to-a-model. Practice was repeated until the examiner was satisfied that the participants were well rehearsed and comfortable with the procedures.
Following the practice session, participants then produced the actual spoken-to-a-model and sung-to-a-model utterances for utilization in the listening study, described below. Each speaker produced six different sentences in spoken-to-a-model and sung-to-a-model modes, resulting in 12 trials (6 sentences × 2 modes) per each speaker. Therefore, each sentence was presented twice, one for spoken-to-a-model and one for sung-to-a-model mode of production. The order of the blocks was alternated across half of the speakers started with the spoken-to-a-model condition followed by the sung-to-a-model condition (spoken– sung), whereas the other half followed the reverse order (sung – spoken). These vocal samples provided recorded vocal excerpts for the listening study, where intelligibility measures were made. See Table 5 for the order of sentence presentation.
Through this procedure, each speaker group provided 144 12 PD speakers × 6 sentences × 2 modes = 144 samples, 12 CD speakers × 6 sentences × 2 modes = 144 samples, 12 control speakers × 6 sentences × 2 modes = 144 samples. PD speakers numbers 1–6 were assigned the same set of utterances as PD speakers numbers 7–12. Similarly, CD speakers numbers 1–6 and numbers 7–12 each received the same set of utterances. Stimulus utterances for control speakers consisted of the same subset of stimuli that were presented to the disordered speaker groups, half taken from utterances prepared for PD speakers (36 utterances) and half from CD speakers (36 utterances).
2) Familiar song production.
Following the elicited production of two vocal modes, PD and CD speakers sang two or three familiar songs of their choice from a list of seven well-known songs. The seven songs were Happy Birthday, Deck the Halls, Silent Night, Amazing Grace, My Country 'Tis of Thee, Jingle Bells, and Eensy Weensy Spider. The songs in the list were also sung by HCs.
Experiment I was designed to assess the percent intelligibility scores of the three vocal spontaneous, spoken-to-a-model, and sung-to-a-model. For the intelligibility listening study, utterances (excerpts) obtained from the three speaker groups, as in the recording protocol given above, were organized for listeners' orthographic transcription.
Listeners for the study were 28 adults (15 men, 13 women) who were native speakers of American English, born and educated in the United States. Listeners ranged from 18 to 58 (M = 31.7, SD = 10.1) years old. They were unaware of the purpose of the experiment and blind to the neurological group identity of the study participants. Inclusionary criteria for listeners were normal or corrected-to-normal vision and hearing with no history of speech or language disorders.
The utterances obtained from the recording procedure were prepared to be presented to the listeners. Eight audio samples (four spontaneous, two spoken-to-a-model, and two sung-to-a-model utterances) were selected from each speaker, generating a total of 288 samples (8 samples × 36 speakers). Specific criteria for how the utterances were selected are discussed in the following section.
1) Selection criteria for spontaneous speech excerpts.
Four utterances were excerpted from each study participants' 3-min sample of spontaneous speech. They were selected to fulfill two (a) constituency and (b) matching in duration and length. The first criterion of constituency required that utterances formed a coherent semantic and syntactic unit (or units), even though they did not necessarily have to be complete sentences. For example, from the original recording “there's 45 people in the cohort from all over the world,” “there's 45 people in the cohort” was extracted.
The second criterion required that that the selected spontaneous utterances were as equal as possible in duration and length (in words) to the spoken-to-a-model and sung-to-a-model counterparts. According to this criterion, the selected utterances for the three modes were roughly equivalent in mean duration and length (see Tables 6 and 7). An additional effort was expended to achieve a balance between utterance duration and number of words due to varying speech rates of PD and CD speakers. It was observed in these cohorts that in PD, speech rate was faster than in HC speakers, whereas speech rate was slower in CD. Therefore, in PD, utterances tended to include more words than in CD over the same amount of time. To accommodate the demands of a listeners' transcription task, a constant mean number of words per excerpt (utterance length) was maintained across the three speaker groups (M = 8.9 words). As a result, a difference in duration of .7 s emerged between the two study groups. Tables 6 and 7 provide the means and SDs of number of words and durations of spontaneous utterances by speaker group.
2) Selection criteria for spoken-to-a-model and sung-to-a-model excerpts.
For the spoken-to-a-model and sung-to-a-model excerpts, the experimenter elicited six original utterances; for the listening study, four were randomly selected from each speaker. Half of the utterances collected were delivered in spoken-to-a-model mode, and the other half in sung-to-a-model mode. Tables 8 and 9 display the means and SDs of number of words and duration for three speaker groups across the two conditions.
3) Creation of stimulus sets.
The collected 288 utterances were divided into two sets of 96 tokens each (Sets A and B) and another two sets of 48 tokens each (Sets C and D). Sets A and B consisted of audio samples produced by dysarthric speakers (i.e., PD and CD), and Set C and D consisted of samples produced by HC speakers. Each set contained utterances of three different modes (spoken-to-a-model, sung-to-a-model, and spontaneous) with a ratio of 1:2. To avoid practice effects in the listening protocol, utterances alternated spoken-to-a-model and sung-to-a-model modes between Sets A and B, and between Sets C and D, such that the two sets mirror each other in terms of modes. For example, utterances no. 1 and 2 were presented in spoken-to-a-model mode in Set A, and the same utterances were presented in sung-to-a-model mode in Set B. In a similar way, utterances no. 3 and 4 were presented in sung-to-a-model mode in Set A, and spoken-to-model mode in Set B. See Tables 10 and 11 for the configuration and details for each stimulus set.
During the intelligibility test, each listener was randomly assigned to receive one of four sets. The intention was to avoid potential practice effects associated with repeated presentations. Given that each spoken-to-a-model and sung-to-a-model sample was produced twice by each speaker, it was possible that the second occurrence of the same sentence might be easier for listeners to understand due to a practice effect. To avoid this issue, the utterances were carefully organized into sets so that they were not repeated for any listener within each set. Having each listener exposed to only one set also ensured that listeners heard each sentence only once. Furthermore, samples within each set were randomized differently for each listener to minimize order effects. The testing samples were presented randomly intermixed within each set, as opposed to separately in different blocks. The order of testing samples was randomized for each listener.
Listeners participated individually in a quiet space. They were told that they would hear a series of speech samples and some of them were not complete sentences. They were asked to listen to each of the samples and write down everything they heard including any repeated syllables and words on the answer sheet, using standard orthographic transcriptions. They were encouraged to make their best guess even when uncertain.
Prior to the experimental task, listeners were given an opportunity to adjust the volume to a comfortable level while listening to a spontaneous speech sample, produced by a healthy person who was not included in the actual test. Once the volume was set, listeners were instructed to keep this setting throughout until the end. Listeners were also given six practice trials to confirm their comprehension of the instructions. Transcriptions of the first two samples were provided for demonstration purposes.
The audio stimuli, excerpts of speech obtained from three groups of study participants, were presented using Praat (Boersma & Weenink, 2019) via headphones (Sony MDR-7502). After a listener finished each item, the experimenter clicked the mouse to advance to the next trial. Each sample was played only once. The listening session lasted between 40 min and 1 hr.
The intelligibility of the utterances was determined by scoring listeners' orthographic transcriptions of the audio samples. Each utterance was given an intelligibility score represented as “percent correct.” Each utterance was scored individually by comparing each written word to those in the target utterances produced by the speakers. Within each transcribed utterance, a score of 1 was assigned to a word if it phonemically matched the word in the target utterance. The number of points earned was tallied for each utterance, divided by the number of points possible, and then multiplied by 100 to yield a percentage of intelligible words. The mean percent intelligibility scores were pooled for each of the samples and then averaged across all listeners for statistical analyses.
Detailed scoring rules for transcriptions are (a) Homonyms (e.g., their and there) and spelling approximating correct English were accepted. (b) Adding or omitting the plural “–s” (e.g., son for sons), the regular third-person singular “–s” (e.g., live for lives), and the regular past tense “–ed” (e.g., travelled for travel) was regarded as correct. (c) A subject and a verb in a contracted form were scored in isolation by assigning 2 points, 1 point for each, so that a credit could still be awarded when only one of them was correct. For instance, if a listener wrote “I'll” instead of the target “She'll,” they would be given 1 point for correctly identifying a verb auxiliary “will.” (d) Contractions that combine a subject pronoun and a verb or verb auxiliary, such as “she'll,” and “you're,” were considered correct if a listener transcribed them in a full (articulated) form. (e) Pause fillers, such as “ah,” “uh,” and “um,” were ignored during scoring. (f) If all words were correctly identified but an additional word or words were inserted, a half point was deducted.
All speech samples contributed by study participants (three groups) were measured in terms of articulation rate, f
o coefficient of variation, intensity SD, and smoothed CPPS using Praat. For each utterance, articulation rate, expressed in syllables per second, was determined by dividing the total number of syllables by the total duration, excluding pause intervals longer than 500 ms (Martel-Sauvageau & Tjaden, 2017). Intensity SD values were measured from each utterance. Coefficient of variation of f
o was the SD of f
o divided by the mean, also measured from each utterance. A smoothed CPPS was measured only from the voice portions of the utterances. The voiced segments were manually detected by an experimenter based on visual inspection of the spectrogram in Praat. The Praat settings used to measure CPPS were adopted from Watts et al. (2017). Using R (R Core Team, 2020), a mixed-model analyses of variance (ANOVAs) was performed on each acoustic parameter, treating etiology (PD, CD, and HC) and vocal mode (spontaneous, spoken-to-a-model, sung-to-a-model) as fixed effects and participants as random effects. Significant interactions were further examined with Tukey's honestly significant difference.
To establish interrater reliability for the intelligibility data, intraclass correlation coefficient (ICC) estimates were calculated for all the trials included in each set. Using the Statistical Package for the Social Sciences (SPSS; IBM Corp, 2021), the two-way mixed-effects model was used with absolute agreement and the mean of multiple raters as the basis of assessment. The intelligibility scores from the seven judges were correlated to obtain the ICC for each set. With seven judges in each of the four sets, there were 28 judges (raters) in total. The ICC estimates were high for the intelligibility measures in Set A (.91), Set B (.94), and Set D (.85), indicating high reliability of measurements across the raters. The estimates were relatively lower in Set C (.6), indicating moderate reliability.
Statistical procedures were completed using R (R Core Team, 2020). Percent correct intelligibility transcription scores were included as a dependent variable in the analyses. A series of mixed-model analyses of covariance (ANCOVAs) were conducted on the percent correct transcription in R with a covariate of articulation rate. Fixed effects included the between-subjects factor of etiology (PD, CD, or HC), the within-subject factor vocal mode (spontaneous, spoken-to-a-model, or sung) and associated interaction. To account for the interindividual variability, speakers (i.e., individual subjects who participated as speakers) were prescribed as a random-effect factor. Post hoc pairwise comparisons, including simple effects, were performed following significant findings using Tukey adjustments. Each participant's four trials for spontaneous mode and two trials for spoken-to-a-model and sung-to-a-model modes were averaged prior to analysis. Comparisons between two sets were performed using SPSS (IBM Corp, 2021).
To examine the balance between the arrangements of the sets as configured for the listening study, the intelligibility measures of the companion data sets (Sets A and B, Sets C and D) were compared using paired-samples t tests. The analyses showed no significant differences in percent correct intelligibility, t(41) = 0.97, p = .34, between Sets A and B. The same results were obtained from the comparison of Sets C and D with respect to percent correct intelligibility, t(20) = −0.86, p = .4. These results confirm that the excerpts were randomly distributed across the sets, minimizing the effects of confounding variables. Based on these findings, the data from the sets were pooled.
The findings of the acoustic analyses revealed that articulation rate alone, among the acoustic variables examined, was relevant to the independent variables of interest in the study. The mean articulation rate of the PD, CD, and HC groups differed significantly as a function of the vocal mode. There were no significant differences on any of the other acoustic measures between the groups and the vocal modes in question. Therefore, in the subsequent intelligibility analyses, the articulation rate was included as a covariate.
The mixed-model ANOVA on articulation rate indicated a significant effect of etiology, F(2, 33) = 15.25, p < .001. Overall, the PD group spoke the fastest (4.37 syllables per second), the CD group spoke the slowest (3.07 syll/s), and the HC group spoke at a rate between the two groups (4.05 syll/s). The articulation rate for PD speakers (M = 4.37, standard error [SE] = 0.18) was 42.3% faster than for CD speakers (M = 3.07, SE = 0.12), t(33) = 5.3, p < .0001. The CD speakers spoke at a rate 24.2% slower than HC speakers (M = 4.05, SE = 0.17), t(33) = −3.99, p = .001.
The effect of mode was significant, F(2, 66) = 62.49, p < .001. The spontaneous samples exhibited the fastest articulation rate (4.29 syll/s), and the sung-to-a-model samples exhibited the slowest (3.01 syll/s), with the spoken-to-a-model samples in the middle (4.19 syll/s). The articulation rate was significantly slower in sung-to-a-model samples (M = 3.01, SE = 0.14) as compared with spoken-to-a-model (M = 4.19, SE = 0.17), t(66) = 9.3, p < .0001, and spontaneous samples (M = 4.29, SE = 0.15), t(66) = 10, p < .0001.
The interaction between etiology and mode was significant, F(4, 66) = 4.87, p < .0001. The simple tests with Tukey adjustment indicated reduced articulatory rates in the sung-to-a-model condition relative to the spoken-to-a-model and spontaneous modes for all speaker groups (all p < .05). Paired-samples t tests further investigating potential differences associated with spoken-to-a-model and spontaneous modes indicated that PD speakers spoke significantly slower in spoken-to-a-model mode than in spontaneous mode, t(11) = −3.07, p = .011, whereas no significant difference in articulation rate was found between spoken-to-a-model and spontaneous modes in CD samples, t(11) = −1.88, p = .087. These differences are illustrated in Figure 1.
Figure 1. Mean articulation rate (syll/s) produced by the Parkinson's disease (PD), cerebellar disease (CD), and healthy control (HC) groups for three modes. The figure illustrates the effect of mode (spontaneous: blue, spoken-to-a-model: red, sung-to-a-model: yellow) for each speaker group. Error bars represent standard errors. Asterisks indicate statistically significant differences as assessed by simple effects tests and paired-samples t tests (*p < .05).
o Coefficient of Variation (f
oCoV; Hz)
Etiology was found to have a main effect on f
oCoV, F(2, 33) = 3.64, p = .037, with f
oCoV highest in CD speakers (M = 17), followed by the HC (M = 17.9), and PD speakers (M = 12.6). There were no significant main effects of mode and interaction. See Figure 2.
Figure 2. Mean f
oCoV values produced by the Parkinson's disease (PD), cerebellar disease (CD), and healthy control (HC) groups for three modes. The figure illustrates the effect of mode (spontaneous: blue, spoken-to-a-model: red, sung-to-a-model: yellow) for each speaker group. Error bars represent standard errors. Asterisks indicate significant differences as assessed by simple effects tests and paired-samples t tests (*p < .05).
For the SD of intensity, the mixed-model ANOVA showed significant main effects of etiology, F(2, 33) = 13.02, p < .001, and of mode, F(2, 66) = 17.83, p < .001. Intensity SD was greater for the CD speakers (M = 12.26, SE = 0.33) than the PD speakers (M = 10.09, SE = 0.31), t(33) = −3.87, p = .001, and HC speakers (M = 9.55, SE = 0.22), t(33) = 4.82, p = .0001.
With respect to the effect of mode, the spoken-to-a-model samples (M = 10.65, SE = 0.34) exhibited significantly larger SDs than the spontaneous samples (M = 9.89, SE = 0.33), t(66) = 3.12, p = .008. The sung-to-a-model samples (M = 11.35, SE = 0.33) had a significantly larger SD of intensity compared with the spoken-to-a-model (M = 10.65, SE = 0.34), t(66) = −2.85, p = .02, and spontaneous (M = 9.89, SE = 0.33), t(66) = −5.97, p < .0001, modes.
A significant interaction effect between etiology and mode was found, F(4, 66) = 4.68, p = .002. The SD intensity differences among modes were particularly prominent for PD speakers. Specifically, for the PD group, the sung-to-a-model mode was associated with a larger SD of intensity relative to both spoken-to-a-model (p = .005) and spontaneous (p < .0001) modes. Paired-samples t tests showed that the intensity of the PD group was significantly more variable in the spoken-to-a-sung mode relative to the spontaneous mode, t(11) = 3.8, p = .003 (see Figure 3). No difference in intensity SD between spoken-to-a-model and spontaneous modes was found in the CD group.
Figure 3. Intensity SD (in dB) derived from the Parkinson's disease (PD), cerebellar disease (CD), and healthy control (HC) groups for three modes. The figure illustrates the effect of mode (spontaneous: blue, spoken-to-a-model: red, sung-to-a-model: yellow) for each speaker group. Error bars represent standard errors. Asterisks indicate significant differences as assessed by simple effects tests and paired-samples t tests (*p < .05).
There was a significant effect of etiology, F(2, 33) = 4.36, p = .02: The CD group (M = 13.04, SE = 0.38) obtained significantly higher CPPS values than the PD group (M = 9.73, SE = 0.49), t(33) = −2.59, p = .04. PD speakers (M = 9.73, SE = 0.49) exhibited significantly lower CPPS values than HC speakers (M = 12.96, SE = 0.87), t(33) = −2.53, p = .04. However, there was no significant difference in CPPS values between CD and HC speakers.
There was also a significant main effect of mode, F(2, 66) = 40.48, p < .001, with the sung-to-a-model samples (M = 14.02, SE = 0.57) exhibiting considerably higher CPPS values relative to the spoken-to-a-model (M = 11.1, SE = 0.61), t(66) = −7.12, p < .0001, and spontaneous (M = 10.61, SE = 0.68) modes, t(66) = −8.32, p < .0001.
A near-significant interaction between etiology and mode, F(4, 66) = 2.46, p = .054, indicated the higher CPPS in the sung-to-a-model mode as compared with the spoken-to-a-model condition for all speaker groups (all p < .05). Paired-samples t tests indicated that for the PD group, the spoken-to-a-model mode was associated with a higher CPPS, t(11) = 3.176, p = .009, as compared with the spontaneous mode, which reflects better voice quality with greater periodic energy in the voice in the presence of external cues. There was, however, no significant difference in CPPS values between spoken-to-a-model and spontaneous modes for the CD group. See Figure 4.
Figure 4. Mean cepstral peak prominence (CPPS; in dB) obtained from the Parkinson's disease (PD), cerebellar disease (CD), and healthy control (HC) groups for three modes. The figure illustrates the effect of mode (spontaneous: blue, spoken-to-a-model: red, sung-to-a-model: yellow) for each speaker group. Error bars represent standard errors. Asterisks indicate significant differences as assessed by simple effects tests and paired-samples t tests (*p < .05).
For percent correct transcription data, three ANCOVAs were conducted to compare two groups (PD-HC, CD-HC, and PD-CD) as a function of three vocal modes (spontaneous, spoken-to-a-model, and sung-to-a-model), with the covariate of articulation rate. A principal finding of this analysis was that transcription intelligibility scores were significantly lower for the spontaneous mode than spoken-to-a-model mode in PD and CD speakers, suggesting that speech tasks that provide an external model facilitate motor speech performance in the two groups. A second main finding was that sung mode conferred significant intelligibility benefits only to PD speakers, not CD speakers, as measured by transcription scores.
Mean percent correct intelligibility scores, as determined from transcriptions, are shown as a function of etiology and mode in Figure 5. The mixed-model ANCOVA with the between-groups factor of group (PD, HC) on percent correct intelligibility revealed a statistically significant main effect of both etiology, F(1, 29) = 76.85, p < .0001, and mode, F(2, 65) = 16.14, p < .0001, as well as an Etiology × Mode Interaction, F(2, 65) = 26, p < .0001, controlling for the effect of articulation rate. Focusing on the results of the PD group, the simple effects due to mode showed that PD speakers were more intelligible in the sung-to-a-model mode than spontaneous mode, t(40) = −6.58, p < .0001. Controlling for the articulation rate, PD speakers' sung-to-a-model utterances were also more intelligible than the spoken-to-a-model, t(53) = −3.98, p = .0006, again providing evidence for the intelligibility boost for sung mode in PD. PD speakers were more intelligible in the spoken-to-a-model mode versus spontaneous speech, t(72) = 6.55, p < .0001, confirming the beneficial effects of external cueing in basal ganglia dysfunction. In HC speakers, none of these differences were significant, suggesting that the speech intelligibility of neurologically healthy individuals was not affected by mode in this test in which ceiling effects were observed for this group.
Figure 5. Mean intelligibility scores as revealed in transcriptions, indexed by percent correct, across modes (blue: spontaneous, spoken-to-a-model, sung-to-a-model) for the Parkinson's disease (PD), cerebellar disease (CD), and healthy control (HC) speaker groups. Error bars represent standard error. Asterisks indicate statistically significant differences as assessed by simple effects tests (*p < .001).
The mixed-model ANCOVA with the between-groups factor of group (CD, HC) on percent correct intelligibility showed significant main effects of etiology, F(1, 67) = 66.41, p < .0001, of mode, F(2, 51) = 43.73, p < .0001, and an interaction between the two factors, etiology and mode, F(2, 55) = 51.18, p < .0001, controlling for the articulation rate. Focusing on the results of the CD group, the simple effects due to mode indicated that there was no significant difference in their intelligibility between the sung-to-a-model and spontaneous mode in CD speakers—in contrast to the results obtained for the PD speakers. CD speakers' intelligibility was worse in the sung-to-a-model mode compared with spoken-to-a-model mode, t(59) = 3.24, p = .0054, suggesting intelligibility boost provided by singing only in PD, not in CD. In line with the results reported for PD speakers, CD speakers were more intelligible during spoken-to-a-model mode than spontaneous mode, t(56) = 7.14, p < .0001, suggesting a significant task effect in CD speakers. Taken together, these findings indicate that the mode of singing had greater impact on the transcription intelligibility of PD speakers than CD speakers, and that a task effect occurs for spoken vocalizations in both groups.
The mixed-model ANCOVA with the between-groups factor of group (PD, CD) on percent correct intelligibility indicated no significant main effect of etiology, F(1, 33) = 2.8, p = .10, after controlling for the articulation rate. There was a significant main effect of mode, F(2, 62) = 38.34, p < .001, and a significant interaction between etiology and mode, F(2, 53) = 4.53, p = .015. An analysis of simple effects using Tukey adjustment revealed that PD speakers were more intelligible during sung-to-a-model mode than spontaneous, t(61) = −6.06, p < .001, and spoken-to-a-model mode, t(74) = −3.21, p = .005. They were more intelligible during spoken-to-a-model mode compared with spontaneous mode, t(62) = 5.33, p < .0001. These findings were consistent with those reported for ANCOVA on PD and HC samples demonstrating the effects of singing and external cues in promoting speech motor performance in the PD group. For CD speakers, sung-to-a-model samples were significantly more intelligible than spontaneous samples, t(71) = −5.95, p < .0001. There were no significant differences in intelligibility between sung-to-a-model and spoken-to-a-model samples, t(76) = −0.18, p = .981. CD speakers' spoken-to-a-model samples were significantly more intelligible than spontaneous samples, t(62) = 7.94, p < .0001.
The two ANCOVA analyses were carried out separately, adding speaker age and dysarthria severity (mild or moderate) as covariates to determine whether speakers' age and dysarthria severity influenced the intelligibility results. The results showed that adding these covariates did not significantly affect the final analysis results.
In order to further examine the extent to which the mode of singing may benefit PD and CD speakers, vocal productions of well-known melodies by the participants were evaluated on the pitch and rhythm dimensions as well as overall competence. These measures directly addressed these sung parameters in the two study groups, PD and CD. Familiar songs were subjected to ratings of pitch and rhythm by persons with musical experience. The competency rating served as a global and holistic judgment of participants' singing abilities. Pitch and rhythm ratings are more analytic measures that decompose singing into specific acoustic parameters. The two kinds of rating scales were used because they were presumed to be sensitive to different elements in participants' singing performance.
Twenty-two musically trained listeners participated in this experiment. All were native speakers of American English with an average of 14.1 years of musical training (range: 5–25, SD = 7.2). They ranged in age from 24 to 66 years (M = 37.5, SD = 13.3) and included 14 males and eight women. No one reported any previous diagnosis of speech or hearing impairment. None of the listeners had participated in the intelligibility portion of the study.
Recorded familiar songs were prepared for playback to musically trained listeners. The songs were trimmed to include only the first half of the lyrics in each song except Happy Birthday, which was included in full. See Appendix D for the corresponding portion of the lyrics used in the study.
For the rating test, recordings of two songs were taken from each speaker, which led to a total of 72 samples (2 songs × 36 speakers [12 PD + 12 CD + 12 HC]). Stimuli were divided into two sets of 36 songs each. The first set, Set A, contained songs produced by half of the speakers (speakers numbers 1–6), one third of whom were speakers with PD, one-third speakers with CD, and one-third HC speakers. The second set, Set B, contained songs produced by the other half of the speakers (speakers numbers 7–12) with the same proportions of PD, CD, and HCs as in Set A and Each listener was randomly assigned to one of two sets such that half of the listeners (10 listeners) received Set A and the other half (10 listeners) received Set B. The testing sample presentation order was randomly intermixed across the three speaker groups within each set and was also determined randomly for each listener.
Listeners performed ratings during individual sessions in a quiet room. Listeners were told that they would hear a series of familiar songs. They were asked to rate each sung sample on two measures, pitch accuracy and rhythm accuracy, utilizing a 5-point scale from poor (1) to excellent (5). Specifically, they were instructed to circle the number on the scale for each parameter after hearing each song. Stimuli were presented only once. Listeners were given six practice trials prior to the actual experimental task. The rating session was approximately 30 min long.
The interrater reliability of the rating was assessed by calculating ICC estimates of pitch and rhythm ratings in familiar songs. The pitch and rhythm ratings of the 11 judges were examined for each of two sets, resulting in 22 scores. An ICC estimate yielded values greater than .90 for all the ratings in each set (pitch rating in Set A: .98, rhythm rating in Set A: .96, pitch and rhythm rating in Set B: .97). These findings indicate a high degree of agreement among the raters.
The dependent measures were goodness ratings of pitch and rhythm in familiar songs. A series of Kruskal–Wallis ANOVAs were conducted on pitch and rhythm ratings with etiology (PD, CD, and HC) as a test variable in SPSS. Significant findings were further examined with post hoc pairwise comparisons with the Bonferroni adjustment. Comparisons between two sets were also performed using SPSS (IBM Corp, 2021).
To determine whether the two sets arranged for the listeners were comparable, a paired-samples t test was conducted. The results indicated that there was no significant difference between the accuracy ratings of pitch and rhythm in Sets A and B, t(65) = 0.95, p = .35, which suggests no systematic differences between the two sets, so that data from the sets were combined.
For the familiar song rating data, the three groups (PD, CD, and HC) were compared on their melody and rhythm ratings, respectively. Table 12 displays the median and interquartile range of ratings of pitch and rhythm of familiar songs produced by the same speakers who were evaluated in the intelligibility tasks. The results obtained from the rating task indicates that CD singing was rated as more impaired than that of PD with respect to both pitch and rhythm dimensions (see Figure 6).
Figure 6. Mean pitch and rhythm rating of accuracy in familiar songs sung by Parkinson's disease PD, cerebellar disease CD, healthy control HC speakers, measured on a 5-point accuracy scale from poor (1) to excellent (5). Error bars represent standard error. All six individual comparisons were statistically significant different (p < .05). PD = Parkinson's disease; CD = cerebellar disease; HC = healthy control.
The Kruskal–Wallis test revealed a statistically significant difference in pitch ratings across three etiology conditions, χ^2^(degrees of freedom [df] = 2, N = 66) = 50.76, p < .001, with HC group being rated higher on the accuracy scale (M = 3.68, SE = 0.31, Mdn = 3.75) at singing in tune, followed by PD (M = 2.70, SE = 0.35, Mdn = 2.71) and CD (M = 1.54, SE = 0.18, Mdn = 1.5) groups. Post hoc pairwise comparisons with Bonferroni adjustments indicated that HC speakers were more accurate in producing correct pitches for the given melody compared with both PD and CD speakers (both p < .005). PD speakers were rated as significantly more accurate than CD speakers with respect to this measure (p < .001; Figure 6).
Listeners also rated the rhythmic features of known melodies. The Kruskal–Wallis test also indicated a statistically significant difference in rhythm ratings across three etiology conditions, χ^2^(df = 2, N = 66) = 50.72, p < .001, with the highest ratings on the accuracy scale from 1 (poor) to 5 (excellent) for the HC group (M = 4.14, SE = 0.23, Mdn = 3.75), followed by the PD (M = 3.56, SE = 0.34, Mdn = 3.42) and CD (M = 2.06, SE = 0.25, Mdn = 2.13) groups. Post hoc pairwise comparisons with Bonferroni adjustments revealed group differences similar to those reported for pitch ratings. HC speakers produced rhythm with greater accuracy than PD and CD speakers (both p < .005). As was the case with the pitch ratings, ratings indicated that PD speakers produced more accurate rhythm on familiar songs than CD speakers (p < .001; see Figure 6).
Pearson product–moment correlation coefficients were computed to evaluate the relationship between (a) pitch ratings and intelligibility of sung-to-a-model mode, and (b) rhythm ratings and intelligibility of sung-to-a-model mode. Specifically, the purpose was to investigate whether speakers with better pitch and rhythm skills demonstrate better intelligibility with sung-to-a-model mode. There was a weak, positive correlation between pitch ratings and intelligibility of sung-to-a-mode mode; however, the relationship was not statistically significant, r(22) = .23, p = .277. Similarly, there was a weak, positive correlation between rhythm ratings and intelligibility of sung-to-a-model mode, but the relationship was not significant, r(22) = .3, p = .15. The pitch and rhythm skills did not appear to be correlated with intelligibility of the sung-to-a-model mode.
Experiment IIb enabled competency ratings of the sung productions (namely, sung-to-a-model phrases and familiar songs) as a function of etiology. This format allowed for a broader examination of competence in singing, as a vocal mode, addressing a key research question in this the effects of neurological impairment on singing ability in the two study groups.
Twenty-eight subjects were recruited for participation in the experiment. The participants averaged 29 years of age (SD = 11 years, 19–61 years) and included 10 males and 18 females. All were high school graduates (mean education = 16.6 years, SD = 1.7) with no history of neurological conditions. All but two participants spoke American English as their first language and were born and raised in the United States. The other two were born in other countries and moved to the United States at 2 and 6 months, respectively. The inclusion criteria excluded participants with a significant history of speech or language disorders (two participants had received speech therapy services in elementary school for remediating articulation disorders).
Sung-to-a-model utterances and familiar songs were both arranged for competence rating. The 72 sung phrases were split into two sets (i.e., Sets A and B) of 36 items, one third of which was composed of samples obtained from PD speakers, another third from CD speakers, and the final third from HC speakers. The 72 familiar song samples were also divided into two sets (i.e., Sets A and B) of 36 items, equally representing the three speaker groups.
The order of the two protocols was counterbalanced to eliminate order effects as a potential confound. As the experimental design involves two sets (A and B) and two tasks (sung-to-a-model and familiar song competence rating tasks), there were four possible (a) sung-to-a-model – Set A, familiar song – Set B; (b) sung-to-a-model – Set B, familiar song – Set A; (c) familiar song – Set A, sung-to-a-model – Set B; and (d) familiar song – Set B, sung-to-a-model – Set A. Participants were assigned to each of the four possible orders randomly. Since the number of participants (n = 28) was a multiple of 4, an equal number of participants (n = 7) completed each of the sequences. The stimuli were randomly ordered within each block independently for each participant.
On all sung-to-a-model rating trials, the participants listened to the audio samples and rated how good each sung sample is on a scale of 1 (very bad) to 7 (very good). The procedure for the familiar song rating protocol was comparable to that of the sung-to-a-model protocol. Raters' responses were recorded in a paper-and-pencil format. Each audio sample was presented only once. Participants completed two practice trials prior to the actual test session.
The dependent measures in both protocols were competence ratings. Using the SPSS software, separate Kruskal–Wallis ANOVAs were conducted for the two rating procedures. Pairwise comparisons with the Bonferroni adjustment were performed following significant findings. Comparisons between two sets were also performed using SPSS (IBM Corp, 2021).
The comparison of the two sets, using paired-samples t tests, revealed that there were no significant differences in competency ratings between the two sets of sung-to-a-model utterances, t(41) = −1.87, p = .072, and familiar songs, t(41) = 1.37, p = .178. This result validates the samples in one data set do not differ from the samples in another data set in any systematic way, controlling for any effects of potential confounding variables, allowing for combining the data for analysis.
Table 13 displays the median and interquartile range of competence ratings obtained for the three groups from each of the two protocols. The Kruskal–Wallis test revealed a statistically significant difference in sung-to-a-model competence ratings across three etiology conditions, χ^2^(df = 2, N = 84) = 39.16, p < .001. Post hoc pairwise comparisons with Bonferroni adjustment indicated that the PD, CD, and HC groups were different from each other in the expected direction. In particular, the HC group received significantly higher ratings than the PD, p = .001, and CD groups, p < .001. The PD group was rated significantly higher than the CD group, p = .018 (see Figure 7). The findings from this experiment suggest that the sung-to-a-model mode yielded greater improvements in singing ability for PD speakers compared with CD speakers.
Figure 7. Mean competence ratings of sung-to-a-model utterances and familiar songs for the Parkinson's disease (PD), cerebellar disease (CD), and healthy control (HC) speaker groups, measured on a 7-point scale from very bad (1) to very good (7). Error bars represent standard error. All six individual comparisons were statistically significant different (p < .05).
As in the sung-to-a-model rating protocol, there was a statistically significant difference in familiar song competence ratings across three etiology conditions, χ^2^(df = 2, N = 84) = 54.64, p < .001. The results of the post hoc comparisons were similar to those observed for sung-to-a-model competence ratings. As expected, the comparisons demonstrated significantly higher competence ratings for the HC group as compared with the PD, p = .001, and the CD group, p < .001. The results also revealed higher competence ratings for the PD group than the CD group, p < .001 (see Figure 7). These findings are compatible with those found in the sung-to-a-model competence ratings, suggesting greater benefits of singing for PD than cerebellar ataxia.
This study examined the perceptual, acoustic, and vocal characteristics of spontaneous, spoken-and sung-to-a-model, and spontaneously sung vocal production in speakers with PD and CD using transcription intelligibility as well as singing accuracy and competency ratings. The results of the Experiment I support our hypothesis that PD speakers benefited from the sung mode in vocal production more than CD speakers. This accords with the results from the Experiment IIa, in which CD speakers demonstrated greater impairment in pitch and rhythm in familiar melodies than PD and HC speaker. Additionally, it accords with the results from the Experiment IIb, in which singing competencies were rated higher in PD than CD speakers. Contrary to expectations, external cueing as represented in the spoken-to-a-model mode elicited higher intelligibility in speakers both with PD and CD, as compared with spontaneous speech.
Our data present evidence for a greater improvement of singing in PD than CD speakers, which was evidenced by increased transcription scores during the sung-to-a-model mode. In particular, under the sung-to-a-model mode, the speakers with PD scored as well as the HCs; on the other hand, the sung-to-a-model productions of CD speakers yielded significantly lower intelligibility scores in comparison to HCs. These findings are compatible with the perceptual rating protocols in Experiments IIa and IIb. The results of the familiar song ratings indicate a greater degree of impairment in producing accurate melodies and rhythm in CD speakers as compared with PD speakers. As pitch and rhythm are fundamental musical building blocks, investigating pitch and rhythm competencies of two clinical populations provides clues as to how basal ganglia and cerebellar dysfunction may influence singing ability.
Recent studies using the brain imaging technique provide corroborating data that indicate a potential cerebellar involvement in singing (Jeffries et al., 2003; Riecker et al., 2000). A positron emission tomography (PET) study conducted by Jeffries et al. (2003) showed relative blood flow increases in the normal cerebellum during singing. The increased blood flow in the cerebellum suggests the engagement of that area during singing. Riecker et al. (2000) reported the differential cerebellar activation patterns between speaking and singing in neurologically healthy participants. In that study, greater activation in the left cerebellum was found for overt singing contrasted with speech, whereas an opposite pattern was found for the overt speech contrasted with the singing greater right cerebellar activation.
Altogether, these findings suggest that the cerebellum may be important for some aspects of vocal motor control required in singing. It seems reasonable to interpret these findings in light of the different levels of vocal demands associated with talking and singing. Singing, where the linguistic information (the lyrics) and music are integrated into one performance, may require a different level or type of motor control and coordination than that typically required for speech. Considering its well-established role in motor coordination, it appears that the cerebellum plays a critical role in modulating coordination associated with the vocal control required in singing. Less competent performance in CD, which is associated with dysfunctional cerebella coupled with intact basal ganglia, alongside preserved singing competence in PD, having dysfunctional basal ganglia with intact cerebella, points toward an important contribution of the cerebellum to motoric aspects in singing.
The present findings of differential effects of singing in speakers with PD and CD raise a key unresolved issue concerning the neural mechanisms subserving talking and Do speech and singing rely on the shared or distinct cerebral systems? The question of one shared or two distinct neural mechanism of talking and singing remains an area of research. The proposition that singing and talking arise from a shared neural system is supported by the behavioral similarities between the two At a physiological level, both involve the same production breath support, vibration of the vocal folds in the larynx (Titze, 1994), and manipulation of oral-pharyngeal structures. Consistent with the idea of a shared system underlying the two vocal modes are findings of neuroimaging studies revealing activation in the overlapping brain areas during overt or covert singing and speaking in healthy individuals (Brown et al., 2006; Callan et al., 2006; Callan et al., 2007).
However, in studies of neurologically impaired individuals, evidence points to a dissociation between the two vocal modes; over the last 2 centuries, it has been consistently reported that the components underlying the ability to sing or talk can be selectively impaired in such individuals (Van Lancker Sidtis et al., 2021), while leaving the complementary vocal mode intact. In addition to the clinical observations and lesion studies reviewed in the introduction, neuroimaging studies provide further support for the distinct neural mechanisms underpinning the two vocal modes (Jeffries et al., 2003; Riecker et al., 2000). Using PET, Jeffries et al. (2003) showed differential activation patterns during reciting or singing words of a familiar increased activation of the left hemisphere during speaking condition, and increased activation in the right hemisphere, medial and dorsolateral prefrontal cortices, and the cerebellum under singing condition. In this study, the data revealed enhanced articulatory vocal output in speakers with PD while singing, but not in speakers with CD, which suggests that talking and singing are governed by different cerebral systems. Although more data and studies are clearly needed on this issue, the results of this study provide empirical evidence obtained from two clinical groups with different lesion loci, supporting distinct cerebral systems underlying the two vocal modes, singing and talking.
This study has shown that transcription intelligibility scores obtained from samples produced by speakers with PD and CD were significantly better under the spoken-to-a-model mode as opposed to the spontaneous mode. Given the study conditions of the spoken-to-a-model mode, where the target utterances were represented in written form and exposed to the participation and thus constitute an external cue, these results suggest a similar effect of external cueing in both clinical groups. These findings extend beyond the earlier view that external cueing was seen to aid speech production primarily in persons with basal ganglia disease, and agree in part with reports of Weir-Mayta et al. (2017). It can be speculated that the positive effects of external cuing on the two clinical groups merit differing explanations, based on their distinctive dysarthric characteristics.
The present findings of enhanced speech motor competence in PD in the presence of external models are in accordance with those reported in the works of Kempler and Van Lancker (2002), Van Lancker Sidtis et al. (2012), and Weir-Mayta et al. (2017). These and related findings contradict a long-held assumption in speech motor literature, which held that speech motor control in the dysarthrias, including in PD and CD, is independent of mode or tasks (Shames & Wiig, 1990; Yorkston et al., 1988). This line of reasoning taught that motor performance in this clinical population is unlikely to vary across spontaneous speech, reading, or repetition (Cannito & Marquardt, 1997; Kent et al., 2000). The finding of a task effect, whereby measures of speech differ depending on vocal modality or task, has generated at least two explanations for the phenomenon. One account concerns a modulatory role of the basal ganglia in planning, initiation, and modulation of movement during various motor gestures including gait, arm reach, and speech (Gurney et al., 2001). Such a view entails that the damaged basal ganglia “benefits” from an external model. When a spoken or written model is provided, the basal ganglia system encounters less demand to prepare, execute, and monitor a motor plan for the verbal gesture. In this study, the spoken-to-a-model mode, PD speakers may have relied on external models that reduced the vocal motor control burden imposed on the basal ganglia system. In contrast, PD speakers' poorer performance under spontaneous mode may be attributable to a poverty of the internal models generated by the compromised basal ganglia. These interpretations are well supported by studies of gait and limb motor gestures in individuals with PD that show improved gait, posture, and movement with the presentation of external cues (Georgiou et al., 1993; Ginis et al., 2018; Morris et al., 1996; Rocha et al., 2014).
As an alternative explanation has been advanced that externally cued tasks require less cognitive–linguistic demands than spontaneous speech (Bohland et al., 2010; Feenaughty et al., 2013; Huber & Darling, 2011; Tourville & Guenther, 2011; Weir-Mayta et al., 2017). While there may be merit in this proposition, putting aside the inherent difficulty in obtaining reliable measures of “cognitive load,” there are caveats with respect to the cognitive load account. This account does not align well with results from studies of gait and limb movement, also like speech, complex movement gestures. For most people, walking and arm reach are highly routinized daily movements. The efficacy of external cueing on gait was shown to be potent in the face of cognitive impairment (Rochester et al., 2009) and general motor control dysfunction (Gräber et al., 2014). Furthermore, the influence of cognitive function on speech measures in multiple sclerosis shows conflicting findings (Feenaughty et al., 2020; Rodgers et al., 2013). Finally, it can be argued that reading, a later acquired skill, imposes a cognitive load in its own right.
Acoustic data in measures of the task effect lend support to the external model account (Kempler & Van Lancker, 2002; Van Lancker Sidtis et al., 2010, 2012). Vowel quality was better, as measured by signal-to-noise ratio, in read and repeated than spontaneous speech (Kempler & Van Lancker, 2002; Van Lancker Sidtis et al., 2010). These acoustic results, organized deeply in the vocal–motor system, appeared to arise not from a role of cognitive effort, but from intrinsic, physiological vocal–laryngeal processes, as mediated by the basal ganglia.
This study found better performance under the spoken-to-a-model mode in CD, compared with the spontaneous mode. This finding also contradicts the previously held believe of consistency of dysarthric features despite mode of speaking. A recent study by Weir-Mayta et al. (2017) also reported that individuals with CD were differentially affected by reading (external cue) as compared with conversation tasks. Based on listener's ratings, speakers with CD were significantly more understandable during the reading task relative to the conversation task. This task-related difference, on the other hand, did not hold for listeners' ratings of No significant differences were found for the naturalness ratings between the two tasks in CD speakers. Again, these results raise the question of CD dysarthric speech characteristics, namely, slow rate and loudness, contributing to the higher intelligibility. In this study, using an objective measure of intelligibility (percent correct in a transcription task), speakers exhibited improvements during the spoken-to-a-model (external cue) mode as compared with the spontaneous mode.
Despite comparable performance in intelligibility between the two groups, it is clear that the dysarthrias in individuals with PD and CD differ in important characteristics, including rate, amplitude, articulatory precision, prosody, and pausing. It is likely that several or all of these differences in dysarthric quality, characteristic of the respective dysarthric profiles in PD and CD, contribute, in part, as suggested by the analysis that covaried rate data. This remains to be investigated in greater detail. Further acoustic analysis may reveal coherent explanations for the observed differences in intelligibility results in the study, uncovering other ways that the basal ganglia and the cerebellum operate differently (Schwartze et al., 2012).
There are several limitations in this study that may influence interpretation of the study findings. First, our study maintained constant recording playback levels, but did not use a reference tone at the time of recording. A calibration tone should be used in future research to carefully manage the examination of intensity in a more orderly way. A second limitation pertains to the fact that estimation of level of dysarthria (mild or moderate) was made by the examiner. It would be advantageous to have a panel of speech-language pathologists experienced in the field of motor speech disorders to determine the level of dysarthria for an unequivocal assessment. Another drawback pertains to the fact that intelligibility was measured using a transcription-based percent score. Scaled intelligibility based on a visual analog scale might provide additional information about listeners' perceptions of dysarthric speech across vocal modes (Stipancic et al., 2016; Tjaden & Wilding, 2011). A minor point is that a small proportion of the study participants (in all three groups) reported amateur choir singing experience, the rest did not. Although it is not likely that these backgrounds played a role in performance, perhaps this factor merits closer control in a future study.
In summary, this article investigated the nature of neural motor control underpinning four different vocal modes in individuals with PD and CD. This study offers a window on how individuals with basal ganglia and cerebellar dysfunction utilize external cues in various speech tasks. The results substantiate the differential production efficiency in spoken-to-a-model mode versus spontaneous mode in the two clinical groups. Future studies are needed to provide more information on how the basal ganglia and cerebellum contribute to different speech measures at various levels.
In addition, this investigation directly addresses the quality of sung production in people with cerebellar dysfunction. The data described here strongly supports a key role of the cerebellum in vocal motor control to produce singing and, hence, offers more perspective on the neural mechanisms underpinning this vocal behavior, supporting the notion of functionally distinct neural mechanisms devoted to speech and singing.
The current results concerning the efficacy of external cueing demonstrate the need to recognize the task effect when clinically evaluating dysarthria. Individuals may vary in the manifestation of their dysarthria (e.g., severity levels) depending on the type of speaking task being used (Sussman & Tjaden, 2012). For the clinician, scaled estimates of intelligibility have been proposed to estimate intelligibility in spontaneous speaking competence from measures of reading (Tjaden & Wilding, 2011). In addition, the finding that singing was more impaired in cerebellar disease than in basal ganglia disease can be of interest in treatment planning. Singing practice may benefit ataxic speakers. In contrast, singing proves a more efficient vocalization in individuals with PD than speaking (Harris et al., 2016; Kempler & Van Lancker, 2002). For all these individuals, treatment with singing has the potential for allowing for vocal practice, which could lead to increased confidence and social participation, thus contributing to overall quality of life.
The data sets generated during this study are available from the corresponding author on request.
This work was completed in partial fulfilment of the requirements for the degree of doctor of philosophy at New York University (NYU) by the first author. It was supported by Doctoral Research and Travel Grant from NYU Steinhardt and a grant from the National Institute of Deafness and Communicative Disorders (R)1 DC 007658. Portions of this work were presented at the 2021 American Speech-Language-Hearing Association convention. The authors extend their sincere appreciation to Maria Grigos, Christina Reuterskiöld, and Aaron Johnson for their insightful comments on this research project.
This work was completed in partial fulfilment of the requirements for the degree of doctor of philosophy at New York University (NYU) by the first author. It was supported by Doctoral Research and Travel Grant from NYU Steinhardt and a grant from the National Institute of Deafness and Communicative Disorders (R)1 DC 007658. Portions of this work were presented at the 2021 American Speech-Language-Hearing Association convention.
The data sets generated during this study are available from the corresponding author on request.