Authors: Mutasim Al-Deaibes, Bassil Mashaqba, Anas Huneety, Mohammed Nour Abu Guba
Categories: Article, Jordanian Arabic, VOT, Stops, Voicing, Laryngeal contrast, Language development
Source: Journal of Psycholinguistic Research
Authors: Mutasim Al-Deaibes, Bassil Mashaqba, Anas Huneety, Mohammed Nour Abu Guba
This study reports the effects of gender and age on the production of Voice Onset Time (VOT) of stop consonants in Rural Jordanian Arabic (RJA). Participants of the study were divided into four age groups, namely children, preadolescents, adolescents, and adults, and were equally stratified according to their gender. They were asked to produce a series of Arabic words beginning with one of the six stop voiceless /tˤ/, /t/, /k/ and voiced /b/, /d/, /ɡ/. The results show that voiceless stops were characterized by a long lag (aspirated, positive VOT), and voiced stops were characterized by a long voicing lead (prevoiced, negative VOT). Across all age bands, the results of the study indicated that females have significantly longer VOT durations than males for both the voiced and voiceless stops. In addition, children had significantly the longest VOT duration as compared to preadolescents, adolescents, and adults across the board. Notably, the VOT duration of both voiced and voiceless stops shortened with increasing age; the younger the age, the longer the VOT is, suggesting that VOT production in RJA is gradually incrementally developing.
The present study reports on an experiment examining the production of stopconsonants by Rural Jordanian Arabic (RJA) speakers. Stop consonants were chosen as a window through which we study the effect of age and gender of four age children, preadolescents, adolescents, and adults in RJA. The acoustic parameter we adopted to evaluate RJA stop productions is voice onset time (VOT), the most reliable acoustic parameter in distinguishing between voiced and voiceless stops in word-initial and word-medial syllables in a wide array of languages. However, the acoustical realization in the two voicing categories has been controversial as whether the laryngeal contrast depends on acoustical features such as voice, aspiration, glottal aperture, or duration of articulatory stricture (Kohler, 1984; Ladefoged & Maddieson, 1996).
VOT is defined as the interval between the release of a stop consonant and the onset of the first glottal pulsing of the following vowel (Cho & Ladefoged, 1999; Li, 2013; Lisker & Abramson, 1964; Oh, 2019, to mention a few). It is an acoustic correlate that is experimentally used to measure the timing between the larynx structures and the oral cavity. Stop consonants are characterized by creating a pressure pulse during the complete occlusion in the vocal tract. There are three phonetic categories of VOT: negative or zero VOT, short positive VOT (< 40 ms), and long positive VOT (also known as long aspiration) (≥ 40 ms). There are two types of VOT classification in the Arabic dialects as reported in the the first group of studies reported that voiced plosives are produced with lead voicing and that the voiceless plosives are produced with short-lag as in Lebanese Arabic (Yeni-Komshian et al., 1977) and Jordanian Arabic (Mitleb, 2001). The second group of studies distinguishes between voiceless unaspirated and voiceless aspirated stops, as in Iraqi Arabic (Al-Ani, 1970) and Najdi Dialect (Alghamdi, 1990). Therefore, in the two groups of studies, it is conspicuous that the voiced stops in all Arabic dialects are truly voiced, with the vocal cords adduction and vibration occurring before the release of the stop, while the voiceless stops have more variation (short lag vs. long lag).
This study is organized as Sect. “Background” provides an overview of VOT variation depending on different factors. Sect. “Materials and Methods” presents the materials used to collect and analyze the data. Sect. “Results and Findings” will lay out the findings of the experiments. Sect. “Discussion and Conclusions” discusses the results and offers a conclusion to the study.
Voice onset time values vary from one language to another based on a number of phonetic and socio-phonetic factors such as voicing (Iverson & Salmons, 1995; Lisker & Abramson, 1964), place of articulation (Cho & Ladefoged, 1999; Kessinger & Blumstein, 1997; Volaitis & Miller, 1992), vowel quality (Fischer & Goberman, 2010; Morris et al., 2008; Pind, 1999; Whiteside et al., 2004), rate of speaking (Beckman et al., 2011; Miller et al., 1986; Volaitis & Miller, 1992; Baum & Ryan, 1993; ), age (Bóna, 2014; Morris et al., 2008; Robb et al., 2005), gender (Li, 2013; Morris et al., 2008; Robb et al., 2005; Ryalls et al., 1997; Whiteside et al., 2004), among others.
Recently, speaker’s gender has been the most studied extralinguistic factor cross-linguistically (Li, 2013). However, most of the studies on the effect of gender on VOT were primarily on adult speakers. Generally, the results on the effect of gender on VOT in adults, including the ones on Arabic dialects have been somewhat conflicting, inconsistent, and far from unanimous. For instance, some studies reported that female speakers tend to have longer VOT values for stops than male speakers (Morris et al., 2008; Robb et al., 2005; Ryalls et al., 1997; Whiteside et al., 2004). Those studies attributed these conclusions to physiological, anatomical, and aerodynamic differences between males and females (Koenig, 2000) or sociophonetic factors such as ‘carefully articulated speech’ (Mattingly, 1966; Whiteside & Irving, 1998). Other studies, however, reported insignificant effects of gender on VOT (Yu et al., 2015) or longer VOTs in males (see Oh, 2019). The latter is ascribed to only sociophonetic factors, excluding the biological and anatomical ones. As for studies on VOT in Arabic dialects, their findings were also inconsistent. For example, Khattab et al. (2006) reported a significant link between VOT values and gender in Lebanese Arabic. However, other studies reported no significant differences between gender VOTs, such as those of Abudalbuh (2010) and AlDamen and Al-Deaibes (2023a, b) for Jordanian Arabic and Almbark (2008) for Syrian Arabic. It is worth pointing out here that the sociophonetic factor of ‘carefully articulated speech’ has been employed by Li (2013) to account for the reason that females produce short VOT while Oh (2019) has used the same exact factor to account for the reason that females produce longer VOTs.
Although most research on gender differences and VOT has been conducted on adults, little is known about gender differences and VOT in children. Investigating VOT in children’s speech is, however, crucial for researchers as it provides us with important data on how this aspect of motor speech skill is being refined (Whiteside et al., 2004), and it helps researchers in clinical phonology to compare normal-hearing children with hearing-impaired children. Most studies have reported that VOT duration and age differences are observed between young people and adults. For example, Yu et al. (2015) reported that gender differences were observed in voiceless stops in children between 8 and 11 years old, where boys had longer VOT durations than girls. However, they found no significant gender differences in VOT durations in adults. Conversely, Whiteside and Marshall (1998) reported significant gender differences in VOTs in children, where girls had longer VOT durations than boys. Bòna (2014) reported age differences in VOTs of voiceless plosives produced by Hungarian young and old subjects. She concluded that with older age, the means of the VOTs of /p/ and /k/ decreased, and the mean of the VOTs of /t/ increased. Also, Whiteside and Marshall (1998) found differences in VOT values depending on age. For example, they reported /p/ and /t/ were shorter with increasing age. However, for bilabial /b/ and /d/, 7-year-old children produced shorter VOT compared to the other age groups namely 9 and 11. This shows no consensus on whether VOT is longer or shorter in young and adult speakers.
As the production of the human voice undergoes ongoing changes throughout our lives, those changes are attributed to physiological and anatomical factors that change the speech mechanism. Therefore, age has been reported in the literature to affect the duration of VOT, and age differences have been observed between young and adult individuals. However, no unanimous results have been reported with regard to the age difference in VOT durations between children and adults. For example, it has been found that four-year-old children have longer VOT durations than adults (Barton & Macken, 1980; Gilbert, 1977; Ma et al., 2018; Menyuk & Klatt, 1975; Smith, 1978). On the other hand, some studies reported that children between 4 and 6 years old produced shorter VOT durations than adults (Bòna, 2014; Kewley-Port & Preston, 1974; Macken & Barton, 1980; Zlatin & Koenigsknecht, 1976). Given that most reported studies focused only on the production of voiceless stops by children, it would be essential to determine whether children, preadolescents, adolescents, and adults of both genders in this dialect can produce the voiced stops with prevoicing. Therefore, this study goes further and reports on both voiced and voiceless stop productions in two genders and four age groups.
In this respect, it is worth pointing out that RJA is an understudied variety of Arabic. Much research has not been done on this variety to document any phonetic realization of VOT or other phonetic phenomena to the best of our knowledge. Jordanian Arabic (JA) is characterized by a variety of sub-dialects. The first is Urban Jordanian Arabic, which is predominantly used by city dwellers who are originally Palestinians or Syrians who migrated to Jordan in the past two centuries due to wars, expulsion, and political unrest. The second is Rural Jordanian Arabic, primarily used by villagers who live in the suburbs and the countryside of northern Jordan, like the districts of Al-Ramtha, Bani Kanana, Bani Ebeid, and Al-Koura and the city of Der’aa and its villages in Syria. This variety is also spoken by city dwellers, originally from the suburbs but have moved to the city. The third is Bedouin Jordanian Arabic, used by desert dwellers (originally descendants of nomadic tribes who currently live in settled communities) in the northern, eastern, and southern parts of Jordan in the districts of Mafraq, Karak, Ma’an, and Al-Badiya Al-Shamaliya (See Al-Deaibes and Rosen, 2022 for information about JA sub-dialects). A major distinction between these three sub-dialects is primarily phonetic, although there appears to be some lexical and syntactic variation to a lesser extent. It is also worth mentioning here that RJA, like other Arabic dialects, has a set of six stops that differ from each other in phonation and places of articulation. These consonants are the voiced /b/, /d/, /ɡ/, and the voiceless /t/, /k/, /tˤ/. As for /tˤ/, it is a pharyngealized voiceless alveolar stop that phonemically contrasts with its plain counterpart /t/. Pharyngealization, an articulatory feature that characterizes Semitic languages, refers to consonants that are articulated with a secondary constriction in the posterior vocal tract and a primary constriction usually in the dental/alveolar oral tract (Jongman et al., 2011; Jaradat et al., 2023; AlDamen and Al-Deaibes, 2023a, b; , among many others).
The importance of this study is twofold. First, it fills the gap in the literature, given that little is known about the phonetic realization of VOT or other phonetic phenomena in this dialect. Second, it investigates the patterns of consonant articulation in RJA across 6 voiced and voiceless stop consonants with respect to age and gender. Therefore, the study provides answers to the following research Does RJA exhibit a laryngeal contrast between prevoiced and voiceless aspirated stops?Do gender and age play a role in affecting the duration of VOT?
Sixty-four speakers participated in the experiment. Sixteen participants (eight males, eight females) were adults aged between 21 and 26 (mean age = 23 years, SD = 1.2), sixteen participants (eight males, eight females) were adolescents aged between 13 and 15 (mean age = 14 years, SD = 2.6), sixteen participants (eight males, eight females) were preadolescents aged between 11–12 (mean age = 12 years, SD = 1.3), and sixteen participants (eight males, eight females) were children aged between 8–9, (mean age = 9, SD = 1.8). All participants were native speakers of RJA who reported no history of speech or hearing impairment. They were born and raised in Jordan, and they (and the children’s parents) reported that they had not lived outside of it except for occasional short vacations and that none had a speech or hearing impairment. To protect the participants’ confidentiality and to meet research ethics, a consent form was given to the participants before collecting the data to explain the general purpose of the study and to seek their permission to record the stimuli of the study.
The stimuli consisted of word-list data comprising 18 word-initial monosyllabic (CVVC) words (see appendix A) for a total of 2304 words (18 tokens × 2 times × 64 speakers = 2304). Upon obtaining the research ethics approval (H22-034), the participants were asked to read the words (provided on a PowerPoint slide) at a normal tempo and how they would use them in everyday communication. The words were written in Standard Arabic (without providing any diacritics), and the participants were advised to familiarize themselves with the words before the recording took place. If participants were not sure about a word, they were advised to ask the researcher. A computer program automatically changed the screen to display one word every three seconds. Each participant repeated the words three times for the recording. Participants’ first two repetitions were analyzed, and their third repetition was analyzed when errors in VOT, such as production errors, were found. This computer-controlled presentation helped have a consistent speech rate between and within participants. The recording was performed using an Audio-Technica AE4100 cardioid microphone and a TubeMP preamp connected directly to a PC. The microphone was placed 25 cm away from the participants’ mouths. The recordings were done as mono sound, and the sampling rate was set to 44,100 Hz with 16-bit digitization and then stored as Waveform Audio File Format.wav files.
Segmentation was performed manually based on Praat (Boersma & Weenink, 2020) waveform, spectrogram representations, and standard segmentation criteria. Measurements were collected using custom-written Praat scripts after all tokens were manually segmented in TextGrids by interpreting the waveforms and the spectrograms. All measurements and segmentation were done by simultaneously evaluating the waveforms and spectrograms. The onset of the initial consonant was taken as the first spike of the burst release. For each stop, VOT was determined as the period between the onset of the release burst (release burst and the first clear glottal pulse) of the stop and the onset of the glottal vibration (Lisker & Abramson, 1964). The onset of the vowel was determined when F1 was clearly visible in the spectrogram, while the vowel’s offset was determined when F2 disappeared from the spectrogram. For each stop, VOT was regarded as the length of time that passes between when a plosive is released and when periodicity begins. If voicing begins after the release, VOT is positive, and if it begins before the release, VOT is negative. The release of the stop was determined based on the appearance of the burst of noise, and the beginning of the voicing was determined at the zero-crossing of the first wave cycle of the vibration that accompanies the production of the vowel . Figure 1 illustrates how the segmentation and the labeling of the VOTs were performed.Fig. 1Segmentation and labeling of the stop /t/in the word [tʰu:t] ‘berry’
The data of the current study consists of one dependent variable (VOT in ms) and three independent target consonant (t, tˤ, k, b, d, ɡ), gender (male, female), and age (child, pre-adolescent, adolescent, adult). Additionally, age and gender were between subjects’ factors while target consonant was a within-subject factor. Therefore, Factorial ANOVA was used to address the study research questions. The statistical analyses were done in SPSS version 22 and R version 4.1.2 facilitated by the packages ggplot2, readxl, and tidyverse. Tukey’s HSD post-hoc analysis was used to demystify the significant age groups. It is important to note that post-hoc analysis was not done on the gender groups since there were less than three gender groups. The results of the data analyses were presented graphically and in tabular form and interpreted into the letter in the results and findings section below. The data analysis and results were done separately for the voiced and voiceless voice stops.
The results emerging from the statistical analysis test showed that voiceless stops in RJA had a positive VOT (mean = 56 ms) and voiced stops had a negative VOT (mean = −46 ms) across the board. The mean and standard deviation were used to summarize the data at first glance. As shown in Table 1 below, the mean VOT (in milliseconds) for females was greater than that of males for both the voiced and the voiceless consonants.Table 1Mean VOTs and standard deviation (in parenthesis) for each consonantRow Labels/t//tˤ//k//b//d//ɡ/**Female64(14)35(11)76(18)−58(25)−64(14)−65(22)**Child70(12)43(9)90(12)−64(26)−71(11)−77(17)PreAdol74(12)46(7)88(6)−64(25)−69(12)−75(17)Adolescent59(11)27(3)72(13)−61(25)−69(14)−57(25)Adult51(10)26(6)56(15)−41(20)−53(14)−50(17)**Male45(11)28(10)48(14)−30(23)−38(16)−24(8)**Child54(6)39(7)56(15)−36(25)−43(17)−20(6)PreAdol52(9)34(4)55(15)−27(28)−41(20)−23(8)Adolescent40(6)21(2)43(9)−32(18)−38(13)−32(7)Adult34(4)18(3)37(7)−23(22)−30(11)−20(6)
Additionally, on average, the VOT duration for the children was longer than that of preadolescents, adolescents, and adults for both females and males across the board. However, to ascertain significant differences in VOT duration based on age and gender, two statistical analyses (Factorial ANOVA) were conducted separately for voiced and voiceless consonants to ensure exhaustive data analysis. These results are also presented graphically in the boxplots in Fig. 2.Fig. 2VOT duration for each age group and gender across the board
Factorial ANOVA analysis, as shown in Table 2 below, suggested that there was a statistically significant difference in the VOT for the voiceless consonants based on the gender of the participants (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ F\left(\mathrm{1,264}\right)=264.1,p<.001,{\eta }^{2}=0.5
\usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=58.381\pm 0.784, SD=22.5429) $$\end{document} having longer VOT durations than males (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=40.372\pm 0.784,SD=14.5809 $$\end{document}). This is further supported by the scatterplot in Fig. 3.Table 2Factorial Anova siting differences in VOT duration based on gender and ageSourceType III Sum of SquaresdfMean SquareFSigPartial Eta SquaredCorrected Model103,080.899^a^234481.77850.692.000.815Intercept702,152.0011702,152.0017941.834.000.968TargetC47,889.550223,944.775270.832.000.672Gender23,349.605123,349.605264.100.000.500Age24,571.47038190.49092.640.000.513TargetC*Gender5323.63622661.81830.107.000.186TargetC*Age752.9766125.4961.419.207.031Gender*Age506.1483168.7161.908.129.021TargetC*Gender *Age687.5156114.5861.296.259.029Error23,340.72026488.412Total828,573.620288Corrected Total126,421.619287^a^R Squared = .815 (Adjusted R Squared = .799)Fig. 3Scatter plot of VOT duration of voiceless stops for females and males Descriptive statistics results in Table 3 show that children had a longer VOT duration (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=58.671\pm 1.108,\mathrm{SD}=20.1551 $$\end{document}) as compared to adolescents ((\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=43.787\pm 1.108, $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ SD=19.5292 $$\end{document}) and adults (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean = 37.112 \pm 1.108, $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {\mathrm{SD}} = 15.6259 $$\end{document}). However, children and preadolescents did not exhibit any significant difference. These results are supported by the scatter plot in Fig. 4, which shows that children had longer VOT stops than all other groups across the board. A statistically significant difference in VOT based on age was also established (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ F\left(\mathrm{3,264}\right)=92.640,p<0.05, {\eta }^{2}=0.513 $$\end{document}).Table 3Summary statistics for VOT duration based on gender and ageMeanStd. errorStd. deviation95% Confidence intervalLower boundUpper bound*Gender*Female58.3810.78422.542956.83859.923Male40.3720.78414.580938.82941.915*Age*Child58.6711.10820.155156.48960.853PreAdol57.9351.10820.067955.75360.117Adolescent43.7871.10819.529241.60645.969Adult37.1121.10815.625934.93139.294Fig. 4Scatter plot of VOTs across all age groups A significant difference in the VOT duration was also observed based on target consonant (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ F\left(\mathrm{2,264}\right)=270.832,p<.001,{\eta }^{2}=0.672 $$\end{document}). Tukey’s HSD post-hoc analysis (See Table 4) indicated that all the consonants were significantly different from each other with velar /k/ having significantly longer VOT duration (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=62.195\pm 0.96, $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ SD=21.5834 $$\end{document}) than both coronal /t/ (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=54.202\pm 0.96, $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ SD=15.5042 $$\end{document}) and /tˤ/ (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=31.733\pm 0.96, $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ SD=10.9740 $$\end{document}).Table 4Tukey’s post-hoc analysis for target consonants of the voiceless stops(I) TargetC(J) TargetCMean Difference (I-J)Std. ErrorSig95% Confidence IntervalLower BoundUpper Boundkt7.992*1.3572.0004.79311.191tˤ30.460*1.3572.00027.26233.659tk−7.992*1.3572.000−11.191−4.793tˤ22.469*1.3572.00019.27025.668tˤk−30.460*1.3572.000−33.659−27.262t−22.469*1.3572.000−25.668−19.270 Further, a statistically insignificant two-way interaction was observed between age and gender (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ F\left(\mathrm{3,264}\right)=1.908, $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ P=0.129, {\eta }^{2}=0.021 $$\end{document}) and target consonant and age (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ F\left(\mathrm{6,264}\right)=1.419, $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ p=0.207,{\eta }^{2}=0.031 $$\end{document}). However, there was a statistically significant two-way interaction between gender and target consonant (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ F\left(\mathrm{2,264}\right)=30.107, $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ p<0.001,{\eta }^{2}=0.186 $$\end{document}). The three-way interaction among gender, target consonant and age was statistically significant (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {F\left(\mathrm{6,264}\right)=1.296),p=0.259,\eta }^{2}=0.029 $$\end{document}). Tukey’s HSD post-hoc analysis in Table 5 indicates that all the age groups were significantly different in VOT (p < 0.05) except for preadolescent and child groups.Table 5Tukey’s post-hoc analysis for the age of the voiceless stops(I) Age(J) AgeMean Difference (I-J)Std. ErrorSig95% Confidence IntervalLower BoundUpper BoundTukey HSDAdolescentAdult6.675*1.56710.0002.62310.727Child−14.883*1.56710.000−18.935−10.832PreAdol−14.147*1.56710.000−18.199−10.095AdultAdolescent−6.675*1.56710.000−10.727−2.623Child−21.558*1.56710.000−25.61−17.507PreAdol−20.822*1.56710.000−24.874−16.77ChildAdolescent14.883*1.56710.00010.83218.935Adult21.558*1.56710.00017.50725.61PreAdol0.7361.56710.966−3.3164.788PreAdolAdolescent14.147*1.56710.00010.09518.199Adult20.822*1.56710.00016.7724.874Child−0.7361.56710.966−4.7883.316 ### Voiced Stops Results Factorial ANOVA conducted on the voiced stops whose results are presented in Table 6 indicated that the VOT duration of the voiced stops differed significantly based on gender (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ F\left(\mathrm{1,264}\right)=226.884,p<.001,{\eta }^{2}=0.462 $$\end{document}) and age (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ F\left(\mathrm{3,264}\right)=10.743,p<.001,{\eta }^{2}=0.109 $$\end{document}).Table 6Factorial Anova siting differences in VOT duration based on gender and ageSourceType III Sum of SquaresdfMean SquareFSigPartial Eta SquaredCorrected Model98,667.700^a^234289.90013.064.000.532Intercept618,744.2811618,744.2811884.197.000.877TargetC3927.42621963.7135.980.003.043Gender74,505.562174,505.562226.884.000.462Age10,583.97533527.99210.743.000.109TargetC*Gender2916.12021458.0604.440.013.033TargetC*Age351.885658.647.179.983.004Gender*Age3044.67731014.8923.091.028.034TargetC*Gender*Age2923.6086487.2681.484.184.033Error86,693.974264328.386Total806,621.140288Corrected Total185,361.673287a. R Squared = .532 (Adjusted R Squared = .492) In terms of gender, females (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=-62.453\pm 1.511, SD=21.2155) $$\end{document} had longer VOT duration as compared to males (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=-30.276\pm 1.5, SD=17.8767 $$\end{document}), as presented and illustrated in Table 7 and Fig. 5.Table 7Summary statistics of voiced VOT duration based on gender and ageMeanStd. errorStd. deviation95% Confidence intervalLower boundUpper bound*Gender*Female−62.4531.51121.2155−65.429−59.478Male−30.2761.5117.8767−33.249−27.302*Age*Child−51.7252.12227.3459−55.903−47.547Preadolescent−49.6652.13628.0026−53.87−45.46Adolescent−47.9962.13622.9195−52.201−43.791Adult−36.0722.15219.9027−40.309−31.835Fig. 5Scatter plot of VOT duration of voiced stops for females and males As can be observed in the summary statistics results in Table 7, the VOT duration of children was significantly longer (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=-51.725\pm 2.122, $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ SD=27.3429 $$\end{document}) than that of preadolescents (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=-49.665\pm 2.136, $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ SD=28.0026 $$\end{document}), adolescents (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ -47.996\pm 2.136,SD=22.9195 $$\end{document}) and adults (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=-36.072\pm 2.152,SD=19.9027 $$\end{document}). Tukey’s HSD post-hoc analysis in Table 8 showed a significant difference between children and adults, preadolescents, and adolescents. The results indicate that as age increases, the VOT duration of voiced stops decreases (See Fig. 6). Further, there was a statistically significant difference in VOT duration based on target consonant (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ F\left(\mathrm{2,264}\right)=5.980,p=0.003,{\eta }^{2}=0.043 $$\end{document}).Table 8Tukey’s Post-hoc analysis for age(I) Age(J) AgeMean difference (I–J)Std. errorSig95% Confidence intervalLower boundUpper boundAdolescentAdult−11.989*3.0309.001−19.825−4.153Child3.8963.0099.567−3.88611.678PreAdol1.6693.0202.946−6.1399.478AdultAdolescent11.989*3.0309.0014.15319.825Child15.885*3.0205.0008.07523.694PreAdol13.658*3.0309.0005.82221.494ChildAdolescent−3.8963.0099.567−11.6783.886Adult−15.885*3.0205.000−23.694−8.075PreAdol−2.2273.0099.881−10.0085.555PreAdolAdolescent−1.6693.0202.946−9.4786.139Adult−13.658*3.0309.000−21.494−5.822Child2.2273.0099.881−5.55510.008Fig. 6Scatterplot of the VOT duration for all age groups Further, Tukey’s HSD post-hoc analysis (see Table 9) indicated voiced target consonants were significantly different from each other were bilabial /b/ and coronal /d/ and coronal /d/ and velar /ɡ/ such that with coronal /d/ had significantly longer VOT duration (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=-51.569\pm 1.851, $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ SD=20.5371 $$\end{document}) than bilabial /b/ (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=-43.375\pm 1.851, $$\end{document}\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ SD=27.9907 $$\end{document}) and velar /ɡ/ (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ Mean=-44.15\pm 1.851,SD=26.5705 $$\end{document}).Table 9Tukey’s Post-hoc analysis for voiced target consonants(I) TargetC(J) TargetCMean difference (I–J)Std. errorSig95% Confidence intervalLower boundUpper boundgb−.5332.6156.977−6.6985.632d7.419*2.6156.0141.25413.584bg.5332.6156.977−5.6326.698d7.952*2.6156.0071.78714.117dg−7.419*2.6156.014−13.584−1.254b−7.952*2.6156.007−14.117−1.787 For the voiced stops, there was a statistically significant interaction between age and gender (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ F\left(\mathrm{3,264}\right)=3.091,p<0.028,{\eta }^{2}=0.034 $$\end{document}) and target consonant and age (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ F\left(\mathrm{2,264}\right)=4.440,p=0.013, {\eta }^{2}=0.033 $$\end{document}) Further, a statistically insignificant two-way interaction was observed between age and target consonant (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ F\left(\mathrm{6,264}\right)=0.179, p=0.983, {\eta }^{2}=0.004 $$\end{document}).The three-way interaction among gender, target consonant, and age was statistically insignificant (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {F\left(\mathrm{6,264}\right)=1.484),p=0.184,\eta }^{2}=0.033 $$\end{document}). ## Discussion and Conclusions Having presented the results, we now turn to answering the research questions posited at the outset of this paper and integrating our results within the wider existing literature. The main objective of the study was to investigate the effects of age and gender on the production of VOT. The results of the study supplement the literature describing different patterns of VOT productions. In this paper, factorial analysis was conducted to examine differences in VOT based on gender and age. The data analysis was divided based on the voiced and voiceless stops. In summary, the data analysis indicated that females have significantly longer VOT duration than males for both the voiced and voiceless stops. In addition, children had significantly longer VOT duration than preadolescents, adolescents, and adults across the board. Notably, the VOT duration of both the voiced and voiceless stops lengthened with decreasing age; the younger the age, the longer the VOT. Also, the results show that voiceless stops were characterized with a long voicing lag (positive), and voiced stops were characterized with a long voicing lead (negative) across the board. This shows that RJA, similarly to what has been reported about Swedish (Beckman et al., 2011), exhibits a different pattern of classifying the laryngeal contrast of stops; that is prevoiced stops with aspirated stops. This result, as reported in the results of this study, has been evident in the speech production of all age groups and genders. Also, this conclusion shows that the laryngeal contrast of RJA is different from what has been reported in the literature where the contrast occurs between lead voicing and short lag in languages such as Lebanese Arabic and Jordanian Arabic (Yeni-Komshian et al. et al., 1977; Mitleb, 2001; French (Caramazza & Yeni-Komshian, 1974)) or those studies which reported a contrast between lead voicing and long lag in languages such as Saudi Arabic ((Ghamdi dialect) Al-Ani, 1970 (Iraqi Arabic); Alghamdi, 1990; Al-Gamdi et al., 2019; and (English) Lisker & Abramson, 1964)). Considering the effects of age and gender, our results show a significant correlation between age and VOT, where female speakers exhibited a longer VOT duration than males. Similarly, the results further reveal a significant correlation between age and VOT, where children had the longest VOTs and adults the shortest. In other words, the younger the participants are, the longer the VOT duration is, and the converse is true. This may lead to the assumption that VOT production in RJA is in incremental development. Although the results of the present study corroborate other studies that linked females’ longer VOTs to anatomical or biological differences, all female speakers in this study had longer VOTs for both voiced and voiceless stops. The results also support other studies that link gender differences to sociophonetic factors such as speaking rate, where women tend to produce more careful speech than men, which results in women having longer VOTs than men (e.g., Kessinger & Blumstein, 1988; Port & Rotunno, 1979; Ryalls et al., 2004; Volaitis & Miller, 1992). However, the question remains why, in some languages, there are no significant differences in VOTs based on gender, as reported by Yu et al. (2015) and Morris et al. (2008). Also, it would be hard to explain why females in other languages produced shorter VOTs than males, as in Ryalls et al. (2004), Li (2013), and Oh (2019). Accordingly, given the results of the study under investigation and the above-mentioned studies, our argument is that the variation of the production of the VOT should be considered language-specific as it varies from one language to another. This, in part, is due to the conflicting results in different languages (whether genetically related or not). These findings suggest that dialectal and language-specific differences affect the VOT variation. This conclusion will rule out the other justifications for the sporadic, conflicting results of other studies, i.e., the sociophonetic and the anatomical. With regards to age differences, our study shows that age plays an integral role in the production of VOT, which is in line with what other studies have reported (e.g., Bóna, 2014 for Hungarian; MacLeod, 2016 for French). Our results demonstrated that VOT duration of both the voiced and voiceless stops was negatively correlated with age. All VOT durations lengthened with decreasing age; the younger the age, the longer the VOT. Since children, both males and females had longer VOTs than preadolescents, preadolescents had longer VOTs than adolescents, and adolescents had longer VOTs than adults, we can conclude that VOT development is incrementally gradual. So, children’s production of VOT gradually becomes close to that of adults as they grow until they reach adult-like proficiency. This gradual development of adult-like proficiency might be ascribed to the anatomical differences between children and adults’ articulators, as suggested by Kewley-Port and Preston (1974). A similar finding was reported by Bóna (2014) for the velar /k/ in Hungarian, where adults produced significantly shorter duration than younger participants. Bóna (2014) ascribed those differences to the slow rate of articulation in children compared to adults, as well as a decrease in muscle tone and the thickness of the vocal folds given the young age of children, as also suggested by Benjamin (1982). Our findings are in agreement with Barton and Macken (1980), Gilbert (1977), Menyuk et al. (1975), Smith (1978), and Ma et al., 2018 in that children produce longer VOTs in voiceless stops. On the other hand, our results differ from Kewley-Port and Preston (1974), Macken and Barton (1980), and Zlatin and Koenigsknecht (1976), who reported that children aged between 4 and 6 years old produced shorter VOT durations than adults.1