Authors: Ziqi Jian, Jingshi Huang, Feng Shi, Yoshihiro Shimomura
Categories: Article, Psychological stress, Heart rate variability, Mental arithmetic task, Cognitive load, Saturation effect, Psychological detachment, Engineering, Mathematics and computing, Neuroscience, Psychology
Source: Scientific Reports
Authors: Ziqi Jian, Jingshi Huang, Feng Shi, Yoshihiro Shimomura
Mental arithmetic task is a classic paradigm for inducing psychological stress and is widely used in heart rate variability research. However, findings across task difficulty are inconsistent, partly due to a lack of standardized difficulty gradation, uncontrolled task order effects, and unconsidered response delays in heart rate variability. We developed a multi-level arithmetic system with low, medium, and high level, combining a serial subtraction task and a Unity-based programmed task. Fifteen healthy graduate students completed the experiment. Electrocardiogram was recorded before and during tasks, and heart rate variability frequency-domain and nonlinear metrics were analyzed as baseline-relative changes. Subjective workload was assessed with NASA-TLX. NASA-TLX ratings and error rates indicated that the difficulty manipulation was effective. Frequency-domain HRV metrics showed higher values under medium- and high-difficulty conditions compared with the low-difficulty condition, while exhibiting no further proportional increases between the medium- and high-difficulty levels, suggesting a possible saturation pattern. In contrast, nonlinear HRV metrics differentiated task difficulty levels more consistently and exhibited response patterns that were more closely aligned with subjective workload ratings. Within the current experimental context, frequency-domain HRV metrics appeared to show limited sensitivity under low and high workload conditions, potentially due to disengagement and saturation effects. By comparison, nonlinear HRV metrics demonstrated greater sensitivity to graded task difficulty and may provide complementary information for characterizing psychological stress.
Psychological stress refers to a series of physiological, psychological, and behavioral responses that occur when individuals face threatening or challenging stimuli. The biological concept of stress was first introduced by Hans Selye, who defined it as a nonspecific response of the body to any demand^1^. During acute psychological stress, the hypothalamic–pituitary–adrenal (HPA) axis and the autonomic nervous system (ANS) are typically activated, leading to sympathetic nerves system (SNS) excitation, parasympathetic nerves system (PNS) inhibition, and a cascade of physiological changes. Heart rate variability (HRV), reflecting fluctuations in R–R intervals driven by the interplay of SNS and PNS inputs, has become a widely used noninvasive marker of autonomic regulation in stress research.
With the growing interest in psychological stress, both frequency-domain and nonlinear analyses of HRV have gained attention. Frequency-domain measures decompose HRV into distinct spectral components, offering physiologically interpretable metric with strong computational stability. Specifically, low-frequency (LF, 0.04–0.15 Hz) power reflects both SNS and PNS contributions but is generally considered to index SNS dominance during stress; high-frequency (HF, 0.15–0.40 Hz) power, closely tied to respiration, is accepted as a robust marker of vagal activity; and the LF/HF ratio is often used to approximate sympathovagal balance^2^. Nonlinear measures, in contrast, capture the complexity and dynamic structure of HRV, offering greater sensitivity to fluctuations always non-stationary in ANS regulation. Approximate entropy (ApEn) and sample entropy (SampEn) quantify the irregularity and unpredictability of RR interval time series, with lower values indicating reduced autonomic flexibility and heightened psychological stress^3,4^. Detrended fluctuation analysis (DFA), particularly the short-term scaling exponent DFA α1, reflects fractal correlation properties of HRV and has been associated with changes in autonomic balance and sympathetic activation under stress conditions. Together, frequency-domain and nonlinear measures provide complementary perspectives and have become essential tools in task-induced stress studies. Yet findings remain while several studies report increased LF and decreased HF or entropy during stress^5,6^, others have found opposite patterns^7,8^. Such discrepancies are likely due less to physiological mechanisms than to task design factors, including difficulty manipulation, task order, and insufficient consideration of HRV response delays.
Cognitive load is regarded as a key psychological mechanism driving stress responses in laboratory paradigms. It refers to the cognitive resources consumed during task execution and depends on the interplay between task complexity and individual capacity. Cognitive load theory posits that when task demands exceed working memory capacity, cognitive overload occurs, triggering SNS activation and stress responses^9–12^. Thus, manipulating task difficulty provides a practical means to regulate cognitive load, allowing researchers to experimentally grade stress intensity and examine HRV responses with greater precision.
Laboratory stress research commonly employs experimental paradigms such as the Trier Social Stress Test (TSST), Stroop task, or emotional picture paradigm. By manipulating task difficulty, time pressure, and feedback, researchers can control cognitive load and thereby induce psychological stress at varying intensities. This not only strengthens the methodological validity of stress paradigms but also provides a framework for examining the psychophysiological coupling between cognitive load and autonomic regulation.
Among experimental paradigms, mental arithmetic tasks are widely adopted due to their simplicity, flexibility, and strong engagement of cognitive resources. These tasks require participants to perform mathematical operations without external aids, engaging core cognitive processes such as working memory, attention, problem solving, and numerical processing^13–15^. These offer high controllability and repeatability, increasing computational complexity reliably elevates cognitive load and elicits stress responses. Early designs, such as the TSST arithmetic component, often used continuous subtractions (e.g., subtracting 13 from 1022) combined with negative feedback to amplify stress^16^, but the fixed difficulty limited adaptability to individual differences^17^. The Montreal Imaging Stress Task (MIST) introduced multi-step operations and time pressure^18^, but its combined cognitive and social stressors, along with a lack of standardized difficulty grading, restricted interpretability. Consequently, developing multi-level mental arithmetic tasks with clear structure, adjustable difficulty, and controlled pacing has become a critical challenge in experimental stress research.
In addition to task difficulty, other design variables can confound HRV responses under psychological stress. For instance, sequential presentation of tasks with different difficulty levels may induce carry-over effects that obscure dynamic HRV changes^19^. Moreover, due to delayed autonomic responses, rapid alternation between task difficulty levels may fail to capture the evolving regulation of the ANS. Li et al. demonstrated that HRV responses to psychological stress typically emerge after ~ 3 min and show linear changes within the subsequent 3–5 min, underscoring the importance of temporal considerations^20^. Similarly, low-difficulty tasks may evoke heightened vigilance at onset due to novelty, but this response often diminishes with adaptation, reducing sustained HRV reactivity^21^. Such findings emphasize that inadequate control of task difficulty, order, and timing can compromise the interpretability and reliability of HRV as a stress marker.
Against this backdrop, the present pilot study systematically investigates the effects of graded mental arithmetic tasks on HRV frequency-domain and nonlinear measures. Given the limited sample size and exploratory scope, this study was designed primarily to evaluate the feasibility of the proposed paradigm and to preliminarily examine the sensitivity of different HRV metrics to graded task difficulty. Specifically, we focus on how difficulty modulation influences HRV responses and elucidates the relationship between cognitive load and autonomic activity. The study addressed three
To validate the effectiveness of the developed arithmetic task system in eliciting psychological stress, thereby assessing its feasibility as an experimental stress paradigm.To examine HRV responses under high cognitive load, with emphasis on the relative sensitivity and dynamic patterns of frequency-domain versus nonlinear measures.To explore temporal changes in stress responses across difficulty levels, evaluating the relationship between task duration and stress intensity, and comparing the time-course trajectories across difficulty levels.
The participants in this study were 15 current graduate students, aged between 24 and 34 years, all with no history of cardiovascular disease. Prior to inclusion in the experiment, each participant underwent a math ability screening test to ensure that individuals with excessively high or low calculation skills were excluded. To ensure the accuracy of the experimental results, all participants were required to provide adequate sleep the night before the experiment and to refrain from smoking or consuming any stimulants, including coffee and tea, for at least 12 h. To further ensure data reliability, all experiments were conducted between 00 p.m. and 00 p.m. Additionally, none of the participants reported any history of heart disease or other chronic conditions. Due to excessive physiological signal artifacts primarily caused by electromagnetic interference and motion-related ECG noise, data from three participants were excluded after preprocessing and visual inspection. These signals exhibited irregular RR intervals and waveform distortions that could not be sufficiently corrected through filtering or artifact removal, resulting in a final sample size of 12 participants (4 females and 8 males; average age = 29 years).
Before the experiment began, the participants were provided with detailed written informed consent forms and were ensured a full understanding of the experimental process, objectives, and potential risks. Each participant received a compensation of 4,000 yen for their involvement. This study was reviewed and approved by the Ethics Committee of Chiba University (Ethics Review No. R4-22), ensuring compliance with ethical standards and regulations.
Given the exploratory nature of the present study and the lack of prior effect size estimates for nonlinear HRV metrics under graded mental arithmetic stress, no a priori power analysis was conducted. The study was therefore designed as a pilot investigation aimed at examining feasibility, sensitivity, and response patterns of multiple HRV indices to task difficulty.
To achieve precise control over task difficulty, we developed a standardized, computerized mental arithmetic task program based on the Unity platform. As shown in Fig. 1, the experimental interface consists of multiple functional the central panel presents a continuous stream of arithmetic problems, each problem composed of multiple calculation units, with each unit containing two single-digit multiplication operations connected by addition. Participants input their answers using the numeric keypad on the right side of the interface and submit responses by clicking the “Confirm” button. The entered answer is displayed in real time in the lower panel, and immediate feedback is provided after submission. A countdown bar at the top indicates the remaining time for the current problem, while the upper-right corner displays the “remaining attempts for answering questions” using red squares to visually represent the number of problems still available, thereby enhancing the sense of psychological stress.
Fig. 1Program interface for mental arithmetic task.
To balance task challenge with participant experience, we implemented a “uniform feedback mechanism,” whereby all submitted answers were displayed as “correct,” regardless of their actual accuracy. This mechanism minimizes negative psychological effects associated with high error rates. To prevent participants from detecting the non-veridical nature of the feedback, a penalty system was if a response was not submitted within the allotted time, the trial was marked as a timeout error, and one “remaining attempt” was deducted. Exceeding the upper limit automatically terminated the task and marked the dataset as invalid, ensuring both data quality and controlled workload.
A researcher interface was also developed to flexibly configure experimental parameters, including task difficulty (ranging from Level 1 to Level 6), baseline and rest durations, maximum number of attempts, auditory prompts, and whether the uniform feedback mechanism is enabled. The system further allows fine-tuned control of time limits and problem counts at each difficulty level, thereby enabling precise adjustment of task pacing and difficulty progression.
Three task conditions were included. The low-level condition employed the classic continuous serial subtraction task (subtracting 7 from 10,000), which has been widely used as a traditional control paradigm in psychological stress research^22^. Participants were instructed to perform the subtraction as quickly and continuously as possible for a fixed duration of 15 min, without receiving any performance-related feedback. The medium- and high-level conditions were delivered through the computerized program, ensuring higher controllability and precision.
The task difficulty was manipulated along two primary dimensions. First, the overall computational load was adjusted by varying the number of calculation units within each problem (i.e., “task complexity”). Second, time pressure was imposed by assigning a fixed response deadline to each problem, thereby inducing stress associated with temporal urgency. Notably, the time limit was determined based on the overall mean response time obtained from all participants during the preliminary mathematical ability screening test. This screening procedure was conducted solely for task calibration rather than participant selection. During the screening, the same participants who later took part in the formal experiment (N = 12) were asked to solve a set of arithmetic problems comparable in structure to those used in the experimental tasks (e.g., multi-digit subtraction and mental calculation problems). The overall mean response time across participants was then used to determine the response deadlines for the medium- and high-difficulty tasks, ensuring that the imposed time pressure was challenging yet achievable. The three task conditions were defined as
Low-level (−7 task): Serial subtraction starting from 10,000, with responses typed in Word. This condition imposed minimal cognitive load and no time pressure.Medium-level (Level 3): Program-based tasks with three multiplication units per problem. Each problem had a 30 s time limit, 85 problems total, and a maximum error allowance of 8.High-level (Level 6): Program-based tasks with six multiplication units per problem. Each problem had a 45 s time limit, 35 problems total, and the same error threshold of 8.
In the present study, only Level 3 and Level 6 were used for analysis, corresponding to the Medium- and High-difficulty conditions, respectively. The Low-difficulty condition employed a traditional continuous subtraction task and was therefore not assigned a software-defined level. For the low-level task, participants were explicitly informed prior to task onset that calculation accuracy was not emphasized, in order to minimize unnecessary performance-related pressure during this control condition.
Because high-level problems required longer solution times, the total number of problems differed between medium and high conditions to ensure a minimum of 15 min task exposure. To avoid fatigue or order effects, each participant completed only one task per day, with at least 24 h between sessions. All participants completed all three conditions. The order of low- and high-level tasks was counterbalanced across participants, while the medium-level task was always administered in the middle to maintain design symmetry.
The experiment was conducted in a soundproof and electromagnetically shielded room maintained at approximately 25 °C. Upon arrival, participants were first familiarized with the laboratory environment, equipment, tasks, potential risks and benefits of the study, and the operation of the mental arithmetic task software. They were then seated comfortably in front of a computer screen and instructed to remain seated and motionless throughout the experiment. As illustrated in Fig. 2, the protocol began with a 15-min seated rest period prior to task onset. This rest period was designed to minimize potential carryover effects from software exercise and to allow participants to reach a relaxed and physiologically stable state. Because the initial minutes of rest may still reflect residual arousal induced by software exercise, the first 5 min were excluded. A stable 5-minute segment was then extracted from the remaining 5–15-minute interval and used as the baseline. This baseline period was followed by at least 15 min of the mental arithmetic task, and then another 15-min rest period. To minimize confounding effects from speech, participants interacted with the system using a mouse and keyboard. This design avoided the potential influence of verbal responses on breathing patterns, which could otherwise obscure vagal withdrawal and alter both the variability and spectral composition of the RR interval series. All task instructions were presented on the computer screen^23,24^.
Fig. 2Experimental procedure.
After completing the experimental procedure, participants were asked to complete the NASA Task Load Index (NASA-TLX). The NASA-TLX is a subjective workload assessment tool developed by the National Aeronautics and Space Administration that evaluates perceived workload across six (1) Mental Demand, reflecting the degree of cognitive activity required (e.g., thinking, calculating, remembering); (2) Physical Demand, reflecting the amount of physical effort required; (3) Temporal Demand, reflecting time pressure and pace of task execution; (4) Performance, reflecting perceived success in accomplishing the task; (5) Effort, reflecting the amount of mental and physical work required to achieve performance; and (6) Frustration Level, reflecting feelings of stress, annoyance, or discouragement experienced during task performance.
Each dimension was rated on a 0–21 scale, with 0 indicating no workload and 21 representing maximum workload. A standard weighting procedure was applied to reflect the relative contribution of each dimension to the overall workload. Participants performed pairwise comparisons among the six dimensions to determine their relative importance. Weighted scores were calculated by multiplying each dimension’s rating by its assigned weight, and the overall workload index was obtained by summing the weighted scores and dividing by the total weight^25^.
Electrocardiogram (ECG) signals were recorded using a Biopac MP160 system with a wireless BioNomadix ECG module. Standard three-lead electrode placement was applied, with electrodes positioned on the right clavicle, left clavicle, and lower left rib cage. Data were sampled at 1000 Hz and acquired using AcqKnowledge 4.2 software. Signals were preprocessed with a five-point smoothing filter, and RR intervals shorter than 0.5 s or longer than 1.5 s were excluded to reduce artifacts.
HRV metric were calculated using Kubios HRV Premium software. Frequency-domain measures included LF, HF, LF normalized units (LF (n.u.)), HF normalized units (HF (n.u.)), and LF/HF. As LF (n.u.) and HF (n.u.) are complementary, only LF (n.u.) was retained. Nonlinear measures included approximate entropy (ApEn), sample entropy (SampEn), and detrended fluctuation analysis α1 (DFA α1)^26^.
To capture temporal dynamics, the 15-minute experimental stage was divided into three consecutive 5-minute segments (EXP5, EXP10, and EXP15) following task onset. The EXP5 segment intentionally included the initial task onset period, during which transient autonomic adjustments are known to occur within the first few minutes of stress exposure, and was therefore conceptualized as reflecting early-stage or initiation-related stress responses rather than a steady-state response. If a task session exceeded 15 min due to individual differences in response speed, only the first 15 min of task-related ECG data following task onset were included in the analysis, and any data beyond this duration were excluded. HRV responses were analyzed as differences relative to baseline to control for inter-individual variability in resting autonomic nervous system activity.
Statistical analyses were conducted using IBM SPSS Statistics 25. Prior to comparing HRV differences across task difficulty levels, the Shapiro–Wilk test was performed to examine the normality of the data. The results indicated that the data did not follow a normal distribution; therefore, subsequent comparisons were performed using a generalized estimating equation (GEE) model to examine the main effects and interaction effects of task difficulty and experimental stage. Because each participant completed tasks of different difficulty levels on separate days, repeated measurements were treated as within-subject observations. Accordingly, participant ID was specified as the clustering variable in the GEE models to account for within-subject correlations across experimental stages and sessions. An exchangeable working correlation structure was adopted. Gender was not included as a covariate in the GEE analyses due to the small and unbalanced sample size, and because the present study was not designed to examine sex-related differences.
In addition, NASA-TLX scores were analyzed using a one-way analysis of variance (ANOVA). To further explore the relationship between subjective workload and physiological responses, Spearman’s rank correlation analysis was conducted to assess the associations between NASA-TLX workload ratings and HRV metrics in both the frequency and nonlinear domains.
As shown in Fig. 3A, descriptive statistics indicated that NASA-TLX scores increased with task difficulty level (High: M = 13.95, SD = 3.41; Medium: M = 10.42, SD = 3.36; Low: M = 7.05, SD = 1.61), consistent with the intended difficulty manipulation. The NASA-TLX workload metric for all three task conditions met the assumption of normality as confirmed by the Shapiro–Wilk test. A one-way ANOVA revealed a significant effect of task difficulty on NASA-TLX workload ratings (F(2, 22) = 16.783, p <.05). Post hoc comparisons showed significant differences between the high- and medium-level tasks (p <.05), high- and low-level tasks (p <.05), and medium- and low-level tasks (p <.05). These results confirm that perceived workload increased with task difficulty.
We compared error rates between the high- and medium-level mental arithmetic tasks. Descriptive statistics showed that participants in the high-level condition had a mean error rate of 33.06% (SD = 15.59), whereas the medium-level condition had a mean error rate of 9.88% (SD = 8.42). As error rates did not meet the assumption of normality, a non-parametric Wilcoxon rank-sum test was applied. As illustrated in Fig. 3B, error rates were significantly higher in the high-level condition than in the medium-level condition (p <.05). For the low-level task, participants were informed prior to the experiment that task accuracy was not emphasized. Therefore, error rates were not included in the analysis for the low-level condition.
For the low-level task, which consisted of continuous subtraction by seven in a Word document, participants could often detect their own mistakes during calculation, potentially leading to unnecessary stress. To avoid additional pressure, participants were explicitly informed prior to the experiment that accuracy was not a concern. Accordingly, error rates were not analyzed for the low-level condition.
Fig. 3Subjective and behavioral measures under mental arithmetic tasks of varying levels of difficulty. (a): NASA-TLX workload. (b): Error rate of mental arithmetic tasks. *p <.05 (Bonferroni corrected).
Table 1Relative baseline differences (M ± SE) of HRV frequency-domain metrics under different levels of mental arithmetic tasks and results of GEE analysis.VariableGroupEXP5 (M ± SE)EXP10 (M ± SE)EXP15 (M ± SE)Group (χ², p)Time (χ², p)Interaction (χ², p)High357.53 ± 465.62427.80 ± 404.34480.83 ± 426.23 LF Medium249.82 ± 609.22304.39 ± 439.33350.24 ± 440.2820.9, 0.001.58, 0.463.99, 0.407Low−286.44 ± 789.18−396.11 ± 637.90−314.15 ± 655.52High0.54 ± 1.750.97 ± 1.321.21 ± 0.96 LF/HF Medium0.60 ± 1.730.95 ± 2.191.23 ± 1.926.37, 0.043.88, 0.144.80, 0.31Low−0.71 ± 2.29−0.61 ± 2.60−0.51 ± 2.60High7.81 ± 19.3813.70 ± 13.0415.35 ± 11.00 LF (n.u.) Medium8.62 ± 16.739.12 ± 16.5412.20 ± 17.617.0, 0.03*3.91, 0.142.69, 0.61Low−0.41 ± 19.35−0.42 ± 22.121.27 ± 20.30High67.85 ± 222.8941.43 ± 251.235.91 ± 215.00 HF Medium64.26 ± 364.3758.07 ± 280.8460.31 ± 275.305.47, 0.0650.80, 0.671.49, 0.83Low−47.51 ± 288.07−66.92 ± 260.77−78.96 ± 227.94
Table 1 presents the relative baseline differences (M ± SE) and GEE results for frequency-domain metrics (LF, LF/HF, LF (n.u.) and, HF). Significant main effects of task difficulty were observed for LF (χ²(2) = 20.9, p <.05), LF/HF (χ²(2) = 6.37, p <.05), and LF (n.u.) (χ²(2) = 7.0, p =.03), whereas HF did not reach significance (χ²(2) = 5.47, p >.05). Neither experimental stage effects nor task difficulty × experimental stage was significant, suggesting that frequency-domain metrics did not exhibit clear temporal dynamics during the 15-min experimental stage.
Figure 4 illustrates the main effects of task difficulty on frequency-domain HRV metrics, with values averaged across experimental stages. For LF, Bonferroni-adjusted pairwise comparisons based on the GEE model revealed that the high-difficulty condition showed significantly higher LF values than the low-difficulty condition (χ² = 19.27, p <.05), and the medium-difficulty condition also exhibited significantly higher LF values than the low-difficulty condition (χ² = 11.76, p <.05). No significant difference was observed between the high- and medium-difficulty conditions (p >.05).
For LF/HF, the high-difficulty condition demonstrated significantly higher values than the low-difficulty condition (χ² = 6.4, p <.05). The comparison between the medium- and low-difficulty conditions approached but did not reach statistical significance (p <.1), while no significant difference was observed between the high- and medium-difficulty conditions (p >.05).
For LF (n.u.), pairwise comparisons indicated that the high-difficulty condition showed significantly higher values than the low-difficulty condition (χ² = 6.66, p <.05), whereas the medium-difficulty condition did not significantly differ from either the high- or low-difficulty conditions (p >.05).
In contrast, HF did not show a significant main effect of task difficulty (χ² = 5.47, p >.05), and none of the pairwise comparisons reached statistical significance (p >.05).
Fig. 4Main effects of task difficulty on frequency-domain HRV metrics. Boxplots depict values of LF, HF, LF/HF, and LF (n.u.) averaged across experimental stages for the low-, medium-, and high-level task difficulty conditions. Statistical significance reflects pairwise comparisons between task difficulty levels based on the generalized estimating equation model. *p <.05, +p <.10 (Bonferroni corrected).
Taken together, the frequency-domain metrics LF, LF/HF, and LF (n.u.) demonstrated robust sensitivity in distinguishing between high- and low-level mental arithmetic tasks, consistently detecting significant differences between these two conditions. However, their discriminative capacity between high- and medium-level tasks was limited, as no significant differences were observed. In contrast, HF exhibited weak sensitivity to task difficulty, with smaller fluctuations overall and no ability to reliably differentiate across difficulty levels.
Fig. 5Scatterplots of correlations between NASA-TLX scores and HRV frequency-domain indices. The red solid line indicates significant correlation, while the gray dashed line indicates non-significant correlation.
As shown in Fig. 5, Spearman’s correlation analysis revealed a significant positive association between NASA-TLX workload scores and LF (ρ = 0.427, p <.05), indicating that higher subjective workload was associated with stronger low-frequency power. In contrast, no significant correlations were observed between NASA-TLX scores and LF/HF, LF in normalized units (LF n.u.), or HF.
Table 2Relative baseline differences (M ± SE) of HRV nonlinear metrics under different levels of mental arithmetic tasks and results of GEE analysis.VariableGroupEXP5 (M ± SE)EXP10 (M ± SE)EXP15 (M ± SE)Group (χ², p)Time (χ², p)Interaction (χ², p)High−0.02 ± 0.10−0.06 ± 0.11−0.07 ± 0.15 ApEn Medium0.03 ± 0.060.05 ± 0.090.02 ± 0.0615.91, 0.000.47, 0.7893.02, 0.56Low0.08 ± 0.070.10 ± 0.070.10 ± 0.08High−0.16 ± 0.28−0.23 ± 0.26−0.28 ± 0.24 SampEn Medium0.00 ± 0.110.06 ± 0.140.01 ± 0.1534.48, 0.003.35, 0.198.54, 0.07Low0.16 ± 0.270.33 ± 0.240.27 ± 0.30High0.17 ± 0.280.23 ± 0.240.21 ± 0.28 DFA α1 Medium0.12 ± 0.190.10 ± 0.220.16 ± 0.227.38, 0.03*1.94, 0.384.81, 0.31Low−0.14 ± 0.23−0.18 ± 0.24−0.13 ± 0.18
Table 2 presents the relative baseline differences (M ± SE) and GEE results for nonlinear metrics (ApEn, SampEn, DFA α1). Significant main effects of task difficulty were found for ApEn (χ²(2) = 15.91, p <.05), SampEn (χ²(2) = 34.48, p <.001), and DFA α1 (χ²(2) = 7.38, p =.03). Neither experimental stage nor task difficulty and experimental stage interactions were significant.
Fig. 6Main effects of task difficulty on nonlinear HRV metrics. Boxplots depict values of ApEn, SampEn, and DFA α1 averaged across experimental stages for the low-, medium-, and high-level task difficulty conditions. Statistical significance reflects pairwise comparisons between task difficulty levels based on the generalized estimating equation model. *p <.05 (Bonferroni corrected).
Figure 6 presents the main effects of task difficulty on nonlinear HRV metrics, with values averaged across experimental stages. For ApEn, a significant main effect of task difficulty was observed (χ²= 15.91, p <.05). Bonferroni-adjusted pairwise comparisons indicated that ApEn values were significantly lower in the high-difficulty condition than in the medium-difficulty condition (χ² = 16, p <.05) and the low-difficulty condition (χ² = 21.81, p <.05). In addition, the medium-difficulty condition exhibited significantly lower ApEn values than the low-difficulty condition (χ² = 8.18, p <.05).
For SampEn, the main effect of task difficulty was also significant (χ² = 34.48, p <.05). Pairwise comparisons revealed significant differences among all three difficulty the high-difficulty condition differed from the medium-difficulty condition (χ² = 39.06, p <.05) and the low-difficulty condition (χ² = 19.01, p <.05), and the medium-difficulty condition also differed significantly from the low-difficulty condition (χ² = 10.82, p <.05).
For DFA α1, a significant main effect of task difficulty was identified (χ² = 7.38, p <.05). Bonferroni-adjusted pairwise comparisons showed that the low-difficulty condition yielded significantly lower DFA α1 values than both the medium-difficulty condition (χ² = 11.42, p <.05) and the high-difficulty condition (χ² = 26.42, p <.05). However, no significant difference was observed between the medium- and high-difficulty conditions (p >.05).
Fig. 7Scatterplots of correlations between NASA-TLX scores and HRV nonlinear indices. The red solid line indicates significant correlation, while the red dashed line indicates marginally significant correlation.
As shown in Fig. 7, Spearman’s correlation analysis indicated significant negative correlations between NASA-TLX workload scores and ApEn (ρ = –0.36, p <.05) as well as SampEn (ρ = –0.48, p <.05), suggesting that higher subjective workload was associated with lower entropy values. Moreover, the correlation between NASA-TLX scores and DFA α1 approached significance (ρ = 0.32, p =.05).
The present study investigated HRV responses to graded mental arithmetic tasks, with particular emphasis on examining and contrasting the responsiveness of frequency-domain and nonlinear metrics to different levels of task difficulty. The study should be regarded as an exploratory (pilot) investigation aimed at assessing the sensitivity of different HRV metrics to graded task difficulty within a controlled experimental paradigm. Based on NASA-TLX workload ratings, the difficulty manipulation of the mental arithmetic tasks was successful. Although NASA-TLX is a subjective self-report measure reflecting perceived workload rather than an objective behavioral index, it provides a well-established indicator of participants’ experienced task demands^25^. As shown in Fig. 3A, perceived workload increased systematically with task difficulty, with the high-level task rated highest, the medium-level task intermediate, and the low-level task lowest. In addition, behavioral performance supported the effectiveness of the difficulty manipulation, as accuracy was significantly lower in the high-level task than in the medium-level task (Fig. 3B), indicating that increasing task difficulty imposed greater cognitive demands. Together, these subjective and behavioral findings confirm that the graded mental arithmetic paradigm successfully elicited differentiated levels of workload, thereby providing an appropriate empirical basis for interpreting the subsequent physiological findings within the exploratory scope of the present study.
As shown in Table 1, the main effects of task difficulty on LF and its related metrics (including LF/HF and LF n.u.) were significant, whereas neither the experimental stage effects nor the task difficulty × experimental stage interaction effects reached significance. Notably, correlation analyses revealed that among these frequency-domain metrics, only LF showed a significant positive correlation with NASA-TLX workload ratings, while LF/HF and LF (n.u.) did not (as illustrated in Fig. 5). In other words, although LF and LF-class metrics reflected differences across task difficulty levels at the group level, only LF was tightly aligned with subjective workload. This dissociation may indicate that LF, although not a pure index of sympathetic activity, is more sensitive to sympathetic modulation under mental stress conditions and has been reported to increase with elevated task demands and perceived workload^27^. By contrast, ratio-based indices such as LF/HF and normalized units (e.g., LF n.u.) are mathematically constrained by their dependence on total power and reciprocal relationships between components. As a result, fluctuations in one component may disproportionately influence the ratio, potentially obscuring physiological interpretation^28^.
In previous research, LF and its related metrics have often been regarded as potential markers of SNS activation, with their increases widely under acute psychological stress^27,29^. In the present study, LF-related metrics showed sustained increases relative to baseline under both medium- and high-difficulty conditions. This pattern of change is consistent with the interpretation that higher task difficulty imposes greater cognitive load, which is accompanied by enhanced autonomic activation and elevated psychological stress, as reflected in HRV alterations observed in prior work investigating autonomic responses to task demands^30^. Although some studies have emphasized that LF is not exclusively governed by SNS activity but rather reflects a combined influence of sympathetic and parasympathetic modulation, and may be more susceptible to baroreflex mechanisms under resting conditions^31,32^, the acute stress paradigm employed in the present experiment featured tightly controlled pacing and feedback, with task difficulty level constituting the primary systematic contrast. Under these conditions, the observed baseline-referenced increases in LF and its related metrics can reasonably be attributed to enhanced sympathetic nervous system involvement in response to heightened task demands, consistent with prior research demonstrating that psychological stress and increased cognitive workload are accompanied by shifts in autonomic regulation reflected in HRV measures^33^.
In contrast, under low-difficulty conditions, LF and its related metrics exhibited sustained decreases relative to baseline. This pattern may reflect a state of “psychological detachment”^34^, in which individuals reduce cognitive and emotional engagement during task performance, allowing previously activated autonomic activity to relax and recover, thereby lowering SNS arousal. Previous studies have suggested that when tasks are overly simple and lack sufficient challenge, individuals are more likely to disengage psychologically^35^. In the present study, the low-difficulty “minus-7” arithmetic task imposed minimal cognitive demands on the graduate student sample, likely promoting routinized and automatic processing. Under such conditions, participants may have reduced sustained attentional and emotional involvement during task execution, resulting in decreases in LF and its related metrics relative to baseline that are consistent with a recovery-like autonomic pattern. This interpretation is further supported by the significantly lower NASA-TLX workload ratings observed under the low-difficulty condition, indicating reduced perceived task demands and lending additional support to the notion of psychological detachment.
As shown in Fig. 4, comparisons at the task-difficulty level revealed no significant differences between the medium- and high-difficulty conditions for LF and its related metrics. This finding indicates that when task difficulty increased from a medium to a high level, LF-class metrics did not exhibit further proportional increases, suggesting a non-linear relationship between workload intensity and frequency-domain HRV responses. Such a pattern is consistent with the “saturation effect” proposed by Goldberger et al. (2001), which posits that once autonomic stimulation reaches a certain threshold, HRV responses plateau and their sensitivity to further increases in stimulation diminishes^36,37^.
The mechanisms underlying the saturation effect may be twofold. On the one hand, it may stem from the dynamic balance of the ANS, under high-load conditions, SNS and PNS branches counteract each other through negative feedback to maintain cardiovascular homeostasis, thereby constraining further HRV fluctuations and preventing overreactions^38^. On the other hand, it may relate to receptor saturation, whereby neurotransmitter concentrations (e.g., acetylcholine, ACh) reach a threshold level beyond which further increases fail to elicit new HRV changes^37^. Scheff et al.’s computational modeling study also indicated that prolonged high-level SNS activation can weaken signal transmission via receptor saturation, leading to reductions in frequency-domain metrics such as LF^36,37^. Within this framework, the lack of further increases in LF and its related metrics when task difficulty increased from a medium to a high level can be interpreted as reflecting an early approach to physiological upper limits under medium load. In other words, frequency-domain HRV metrics may already operate near their effective response ceiling at moderate workload levels, thereby limiting their sensitivity to additional increases in task difficulty.
In conjunction with the NASA-TLX workload ratings, subjective scores were able to significantly distinguish between medium- and high-difficulty tasks, indicating that participants perceived meaningful differences in workload intensity across these conditions. By contrast, although LF-related metrics showed elevated values at the group level with increasing task difficulty, correlation analyses did not reveal significant associations between these indices and NASA-TLX ratings. This dissociation may reflect inherent limitations of ratio- and normalized frequency-domain metrics under high-load conditions. In particular, once sympathetic activation approaches elevated levels, LF/HF and LF (n.u.) may enter a response plateau or become clearly susceptible to denominator fluctuations, thereby reducing their capacity to linearly track subtle variations in subjective workload^39^. Such characteristics are consistent with the saturation framework and further suggest that normalized frequency-domain indices may be less sensitive to fine-grained differences in perceived workload at higher levels of cognitive demand.
As shown in Table 1, none of the LF-related metrics exhibited significant main effects of experimental stage. This absence of a clear time effect may be attributable to two complementary factors. First, the early phases of task performance are often accompanied by transient, non–task-specific autonomic responses associated with task onset, such as novelty-related arousal or brief overactivation. These short-lived responses may temporarily obscure workload-related HRV modulation^40^. As task engagement is sustained and autonomic responses gradually stabilize, workload-related modulation of HRV may become more clearly expressed. Consistent with previous findings, HRV responses to acute psychological stress typically require a period of sustained engagement before more interpretable regulatory patterns emerge^20^. Second, under high cognitive load, sustained sympathetic activation often develops only after prolonged task exposure, commonly exceeding 20–30 min, before producing reliable and interpretable changes in HRV indices^41^. Accordingly, the relatively short task duration in the present study (15 min) may have limited the emergence of pronounced time-dependent effects.
As shown in Fig. 4, the HF indicator exhibited numerical variability across the three task conditions, however, statistical analyses did not reveal significant main effects of task difficulty or experimental stage, nor a significant task difficulty × stage interaction. Correlation analyses further indicated that HF was not significantly associated with NASA-TLX workload scores, suggesting that HF showed limited sensitivity to subjective workload in the present mental arithmetic paradigm. This pattern may be attributable to two factors. First, mental arithmetic tasks predominantly engage sympathetic activation, whereas parasympathetic responses are relatively less pronounced, resulting in modest modulation of HF^33,42^. Second, HF is strongly influenced by respiratory patterns, and the absence of explicit respiratory control in the present study may have introduced additional variability, thereby reducing its ability to reliably reflect subjective workload^43^.
This study employed three nonlinear HRV metrics—ApEn, SampEn, and DFA α1—to evaluate the effects of mental arithmetic tasks of varying difficulty on psychological stress. Among them, ApEn and SampEn are classic measures of the short-term complexity and irregularity of heart rate time series. Higher values indicate greater relaxation and more flexible, diverse autonomic modulation of heart rate, whereas lower values suggest elevated psychological stress, reduced flexibility, and more rigid or monotonous heart rate dynamics^3,4,44^.
As shown in Table 2; Fig. 6, ApEn and SampEn exhibited a clear ordering associated with task difficulty levels at the descriptive level. Specifically, ApEn values were lowest under the high-difficulty condition, intermediate under the medium-difficulty condition, and highest under the low-difficulty condition, forming a stable High < Medium < Low pattern. SampEn demonstrated a similar but more pronounced graded ordering across difficulty levels. Together, these findings indicate that entropy-based nonlinear metrics are sensitive to graded changes in task difficulty, with increasing cognitive load associated with progressively reduced cardiac signal complexity.
This interpretation is further supported by correlation analyses, which revealed that both ApEn and SampEn were significantly negatively associated with NASA-TLX workload ratings. These findings reinforce the notion that higher subjective workload is accompanied by lower HRV complexity, reflecting reduced autonomic flexibility under psychological stress. At the aggregate level, entropy-based HRV metrics exhibited an overall correspondence with subjective workload conditions associated with higher perceived workload tended to show lower complexity, whereas conditions with lower perceived workload were characterized by higher HRV complexity (Fig. 6B). This concordance suggests that task difficulty–related increases in subjective workload are consistently accompanied by reductions in autonomic complexity, while lower task demands are associated with preserved or enhanced autonomic adaptability.
Such consistency between psychological and physiological measures underscores the validity of entropy-based HRV nonlinear metrics in capturing task difficulty and psychological stress intensity. In the present study, ApEn and SampEn demonstrated clear task difficulty–related main effects and showed robust associations with subjective workload, indicating a high sensitivity to variations in cognitive load. Compared with frequency-domain metrics, in the present dataset, entropy-based measures appeared less susceptible to early-stage insensitivity and saturation effects observed under medium-to-high workload conditions. Rather than relying on absolute or ratio-based power changes, ApEn and SampEn reflect the structural complexity of heart rate dynamics, allowing them to capture stress-related alterations in autonomic regulation across a broad range of task demands^45^. Taken together, these findings suggest that entropy-based nonlinear HRV metrics may offer a more sensitive characterization of psychological stress than traditional frequency-domain measures when the primary aim is to differentiate levels of cognitive workload.
In addition to entropy-based metrics, the present study also employed DFA to further examine the impact of task difficulty on heart rate dynamics. Originally introduced by Peng et al. (1995), the short-term scaling exponent DFA α1 primarily reflects correlation properties of HRV at short time scales (on the order of minutes), with higher values commonly interpreted as indicating enhanced sympathetic modulation^46,47^. Correlation analyses revealed a weak positive association between DFA α1 and NASA-TLX workload ratings. This pattern is consistent with the notion that DFA α1 is to some extent sensitive to variations in subjective workload, although its association appears less robust than that observed for entropy-based metrics. These findings suggest that while DFA α1 provides complementary information regarding autonomic regulation under cognitive stress, its sensitivity to graded workload differences may be comparatively limited.
As shown in Table 2; Fig. 7, DFA α1 exhibited a significant main effect of task difficulty, with lower values under the low-difficulty condition than under the medium- and high-difficulty conditions, while no reliable difference was observed between the latter two. Unlike entropy-based nonlinear metrics that exhibited graded sensitivity across all difficulty levels, the short-term scaling exponent DFA α1 has shown dynamic changes across broad intensity domains but tends to display a compressed response range at higher load levels, suggesting that it is particularly effective at differentiating low-load from moderate workload states, yet may be less sensitive in finely discriminating between medium and high workload conditions. Such patterns have been noted in studies linking DFA α1 responses to graded physiological demands, where its dynamic range corresponded with broad intensity transitions but flattened within higher intensity bands, making it more state-dependent than a precise fine-grained marker^48,49^.
Several limitations should be acknowledged. First, the inclusion of multiple HRV indices spanning frequency-domain and nonlinear domains increases the possibility of detecting statistically significant effects by chance. Although correction procedures were applied in pairwise comparisons, the exploratory (pilot) nature of this study warrants cautious interpretation of isolated significant findings, particularly when effect sizes were modest. Future research employing preregistered hypotheses, multivariate modeling approaches, or dimensionality-reduction techniques may help mitigate potential inflation risks and clarify the unique versus overlapping contributions of different HRV metrics in characterizing cognitive workload. Second, the relatively small sample size may have constrained statistical power, and the short task duration (15 min) was insufficient to fully capture potential time-on-task effects. Under conditions of sustained cognitive demand, autonomic responses may evolve over longer exposure periods, which were not examined in the present design. Third, because the medium-difficulty condition was consistently administered in the middle session of the multi-day protocol, potential session-order or motivational effects specific to the intermediate stage cannot be fully ruled out. Although this design choice was intended to minimize fatigue and maintain a symmetrical experimental structure, future studies should adopt fully counterbalanced task orders to more rigorously disentangle task difficulty from order-related influences. Furthermore, although a uniform feedback mechanism was employed to minimize performance-related emotional fluctuations, no formal manipulation check was conducted to verify whether participants consistently perceived the feedback as valid, particularly under high-difficulty conditions with elevated error rates. It therefore cannot be excluded that discrepancies between subjective performance and displayed feedback may have influenced autonomic responses. Finally, the absence of multimodal physiological measures—such as cortisol assessment or respiratory monitoring—limits the comprehensiveness of the autonomic interpretation. Future investigations should incorporate multimodal monitoring, extend task exposure duration, and employ larger samples and longitudinal designs to further validate the robustness and ecological applicability of nonlinear HRV metrics in both experimental and real-world workload assessment contexts.
This study examined a multi-level mental arithmetic paradigm designed to elicit graded psychological stress. NASA-TLX and behavioral measures indicated that the task difficulty manipulation was effective. HRV analyses showed that LF-related metrics differentiated low from medium/high workload conditions but exhibited limited additional changes between the medium and high levels, consistent with a saturation-like pattern. Under low workload, reduced LF-related responses were observed, which may reflect task disengagement or psychological detachment.
In contrast, nonlinear HRV metrics (ApEn, SampEn, and DFA α1) demonstrated greater sensitivity to differences in task difficulty and showed response patterns more closely aligned with subjective workload ratings. Entropy-based measures tended to decrease with increasing perceived workload, a pattern consistent with reduced autonomic complexity under higher stress.
From a methodological perspective, the precisely graded, unified-feedback arithmetic paradigm provides a controlled experimental framework for investigating stress-related HRV responses under varying cognitive loads. From an applied perspective, nonlinear HRV metrics may offer complementary and potentially sensitive indicators for psychological stress assessment in laboratory and health-monitoring contexts, warranting further validation in larger and more diverse samples.
It is important, however, to contextualize these findings within the constraints of a controlled laboratory paradigm. The present study employed fixed pacing, standardized feedback, minimal physical movement, and relatively stable respiratory patterns, thereby maximizing internal validity and signal stability. In contrast, real-world workload monitoring—such as in occupational, educational, or clinical environments—typically involves spontaneous respiration, posture changes, environmental variability, and concurrent emotional influences. Under such ecological conditions, the relative performance characteristics of frequency-domain and nonlinear HRV metrics may differ substantially. Therefore, while entropy-based measures appeared relatively more sensitive within the present controlled setting, caution is warranted when generalizing these sensitivity patterns to ambulatory or field-based workload monitoring contexts. Future studies directly comparing laboratory-induced workload manipulation with ecologically valid monitoring paradigms would help clarify the translational implications of these findings.