Authors: You-Shyang Chen, Ching-Hsue Cheng, Su-Fen Chen, Jhe-You Jhuang
Categories: Article, Framingham risk attributes, Framingham risk score (FRS), applications of medicine, cardiovascular disease, entropy-based rule model, machine learning techniques
Source: Entropy
Doi: 10.3390/e22121406
Since 2001, cardiovascular disease (CVD) has had the second-highest mortality rate, about 15,700 people per year, in Taiwan. It has thus imposed a substantial burden on medical resources. This study was triggered by the following three factors. First, the CVD problem reflects an urgent issue. A high priority has been placed on long-term therapy and prevention to reduce the wastage of medical resources, particularly in developed countries. Second, from the perspective of preventive medicine, popular data-mining methods have been well learned and studied, with excellent performance in medical fields. Thus, identification of the risk factors of CVD using these popular techniques is a prime concern. Third, the Framingham risk score is a core indicator that can be used to establish an effective prediction model to accurately diagnose CVD. Thus, this study proposes an integrated predictive model to organize five notable the rough set (RS), decision tree (DT), random forest (RF), multilayer perceptron (MLP), and support vector machine (SVM), with a novel use of the Framingham risk score for attribute selection (i.e., F-attributes first identified in this study) to determine the key features for identifying CVD. Verification experiments were conducted with three evaluation criteria—accuracy, sensitivity, and specificity—based on 1190 instances of a CVD dataset available from a Taiwan teaching hospital and 2019 examples from a public Framingham dataset. Given the empirical results, the SVM showed the best performance in terms of accuracy (99.67%), sensitivity (99.93%), and specificity (99.71%) in all F-attributes in the CVD dataset compared to the other listed classifiers. The RS showed the highest performance in terms of accuracy (85.11%), sensitivity (86.06%), and specificity (85.19%) in most of the F-attributes in the Framingham dataset. The above study results support novel evidence that no classifier or model is suitable for all practical datasets of medical applications. Thus, identifying an appropriate classifier to address specific medical data is important. Significantly, this study is novel in its calculation and identification of the use of key Framingham risk attributes integrated with the DT technique to produce entropy-based decision rules of knowledge sets, which has not been undertaken in previous research. This study conclusively yielded meaningful entropy-based knowledgeable rules in tree structures and contributed to the differentiation of classifiers from the two datasets with three useful research findings and three helpful management implications for subsequent medical research. In particular, these rules provide reasonable solutions to simplify processes of preventive medicine by standardizing the formats and codes used in medical data to address CVD problems. The specificity of these rules is thus significant compared to those of past research.
Keywords: applications of medicine, cardiovascular disease, Framingham risk score (FRS), Framingham risk attributes, entropy-based rule model, machine learning techniques
This section explores the research background and problem in the relevant medical domains, the research gaps and motivations, and the study goals and research objectives.
Cardiovascular disease (CVD) is one of the main causes of death [1] in most countries and likely results in related problems in the blood vessels or the heart, such as cerebrovascular disease (i.e., stroke), congenital heart disease, coronary heart disease (CHD) [2], heart failure, peripheral artery disease, raised blood pressure (i.e., hypertension), and rheumatic heart disease. It is estimated that 25 million new cases of heart disease each year are diagnosed. Thus, CVD has been a key cause of death by a serious illness in recent years [1,3]. Effective identification of the risk factors of CVD is therefore highly important for clinical research for long-term therapy and prevention to reduce wastage of medical resources. Based on past studies, the main risk factors of CVD include an unhealthy diet, harmful use of alcohol and tobacco, and physical inactivity. According to a report from the World Health Organization [4], in 2015, there were 17.7 million deaths related to CVD, representing 31% of global deaths and a higher mortality risk than that of general diseases. Among these deaths, 6.7 million died from a stroke, and 7.4 million died from CHD. In 2015, 82% of premature deaths (under age 70) were caused by non-communicable diseases (NCDs) in low- and middle-income countries (LMICs), 37% of which resulted from CVD. In recent decades, the mortality rate of CVD has declined in high-income countries (HICs). Conversely, it has increased in LMICs surprisingly rapidly. However, it is possible that efforts to improve healthcare interventions against risk factors of CVD result in a significant reduction of deaths and medical resources and socioeconomic burdens. Thus, identification of the CVD problem is first emphasized in this study. This problem poses significant challenges and an opportunity in advanced preventive medicine for medical application research.
Given the above characterizations, it is clear that the issue of identifying CVD is urgent. Thus, the related issue of the CVD problem is a key research concern and is one of the core research goals of this study. The benefits of this study are an exploration of the risk factors of CVD fatal diseases and effective methods for their identification. In the context of past studies, these are novel findings.
Table 1 lists the top five causes of death in 2015 in Taiwan, as reported by the Ministry of Health and Welfare [5]. It is clear that CVD has been the second most common cause of death since 2001. CVD not only results in higher medical resource expenses but also greater long-term expenditure for countries and individual families. To tackle this severe problem, developing countries have dedicated funds to CVD prevention and treatment through national education and training resources, thereby reducing incidence and mortality rates [6]. However, due to changes in lifestyle and eating habits in recent decades, incidence rates of CVD among younger people have experienced a growing trend. Thus, early detection and prevention of CVD and closing this research gap related to younger people have become critical issues. From the perspective of preventive medicine, the design of an effective predictive (classification) model to help doctors in the early and accurate diagnosis of CVD is required, and the effectiveness and quality of prevention and treatment need to be improved. Thus, this study is motivated by the following. First, the classification function of machine learning techniques for mining meaningful CVD rules (knowledge) from large amounts of data is a valuable approach. Notably, classification techniques have been successfully applied in medical fields, such as breast cancer [7,8] and heart diseases [9,10], by studying past research. Furthermore, Boursalie et al. [11] efficiently utilized a support vector machine (SVM) classifier to monitor CVD via effective features and characteristics. Second, it is clear from a limited literature review that using effective risk factors (e.g., Framingham risk attributes) to mine CVD data is efficient and effective [1]. Although past studies have used features of Framingham risk to help prevent CVD [12], they have rarely used a hybrid model to integrate the Framingham risk score and classification techniques to address the issue of CVD for doctors and patients. Finally, the Framingham risk score is an effective instrument that can be used to build a hybrid prediction model to accurately diagnose CVD. Thus, this study is beneficial both theoretically and practically.
Based on the above descriptions, serious diseases (e.g., CVD) require greater healthcare and the identification of potential preventable materials or approaches at both the country and individual family levels. The challenge related to CVD has thus attracted significant attention from practitioners and academics. Therefore, the study examines the noteworthy and important issue of effectively identifying the risk factors of CVD.
The CVD problem reflects an urgent issue. A high priority has been placed on long-term therapy and prevention to reduce the wastage of medical resources via the greater use of preventive medicine, particularly for developed countries. Furthermore, popular data-mining methods have been well learned and studied, with excellent performance in medical fields. Thus, identification of the risk factors of CVD using popular data-mining techniques is also a major goal of this study. In addition to addressing the issue of CVD, the study also aims to reduce the wastage of medical resources and financial expenditure.
In data-mining fields, various emerging machine learning models have been found to be superior in their application to healthcare issues [13]. Thus, in this study, a hybrid predictive model was first built to identify CVD using five well-performing classifiers in the medical domain, namely, decision tree (DT), multilayer perceptron (MLP), random forest (RF), rough set theory (RST), and SVM, and simultaneously measure the performance of the Framingham risk score with full attributes. In practice, the study used two real datasets collected from a Taiwan teaching hospital with CVD cases and a public Framingham dataset from the Internet for the further benefit of preventive treatment. To summarize, this study has three research (1) construction of a hybrid predictive model of CVD based on diverse classifiers and the Framingham risk score; (2) identification of the key determinants of CVD by attribute selection methods and extraction of comprehensible entropy-based rule sets based on DT; and (3) provision of analytical results and research findings with management implications to healthcare providers, physicians, and patients as a useful medical reference. The entropy-based method is beneficial for information gain [14].
The remainder of this paper is organized as A literature review of the study issues, including CVD, the Framingham risk score, and the five classification techniques, is provided in Section 2. The research concepts justifying the procedure of the proposed method are presented in Section 3. The analysis results and the core research findings from the experiments are provided in Section 4. Finally, Section 5 concludes and suggests subsequent research.
The literature and concepts related to the identification of CVD, the Framingham risk score, and the five classification techniques are introduced and explored in this section.
CVD is a disease related to the blood vessels and heart, and coronary artery diseases (CAD, e.g., heart attacks or angina). Its major clinical symptoms have various levels of severity and include cardiomyopathy, congenital heart disease, hypertensive heart disease, peripheral artery disease, rheumatic heart disease, and stroke [15]. In clinical practice, CVD and cerebrovascular disease have a cause-and-effect relationship [16]; thus, CVD is a sign of a serious illness because CVD patients usually also have various extensions of serious chronic diseases (e.g., hypertension, stenocardia, hyperlipidemia, hyperuricemia, diabetes mellitus, and obesity). Similarly, other heart diseases, such as coronary sclerosis, valvar heart disease, and ventricular fibrillation arrhythmia, and serious conditions, such as cerebral embolism and myocardial infarction, are complicated and exacerbated by the negative effect of CVD. Henriksson et al. [17] provided evidence that cardiorespiratory fitness, body mass index, and muscular strength in adolescents can lead to later chronic incapacity and CVD disability, including cerebrovascular disease, heart failure, and ischemic heart disease.
The risk factors of CVD must first be explored and identified. These risk factors are classified into two categories. The main risk factors include diabetes mellitus, an older age, family history, hyperlipidemia, hypertension, and obesity. Subordinate risk factors include high uric acid, narrow heart disease, myocardial infarction, arrhythmia, cardiogenic shock, peripheral vein disease, and smoking. The three main risk factors of CVD disease are as
In addition to these three main risk factors, other potential factors need to be further examined and identified effectively to enable the results to be compared to those of the literature.
Identifying and preventing the major modifiable risk factors of CVD in advance can lower the prevalence rate for related heart disorders [1]. In 1948, the Framingham Heart Study (FHS), directed by the National Heart Institute, USA, began studying cardiac disease (heart disease) for health research purposes. At that time, the general cause–effect relationship between heart disease and stroke was not known. Subsequently, however, the mortality of CVD has steadily increased, reaching epidemic levels in the USA at the beginning of the current century [23]. The main aims of this project were to identify the common characteristics and risk factors of CVD patients by monitoring participants who had not yet suffered strokes and heart attacks due to CVD. The FHS investigated the risk factors of CVD in well-designed experimental groups, which consisted of 5209 people from Framingham, Massachusetts, USA, between the ages of 30–62 and free of CVD [23].
Investigators of the FHS first developed the risk equation of CHD [24] for clinicians to predict coronary disease from 1950 to the mid-1960s, reflecting early efforts from investigators. The long-term studies of the FHS found data with respect to related heart diseases. In addition to age and gender, the risk factors of CVD were found to include high SBP, serum cholesterol, glucose intolerance (e.g., diabetes), and left ventricular hypertrophy (LVH); furthermore, the FHS further identified lifestyle factors, such as sedentary lifestyle, eating an unhealthy diet, and cigarette smoking. In addition, the influence of HDL-C [25] was measured and is reflected in the equations. Measurement of total HDL-C was also found to be superior to that of serum cholesterol as a predictor of CVD. Based on the above, high blood cholesterol, high blood pressure, diabetes, obesity, physical inactivity, and smoking were defined as the main risk factors of CVD. Other characteristics, such as age, blood triglycerides, gender, HDL-C level, and psychosocial issues, have also been determined as relevant factors over the years [26]. More important, behavioral mechanisms for lowering the main risk factors of CVD showed a positive relationship with modern medical trends, which resulted in an emerging treatment and prevention strategy for the practice of clinical medicine.
The National Heart, Lung, and Blood Institute (NHLBI) and Framingham investigators continue to research genetic factors (genetic material or genetic patterns) of CVD using new diagnostic technologies. Hojat et al. [1] compared the risk factors of CVD among male and female nurses and found that men have a higher risk and more complications of CVD than women. According to the FHS conversion standard, the index factors of Framingham risk are age, blood pressure, diabetes status (e.g., fasting glucose), gender, HDL-C level, smoking status, and total cholesterol level [27,28]. Consequently, the transformed scores of the above factors are listed in Table 2, Table 3 and Table 4 below.
This subsection uses advanced classification machine learning technology to highlight the purpose of dealing with large amounts of medical data from clinical databases. This study applied five noted classification technologies, namely, RST, DT, RF, MLP, and SVM, which were selected because they were previously used in medical classification and prediction of healthcare problems and found to have superior performance.
RST was developed initially by Pawlak [29]. It classifies by analyzing incomplete, unspecific, or imprecise information and extracts decisional rule-sets. In the data analytics community, RST is used for non-statistical techniques. Its essential conception is to use β-lower and β-upper approximations for a given set, and the β-lower and β-upper approximations form formal classification information in a specific domain. Objects shape a part of a subset, and objects that possibly shape a part of a subset characterize the subset of the β-upper approximations and the subset generated by the β-lower approximations. Each subset determined by β-lower and β-upper approximations is called a rough set [30,31,32]. RST is widely used today in medical research, process control, and database attribute analyses. Other fields include heart disease [33], medical data diagnosis [34], multiobjective attribute reduction [35], the safe use of medical devices [36], short text data [37], textile applications [38], and the supply chain of auto spare parts [39].
RST provides four benefits when processing classification (1) It is unnecessary to use external information; (2) it is appropriate for studying qualitative and quantitative data; (3) hidden information from given data is discovered through decisional rule-sets and is thus easier to interpret; and (4) extracted decisional rules eliminate redundant or useless information from the given dataset. In further exploration, RST is implemented in six steps [40]:
The literature suggests that RST is suitable for use in medical fields [45], and it was thus selected for this study.
DT is a regression or classification approach for forming a tree-based rule structure from a supervised algorithm. DT uses supervised learning technology and has a classification (or prediction) function for well-known machine learning techniques in various application domains. Thus, DT results in fast formations and produces an entropy-based decision rule of easy if–then–else interpretations in a variety of fields; it has become a common application technique among classification approaches [46,47]. DT algorithms contain two developing a tree and pruning a tree. For the first phase, DT selects a better suitable attribute from a split training dataset. The final outcome only applies to that data, and the common training splits of the sub-dataset target a specific class. The repetition or recursion functions concerning the selection attribute and splitting-set construct internal nodes and a root node. In the pruning-tree phase, some data from the trained sub-dataset may have an inappropriate branch after the DT structure is built, i.e., overfitting, and these inappropriate branches must be eliminated to improve the classification performance [48,49].
In the DT community field, a core algorithm for employing a top-down structure is called iterative dichotomizer 3 (ID3) [50]. ID3 uses information theory to build an entropy-based tree rule for information gain and entropy. The key concept of quantifying information is called entropy and is mainly used for calculating the homogeneity of the given data by the ID3 algorithm. The entropy will be zero if the data sample is homogeneous; conversely, it will be one if the data sample is completely heterogeneous. Information gain is used for training a decisional tree and is based on the reduction in entropy when a given dataset is converted. A calculation is made by comparing the entropy of this dataset before and after conversation; thus, the highest information gain will be used to construct a decisional tree in the process of finding attributes. The entropy is defined as the following Equation (1) by using a frequency table of one
where p is the probability of one class.
A collection set S of outcomes is determined, and the entropy is formatted
where pi is the proportion of S, which belongs to class i?
Gain (S, A) is the information gain of a case set S on attribute A and is formatted
where v is a value of A, Sv is a subset of S, |Sv| is the number of elements on Sv, and |S| is the number of elements on S.
In previous studies, the related DT algorithms have typically been applied to functions of the relevant areas of specialization, such as churn prediction [51], forecasting corporate credit ratings [52], object classification in autonomous driving [53], and the identification of a route selection strategy in classification [49]. Thus, it is an effective method for the selection of medical research.
RF is a novel and highly effective machine learning technique. It is a type of ensemble-learning approach [54] that combines predictors of trees such that trees are highly dependent on random vector values for independent samples with the same classification as those in decisional trees in a constructed forest. The basic concept of RF is to assemble multiple DTs, bag, and bootstrap [7], and identify the classification results with negative or positive information voted using individual decision trees. It is capable of handling large datasets and process a large number of input variables without first deleting variables. RF classifiers can provide two measurements of randomness for a comprehensive view of a tree the first is for data, and the second is for features.
RF can be applied in diverse domains due to its features, and particularly in clinical diagnosis [55]. Thus, RF was used to processing the CVD problem in this study.
MLP works on a supervised neural network model that uses a simple perceptron learning algorithm and an original backpropagation for training. MLP is adopted in most series forecasting approaches [56]. It usually has three types of layers of the input layer (i.e., for some source nodes), hidden layers (i.e., for computational nodes), and the output layer, which highlights the learning performance and the MLP classifier [57]. Some hidden nodes possess the ability of MLP learning classifiers with a minimum of hidden nodes; hence, the network is the most recommended and selected.
Regarding the layers for the MLP, each layer has a plurality of processing units, and each unit has subsequent layers in unit-weighted links. An MLP with inputs from a plurality of nonlinear variables to an output [58] is depicted in Figure 1.
Figure 1 Schematic diagram for multilayer perceptron (MLP).
If an input vector of an input layer is acquired from an output value bnj of hidden nodes, the following equation is
where φ refers to an activation function from the hidden nodes, θM represents a threshold from the hidden nodes, wij refers to a connection weight between the hidden layers and the input layers, and xi is used as the input vector from these input nodes. The only output from these output nodes is acquired from the nodes of a hidden layer in a similar
The activation function used is sigmoid and defined
The training of an MLP network is supervised. MLP can use the gradient descent algorithm [59] for training or use the conjugated gradient algorithm with a nonlinear optimal algorithm. It is possible to accelerate the convergence speed of the weights with respect to the gradient descent algorithm.
Due to its excellent prior performance in medical applications [60], MLP was also selected for comparative analysis of the proposed study model.
SVM, invented initially by Vapnik, is one of the most common machine learning approaches used, with excellent performance in various applications [61]. SVM is a supervised learning method [62] with the main functions of regression and classification; it separates linearly separable input data into preset classes using the hyperplanes from an n-dimensional space. Using an optimal separated hyperplane maximizes the margins between data points that are nearest to each other within the different known classes. SVM was used in many prior studies [63,64,65] and has a good professional reputation due to its ability within a binary classification model.
Based on its strong positive reputation for medical research [66], SVM was thus selected and used for one of the proposed models in this study.
To achieve the study goal, another key function of a classification technique is to classify an unknown category of given data objects corresponding to the known category for prediction purposes. The most challenging task of classification techniques is to learn how to select a suitable technique to improve the accuracy in the medical field. Previous studies of Framingham risk attributes have used statistical methods conducted with the SPSS software, in addition to the area under the curve (AUC) of the receiver operating characteristic (ROC) curve, sensitivity, and specificity, to show the prediction power for addressing CVD problems during the past 10 years [12]. However, they did not apply the Framingham risk score to intelligent machine learning classification techniques for identifying CVD and, in particular, for the extraction of knowledgeable rules. Thus, the above gaps should be bridged using well-defined methods of techniques used in medical research applications, which is the focus of this study. This study proposed using the Framingham risk score to filter attributes and then employing five classification techniques to build a hybrid model for identifying CVD. Two CVD datasets were compiled for the verification of the proposed model, and the performances of the Framingham risk score, Framingham attributes, and the full attributes in five classifiers were compared.
The proposed method has five compile the dataset → preprocess the data → select the attributes → build the model → evaluate the results. Figure 2 shows a flow diagram of the algorithm proposed in this method, and detailed information about the proposed algorithms is provided in the following.
Figure 2 Research procedure of the proposed method.
The first CVD dataset is employed to illustrate the proposed method. Based on a literature review, there are eight types of related CVDs, such as arrhythmia, cardiogenic shock, and diabetes mellitus. Twenty original attributes are reduced to 13 according to the expert opinion of physicians. Consequently, the 13 attributes include five physical exam attributes, seven blood test attributes, and one decision-attribute of the class with the eight CVD names noted above. There are two categories in the class, that is, Y (Yes): Have at least one CVD and N (No): None. These are listed in Table 5 below. Next, the second dataset has all of the attributes, including age, sex, serum cholesterol, DBP, SBP, Metropolitan relative weight, smoke, and CVDs (class), which is based on the expert recommendation of doctors. For further understanding of these datasets, their properties are shown in Table 6.
Step Select attributes. First, the eight attributes of the Framingham risk score are age, gender, DBP, SBP, LDL-C, HDL-C, fasting glucose, and smoking, which are named the Framingham attributes. Next, these factors are selected to calculate and transform their score values based on the conversion method of the Framingham risk score discussed in Section 2.2 [28] and shown in Table 2, Table 3 and Table 4 from the CVD dataset. Finally, all 13 attributes of the CVD dataset are selected and called the full attributes, which additionally include body mass index, waistline, red blood cells, and white blood cells.
Step Build a model. This step designs a hybrid model of five classifiers (i.e., RST, DT, RF, MLP, and SVM) with various attribute components, which include the Framingham score attributes, the Framingham attributes, and the full attributes, to highlight and differentiate the performance of the proposed method based on a commonly used 67–33% training–testing ratio. This step can be divided into five sub-steps and is executed using different software packages. The procedure of the five sub-steps is as First, the selected attributes are used as input variables. Second, for the percentage-split data method, the two training–testing sub-datasets are formed using the common 1 ratio to achieve a good and reasonable result in practice. Thus, 67% of the data is used as a training sub-dataset, and the remaining 33% of the data is used as the testing sub-dataset. Third, all of the default parameters are defined to implement each of the above classifiers. Fourth, RST is applied using the rough set exploration system (RSES) [29], and DT, RF, MLP, and SVM are applied separately. Fifth, comprehensive knowledge-based rule sets are created using DT. For further details, the pseudo-code of the construction of the hybrid model is shown in Algorithm 1.
Step Evaluate the results. To evaluate machine learning, most researchers use three common accuracy, sensitivity, and specificity. These are defined in Table 7 of the confusion matrix and Equations (7)–(9). In Table 7, a true positive (TP) means the truth is positive, and the prediction is positive; a false positive (FP) means the truth is negative, but the prediction is positive; a false negative (FN) means the truth is positive, but the prediction is negative, and a true negative (TN) means the truth is negative, and the prediction is negative.
This section verifies the proposed method and its algorithms with two medical datasets and compares the listing models to further evaluate the classification performance.
To further explore the experimental CVD dataset, Table 8 shows the descriptive statistics for the first CVD dataset using the chi-squared test to summarize and compare the baseline characteristics of the continuous attributes and the categorical attributes.
The extracted entropy-based decision rules, core attributes, accuracy rate, sensitivity, and specificity were calculated from the two medical datasets after the experiments.
After data preprocessing and attribute selection, the Framingham attributes and full attributes were identified, and the five different classifiers were used and compared with various evaluation standards for the overall accuracy rate using a training–testing ratio of 67–33% to measure the performance of the proposed method.
Consequently, the visualized tree structure of if–then–else control statements in the DT model was used to identify future CVD issues. Figure 3 lists the entropy-based rule results of the CVD dataset in a visualized tree structure. As shown in Figure 3, one case is exemplified and highlighted in red and green, and one key result of the core attributes is defined.
Figure 3 Entropy-based rule results of the cardiovascular disease (CVD) dataset in a visualized tree structure.
Accordingly, Table 9 shows the comparative results of various evaluation standards on the Framingham score attributes, the Framingham attributes, and the full attributes of the CVD dataset with five classifiers. In Table 9, it is clear that SVM outperforms the other four classifiers in terms of accuracy (99.67%), sensitivity (99.93%), and specificity (99.71%) for the three aspects of the Framingham score attributes, the Framingham attributes, and the full attributes. This implies that SVM is more suitable for identifying and addressing CVD than the other methods examined in this study.
The second Framingham dataset with 18 original attributes was also considered. Similarly, after the preprocessing and identification of the Framingham attributes and the full attributes, this dataset was used with the five different classifiers to assess the evaluation performance of the proposed method in terms of average accuracy rate for 10 repetitions, also using a training–testing ratio of 67–33%.
DT was also used to visualize the tree structure of if–then–else control statements for identifying CVD. Figure 4 shows the entropy-based rule results of the Framingham dataset in a visualized tree structure. In Figure 4, two key points can be determined. First, a case from the tree-like structure is highlighted in red and green. Next, key attributes are accordingly identified and determined.
Figure 4 Entropy-based rule results of the Framingham dataset in a visualized tree structure.
Table 10 lists the analytical results for the Framingham attributes and the full attributes in the Framingham dataset after the experiments. In Table 10, the five classifiers are compared with the three evaluation standards in terms of accuracy, sensitivity, and specification. It is clear that the RS method has the highest performance in terms of accuracy (85.11%), sensitivity (86.06%), and specificity (85.19%) in all Framingham score attributes, the Framingham attributes, and the full attributes for the Framingham dataset. This information implies that the rough set theory is more suitable for identifying CVD in the Framingham dataset than the other four classifiers examined in this study.
Three key findings and management implications follow from the experimental
This study proposes a hybrid method to integrate and model Framingham risk attributes and five novel classification techniques—RST, DT, RF, MLP, and SVM—for the identification of key attributes that influence CVD and to highlight preventive practices in healthcare services. The study’s contribution consists of calculating the score of Framingham risk attributes and identifying CVD using a suitable classifier for different datasets in hybrid medical applications. This contribution differs from those of previous studies [67]. For verification using the three evaluation criteria (accuracy, sensitivity, and specificity), 1190 instances in the CVD dataset available from Taiwan’s regional teaching hospital and 2019 examples from the public Framingham dataset were used. SVM showed the best performance in terms of accuracy (99.67%), sensitivity (99.93%), and specificity (99.71%) in all of the F-attributes in the CVD dataset, and RS showed the best performance in the accuracy (85.11%), sensitivity (86.06%), and specificity (85.19%) in most of the F-attributes for the Framingham dataset. Consequently, three main points can be made regarding the contribution, specificity, and novelty of this (1) Regarding its contribution, this study supports meaningful entropy-based knowledgeable rules for visualizing a tree structure and differentiates the classifiers from the two datasets, resulting in three useful research findings and three helpful management implications for subsequent medical research and other interested parties. (2) Regarding its novelty, the study results provide novel evidence that indicates no classifier or model is suitable for all practical datasets of medical applications. Thus, finding an appropriate classifier to address specific medical data is highly important. Furthermore, this study is the first to calculate and identify the use of key Framingham risk attribute scores, integrated with the five classification techniques noted above and with the DT technique, to produce entropy-based decision rules of knowledge sets. This has not been achieved in previous studies. (3) Regarding the specificity of the study, the knowledgeable rule sets created by DT provide reasonable solutions to simplifying the processes of preventive medicine by standardizing the formats and codes used in medical data to address CVD problems. The specificity of these rules is thus significant compared to those of past research.
Five conclusive and important directions are indicated by the analytical results of the experiments using the two medical
The authors would like to cordially express many thanks for the above financial support of this paper and also really appreciate the editors of Entropy and the anonymous reviewers for providing their constructive suggestions for the improvement of paper quality.
Conceptualization, J.-Y.J.; methodology, C.-H.C. and Y.-S.C.; software, S.-F.C.; validation, Y.-S.C. and S.-F.C.; formal analysis, S.-F.C.; investigation, Y.-S.C. and C.-H.C.; resources, Y.-S.C.; data curation, S.-F.C.; writing—original draft preparation, J.-Y.J.; writing—review and editing, Y.-S.C. and C.-H.C.; visualization, Y.-S.C. and C.-H.C. All authors have read and agreed to the published version of the manuscript.
This research was supported and funded by the Ministry of Science and Technology, Taiwan, grant number MOST 109-2221-E-146-003.
The authors declare that there are no conflict of interest.