Authors: Lianbaichao Liu, Zhanping Song, Ping Zhou, XinHe He, Liang Zhao
Categories: Article, Civil engineering, Geology, Lithology identification, Neural network, Rock strength, Tunnel face, Weathering degree
Source: Scientific Reports
Authors: Lianbaichao Liu, Zhanping Song, Ping Zhou, XinHe He, Liang Zhao
In geological engineering and related fields, accurately and quickly identifying lithology and assessing rock strength are crucial for ensuring structural safety and optimizing design. Traditional rock strength assessment methods mainly rely on field sampling and laboratory tests, such as uniaxial compressive strength (UCS) tests and velocity tests. Although these methods provide relatively accurate rock strength data, they are complex, time-consuming, and unable to reflect real-time changes in field conditions. Therefore, this study proposes a new method based on artificial intelligence and neural networks to improve the efficiency and accuracy of rock strength assessments. This research utilizes a Transformer + UNet hybrid model for lithology identification and an optimized ResNet-18 model for determining rock weathering degrees, thereby correcting the strength of the tunnel face surrounding rock. Experimental results show that the Transformer + UNet hybrid model achieves an accuracy of 95.57% in lithology identification tasks, while the optimized ResNet model achieves an accuracy of 96.13% in rock weathering degree determination. Additionally, the average relative error in tunnel face strength detection results is only 9.33%, validating the feasibility and effectiveness of this method in practical engineering applications. The multi-model neural network system developed in this study significantly enhances prediction accuracy and efficiency, providing robust scientific decision support for tunnel construction, thereby improving construction safety and economy.
During tunnel construction, assessing the strength of the rock at the tunnel face is crucial due to the complex and variable geological conditions, which pose significant challenges for accurate evaluation. Including visuals such as images of tunnel face conditions and rock samples can highlight these challenges and underscore the importance of the proposed AI-based methods in improving assessment accuracy and construction safety.
Traditional methods primarily rely on on-site sampling and laboratory testing, such as uniaxial compressive strength (UCS) tests and velocity tests. While these methods provide relatively accurate rock strength data, they are complex and time-consuming^1–3^. More importantly, traditional methods cannot reflect real-time changes in on-site conditions. During tunnel construction, geological conditions are complex and variable, and the physical and mechanical properties of the rock can change significantly with construction progress and external environmental changes^4–7^. The results of traditional tests often lag, making it difficult to reflect the current state of rock strength in a timely manner^8,9^. This lag not only reduces the practical application value of the test results but also potentially increases safety hazards during construction^10–14^.
Given these issues, the application of existing methods for obtaining rock strength in modern tunnel construction faces many challenges^15–17^. To improve the efficiency and accuracy of rock strength assessment, there is an urgent need to explore new technologies and methods. In this context, artificial intelligence technology, particularly neural network methods, has gradually attracted the attention of researchers. Research on the application of neural network methods in geological engineering is increasing^18–21^. One of the most popular methods is UNet, designed by Olaf Ronneberger, Philipp Fischer, and Thomas Brox in 2015 for medical imaging. The unique feature of UNet is its U-shaped architecture, comprising two one for down-sampling (reducing the image) and the other for up-sampling (expanding the image). The first part identifies objects in the image, while the second part makes the image clearer. Due to UNet's excellent performance in image object extraction, researchers have applied it to various object extraction tasks, such as lung nodule and tumor detection in CT images^22,23^, building extraction in satellite images^24,25^, and road debris classification and identification^26^. In the construction field, researchers have also utilized UNet for road detection and metro tunnel leakage inspection^27–29^.
The Transformer model was introduced by Vaswani et al. in 2017 at Google Brain^30^. It is faster and more efficient than traditional models (such as RNNs and CNNs) because it employs a self-attention mechanism. It has gained popularity in natural language processing tasks, such as machine translation and language modeling. Google's BERT model is an example of the Transformer architecture, achieving outstanding results in many NLP tasks^30^. The Transformer model has many advantages, such as parallel computing and capturing long-distance dependencies, but it can also be complex and sensitive to sequence length variations. Researchers have begun exploring the application of Transformer-based methods in 2D image segmentation tasks. In medicine, it has been used for kidney stone detection^31,32^. In aerial images, Bi et al. employed ViT for object classification^33^, and some studies have applied it to forest fire segmentation^34^. In tunnel construction, Transformer has been used for similar tasks, such as crack detection^35–38^, electronic detonator misfire detection^39^, automatic low-resolution borehole image stitching, and improving GPR surveys in tunnel construction^40,41^.
ResNet (Residual Network) is a deep convolutional neural network structure proposed by He et al. at Microsoft Research in 2015. The innovation of this model lies in the introduction of residual blocks, which significantly alleviate the problem of vanishing and exploding gradients as network depth increases^42^. The ResNet structure can be easily extended to deeper networks, such as ResNet-50, ResNet-101, and ResNet-152, while maintaining good performance as depth increases. ResNet has been applied in various aspects of construction, including detecting cracks on the surfaces of tunnels and bridges^43,44^, TBM vibration analysis prediction, and EPB utilization coefficient prediction accuracy^45,46^.
This study innovatively applies the Transformer + UNet hybrid model for lithology identification in tunnel construction, enhancing segmentation accuracy through superior global contextual information capture. Additionally, the ResNet-18 model is utilized to distinguish weathering degrees, significantly improving the precision of rock strength evaluation. This dual-model approach offers substantial theoretical and practical advancements in rock strength assessment.For instance, by processing tunnel face images, ViT can efficiently identify rock types and fracture distributions, providing reliable data support for construction^47^. Additionally, UNet is used in geotechnical engineering for geological profile segmentation, helping engineers better understand stratigraphy and geotechnical properties^48^. ResNet, through training on a large number of rock images, can automatically classify different rock types and identify the degree of weathering, providing scientific basis for engineering decisions^49,50^.
By analyzing real-time construction site image data, AI systems can timely detect potential geological hazards and issue warnings to construction personnel^51^ . Moreover, AI technology can be used for data analysis and optimization during construction, analyzing historical construction data through machine learning algorithms to summarize optimal construction parameters and operational processes, thereby improving overall construction efficiency and quality^52^.
In summary, exploring and developing AI and neural network-based methods for rock strength assessment has become a key direction for addressing this issue. This paper proposes an innovative method that identifies lithology through a Transformer + UNet image segmentation approach, uses ResNet18 to distinguish weathering degrees, and corrects rock strength based on weathering degree. This research has significant theoretical value and broad prospects for practical engineering applications.
Weathering is a geological process in which rocks at or near the earth's surface decompose and alter under the influence of atmospheric, hydrological, and biological factors. Weathering is generally categorized into three basic physical weathering, chemical weathering, and biological weathering, each affecting the structure and composition of rocks through different mechanisms^53^. Physical weathering involves the fragmentation of rocks due to temperature changes, freeze–thaw cycles, or salt crystal growth. Chemical weathering occurs when chemical substances in water and the atmosphere react with rock minerals, altering their mineralogical properties. Biological weathering results from biological activities, such as plant root growth or microbial metabolism, causing structural or chemical changes in rocks.
Studies have shown that weathering significantly impacts the uniaxial compressive strength (UCS) of rocks. As the degree of weathering increases, the UCS of rocks typically decreases markedly. This is due to the destruction of the internal structure of rocks during weathering, such as the increase in fractures, expansion of pores, and weakening of cohesive forces, all of which reduce the rock's load-bearing capacity^54^. For instance, an increase in moisture content, often associated with chemical weathering, not only alters the physical state of the rock but may also cause hydrolysis and dissolution of certain minerals, further reducing the rock's compressive strength.
Existing literature indicates that the reduction ratios of UCS for sedimentary, igneous, and metamorphic rocks under various degrees of weathering are as shown in Table 1. Table 1UCS reduction ratios for different rock types under various weathering grades.Rock typeSlightly weathered (%)Moderately weathered (%)Highly weathered (%)Completely weathered (%)Volcanic^55^ − 20− 40− 55− 75Limestone^56^ − 33.53− 41.28− 63.88− 100metamorphic^57^ − 43.37 to 51.57− 50.82 to 61.14− 100− 100
From Table 1, it is evident that different types of rocks exhibit significant differences in UCS changes under weathering. To obtain a general attenuation coefficient for UCS, further summarization of relevant literature was conducted, and the UCS reduction curve was plotted as shown in Fig. 1.Figure 1Effect of weathering on UCS reduction^58–68^.
As illustrated in Fig. 1, the higher the degree of weathering, the weaker the structural integrity and mechanical strength of the rock. The empirical attenuation coefficients (Fw) for general rock strength under different weathering degrees are as Fresh: Fw ≈ 1.0; Slightly Weathered: Fw ≈ 0.8; Moderately Weathered: Fw ≈ 0.6; Highly Weathered: Fw ≈ 0.3; Completely Weathered: Fw ≈ 0.
This implies that the compressive strength of rocks needs to be adjusted according to their weathering degree. The application of correction coefficients is a process of adjusting the original compressive strength of the rock based on its weathering condition to obtain a rock strength value that better reflects the actual conditions.
The final rock strength (RC value) can be calculated using the following 1\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {\text{RC}} = {\text{Rt}} \times {\text{Fw}}
This results in a more accurate rock strength value considering the effects of weathering. This outcome is significant for guiding tunnel design and construction, helping engineers select appropriate construction methods and support structures to ensure the safety and reliability of tunnel construction. This flowchart in Fig. 2 outlines the methodology used in the manuscript for determining rock strength through image processing neural networks, aiming to facilitate the readers' understanding. The methodology involves the following Building Database: Establish a standard rock database to obtain image data and rock strength data.Image Processing Neural Network: Utilize the Transformer + UNet model for lithological image segmentation and the ResNet-18 model for determining the degree of lithological weathering.Acquisition of Tunnel Face Rock Strength Values: Integrate rock strength data, lithology determination results, and weathering degree determination results to correct and obtain the final rock strength values.Figure 2The flowchart of acquisition of tunnel face strength values using image processing neural networks. The process of applying correction coefficients and determining rock strength demonstrates the advantage of combining traditional geological engineering experience with modern neural network technology. ### Database establishment #### Rock data collection To ensure comprehensive coverage of rock characteristics and fully account for their geological background and engineering applications, we collected image data encompassing igneous, sedimentary, and metamorphic rocks. The sources of rock image data are one part comes from online rock databases from various global regions, including databases from authoritative organizations such as the National Mineral Rock Fossil Specimen Resource Sharing Platform of China (NIMRF), the United States Geological Survey (USGS), and the British Geological Survey (BGS). These databases have established extensive rock databases, providing detailed information on the composition, structure, formation environment, and geographical location of rocks, as well as high-resolution images of various rock types. The other part of the data comes from on-site collection during tunnel, slope, and highway construction projects both domestically and internationally. This includes rock images captured during the construction process and the corresponding rock strength data. Currently, the data collection covers image data from construction projects in 12 provinces in China, as well as several tunnel and highway construction projects in Central Asia, West Asia, and parts of Africa. The engineering team conducted detailed mechanical tests on the collected rock samples. In tunnel construction, obtaining clear photos of the tunnel face is crucial for image training. We used high-performance DSLR cameras or high-resolution smartphones for photography, adjusting camera parameters to account for the low light and high dust environments typical in tunnels. The optimal time for capturing images is usually after blasting when the dust has settled and before the commencement of preliminary support work, as shown in Fig. 1. This approach effectively captures high-quality images of the tunnel face, providing an accurate data foundation for subsequent deep learning analysis. Figure 3 showcases examples of images from the two data sources.Figure 3Examples of images from online databases and on-site collection. #### Data preprocessing ##### Data augmentation methods When conducting machine learning training tasks, we need to preprocess the image data in the database to enhance data quality and consistency, thereby improving the accuracy and reliability of the analysis results. The first step in preprocessing is to crop the images to remove irrelevant details, ensuring the focus remains on the rock samples, which maintains dataset uniformity. Next, we employ various data augmentation methods to diversify the training dataset and enhance model robustness. The specific techniques Rotation: Images are rotated at various angles (e.g., 90°, 180°, 270°) to simulate different orientations of rock samples.Scaling: The images are resized by different scaling factors to mimic variations in distance and size.Horizontal and Vertical Flipping: Images are flipped horizontally and vertically to introduce variance in rock presentation.Brightness Adjustment: Brightness levels are randomly altered to account for varying lighting conditions in the tunnel environment.Noise Addition: Gaussian noise is added to the images to improve the model’s ability to handle real-world imperfections and noise. This augmentation process is crucial for enhancing the model's robustness and helps expand the dataset, thereby improving model performance. Each augmentation method contributes to creating a diverse set of training images, which helps in reducing overfitting and improving generalization. ##### Data annotation methods For the lithology segmentation and recognition part of this study, we accurately annotate rock lithology images based on source information, covering rock attributes such as porphyrite, granite, loess clay, fault, and background. Using the polygon annotation method in the LabelMe annotation software, we perform pixel-level annotation of all rock properties and backgrounds in the images, as shown in Fig. 4. The upper part shows the image data with red boundary lines indicating the annotated areas; the lower part shows the corresponding images after polygon annotation.Figure 4Lithology annotation method in rock images. For the weathering degree determination study, we identify the weathering degree by closely observing changes in rock structure, mineral composition, and color, as illustrated in Fig. 5. In an unweathered state, rocks maintain their original properties, with little to no change in structure and color. In the slightly weathered stage, the rock structure begins to deteriorate, with noticeable changes in color and mineral composition. Moderate weathering shows more significant changes, with intensified weathering on fracture surfaces. In the highly weathered stage, the rock structure is completely destroyed, turning into loose soil or sand-like material, with all minerals except quartz transforming into secondary minerals.Figure 5Schematic diagram of rock weathering degrees. Using classification labels for weathering determination, rock images are categorized into four unweathered (0), slightly weathered (1), moderately weathered (2), and highly weathered (3). Each rock image is assigned a corresponding label based on its degree of weathering, facilitating neural network training for classification. This method is straightforward and helps the model understand and distinguish between different weathering levels of rock images. ### Lithology identification model based on transformer and UNet Processing tunnel face images for rock lithology segmentation encounters various specific challenges due to its complexity. Firstly, the heterogeneity and diversity of surrounding rock lead to significant differences in the texture, color, and morphology of rocks, posing challenges for image segmentation. Secondly, lighting variations and noise interference in the tunnel environment affect image quality, further increasing the difficulty of image processing. Against this backdrop, traditional convolutional neural network (CNN)-based UNet models face several issues when processing tunnel face images. While the UNet model excels in fields like medical image segmentation, its limitations become apparent when dealing with tunnel face images. Firstly, UNet relies on local convolution operations, primarily focusing on capturing local features, and struggles to fully utilize global contextual information. In complex tunnel face images, global information is crucial for accurate lithology segmentation. Additionally, UNet can struggle with segmentation accuracy when dealing with complex backgrounds and blurred boundaries. To address these issues, the attention mechanism of Transformers shows excellent performance in tunnel face image segmentation. Transformers can effectively capture global contextual information through self-attention mechanisms, overcoming the limitations of traditional CNNs in global feature extraction. Compared to the UNet model, Transformers handle images with complex backgrounds and multi-scale features more accurately for segmentation and recognition. Therefore, combining Transformers with UNet to form a hybrid model can leverage the strengths of both, improving lithology segmentation performance. As shown in Fig. 6, the new hybrid model architecture replaces the encoder part of the original UNet with a Transformer (highlighted in the red box), while retaining the UNet's decoder part and the skip connections between the encoder and decoder (highlighted in the green box).Figure 6Schematic diagram of transformer + UNet operation. The Transformer + UNet deep learning model process Input Image: An input image with a size of 448 × 448 pixels is fed into the network. This resolution ensures a balance between detail and computational efficiency.Encoder: The input image passes through the encoder, which consists of multiple self-attention layers and convolutional layers. The encoder compresses the input image into a set of feature maps containing essential information about the image. The size of the encoder's feature maps is generally smaller, depending on the reduction in channels and spatial resolution. For example, if the encoder reduces the spatial resolution by half and doubles the number of channels at each layer, the feature map size might be 224 × 224 × 64.Bottleneck Layer: The feature maps from the encoder then pass through a bottleneck layer, which further compresses the information and reduces the number of channels. The size of the bottleneck layer's feature maps depends on the specific architecture but is typically smaller than the encoder's feature maps. For example, if the bottleneck layer reduces the number of channels by four times, the feature map size might be 112 × 112 × 256.Decoder: The compressed information from the bottleneck layer is passed through the decoder, which consists of multiple convolutional layers and transpose convolutional layers. The decoder uses transpose convolutions to upsample the feature maps, gradually restoring the image's original spatial resolution. For example, if the decoder doubles the spatial resolution at each layer and halves the number of channels, the feature map size might be 224 × 224 × 128.Skip Connections: The UNet architecture includes skip connections, allowing information from the encoder path to be directly passed to the decoder path. These skip connections help preserve spatial information, avoiding the loss of details during the upsampling process. The size of the feature maps in each skip connection typically matches the size of the corresponding encoder feature maps.Segmentation Mask: The decoder's output is a set of feature maps used to generate the segmentation mask. The segmentation mask is a binary image identifying the regions of interest in the input image. The segmentation mask size matches the input image size, which in this example is 448 × 448 pixels.Loss Function: Focal Loss is used as the loss function. It is designed to address the issue of class imbalance, particularly useful for tasks where some classes have relatively few samples. Focal Loss adjusts the weight of hard and easy samples through a parameter γ, reducing the contribution of easy samples and increasing the weight of hard samples, improving the model's performance on minority classes. The formula is as 2\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {\text{FL}}({\text{p}}_{{\text{t}}} ) = - {\upalpha }_{{\text{t}}} {\text{(1 - p}}_{{\text{t}}} {)}^{\gamma } {\text{log}}({\text{p}}_{{\text{t}}} ) $$\end{document}FL(pt)=-αt(1 - pt)γlog(pt)where pt represents the probability of the model classifying the sample, α~t~ represents the weight coefficient of each class, and γ is an adjustment parameter.The loss function adjusts the weight of hard and easy samples through γ, reducing the contribution of easy samples and increasing the contribution of hard samples. This allows Focal Loss to focus more on difficult-to-classify samples, improving the model's performance on minority classes.Output: The final output of the network is the segmentation mask, which can be visualized as an image or used for further analysis. The output size matches the input image size, which is 448 × 448 pixels. Our proposed Transformer and UNet hybrid network combines the Transformer’s ability to capture global contextual information in rock images with the UNet’s capability of restoring lost spatial details through upsampling and convolution blocks. This integration of multi-scale encoder features and skip connections at matching resolutions allows the transfer of fine-grained local information to the decoder. This multi-resolution representation capability enables the model to produce highly accurate segmentation masks. ### Weathering degree determination base on ResNet18 Due to the requirement for image classification methods to directly output the weathering degree category represented by each image (e.g., unweathered, slightly weathered, moderately weathered, and highly weathered), this study adopts an image classification approach for the weathering degree determination. Dr. Kaiming He provided various depth ResNet models in his 2016 paper, such as ResNet18, ResNet34, ResNet50, ResNet101, and ResNet152. Unlike UNet-based segmentation algorithms, ResNet networks extract image features in a hierarchical manner. For rock images, the complex textures and details can be effectively captured at the shallower layers, while the deeper layers can extract more abstract features, such as macroscopic weathering patterns. This hierarchical feature extraction helps to comprehensively analyze the weathering conditions on the rock surface. Figure 7 illustrates the ResNet-18 network architecture and its process in determining weathering degrees.Figure 7ResNet-18 network architecture and weathering degree determination process. By using the ResNet-18 model, we leverage its hierarchical feature extraction capability to accurately determine the weathering degree of the tunnel face surrounding rock. The model's architecture allows for efficient learning and representation of both detailed and abstract features, providing a robust solution for weathering degree classification. The outputs from this model can be visualized and further analyzed to support engineering decisions in tunnel construction, ensuring safety and reliability. In summary, integrating the ResNet-18 model for weathering degree determination complements the lithology segmentation model, forming a comprehensive framework for analyzing tunnel face images. This dual-model approach enhances the overall accuracy and efficiency of rock strength assessment in tunnel construction. ## Model results comparative analysis ### Lithology segmentation implementation of rock images #### Model training strategy The Transformer + UNet model was executed on a computer equipped with an Intel(R) Core(TM) i7-10,700 CPU @ 2.90GHz processor and an NVIDIA 2060 graphics card to ensure efficient training and evaluation. We used the PyTorch deep learning framework for experiment management and reproducibility. The batch size was set to 4, the optimization method used was stochastic gradient descent (SGD), with a minimum learning rate of 0.01 and a momentum of 0.9. #### Model training and comparative analysis results During the model training process, both the training loss and validation loss gradually decreased over 500 epochs, as shown in Fig. 8. The smoothed training loss and validation loss displayed similar trends, gradually decreasing and stabilizing around 450–500 epochs. This indicates that the model is learning and improving its ability to make accurate predictions.Figure 8Training loss and validation loss over epochs. Training loss measures the error on the training data, while validation loss evaluates the model's performance error on an independent dataset. The significant decrease in both loss values indicates the model's excellent generalization capability, effectively handling new data without overfitting. The steady decline in both training and validation loss with increasing data volume further demonstrates the model's strong ability to improve prediction accuracy under data-driven conditions. Observing the smoothly declining curves of training and validation loss, it is evident that the model performance is steadily improving and stable. The results fully demonstrate that our model is well-suited for this task, making accurate and reliable predictions on both training and validation data. As shown in Fig. 9, our model achieved excellent performance on the test set, with high evaluation a Dice Similarity Coefficient (DSC) of 95.43%, a mean Intersection over Union (mIoU) of 91.29%, an MPa (mean pixel accuracy) of 95.57%, and precision and recall both stabilized at 95.31% and 95.57%. This comprehensive evaluation result fully verifies the model's excellent capability in fine-grained geological structure and defect segmentation and its strong adaptability to various complex field conditions.Figure 9Segmentation results for each metric by transformer + UNet. To better understand the progress of the algorithm, additional models were trained for comparison, including DeeplabV3, DeeplabV3 + , FPN, Linknet, PAN, Pspnet, and UNet + + alongside Transformer + UNet. The training strategies for all models were kept consistent to ensure a fair comparison. Specifically, all models were trained on the same dataset, using identical batch sizes, learning rates, and optimization methods. This consistency allows for a direct comparison of the models' performance in image segmentation evaluation. Table 2 and Fig. 9 provide a comprehensive comparison of different models in this context, which is a critical process in computer vision. Table 2Comparison of results for each model in different evaluations.NoModelsDice (%)mIoU (%)MPa (%)mRecall (%)mPrecison (%)1Deeplabv383.8381.5581.5584.4082.782Deeplabv3plus84.2883.8483.8480.4076.563FPN72.4070.5970.5973.8871.054Linknet68.0876.3576.3579.1676.015PAN83.9785.9985.9984.4786.016pspnet71.8075.4575.4585.1084.297UNet + + 82.6382.3482.3480.4580.458Transform + UNet91.2995.5795.5795.3195.43 The models were trained on the same batch database and evaluated using several metrics. The models listed in the table are arranged in descending order of their Dice coefficient performance. The best-performing model is Transformer + UNet, with a Dice score of 95.43%, mIoU of 91.29%, MPa of 95.57%, mRecall of 95.57%, and mPrecision of 95.31%. This model combines the architectures of Transformer and UNet, enabling it to effectively capture spatial and contextual information. The "PAN" model is the second-best performer with a score of 86.01%, and "DeeplabV3" is the third-best performer with a score of 82.78%. The results show that the Transformer + UNet model's success rate is as high as 95.57%, surpassing other popular models such as DeepLabV3, DeepLabV3 + , FPN, Linknet, PSPNet, PAN, and UNet + + . This highlights the proposed model's effectiveness in accurately detecting tunnel face lithology. ### Weathering degree determination implementation of rock images #### Model training details In this study, we used PyTorch as the primary deep learning framework to investigate the identification of rock weathering degrees, selecting high-performance hardware and software configurations to ensure model training efficiency and performance. The experimental environment is detailed as we used Python 3.8 and Windows 11 OS with the PyTorch framework. Hardware includes a 13th Gen Intel(R) Core(TM) i5-13600KF @ 3.50 GHz processor, 63.8GB of RAM, and an NVIDIA GeForce RTX 4070Ti GPU with 43.9GB of VRAM. The CUDA version is 12.1, and cuDNN version is 8.9.6. For model training and optimization, we set 50 epochs, a learning rate of 0.05, weight decay of 5e-4, momentum of 0.9, and used stochastic gradient descent (SGD) as the optimizer. This configuration provides a stable and efficient platform for smooth deep learning model training. Additionally, we optimized the ResNet-18 model (ResNet-18opt) by setting the learning rate to 0.1 and employing a cosine annealing method, which adjusts the learning rate according to the cosine decay schedule, dynamically reducing it to prevent overfitting. This method updates the learning rate according to the decay cycle of a cosine wave, decreasing from the maximum value to the minimum value in the first half of the cycle and increasing from the minimum value to the maximum value in the second half. #### Model training and comparative analysis results In this study, we constructed and trained ResNet series models, DenseNet-121, and Inception ResNetV2 models within the PyTorch environment. The necessary components for the ResNet models, such as conv2d, BatchNorm2d, and ReLU, are provided by the torch.nn library. The image datasets were input into the ResNet models, trained with pre-set hyperparameters, and monitored using TensorBoard. Additionally, we optimized the ResNet-18 model by setting the learning rate to 0.1 and employing a cosine annealing method for dynamic learning rate adjustment. This method updates the learning rate according to the decay cycle of a cosine wave, decreasing from the maximum value to the minimum value in the first half of the cycle and increasing from the minimum value to the maximum value in the second half. Figure 10 and Table 3 present a comparison of the training results of the optimized ResNet-18 opt model with the ResNet series models, DenseNet-121, and Inception ResNetV2 models.Figure 10Training results comparison.Table 3Training results of networks comparison.NetworkBest training set accuracy (%)Best test set accuracy (%)Minimum cross-entropyResNet-1895.1388.530.049ResNet-3494.3487.090.113ResNet-5072.8971.120.296ResNet-10181.9476.510.247ResNet-15272.2066.630.474ResNet-18-opt96.1395.950.045DenseNet-12175.4367.540.173Inception ResNetV295.9395.900.076 Figure 10a shows the best training set accuracy, indicating that ResNet-18-opt performed significantly better than other models. Figure 10b displays the accuracy variation on the training set during training, revealing a fluctuating upward trend typical of deep learning network training. Figure 10c presents the accuracy variation on the test set, showing that ResNet-18-opt performed the best on the validation set when all model hyperparameters remained constant. Figure 10d reflects the cross-entropy changes, indicating ResNet-18-opt superior performance in the task of determining rock weathering degrees. Overall, the ResNet-18 model is the optimal model. ## Case study and evaluation of method effectiveness ### Project introduction The data for this project comes from the construction of a highway tunnel project in Georgia. The Ubisa-Shorapani (F3) project route on the E60 highway in Georgia has a total length of 13.04 km, with a design speed of 100 km/h. The project features a two-way, four-lane cement concrete pavement, with a road width of 27.6 m and a lane width of 3.75 m, totaling 7.5 m for one-way lanes. This case study was applied to three tunnels in this construction area, where the geological conditions are complex. According to the geological survey report during the tunnel design phase, the primary rock types in this area are porphyrite and granite, with the strength of these rocks identified as 125 MPa and 175 MPa, respectively. Therefore, this experiment was conducted based on database data of porphyrite and granite, along with some on-site image data. ### Model integration and on-site application In the multimodal method proposed in this paper, we adopted an integrated module design to achieve lithology identification and weathering degree determination in the rock strength assessment process. Users can select lithology identification result files and weathering degree identification result files through a graphical user interface (GUI) to ensure the accuracy and diversity of the input data. When the user clicks the "Calculate" button, the program loads the file contents and performs a comprehensive analysis. The system calls pre-trained neural network models to process the input image data, identifying rock types and weathering degrees, and calculates the corrected rock strength values by combining these results. The calculation results are immediately displayed in the GUI window, including specific categories of lithology identification, weathering degree classification, and corrected rock strength values. Users can clearly view and analyze the result data through the interface, facilitating construction decision-making and safety assessments. This integrated method, combined with the Python PySimpleGUI library, provides a simple and user-friendly interface, allowing users to complete complex rock strength assessments without modifying the code. Figure 11 shows the real-time display of Tunnel 1 in the GUI window. This method not only improves operational convenience but also provides solid technical support for on-site rock strength evaluation practices through the integration of multimodal data. An additional aspect worth considering is the cost–benefit analysis of the proposed AI-based method compared to traditional rock strength assessment techniques. Traditional methods, such as uniaxial compressive strength (UCS) tests and velocity tests, often require extensive field sampling and laboratory testing, which can be time-consuming and expensive. In contrast, the AI-based approach significantly reduces the time and cost associated with data collection and analysis by leveraging automated image processing and neural networks. While the initial setup of AI models and the acquisition of high-quality image data may involve higher upfront costs, the long-term benefits, including improved accuracy, real-time assessment capabilities, and reduced labor costs, can result in substantial economic advantages. A detailed cost–benefit analysis would provide valuable insights into the financial implications of adopting AI-based methods, highlighting the potential for cost savings and efficiency gains in geological engineering projects.Lithology Segmentation Results: The lithology of the current tunnel face was identified, with the following Granite (red): 53.83%, area 2.67, strength 125 MPaFault zone (green): 23.51%, area 1.18, strength 0 MPaLoess clay (yellow): 8.62%, area 0.431, strength 0 MPaPorphyrite (blue): 7.46%, area 0.373, strength 175 MPaBackground (black): 6.58%, area 0.329, strength 0 MPaStandard Strength Values: Subsequently, the weathering identification results were obtained, indicating slightly weathered and moderately weathered conditions.Strength Calculation: The strength value of the lithology with the largest area proportion (Granite, 175 MPa) was selected and corrected using the largest weathering degree (moderately weathered, correction factor 0.6). The final rock strength value of the tunnel face was calculated \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {\text{Corrected}}\;{\text{Strength}} = {\text{Granite\;Strength}} \times {\text{Weathering}}\;{\text{Correction}}\;{\text{Factor}} = 125\;{\text{MPa}} \times 0.6 = 105\;{\text{MPa}} $$\end{document}CorrectedStrength=Granite\;Strength×WeatheringCorrectionFactor=125MPa×0.6=105MPaFigure 11Example of RC value acquisition. As shown in Fig. 12, the on-site engineering team conducted laboratory tests on rock samples collected from the field. The laboratory test results were compared with the RC values predicted using the correction factor method. The test results showed that the actual compressive strength of the porphyrite was 80 MPa, which is very close to the predicted RC value of 75 MPa, indicating that the correction factor method has high reliability in predicting rock compressive strength.Figure 12Field sampling and laboratory images. Using this method, we conducted experimental comparisons on continuous tunnel face surrounding rock data for 5 groups in each of the three tunnels in the project. Figure 13 shows the first tunnel face surrounding rock images and image processing results for Tunnel 2 and Tunnel 3. Through image processing, the structure and potential weaknesses of the surrounding rock can be clearly identified, providing important reference information for subsequent construction.Figure 13Tunnel face surrounding rock images and image processing results. The results, shown in Table 4, indicate that the average error values for Tunnel 1, Tunnel 2, and Tunnel 3 are 9.772%, 8.844%, and 9.58%, respectively. The overall average error across all tunnels is 9.33%. These results demonstrate a high consistency between laboratory testing and on-site identification methods in determining the strength of the tunnel face surrounding rock, validating the effectiveness of the correction factor method. This method allows the construction team to more accurately assess the strength of the surrounding rock, thereby optimizing construction plans and improving construction safety and efficiency. Table 4Continuous tunnel face strength experimental comparison results for three tunnels.Tunnel 1Tunnel 2Tunnel 3NoLaboratory test strength (MPa)On-site identified strength (MPa)Error (%)NoLaboratory Test Strength (MPa)On-site identified strength (MPa)Error (%)NoLaboratory test strength (MPa)On-site identified strength (MPa)Error (%)1114.51059.071128.51408.211109.51009.502115.310015.302190.41758.802134.21257.36391.51008.503154.317511.833113.71259.844135.01258.004130.21407.00468.0759.335151.21407.995113.81058.38582.4759.87Average 9.772Average 8.844Average 9.58Overall Average Error9.33 This case application demonstrates the feasibility and accuracy of integrating rock type identification, weathering degree assessment, and correction factor application in practical engineering. The method not only enhances the precision of rock strength prediction but also provides a reliable scientific basis for tunnel construction design and support structure selection, thereby improving the safety and economy of the project. Additionally, this case highlights the advantages of combining modern neural network technology with traditional geotechnical engineering knowledge, showcasing the importance of technological innovation in engineering practice. Despite its advantages, the proposed method may face limitations in different tunnel construction environments. Varying geological conditions, diverse rock types, and environmental factors can affect its generalizability. Unusual mineral compositions or highly heterogeneous rock structures might challenge accurate image segmentation and classification. Additionally, input image quality, influenced by lighting, dust, or water presence, can impact performance. To address this, our rock database includes a wide range of rock types to enhance adaptability. Continuous updates and further validation in diverse environments are essential to ensure robust performance. ## Conclusion This study explores the application of various neural network models for assessing tunnel face rock strength during construction. By integrating a Transformer + UNet hybrid model and a ResNet18 model, we developed an innovative multi-neural network-based rock strength assessment method using an image dataset. The conclusions drawn are as Through a statistical analysis of the literature on the impact of increased weathering on UCS, the attenuation coefficients for UCS reduction due to increased weathering were Fresh: Fw ≈ 1.0; Slightly Weathered: Fw ≈ 0.8; Moderately Weathered: Fw ≈ 0.6; Highly Weathered: Fw ≈ 0.3; Completely Weathered: Fw ≈ 0.In terms of lithology image segmentation and identification, the Transformer + UNet model performed excellently in both training and validation loss, gradually decreasing and stabilizing. On the test set, the model also achieved significant results, with a Dice similarity coefficient of 95.43%, mIoU of 91.29%, mean pixel accuracy of 95.57%, and precision and recall rates of 95.31% and 95.57%, respectively. These results indicate that the Transformer + UNet model has strong adaptability and accuracy in fine-grained geological structure and defect segmentation.In the determination of rock weathering degrees, the ResNet18-opt model also performed excellently, with training and test set accuracies of 96.13% and 95.95%, respectively, and the lowest cross-entropy loss reaching 0.045. This fully demonstrates the predictive accuracy and robustness of the ResNet18-opt model under complex geological conditions.By processing and analyzing tunnel face images, we were able to clearly identify rock types and their weathering degrees and calculate the corrected rock strength values accordingly. Compared to on-site rock strength experimental results, the predicted rock strength values from our method had an error of only 9.33% compared to laboratory test values, fully validating the feasibility and accuracy of this method. In summary, the rock strength assessment method based on Transformer + UNet and ResNet18-opt proposed in this study significantly improves assessment accuracy and efficiency. By analyzing construction site image data in real-time, the neural network system can promptly detect potential geological hazards and issue warnings. Additionally, this method demonstrates superior performance in data analysis and optimization, helping to determine the best construction parameters and procedures, thereby enhancing overall construction efficiency and quality. The approach holds the potential to be generalized to other geological settings and construction projects, offering a robust framework for diverse engineering applications. Combining modern neural network technology with traditional geotechnical engineering knowledge improves rock strength prediction accuracy and provides a reliable scientific basis for tunnel construction design and support structure selection, thereby enhancing the safety and economy of engineering projects.