Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Mode-Aware Hybrid Machine-Learning Framework for Full-Field Warpage Prediction of Fan-Out Panel-Level Packaging After Debonding

1
Department of Power Mechanical Engineering, National Tsing Hua University, Hsinchu 300044, Taiwan
2
College of Semiconductor Research, National Tsing Hua University, Hsinchu 300044, Taiwan
*
Author to whom correspondence should be addressed.
Materials 2026, 19(16), 3500; https://doi.org/10.3390/ma19163500
Submission received: 12 July 2026 / Revised: 7 August 2026 / Accepted: 17 August 2026 / Published: 18 August 2026
(This article belongs to the Special Issue Advances in Modeling and Analysis of Materials Processing)

Abstract

Fan-Out Panel-Level Packaging (FO-PLP) enables high area utilization and manufacturing efficiency, but process-induced warpage caused by the coefficient of thermal expansion (CTE) mismatch and polymer shrinkage remains a major challenge. This study presents a classifier-gated hybrid machine-learning framework for the rapid and accurate FO-PLP warpage prediction using a database generated from a validated three-dimensional finite element process model. A Random Forest classifier first estimates the probability of each global warpage mode, while cluster analysis reduces the spatial training dataset. Two mode-specific artificial neural networks are then combined through probability-weighted fusion to predict the full warpage field and enable warpage prediction for previously unseen geometry layouts. The framework was evaluated on 16 independent finite element designs spanning both warpage modes. Compared with an equivalent single-network model, the proposed approach consistently achieved lower mean and maximum prediction errors across all designs, with the greatest improvements at the panel edges and corners where the prediction is most challenging. In addition, the clustering strategy reduced the training-set size and computational cost. These results demonstrate that integrating warpage-mode classification with mode-specific learning improves both the prediction accuracy and training efficiency, providing a practical tool for the fast warpage assessment of new FO-PLP layout designs.

1. Introduction

The Fan-Out Panel-Level Package (FO-PLP) discussed in this study, compared to wafer-level packaging (WLP), offers a large-area format with a high utilization efficiency and a broader distribution of I/O connections. However, due to its larger footprint, panel-level packaging is more susceptible to significant warpage during manufacturing processes [1].
FO-PLP can be implemented using the Die-first face-down, Die-first face-up, or Die-last process flows [2]. This study focuses on the Die-first face-down route, with compression molding and post-mold curing at 175 °C, followed by laser debonding at 185 °C. Although panel-level processing improves the substrate utilization and reduces the packaging cost compared with wafer-level packaging [2,3], the larger panel size amplifies the out-of-plane deformation caused by a CTE mismatch, increasing the challenges in handling and lithographic alignment. Warpage is primarily driven by EMC chemical shrinkage during curing and the CTE mismatch among the EMC, dies, and glass carrier during thermal processing [1,4]. Consequently, an accurate warpage prediction requires a rheological representation of the EMC behavior.
Thermosetting polymer materials exhibit complex nonlinear behavior during heating and cooling cycles. Therefore, in this study, a nonlinear rheological model is adopted to describe the mechanical properties and material characteristics of such materials [5,6].
The FEM is frequently used to simulate and analyze engineering problems [7,8,9,10,11,12]. However, 3D FEM simulations are often time-consuming, and the results may be affected by modeling assumptions, including the mesh strategy, interface and boundary condition definitions, and material parameter selection. To address these limitations, this study introduces a machine-learning approach to predict the post-process warpage behavior of FO-PLP. The validated FEM is employed to construct both training and testing datasets, and machine-learning techniques are then applied to analyze correlations within the data, ultimately predicting the overall warpage distribution of the package. While ensemble learning frameworks are robust for predicting the electronic packaging reliability [13,14,15,16,17,18,19,20], accurately capturing the complex warpage of FO-PLP remains a challenge. Previous studies from our laboratory have explored warpage mechanisms using equivalent CTE methods [4], but these often simplify the nonlinear material behavior during processing.
The advantages of machine learning lie in its high prediction accuracy and strong computational capability. Different engineering problems require different algorithms, and the choice of algorithm should be based on the data’s characteristics. Since panel-level packaging warpage can exhibit multiple modes and the warpage values in adjacent regions are often very close, this study designs a modeling approach tailored to these characteristics.
First, a Random Forest classifier is employed to determine the probability of each warpage mode. Next, clustering analysis, an unsupervised learning technique, is used to select representative points from the dataset as training samples, thereby reducing the computational cost of model training. Finally, an ANN is applied to predict panel-level warpage values, with warpage mode probabilities incorporated into an ensemble learning framework to enhance the prediction accuracy further. Some basic and primary concepts of this framework have been shown in our previous works [13,19,21,22,23].

2. Fundamental Theory

2.1. Nonlinear Rheological Viscoelastic–Plastic Model

In this study, the nonlinear viscoelastic–plastic rheological material model [24,25,26] is employed to simulate the warpage behavior of thermosetting polymers subjected to multiple thermal cycles. This model is also known as Burger’s rheological model.
The nonlinear rheological model is a mathematical representation composed of serially connected components, including the Maxwell model [27] and the Kelvin–Voigt model [28], as illustrated in Figure 1. Each element within the model serves a distinct function: a single spring represents the elastic behavior of the material; the parallel combination of a spring and a damper characterizes the time-dependent viscous behavior; and the individual damper describes the plastic deformation following the elastic strain.
Figure 1. Nonlinear rheological model: (a) the Burgers model; and (b) the two elementary elements. Maxwell-element parts are drawn in blue, Kelvin–Voigt-element parts in orange.
In the Burgers model, the total strain comprises the instantaneous elastic deformation from the Maxwell spring, the delayed elastic deformation from the Kelvin–Voigt element, and the irreversible viscous deformation from the dashpot. The model parameters, including the spring constants Em and Ek and the viscosities ηm and ηk, were identified by fitting the model to EMC stress-relaxation data. Because the Burgers viscoelasticity is not directly available in ANSYS (version 2024 R2, ANSYS, Inc., Canonsburg, PA, USA), the fitted relaxation response was converted to an equivalent Prony series through the Laplace-domain transformation and term-by-term matching of the relaxation modulus [29]. The resulting Prony coefficients were then implemented in the finite element model. The Prony series formulation is the standard representation of the linear viscoelastic relaxation behavior in commercial finite element solvers [27,30].

2.2. Random Forest

Random Forest [31,32,33] is an algorithmic framework composed of multiple decision trees. It begins by generating multiple bootstrap subsets of the dataset, each used to train a distinct decision tree. The final prediction is obtained by aggregating the outputs of all trees: classification tasks use majority voting, while regression tasks use the average of the predictions.

2.3. K-Means Clustering

Cluster analysis is an unsupervised learning technique that identifies the inherent patterns and correlations within data without prior knowledge of its structure [34,35,36]. Determining the appropriate number of clusters, K, requires evaluating multiple candidate values [37,38]. The selected K should provide a sufficient representation of the dataset variability while maximizing the computational efficiency through training data reduction.

2.4. Artificial Neural Network

Artificial neural networks (ANNs) are composed of interconnected neurons arranged in the input, hidden, and output layers, as illustrated in Figure 2 [39,40,41]. During training, connection weights are iteratively updated through backpropagation to minimize a loss function representing the prediction error. The predictive performance of an ANN depends strongly on the selection and optimization of key hyperparameters.
Figure 2. Schematic diagram of an artificial neural network.

3. Finite Element Analysis of FO-PLP

In this study, a Fan-Out Panel-Level Packaging (FO-PLP) model was developed using the FEM in ANSYS®. By integrating process modeling technology, the simulation better reflects actual manufacturing conditions.

3.1. Finite Element Model

The material parameters considered in model construction include the CTE, Young’s modulus, Poisson’s ratio, and material density. Among these, the mismatch in CTE between different materials is the primary cause of package warpage, making CTE one of the most critical parameters. The material properties are summarized in Table 1 [22]. In addition, the CTE of the EMC exhibits temperature-dependent behavior, as shown in Table 2 [22].
Table 1. Material parameters [22].
Table 2. Variation of the CTE of EMC with temperature [22].
Due to the complex and variable structure of the EMC, a nonlinear rheological model is adopted in this study to describe its influence on the overall warpage during multiple thermal cycles in the manufacturing process. The parameters of the nonlinear rheological model components are listed in Table 3. These component parameters are then converted into a Prony series for finite element analysis, as shown in Table 4.
Table 3. Constitutive parameters of the rheological model.
Table 4. Parameter of Prony series.
The finite element settings are summarized in Table 5; the element types and mesh density follow the mesh size control study previously reported for FO-PLP panel models [4].
Table 5. Summary of the finite element model.
In this study, the Fan-Out Panel-Level Packaging (FO-PLP) structure with dimensions of 320 mm × 320 mm, as proposed by Su et al. [7], is adopted as the basis for model development. A full-chip is designed in both the x and y directions. Additionally, the feature sizes of each material layer, such as chip thickness, chip size, chip spacing, EMC thickness, and carrier thickness, are defined based on our previous study to construct the finite element model, as illustrated in Figure 3 and Figure 4 [23].
Figure 3. Top view of the FO-PLP model [23].
Figure 4. Side view of the FO-PLP model [23].
The viscoplastic Burgers model and the process-modeling procedure adopted here have been validated against experimental measurements in our previous work, reproducing measured warpage to within 4.5% [22] and 0.28–0.52% [4].

3.2. Boundary Condition and Loading Setting

The constructed model represents an FO-PLP. To prevent rigid-body motion, all degrees of freedom are constrained at the panel’s central node. Gravitational effects are also considered during the process simulation, and, thus, gravity is applied to the model.
After the Compression Molding and PMC processes, the simulation proceeds to the debonding stage. In practical manufacturing, laser debonding is conducted in a vacuum chamber, with the parts released instantaneously. To replicate this behavior, the present study simulates the vacuum adsorption and release mechanisms by incorporating multiple spring elements.
These springs are created by connecting the upper nodes of the EMC layer to the top nodes of the spring elements. The spring element type used is COMBIN14, which is massless and, therefore, does not require gravity consideration. The configuration of the spring connections is illustrated in Figure 5.
Figure 5. Illustration of the connection between spring elements and the FO-PLP model: ① the upper surface of the EMC; ② the spring elements; ③ the vacuum chuck; and ④ the panel cross-section.
Before the simulation begins, the spring elements are deactivated. During debonding, these elements are activated to emulate the release behavior.
The process simulation procedure is summarized in Table 6. Initially, the entire model is constructed, including the glass carrier, chips, EMC, and spring elements. In the early stages of the simulation, only the glass carrier and chips are active, while the EMC and spring elements are deactivated using element death techniques.
Table 6. Process modeling technology.
The simulation begins by applying a temperature load that ramps from room temperature to the molding temperature. Once the molding temperature is reached, the EMC elements are reactivated to simulate the molding process. The temperature is held at the molding temperature for a specified time, then cooled back to room temperature.
The simulation then proceeds to the PMC stage, where the temperature is held constant for a period. After another cooling phase to room temperature, the debonding process begins.
At the start of the debonding stage, the spring elements are reactivated to simulate vacuum adsorption. The spring stiffness is set to a very high value to ensure strong adhesion. During heating to the debonding temperature, care is taken to prevent excessive deformation that may lead to convergence issues in the finite element simulation.
In the vacuum release phase, the carrier is gradually peeled off from both sides to avoid sudden warpage changes that could cause numerical instability. After the debonding is completed, the model is cooled back to room temperature. The spring stiffness is then gradually reduced, and, once sufficiently small, all spring elements are deactivated to simulate the vacuum release behavior.

4. Hybrid Machine-Learning Framework

In this study, a finite element model of FO-PLP was developed, and a comprehensive database was successfully constructed for various feature sizes. A hybrid machine-learning framework is proposed to predict the warpage behavior based on this dataset.
The method begins by using a Random Forest to estimate the probability of each possible warpage mode (e.g., saddle-shaped, and bowl-shaped), enabling the classification of the warpage type. These probabilities are then used as weights to incorporate the likelihood of each mode, thereby improving the subsequent ANN model’s prediction accuracy.
To reduce the data volume and training time, clustering analysis is applied to the spatial warpage field to select representative locations (spatial points) on the panel, which are then used as training data for the ANN. This approach effectively captures critical information while improving the computational efficiency.
By integrating multiple machine-learning techniques, the proposed method leverages the strengths of each to enhance the prediction accuracy and reduce the training time. The overall framework of the proposed model is illustrated in Figure 6.
Figure 6. Framework of the hybrid machine-learning method.

4.1. Database

Several key feature sizes in FO-PLP are closely related to warpage behavior. In this study, five geometric parameters of the package structure are selected: the chip thickness, chip size, chip-to-chip gap, EMC thickness above the chips, and glass carrier thickness, as summarized in Table 7. The geometric features are illustrated in the top and side views of the package shown in Figure 7 and Figure 8, respectively [23].
Figure 7. Top view of the geometric dimensions [23].
Figure 8. Side view of the geometric dimensions [23].
Table 7. Five feature sizes.
After defining the five geometric features, a dataset comprising multiple design samples (each representing a unique combination of geometric parameters) was constructed. For each design sample, spatial warpage data points were extracted using the Full Array Method. The 320 mm × 320 mm panel was divided at 10 mm intervals in both the x and y directions, resulting in 1089 spatial points per design sample, as illustrated in Figure 9 [23]. Each point corresponds to specific X and Y coordinates on the package; the Z coordinate represents the associated warpage value. The training and testing datasets are summarized in Table 8 and Table 9, respectively. Since each design sample contains 1089 spatial points, the training dataset contains 264,627 pointwise warpage entries. Two further designs, Test Sample 1 and Test Sample 2, were used during development to determine the number of clusters and to optimize the model hyperparameters, and are reported separately in Section 5.1 and Section 5.2 as illustrative cases. Sixteen additional designs, disjoint from both the training grid and the two development samples, were held out from every stage of the model selection and are used for all subsequent evaluations.
Figure 9. Full array method [23]. Red markers indicate the 33 × 33 sampling points, spaced 10 mm apart.
Table 8. Training dataset.
Table 9. Testing dataset.

4.2. Warpage Mode

Compared to wafer-level packaging, panel-level packaging offers the advantage of a higher utilization efficiency. However, it is often accompanied by severe warpage during manufacturing. Depending on the degree of warpage and the influence of various geometric features, panel-level packages may exhibit different warpage modes. In the absence of geometric and material uncertainties [42], the warpage behavior tends to follow uniform deformation modes.
The database constructed in this study includes two primary warpage modes: bowl-shaped and saddle-shaped, as illustrated in Figure 10 [21]. Bowl- and saddle-type warpage are the two predominant deformation modes reported for panel-level packages in the literature [8,9].
Figure 10. The two primary warpage modes: (a) bowl-shaped; and (b) saddle-shaped [21].
Bowl-shaped warpage refers to a deformation mode in which all four corners deform in the same direction relative to the center, resulting in a symmetric concave or convex shape. This type accounts for approximately 37% of the dataset.
Saddle-shaped warpage refers to a deformation mode in which opposite corners deform in opposite directions relative to the center. This type accounts for approximately 63% of the dataset. In this study, the warpage mode label was determined based on the sign distribution of the out-of-plane displacement at the four panel corners.
Due to the variations in warpage modes, warpage values along the four edges of the panel can fluctuate significantly, severely impacting the prediction accuracy at the periphery of panel-level packages. To address this issue, this study employs a Random Forest to identify the probability of each warpage mode. These probabilities are then used as effective weights for the ANN, thereby enhancing the overall prediction accuracy.

4.3. Point Selection

Cluster analysis is a data reduction technique that groups data based on shared characteristics. Within each cluster, the data points exhibit similar features. By calculating either the average of the points within a cluster or the distances among them, a representative point can be identified. These representative points capture the underlying characteristics of their respective clusters and, when considered collectively, can reflect the overall warpage behavior of the FO-PLP. This approach reduces the number of training data points required and effectively lowers the computational cost of model training. A schematic of the representative point selection is shown in Figure 11 [23].
Figure 11. Schematic illustration of clustering results [23]. Red markers indicate the points retained after K-means reduction.
Since clustering is an unsupervised learning method, there is no direct ground-truth label for evaluating the clustering accuracy. Therefore, the selection of K is mainly determined by the stability of the prediction performance after integrating clustering with ANN training.

5. Results and Discussion

5.1. Single ANN Model

This study first investigates the performance of a single ANN in warpage prediction. To optimize the network architecture, a comprehensive grid search was conducted, with the hyperparameter search space established based on our prior methodology [23]. The hyperparameter search space is summarized in Table 10, and the selected optimal combination is listed in Table 11. Two sets of test samples were used for the evaluation, as illustrated in Figure 12.
Figure 12. The two test designs used for the detailed illustration: (a) Test Sample 1, bowl-dominant; and (b) Test Sample 2, saddle-dominant.
Table 10. Grid search range of ANN hyperparameters.
Table 11. Single ANN hyperparameters.
The prediction was first performed on Test Sample 1, with the results presented in Table 12 and Figure 13. MAE and MAPE denote the mean absolute and mean absolute percentage error, each averaged over the N = 1089 spatial points of a design; the maximum error and its percentage refer to the single worst point. While the mean testing error was relatively low, the maximum testing error reached 10.69%, corresponding to 265.07 µm. The error surface in Figure 13b shows that the largest deviations occur at the outer regions of the package.
Figure 13. Test Sample 1 predicted by the ungated baseline: (a) predicted warpage field; and (b) absolute error surface. Axes in mm, warpage and error in µm.
Table 12. Prediction performance of the single ANN model on Test Sample 1 and Test Sample 2.
Subsequently, warpage prediction was performed for Test Sample 2. The prediction results for Test Sample 2 are given in Table 12 and Figure 14, in which Figure 14a shows the predicted warpage field and Figure 14b the corresponding error surface [21]. As in the previous case, larger prediction errors are observed in the outer regions of the package.
Figure 14. Test Sample 2 predicted by the ungated baseline: (a) predicted warpage field; and (b) absolute error surface. Axes in mm, warpage and error in µm.
The overall prediction results, along with observations from the training dataset, indicate that using a single ANN for warpage prediction may lead to significant errors, particularly at the outer regions of the FO-PLP. This is primarily because the relationship between geometric parameters and the spatial warpage field differs significantly between the bowl-shaped and saddle-shaped deformation modes. As a result, a single ANN trained on mixed modes tends to average these behaviors, leading to larger prediction deviations, particularly near the panel edges.
To address this issue, this study proposes a warpage prediction approach based on Ensemble Learning. The predicted probabilities of different possible warpage modes, obtained using a Random Forest, are used as weights to integrate the outputs of multiple ANN models, thereby improving the overall prediction accuracy.

5.2. Hybrid Machine Learning

To address the large prediction errors observed in the outer regions of the package, this study introduces a hybrid warpage prediction approach. Prior to predicting the warpage magnitude, a Random Forest is employed to estimate the probability of each warpage mode for a given sample. These probabilities are then used as weights in the subsequent prediction stage to guide the ANN models.
Two separate ANN models are trained using datasets corresponding to different warpage modes. Before ANN training, cluster analysis is applied to reduce the number of data points in the training set, thereby simplifying the dataset and improving the computational efficiency. Each sample is then predicted using both ANN models.
Finally, the outputs of the two ANN models are combined using probability-weighted integration based on the warpage mode probabilities predicted by the Random Forest. Since the predicted probabilities satisfy PB + PS = 1, the final output can be interpreted as the expected warpage field conditioned on the given geometric parameters. The overall workflow of the proposed hybrid model is illustrated in Figure 15. It consists of three stages: the warpage mode probability estimation, data reduction combined with warpage prediction, and weighted result integration.
Figure 15. Workflow of the hybrid machine learning method. Blue for pattern probability, green for reducing points, purple for warpage prediction, and orange for weight distribution.
The first stage of the proposed workflow is the estimation of the warpage mode probability. A Random Forest model is trained using five geometric features. The hyperparameters used in the Random Forest training are listed in Table 13.
Table 13. Hyperparameters of Random Forest.
To validate the classification capability of the framework, the Random Forest classifier was evaluated under stratified five-fold cross-validation. The classifier achieves an accuracy of 95.9%, an area under the ROC curve of 0.995, and an out-of-bag accuracy of 98.8%, as detailed in Table 14. Furthermore, it correctly identified the expected mode in every one of the sixteen independent test designs evaluated in this study.
Table 14. Cross-validated classification performance of the Random Forest over the 243 labeled designs.
After training and validation, the model is used to predict the probability of possible warpage modes for specific test samples. Test Sample 1 is first input into the trained Random Forest model. As shown in Table 15, the predicted probability for the bowl-shaped pattern is 1, while that for the saddle-shaped pattern is 0, which is consistent with the ground truth. Similarly, Test Sample 2 is then input into the model. The predicted probability for the saddle-shaped pattern is 0.97, also matching the true label. Meanwhile, the bowl-shaped pattern is predicted with a probability of 0.03.
Table 15. Predicted warpage-mode probabilities for the two representative test designs.
Compared to directly outputting a single class label, providing classification probabilities preserves the model’s confidence in each class, enabling more nuanced and informative decision-making in the subsequent prediction stage. Upon completing the warpage mode probability prediction, the next steps involve the data reduction and warpage prediction stage, followed by the weighted integration stage.
Since different values of K can affect the clustering quality and data representation, comparing multiple K values provides a more effective reflection of the underlying data characteristics. As noted by Pham et al. [38], when the ratio of K to N (K/N) is too small, the clustering structure may be insufficient; conversely, a large K/N ratio may undermine the purpose of data reduction. In this study, this principle is applied to the spatial warpage field, where N denotes the number of spatial data points per sample (1089 points from the 33 × 33 full array).
Therefore, the K/N ratio is set between 3.33% and 20%, corresponding to K values of approximately 30–200. For each fixed K value, a grid search is conducted to determine the optimal hyperparameters for the corresponding ANN, enabling a comprehensive comparison of the prediction performance across different K values.
The model architecture consists of two distinct neural networks: ANN model 0 and ANN model 1. ANN model 0 is trained using samples with bowl-shaped warpage, while ANN model 1 is trained using samples with saddle-shaped warpage. Accordingly, the impact of K is analyzed and compared independently for each model. The workflow combining clustering analysis with neural network prediction is illustrated in Figure 16.
Figure 16. Cluster analysis combined with the two mode-specific networks: (a) ANN model 0; and (b) ANN model 1. Colors and shapes follow the convention of Figure 15.
ANN model 0 was initially evaluated at K=30, 40, and 50, with the analysis subsequently extended to K=100 and higher; the complete results are summarized in Table 16. The mean testing error stabilized below 6% for K40, while the maximum testing error stabilized for K50, with only marginal changes in the prediction performance at larger K values. These results indicate that the predictive accuracy of ANN model 0 converges as K exceeds 50. Accordingly, K=50 was selected as the minimum viable number of clusters for ANN model 0.
Table 16. Performance of ANN model 0 with varying K values (K = 30 to 100, 150, and 200).
ANN model 1 was evaluated over the same range, first at K = 30, 40, and 50, and then up to K = 100 and beyond; the complete sweep is summarized in Table 17. The mean testing error approaches 3% once K exceeds 50, but both the mean and the maximum errors continue to fluctuate until K = 70, beyond which they remain essentially unchanged. The prediction accuracy of ANN model 1 therefore stabilizes above K = 70, which is taken as the minimum stable K for this model. Selecting the smallest K at which the accuracy has stabilized avoids unnecessary computation without sacrificing the prediction quality.
Table 17. Performance of ANN model 1 with varying K values (K = 30 to 100, 150, and 200).
The results indicate that, once K exceeds a threshold, the prediction accuracy of K-means combined with the ANN stabilizes. Above that threshold, the remaining variation is not systematic. Repeating this evaluation at every K in the grid search gives a mean test error of 5.28 ± 0.39% for ANN model 0 and 3.74 ± 0.19% for ANN model 1 above the minimum stable K, as shown in Table 18. This confirms the reproducibility of the selected operating point, and that further increases in K yield no systematic improvement in accuracy. This phenomenon reflects the diminishing-return regime [43,44,45], which motivates selecting the smallest stable K. The point at which the accuracy first stabilizes is therefore adopted as the minimum stable K, which maximizes the data reduction while keeping the prediction accuracy within an acceptable range and balancing performance against computational efficiency.
Table 18. Dispersion of the test error over independently optimized configurations.
For each model, the selected K value corresponds to its minimum viable value. The clustering hyperparameters used are listed in Table 19. For ANN model 0 and ANN model 1, the minimum viable K values are 50 and 70, respectively. The optimal hyperparameter settings for each ANN model, obtained through a grid search, are presented in Table 20 and Table 21.
Table 19. Hyperparameter of K-means.
Table 20. Optimal hyperparameters of ANN model 0 at K = 50.
Table 21. Optimal hyperparameters of ANN model 1 at K = 70.
By employing this data selection strategy, the training time was significantly reduced from 443 s required by the original single neural network to approximately 42 s in total (19 s for ANN model 0 and 23 s for ANN model 1). The additional computational cost of Random Forest classification and K-means clustering is relatively minor compared to ANN training, achieving an overall reduction of about 90% in training cost. The training time scales linearly with the number of clusters K over the full range tested, as summarized in Table 22 and illustrated in Figure 17. At the selected K, the training set falls from 264,627 to 15,190 pointwise entries, a 94.3% reduction, consistent with the 90.4% reduction in network training time.
Figure 17. Measured training time of the two mode-specific networks against the number of clusters K. The vertical dashed lines represent K = 50 (blue) and K = 70 (red).
Table 22. Training points, iterations, and measured training time versus the number of clusters.
Next, all test samples are individually input into the two artificial neural network models within the ensemble learning framework. The prediction results from the two models are integrated using a weighted calculation, as defined in Equation (1).
The prediction from ANN model 0, trained on bowl-shaped warpage samples, is denoted as W0, while the prediction from ANN model 1, trained on saddle-shaped warpage samples, is denoted as W1. The corresponding Random Forest probabilities are PB and PS, respectively. The final warpage prediction W is obtained as follows:
W = PB × W0 + PS × W1,
where PB + PS = 1.
The results for Test Sample 1 are presented in Table 23. ANN model 0 effectively predicts the warpage due to its training on bowl-shaped warpage data, while ANN model 1, trained exclusively on saddle-shaped warpage, exhibits larger prediction deviations.
Table 23. Performance of the mode-specific networks and of the fused prediction on the two representative test designs.
By applying ensemble learning, the predictions from both models are weighted according to their respective warpage mode probabilities and then summed. Since the Random Forest predicted PB = 1 and PS = 0 for Test Sample 1, the ensemble result is mathematically equivalent to the prediction of ANN model 0 alone. The resulting integrated prediction demonstrates an improved accuracy compared to the single ANN baseline, with the average testing error reduced from 6.45% to 4.77%, and the maximum testing error reduced from 10.69% to 5.76%. The predicted warpage distribution and corresponding surface error map are shown in Figure 18.
Figure 18. Test Sample 1 predicted by the proposed framework: (a) predicted warpage field; and (b) absolute error surface. Axes in mm, warpage and error in µm.
The prediction results for Test Sample 2 are also presented in Table 23. Since ANN model 0 was trained solely on bowl-shaped warpage data, its prediction deviated significantly from the actual values. In contrast, ANN model 1 accurately captured the sample’s warpage trend. After weighted integration, the overall prediction performance improved compared to using a single ANN model. The average testing error was reduced to 3.49%, and the maximum testing error was improved to 3.82%. The warpage prediction and corresponding surface error maps are shown in Figure 19.
Figure 19. Test Sample 2 predicted by the proposed framework: (a) predicted warpage field; and (b) absolute error surface. Axes in mm, warpage and error in µm.
To further evaluate the proposed framework across the design space, 16 held-out designs, including five bowl-dominant and eleven saddle-dominant cases, not represented in the training set, were analyzed. Table 24 presents the classifier probability of the saddle-shaped mode and the corresponding prediction errors of both the ungated baseline and the proposed framework, while Table 25 summarizes the overall performance. As shown in Figure 20, the proposed framework achieved lower mean and maximum absolute errors than the baseline for all 16 designs, yielding average reductions of 38.5% and 31.9%, respectively. Moreover, applying a network trained for the opposite warpage mode increased the mean absolute error by a factor of 5.9, from 30.71±5.24 µm to 181.17±64.49 µm, demonstrating the effectiveness of mode-specific network specialization.
Figure 20. Per-design comparison of the ungated baseline and the proposed framework over the sixteen held-out test designs: (a) mean absolute error; and (b) maximum absolute error.
Table 24. Ungated baseline and proposed framework on sixteen held-out designs.
Table 25. Aggregate performance over the sixteen designs (mean ± standard deviation).
An ablation study compared the probability-weighted fusion against a hard arg-max switch over the eleven non-degenerate designs. As detailed in Table 26, the soft gate gave the lower mean error in seven of eleven cases, by 2.82% on average. This benefit is concentrated near the classifier’s decision boundary rather than in the average accuracy for confidently classified designs, as illustrated in Figure 21. Because using the wrong expert is 5.9 times worse, replacing the hard gate with the classifier’s cross-validated accuracy of 95.9% would raise the expected mean error by 20.1%. Thus, retaining the probabilities buys robustness at a negligible cost where the classification is already confident.
Figure 21. Behavior of the two gating schemes: (a) mean absolute error with the correct versus the wrong mode-specific network; and (b) prediction error as a function of the classifier probability of the true mode. The vertical dashed line represents the decision threshold where the classifier probability equals 0.5.
Table 26. Probability-weighted fusion compared with hard switching over the eleven non-degenerate designs.

6. Conclusions

A classifier-gated hybrid machine-learning framework was developed to predict the full warpage field of Fan-Out Panel-Level packages directly from geometric design parameters. The framework employs a Random Forest classifier to estimate the probabilities of the two dominant warpage modes, bowl-shaped and saddle-shaped, while cluster analysis reduces the spatial training data within each mode to a set of representative points. Two mode-specific artificial neural networks are then trained on the reduced datasets and combined through probability-weighted fusion to generate the complete warpage-field prediction. This mode-aware architecture enables specialization for distinct deformation patterns while substantially reducing the training data requirements and computational cost.
The evaluation on 16 held-out designs spanning the design space showed that the proposed framework consistently outperformed an equivalent ungated single-network model, reducing the mean and maximum absolute prediction errors by 38.5% and 31.9%, respectively. The largest improvements occurred at the panel edges and corners, where the warpage is most severe and strongly influenced by the deformation mode. Moreover, applying a network trained for the incorrect warpage mode increased the mean absolute error by a factor of 5.9, confirming the importance of mode-specific specialization. By reducing the effective training dataset by more than 94%, the representative-point strategy also lowered the network training cost by 90.4% without compromising the prediction accuracy, making the framework substantially more efficient than a conventional single-network approach.
The trained framework provides a practical tool for the rapid warpage prediction of new FO-PLP layouts. Once trained, it predicts the full warpage field directly from the geometric design parameters, eliminating the need for repeated finite element simulations and enabling efficient design verification and optimization. Although the current AI model is limited to a fixed material set, it can be readily retrained using the same methodology when the material properties change. Future work could focus on incorporating material systems, including alternative encapsulation materials, to improve the model generalizability across a broader FO-PLP design space.

Author Contributions

Conceptualization, Y.-T.S. and K.-N.C.; methodology, Y.-T.S. and K.-N.C.; software, M.-C.H. and Y.-T.S.; validation, M.-C.H., Y.-T.S. and K.-N.C.; formal analysis, M.-C.H. and Y.-T.S.; investigation, K.-N.C.; resources, K.-N.C.; data curation, M.-C.H. and Y.-T.S.; writing—original draft preparation, Y.-T.S.; writing—review and editing, K.-N.C.; visualization, Y.-T.S.; supervision, K.-N.C.; project administration, K.-N.C.; funding acquisition, K.-N.C. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by National Tsing Hua University and National, Science and Technology Council, Taiwan, grant number NSTC 114-2221-E-007-025-MY3.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Nomenclature

The following symbols are used in this manuscript:
EmMaxwell spring stiffness[N/mm]
EkKelvin–Voigt spring stiffness[N/mm]
ηmMaxwell dashpot damping coefficient[N·s/mm]
ηkKelvin–Voigt damping coefficient[N·s/mm]
GmProny relative modulus, first term (0.74)[–]
GkProny relative modulus, second term (0.26)[–]
τ1Prony relaxation time, first term (6229.821)[s]
τ2Prony relaxation time, second term (58.303)[s]
PBPredicted probability of the bowl-shaped mode[–]
PSPredicted probability of the saddle-shaped mode[–]
WFused predicted warpage field[µm]
W0Warpage predicted by ANN model 0[µm]
W1Warpage predicted by ANN model 1[µm]
NNumber of spatial points per design (1089)[–]
KNumber of clusters in K-means[–]
CTChip thickness[mm]
CSChip size[mm]
GSGap between dies[mm]
ETExtra molding thickness[mm]
CRTCarrier thickness[mm]

Abbreviations

The following abbreviations are used in this manuscript:
ANNArtificial neural network
CTECoefficient of thermal expansion
EMCEpoxy molding compound
FEM/FEAFinite element method/analysis
FO-PLPFan-Out Panel-Level packaging
MAEMean absolute error
MAPEMean absolute percentage error
MSEMean squared error
PMCPost-mold cure
RDLRedistribution layer
ReLURectified linear unit
RFRandom Forest
WLPWafer-level packaging

References

  1. Braun, T.; Becker, K.F.; Hoelck, O.; Kahle, R.; Wöhrmann, M.; Boettcher, L.; Töpper, M.; Stobbe, L.; Zedel, H.; Aschenbrenner, R.; et al. Panel Level Packaging—A View Along the Process Chain. In Proceedings of the 2018 IEEE 68th Electronic Components and Technology Conference (ECTC), San Diego, CA, USA, 29 May–1 June 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 70–78. [Google Scholar] [CrossRef] [Scilit]
  2. Praful, P.; Bailey, C. Warpage in wafer-level packaging: A review of causes, modelling, and mitigation strategies. Front. Electron. 2025, 5, 1515860. [Google Scholar] [CrossRef] [Scilit]
  3. Lau, J.H. Recent Advances and Trends in Advanced Packaging. IEEE Trans. Compon. Packag. Manuf. Technol. 2022, 12, 228–252. [Google Scholar] [CrossRef] [Scilit]
  4. Tsai, C.H.; Liu, S.W.; Chiang, K.N. Warpage Analysis of Fan-Out Panel-Level Packaging Using Equivalent CTE. IEEE Trans. Device Mater. Reliab. 2020, 20, 51–57. [Google Scholar] [CrossRef]
  5. Perzyna, P. Fundamental Problems in Viscoplasticity. In Advances in Applied Mechanics; Chernyi, G.G., Dryden, H.L., Germain, P., Howarth, L., Olszak, W., Prager, W., Probstein, R.F., Ziegler, H., Eds.; Elsevier: New York, NY, USA, 1966; Volume 9, pp. 243–377. [Google Scholar] [CrossRef] [Scilit]
  6. Perzyna, P. The thermodynamical theory of elasto-viscoplasticity for description of nanocrystalline metals. Eng. Trans. 2010, 58, 15–74. [Google Scholar] [CrossRef]
  7. Su, M.; Cao, L.; Lin, T.; Chen, F.; Li, J.; Chen, C.; Tian, G. Warpage simulation and experimental verification for 320 mm × 320 mm panel level fan-out packaging based on die-first process. Microelectron. Reliab. 2018, 83, 29–38. [Google Scholar] [CrossRef] [Scilit]
  8. Chen, C.; Su, M.; Ma, R.; Zhou, Y.; Li, J.; Cao, L. Investigation of Warpage for Multi-Die Fan-Out Wafer-Level Packaging Process. Materials 2022, 15, 1683. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Wu, M.L.; Lan, J.S. Simulation and Experimental Study of the Warpage of Fan-Out Wafer-Level Packaging: The Effect of the Manufacturing Process and Optimal Design. IEEE Trans. Compon. Packag. Manuf. Technol. 2019, 9, 1396–1405. [Google Scholar] [CrossRef] [Scilit]
  10. Lee, C.C.; Chang, C.P.; Chen, C.Y.; Lee, H.C.; Chen, G.C.F. Warpage Estimation and Demonstration of Panel-Level Fan-Out Packaging with Cu Pillars Applied on a Highly Integrated Architecture. IEEE Trans. Compon. Packag. Manuf. Technol. 2023, 13, 560–569. [Google Scholar] [CrossRef] [Scilit]
  11. Cheng, H.C.; Wu, Z.D.; Liu, Y.C. Viscoelastic Warpage Modeling of Fan-Out Wafer-Level Packaging During Wafer-Level Mold Cure Process. IEEE Trans. Compon. Packag. Manuf. Technol. 2020, 10, 1240–1250. [Google Scholar] [CrossRef] [Scilit]
  12. Che, F.X.; Yamamoto, K.; Rao, V.S.; Sekhar, V.N. Panel Warpage of Fan-Out Panel-Level Packaging Using RDL-First Technology. IEEE Trans. Compon. Packag. Manuf. Technol. 2020, 10, 304–313. [Google Scholar] [CrossRef] [Scilit]
  13. Lau, J.; Chiang, K.N. Cu-Interconnects, Glass, and AI-Assisted Simulation for Chiplets and Heterogeneous Integration; Springer: Singapore, 2026. [Google Scholar] [CrossRef] [Scilit]
  14. Yu, C.F.; Peng, J.W.; Hsiao, C.C.; Wang, C.H.; Lo, W.C. Development of GUI-Driven AI Deep Learning Platform for Predicting Warpage Behavior of Fan-Out Wafer-Level Packaging. Micromachines 2025, 16, 342. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Su, Q.; Chiang, K.N. Multi-Algorithm Ensemble Learning Framework for Predicting the Solder Joint Reliability of Wafer-Level Packaging. Materials 2025, 18, 4074. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Panigrahy, S.K.; Che, F.X.; Ong, Y.C.; Ng, H.W.; Kumar, G. Deep Learning Study on Memory IC Package Warpage Using Deep Neural Network and Finite Element Simulation. Chips 2025, 4, 35. [Google Scholar] [CrossRef] [Scilit]
  17. Prada Parra, D.; Ferreira, G.R.B.; Díaz, J.G.; Gheorghe de Castro Ribeiro, M.; Braga, A.M.B. Supervised Machine Learning Models for Mechanical Properties Prediction in Additively Manufactured Composites. Appl. Sci. 2024, 14, 7009. [Google Scholar] [CrossRef] [Scilit]
  18. Su, Q.; Chiang, K.N. Predicting Wafer-Level Package Reliability Life Using Mixed Supervised and Unsupervised Machine Learning Algorithms. Materials 2022, 15, 3897. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Panigrahy, S.K.; Tseng, Y.C.; Lai, B.R.; Chiang, K.N. An Overview of AI-Assisted Design-on-Simulation Technology for Reliability Life Prediction of Advanced Packaging. Materials 2021, 14, 5342. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  20. Dietterich, T.G. Ensemble Methods in Machine Learning. In Proceedings of the Multiple Classifier Systems, Cagliari, Italy, 21–23 June 2000; IEEE: Piscataway, NJ, USA, 2025; pp. 1–15. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Huang, M.C.; Chen, Y.C.; Chiang, K.N. Hybrid Machine Learning-Based Warpage Prediction for Fan-Out Panel Level Packaging. In Proceedings of the 2025 20th International Microsystems, Packaging, Assembly and Circuits Technology Conference (IMPACT), 21–24 October 2025; IEEE: New York, NY, USA, 2025; pp. 181–184. [Google Scholar] [CrossRef] [Scilit]
  22. Tseng, Y.H.; Huang, M.C.; Chiang, K.N. Estimating Warpage of Fan-Out Panel-Level Packages Under Different Manufacturing Processes Using a Viscoplastic Model. In Proceedings of the 2024 19th International Microsystems, Packaging, Assembly and Circuits Technology Conference (IMPACT), Taipei, Taiwan, 22–25 October 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 154–157. [Google Scholar] [CrossRef] [Scilit]
  23. Huang, M.C.; Cheng, H.C.; Chiang, K.N. Integrating Cluster Analysis and Artificial Neural Networks to Predict Warpage in Fan-out Panel-Level Packaging. In Proceedings of the 2025 26th International Conference on Thermal, Mechanical and Multi-Physics Simulation and Experiments in Microelectronics and Microsystems (EuroSimE), Utrecht, Netherlands, 6–9 April 2025; IEEE: Piscataway, NJ, USA, 2025; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  24. Adibeig, M.R.; Hassanifard, S.; Vakili-Tahami, F. Optimum creep lifetime of Polymethyl Methacrylate (PMMA) tube using rheological creep constitutive models based on experimental data. Polym. Test. 2019, 75, 107–116. [Google Scholar] [CrossRef] [Scilit]
  25. Li, Y.; Kessler, M.R. Creep-resistant behavior of self-reinforcing liquid crystalline epoxy resins. Polymer 2014, 55, 2021–2027. [Google Scholar] [CrossRef] [Scilit]
  26. Majda, P.; Skrodzewicz, J. A modified creep model of epoxy adhesive at ambient temperature. Int. J. Adhes. Adhes. 2009, 29, 396–404. [Google Scholar] [CrossRef] [Scilit]
  27. Roylance, D. Engineering Viscoelasticity; Department of Materials Science and Engineering, Massachusetts Institute of Technology: Cambridge, MA, USA, 2001; pp. 1–37. [Google Scholar]
  28. Voigt, W. Ueber innere Reibung fester Körper, insbesondere der Metalle. Ann. Phys. 1892, 283, 671–693. [Google Scholar] [CrossRef] [Scilit]
  29. Bharadwaj, M.V.; Claramunt, S.; Srinivasan, S. Modeling Creep Relaxation of Polytetrafluorethylene Gaskets for Finite Element Analysis. Int. J. Mater. Mech. Manuf. 2017, 5, 123–126. [Google Scholar] [CrossRef] [Scilit]
  30. Ferry, J.D. Viscoelastic Properties of Polymers, 3rd ed.; John Wiley & Sons: New York, NY, USA, 1980. [Google Scholar]
  31. Salman, H.A.; Kalakech, A.; Steiti, A. Random forest algorithm overview. Babylon. J. Mach. Learn. 2024, 2024, 69–79. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  32. Hsiao, H.Y.; Chiang, K.N. AI-assisted reliability life prediction model for wafer-level packaging using the random forest method. J. Mech. 2021, 37, 28–36. [Google Scholar] [CrossRef] [Scilit]
  33. Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
  34. MacKay, D.J.C. Information Theory, Inference and Learning Algorithms; Cambridge University Press: Cambridge, UK, 2003. [Google Scholar]
  35. MacQueen, J. Some Methods for Classification and Analysis of Multivariate Observations. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability; University of California Press: Berkeley, CA, USA, 1967; pp. 281–297. [Google Scholar]
  36. Lloyd, S. Least squares quantization in PCM. IEEE Trans. Inf. Theory 1982, 28, 129–137. [Google Scholar] [CrossRef] [Scilit]
  37. Jain, A.K. Data clustering: 50 years beyond K-means. Pattern Recognit. Lett. 2010, 31, 651–666. [Google Scholar] [CrossRef] [Scilit]
  38. Pham, D.T.; Dimov, S.S.; Nguyen, C.D. Selection of K in K-means clustering. Proc. Inst. Mech. Eng. Part C J. Mech. Eng. Sci. 2005, 219, 103–119. [Google Scholar] [CrossRef] [Scilit]
  39. McCulloch, W.S.; Pitts, W. A logical calculus of the ideas immanent in nervous activity. Bull. Math. Biol. 1943, 5, 115–133. [Google Scholar] [CrossRef] [Scilit]
  40. Hornik, K.; Stinchcombe, M.; White, H. Multilayer feedforward networks are universal approximators. Neural Netw. 1989, 2, 359–366. [Google Scholar] [CrossRef] [Scilit]
  41. Rumelhart, D.E.; Hinton, G.E.; Williams, R.J. Learning representations by back-propagating errors. Nature 1986, 323, 533–536. [Google Scholar] [CrossRef] [Scilit]
  42. Chen, H.L.; Chiang, K.N. The Effect of Geometric and Material Uncertainty on Debonding Warpage in Fan-Out Panel Level Packaging. In Proceedings of the 2023 24th International Conference on Thermal, Mechanical and Multi-Physics Simulation and Experiments in Microelectronics and Microsystems (EuroSimE), 16–19 April 2023; IEEE: New York, NY, USA, 2023; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  43. Thompson, N.C.; Greenewald, K.; Lee, K.; Manso, G.F. Deep Learning’s Diminishing Returns: The Cost of Improvement is Becoming Unsustainable. IEEE Spectr. 2021, 58, 50–55. [Google Scholar] [CrossRef] [Scilit]
  44. LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. Hastie, T.; Tibshirani, R.; Friedman, J. The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd ed.; Springer: New York, NY, USA, 2009. [Google Scholar] [CrossRef]
Figure 1. Nonlinear rheological model: (a) the Burgers model; and (b) the two elementary elements. Maxwell-element parts are drawn in blue, Kelvin–Voigt-element parts in orange.
Materials 19 03500 g001
Figure 2. Schematic diagram of an artificial neural network.
Materials 19 03500 g002
Figure 3. Top view of the FO-PLP model [23].
Materials 19 03500 g003
Figure 4. Side view of the FO-PLP model [23].
Materials 19 03500 g004
Figure 5. Illustration of the connection between spring elements and the FO-PLP model: ① the upper surface of the EMC; ② the spring elements; ③ the vacuum chuck; and ④ the panel cross-section.
Materials 19 03500 g005
Figure 6. Framework of the hybrid machine-learning method.
Materials 19 03500 g006
Figure 7. Top view of the geometric dimensions [23].
Materials 19 03500 g007
Figure 8. Side view of the geometric dimensions [23].
Materials 19 03500 g008
Figure 9. Full array method [23]. Red markers indicate the 33 × 33 sampling points, spaced 10 mm apart.
Materials 19 03500 g009
Figure 10. The two primary warpage modes: (a) bowl-shaped; and (b) saddle-shaped [21].
Materials 19 03500 g010
Figure 11. Schematic illustration of clustering results [23]. Red markers indicate the points retained after K-means reduction.
Materials 19 03500 g011
Figure 12. The two test designs used for the detailed illustration: (a) Test Sample 1, bowl-dominant; and (b) Test Sample 2, saddle-dominant.
Materials 19 03500 g012
Figure 13. Test Sample 1 predicted by the ungated baseline: (a) predicted warpage field; and (b) absolute error surface. Axes in mm, warpage and error in µm.
Materials 19 03500 g013
Figure 14. Test Sample 2 predicted by the ungated baseline: (a) predicted warpage field; and (b) absolute error surface. Axes in mm, warpage and error in µm.
Materials 19 03500 g014
Figure 15. Workflow of the hybrid machine learning method. Blue for pattern probability, green for reducing points, purple for warpage prediction, and orange for weight distribution.
Materials 19 03500 g015
Figure 16. Cluster analysis combined with the two mode-specific networks: (a) ANN model 0; and (b) ANN model 1. Colors and shapes follow the convention of Figure 15.
Materials 19 03500 g016
Figure 17. Measured training time of the two mode-specific networks against the number of clusters K. The vertical dashed lines represent K = 50 (blue) and K = 70 (red).
Materials 19 03500 g017
Figure 18. Test Sample 1 predicted by the proposed framework: (a) predicted warpage field; and (b) absolute error surface. Axes in mm, warpage and error in µm.
Materials 19 03500 g018
Figure 19. Test Sample 2 predicted by the proposed framework: (a) predicted warpage field; and (b) absolute error surface. Axes in mm, warpage and error in µm.
Materials 19 03500 g019
Figure 20. Per-design comparison of the ungated baseline and the proposed framework over the sixteen held-out test designs: (a) mean absolute error; and (b) maximum absolute error.
Materials 19 03500 g020
Figure 21. Behavior of the two gating schemes: (a) mean absolute error with the correct versus the wrong mode-specific network; and (b) prediction error as a function of the classifier probability of the true mode. The vertical dashed line represents the decision threshold where the classifier probability equals 0.5.
Materials 19 03500 g021
Table 1. Material parameters [22].
MaterialCTE
[ppm/°C]
Young’s Modulus
[GPa]
Poisson’s RatioDensity
[g/cm3]
Stress-Free Temperature [°C]
Si2.81310.282.32925
EMC7.79–35.927.350.351.8175
Glass carrier4.573.60.232.325
Table 2. Variation of the CTE of EMC with temperature [22].
Temperature [K]CTE [ppm/°C]
2987.79
37310.92
Table 3. Constitutive parameters of the rheological model.
ConstantSymbolValueUnit
Maxwell dashpot damping coefficientηm20,506.158N·s/mm
Kelvin–Voigt damping coefficientηk1016.199N·s/mm
Maxwell spring stiffnessEm4.435N/mm
Kelvin–Voigt spring stiffnessEk12.940N/mm
Table 4. Parameter of Prony series.
Creep ModulusValue
τ16229.821
τ258.303
Gm0.74
Gk0.26
Table 5. Summary of the finite element model.
ItemSetting
Software/analysis typeANSYS®; large-deflection static analysis, full Newton–Raphson
Panel dimensions320 × 320 mm; quarter geometry generated and mirrored about both axes
Element type (EMC, die, glass carrier)SOLID185 (enhanced-strain formulation)
Element type (extra-molding thin layer)SHELL181 with defined shell section (EMC, 3 integration points)
Element type (vacuum chuck)COMBIN14 spring elements (massless)
In-plane element size1.5 mm
Through-thickness discretization1 element for the die layer, 1 element for the carrier
Number of equations (DOF)629,136 for the representative design
Approximate element/node count≈1.0 × 105 SOLID185, 5.2 × 104 SHELL181; ≈1.6 × 105 nodes
Equation solverSparse direct solver
Time steppingAutomatic, maximum 100 substeps per load step
Convergence criterionDisplacement-based, tolerance 5%
Solution size182 load steps; 3288 equilibrium iterations
Process modelingElement birth and death
ConstraintAll degrees of freedom fixed at the panel center node
Body loadGravity throughout; sign reversed when the panel is inverted before debonding
Thermal historyMolding 175 °C for 1 h, cooling to 25 °C, post-mold cure 175 °C for 4 h, cooling to 25 °C, debonding at 185 °C
Vacuum releaseProgressive from the panel edges inward, applied in discrete sub-steps
Table 6. Process modeling technology.
StepProcess Description
1Deactivate the EMC and spring elements. Set the initial temperature of the chip and glass carrier to 25 °C.
2Heat the active model to the molding temperature of 175 °C.
3Reactivate the EMC elements at 175 °C to simulate the encapsulation process.
4Cool the entire model down to 25 °C.
5Reheat the model to the post-mold cure (PMC) temperature of 175 °C.
6Cool the model down to 25 °C again.
7Activate the spring elements (for vacuum simulation), invert the model, and heat it to the debonding temperature of 185 °C.
8Gradually deactivate the carrier elements from the edges toward the center, then cool the model to 25 °C.
9Progressively reduce the spring stiffness and deactivate the springs to simulate the final release.
Table 7. Five feature sizes.
No.Feature NameSymbol
1Chip thicknessCT
2Chip sizeCS
3Gap between chipsGS
4Extra molding thicknessET
5Glass carrier thicknessCRT
Table 8. Training dataset.
Feature NameSize [mm]
Chip size (CS)6, 7, 8
Gap between chips (GS)1, 1.4, 2
Chip thickness (CT)0.35, 0.4, 0.45
Glass carrier thickness (CRT)1.1, 1.3, 1.5
Extra molding thickness (ET)0.028, 0.042, 0.056
Total data243
Total data point243 × 33 × 33 = 264,627
Table 9. Testing dataset.
Feature NameTest Sample 1 [mm]Test Sample 2 [mm]
Chip size (CS)6.56.5
Gap between chips (GS)1.251.9
Chip thickness (CT)0.3750.36
Glass carrier thickness (CRT)1.21.2
Extra molding thickness (ET)0.0350.035
Total data11
Total data point1 × 33 × 33 = 10891 × 33 × 33 = 1089
Table 10. Grid search range of ANN hyperparameters.
ANN ModelSetting
PreprocessorStandardScaler
Random stateFixed
Hidden layers4
Neuron100, 150, 200
Activation functionReLU
Loss functionMSE
SolverAdam
Learning rate0.001, 0.0001
Batch size150, 300, 450
Epoch100, 200, 300
Table 11. Single ANN hyperparameters.
ANN ModelSetting
PreprocessorStandardScaler
Random stateFixed
Hidden layers4
Neuron100
Activation functionReLU
Loss functionMSE
SolverAdam
Learning rate0.0001
Batch size300
Epoch300
Table 12. Prediction performance of the single ANN model on Test Sample 1 and Test Sample 2.
MetricTrainingTest Sample 1Test Sample 2
Mean absolute error [µm]7.8665.0435.20
Mean absolute error [%]1.346.455.08
Maximum absolute error [µm]59.65265.07176.48
Maximum absolute error [%]2.6510.695.96
Reference value at maximum error [µm]24802960
Training time [s]443.32
Table 13. Hyperparameters of Random Forest.
Random ForestSetting
Random stateFixed
n_estimators100
max_depthNone
Table 14. Cross-validated classification performance of the Random Forest over the 243 labeled designs.
MetricBowl-ShapedSaddle-ShapedOverall
Support [designs]91152243
Precision0.9660.9550.959
Recall0.9230.9800.959
F1-score0.9440.9680.959
Confusion matrix84/73/149
ROC-AUC0.995
Out-of-bag accuracy0.988
Table 15. Predicted warpage-mode probabilities for the two representative test designs.
Warpage ModeTest Sample 1Test Sample 2
Bowl-shaped, PB1.000.03
Saddle-shaped, PS0.000.97
Table 16. Performance of ANN model 0 with varying K values (K = 30 to 100, 150, and 200).
KMean Test
Error [µm]
Mean Test
Error [%]
Maximum Test Error [µm]Reference Value at Maximum Error [µm]Maximum Test Error [%]
3039.346.35135.7415208.93
4035.985.79125.3916807.46
5035.234.77112.9619605.76
6042.785.66127.3620506.21
7045.665.83144.7924805.84
8044.145.52133.0724805.37
9036.675.29134.1624805.41
10038.724.87121.8520505.94
15039.824.94148.2324805.98
20041.055.35135.3022506.01
Table 17. Performance of ANN model 1 with varying K values (K = 30 to 100, 150, and 200).
KMean Test
Error [µm]
Mean Test
Error [%]
Maximum Test Error [µm]Reference Value at Maximum Error [µm]Maximum Test Error [%]
3030.505.84161.9916809.64
4032.436.21159.2416809.48
5024.164.46113.5716806.76
6029.674.16168.0829605.68
7022.503.49138.6229604.68
8021.803.83116.7129603.94
9026.743.98141.3729604.78
10022.753.91119.3429604.03
15023.453.5894.9729603.21
20022.843.67108.2329603.66
Table 18. Dispersion of the test error over independently optimized configurations.
ModelRegionnMean Error [%]Maximum Error
[%]
Mean Error [µm]Maximum Error [µm]
ANN model 0K < 5026.07 ± 0.408.20 ± 1.0437.66 ± 2.38130.56 ± 7.32
ANN model 0K ≥ 5085.28 ± 0.395.82 ± 0.2940.51 ± 3.61132.22 ± 11.53
ANN model 1K < 7045.17 ± 1.017.89 ± 1.9829.19 ± 3.55150.72 ± 25.04
ANN model 1K ≥ 7063.74 ± 0.194.05 ± 0.6023.35 ± 1.75119.87 ± 17.77
Table 19. Hyperparameter of K-means.
K-MeansSetting
initK-means++
n_cluster50 (ANN model 0)
70 (ANN model 1)
Max_iterDefault (300)
tolDefault (1 × 10−4)
Table 20. Optimal hyperparameters of ANN model 0 at K = 50.
ANN Model 0Setting
PreprocessorStandardScaler
Random stateFixed
Hidden layers4
Neuron150
Activation functionReLU
Loss functionMSE
SolverAdam
Learning rate0.001
Batch size150
Epoch300
Table 21. Optimal hyperparameters of ANN model 1 at K = 70.
ANN Model 1Setting
PreprocessorStandardScaler
Random stateFixed
Hidden layers4
Neuron100
Activation functionReLU
Loss functionMSE
SolverAdam
Learning rate0.001
Batch size300
Epoch300
Table 22. Training points, iterations, and measured training time versus the number of clusters.
KModel 0 PointsModel 0
Iterations
Model 0 Time [s]Model 1 PointsModel 1
Iterations
Model 1 Time [s]
504550930019.06760010,20020.69
100910018,30031.3915,20020,40036.72
15013,65027,30043.3122,80030,60055.12
20018,20036,60056.8330,40040,80070.77
25022,75045,60070.0738,00051,00084.97
30027,30054,60081.0945,60061,200102.66
35031,85063,90093.6453,20071,400120.85
40036,40072,900107.7260,80081,600136.18
45040,95081,900118.1668,40091,800151.89
50045,50091,200132.2976,000102,000169.21
Table 23. Performance of the mode-specific networks and of the fused prediction on the two representative test designs.
QuantityANN Model 0 (K = 50)ANN Model 1 (K = 70)Fused Prediction
Training mean error [µm/%]15.69/2.078.65/1.21
Training max error [µm/%]115.61/5.5687.74/4.22
Training time [s]19.4323.3042.73
Test Sample 1, mean error [µm/%]35.23/4.77158.91/16.6835.23/4.77
Test Sample 1, maximum error [µm/%]112.96/5.76550.25/37.43112.96/5.76
Test Sample 1, reference at maximum error [µm]196014701960
Test Sample 2, mean error [µm/%]344.99/36.2222.50/3.4919.20/3.49
Test Sample 2, maximum error [µm/%]1109.01/43.32138.62/4.68112.99/3.82
Test Sample 2, reference at maximum error [µm]256029602960
Table 24. Ungated baseline and proposed framework on sixteen held-out designs.
DesignModePSSingle MAE [µm]Single Max [µm]Framework MAE [µm]Framework Max [µm]
CT0.375CS6.5GS1.25ET0.035CRT1.4bowl0.0156.14222.4532.94142.15
CT0.375CS6.5GS1.25ET0.049CRT1.2bowl0.0076.73280.2933.44107.75
CT0.375CS6.5GS1.25ET0.049CRT1.4bowl0.0278.69257.7042.63159.04
CT0.425CS6.5GS1.25ET0.049CRT1.2bowl0.0341.18182.9235.15142.56
CT0.425CS6.5GS1.25ET0.049CRT1.4bowl0.0248.04180.0234.08140.81
CT0.375CS6.5GS1.75ET0.035CRT1.2saddle0.9939.59175.0624.85124.45
CT0.375CS6.5GS1.75ET0.035CRT1.4saddle0.9636.82129.6220.10106.43
CT0.375CS6.5GS1.75ET0.049CRT1.2saddle0.9935.70182.5325.09113.34
CT0.375CS6.5GS1.75ET0.049CRT1.4saddle0.9933.15126.9725.78114.71
CT0.375CS7.5GS1.25ET0.049CRT1.2saddle0.9958.01193.2441.63135.41
CT0.375CS7.5GS1.25ET0.049CRT1.4saddle1.0043.67144.4130.30102.96
CT0.375CS7.5GS1.75ET0.049CRT1.2saddle1.0052.98162.0826.3294.74
CT0.375CS7.5GS1.75ET0.049CRT1.4saddle1.0057.40155.9726.7299.33
CT0.425CS6.5GS1.75ET0.049CRT1.2saddle0.9938.02154.9522.96119.79
CT0.425CS7.5GS1.75ET0.049CRT1.2saddle0.9957.14191.3731.88107.89
CT0.425CS7.5GS1.75ET0.049CRT1.4saddle1.0059.22175.7930.57116.49
Table 25. Aggregate performance over the sixteen designs (mean ± standard deviation).
GroupnSingle MAE [µm]Single Max [µm]Framework MAE [µm]Framework Max [µm]Mean Error GainMaximum Error Gain
All designs1650.78 ± 13.82182.21 ± 41.8930.28 ± 6.33120.49 ± 18.5038.5%31.9%
Bowl-dominant560.15 ± 16.89224.68 ± 44.5235.65 ± 3.99138.46 ± 18.7337.5%36.0%
Saddle-dominant1146.52 ± 10.42162.91 ± 22.9827.84 ± 5.73112.32 ± 11.7539.0%30.1%
Table 26. Probability-weighted fusion compared with hard switching over the eleven non-degenerate designs.
GroupnSoft Better on Mean ErrorGain in Mean ErrorSoft Better on Maximum ErrorGain in Maximum Error
All non-degenerate117 (64%)+2.82%6 (55%)+1.72%
Saddle-dominant77 (100%)+6.41%6 (86%)+3.90%
Bowl-dominant40 (0%)−3.45%0 (0%)−2.08%
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Huang, M.-C.; Su, Y.-T.; Chiang, K.-N. A Mode-Aware Hybrid Machine-Learning Framework for Full-Field Warpage Prediction of Fan-Out Panel-Level Packaging After Debonding. Materials 2026, 19, 3500. https://doi.org/10.3390/ma19163500

AMA Style

Huang M-C, Su Y-T, Chiang K-N. A Mode-Aware Hybrid Machine-Learning Framework for Full-Field Warpage Prediction of Fan-Out Panel-Level Packaging After Debonding. Materials. 2026; 19(16):3500. https://doi.org/10.3390/ma19163500

Chicago/Turabian Style

Huang, Ming-Ching, Yu-Ting Su, and Kuo-Ning Chiang. 2026. "A Mode-Aware Hybrid Machine-Learning Framework for Full-Field Warpage Prediction of Fan-Out Panel-Level Packaging After Debonding" Materials 19, no. 16: 3500. https://doi.org/10.3390/ma19163500

APA Style

Huang, M.-C., Su, Y.-T., & Chiang, K.-N. (2026). A Mode-Aware Hybrid Machine-Learning Framework for Full-Field Warpage Prediction of Fan-Out Panel-Level Packaging After Debonding. Materials, 19(16), 3500. https://doi.org/10.3390/ma19163500

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Citations

No citations were found for this article, but you may check on Google Scholar

Article Access Statistics

Created with Highcharts 4.0.4Chart context menuArticle access statisticsArticle Views18. Aug19. Aug20. Aug21. Aug22. Aug23. Aug24. Aug25. Aug050100150200250
For more information on the journal statistics, click here.
Multiple requests from the same IP address are counted as one view.
Back to TopTop