Optimal Biasing Parameter Selection for Shrinkage Estimators in High-Dimensional Multicollinear Systems

Authors and Affiliations

  • Afeez Abolaji Lawal Department of Statistics, Faculty of Science, University of Ibadan
  • Oluwayemisi Oyeronke Alaba Department of Statistics, Faculty of Science, University of Ibadan
  • Ibraheem Sanusi Department of Statistics, Faculty of Science, University of Ibadan

About this article

Download PDF

Keywords:

Cross-Validation; Data-Driven Methods; Machine Learning; Multicollinearity

Abstract

Severe multicollinearity remains a challenge in regression analysis, leading to reduced estimator efficiency. A biased parameter estimator is a popular approach that provides various ‎remedies for the problem, but the performance of this estimator will depend on the selection of ‎the biasing parameter. This paper explores and compares different biasing parameter selection ‎strategies, including one-and two-parameter biased estimators, emphasizing fixed, data-driven, ‎cross-validation, and machine learning-based solutions in different sample sizes and predictor ‎degree relationships. The study used a Monte Carlo simulation that generated ‎data at the conditions of strong multicollinearity (ρx ranging from 0.7, 0.75, 0.8, 0.85, 0.9, 0.95, and ‎‎0.99), varying error variance (σ = 3, 5, 10, 15, and 20), and sample sizes (n = 100, 500, 1000, ‎‎5000, and 10,000). The performance of the estimators was analyzed with MSE, which allows the ‎systematic comparison of different methods for selection of biasing parameters under different ‎types of multicollinearity and sample sizes. The results consistently indicate that the machine ‎learning–based selection provides the lowest MSE in all cases and is increasingly advantageous as ‎increasing multicollinearity is encountered and sample size reduces. Data-driven methods provide ‎competitive outcomes and clearly improve over fixed parameter and cross-validation. Cross-validating ‎methods experience inefficiency under more complex multicollinearity scenarios. To validate the simulation studies, a real-world data setting was applied using the Ames ‎housing dataset. The machine learning approach also produced the best predictive performance, ‎recording the lowest MSE, substantially outperforming the fixed, data-driven, and cross-validation ‎methods in the real-data application. Thus, in the presence of severe ‎multicollinearity, practitioners and researchers should go beyond applying fixed or only cross-‎validation–based rules but use machine learning or data-driven guided approaches‎.

References

[1] Akay, K. U., Ertan, E., and Erkoç, A. (2023). A New biased estimator and variations based on the Kibria Lukman Estimator. Istanbul Journal of Mathematics, 1(2), 74-85. https://doi.org/10.26650/ijmath.2023.00009.

[2] Akhtar, N., and Alharthi, M. F. (2025). Enhancing accuracy in modelling highly multicollinear data using alternative shrinkage parameters for ridge regression methods. Scientific Reports, 15, 10774. https://doi.org/10.1038/s41598-025-94857-7.

[3] Ayinde, K., Alabi, O. O., and Nwosu, U. I. (2021). Solving multicollinearity problem in linear regression model: The review suggests new idea of partitioning and extraction of the explanatory variables. Journal of Mathematics and Statistics Studies, 2(1), 12-20. https://doi.org/10.32996/jmss.2021.2.1.2.

[4] Dawoud, I., Abonazel, M. R., and Awwad, F. A. (2022). Generalized Kibria-Lukman estimator: method, simulation, and application. Frontiers in Ap-plied Mathematics and Statistics, 8, 880086. https://doi.org/10.3389/fams.2022.880086.

[5] De Cock, D. (2011). Ames, Iowa: Alternative to the Boston housing data as an end-of-semester regression project. Journal of Statistics Educa-tion, 19(3). https://doi.org/10.1080/10691898.2011.11889627.

View more references (8)

[6] El-doakly, W. S. H. (2024). A comparative study of some two-parameter ridge-type and Liu-type estimators to combat the multicollinearity problem in regression models: Simulation and application. Journal of Statistics and Econometrics, 54(4), 105–136. https://doi.org/10.21608/jsec.2024.395938.

[7] Hoerl, A. E., and Kennard, R. W. (1970). Ridge regression: Biased estimation for nonorthogonal problems. Technometrics, 12(1), 55–67. https://doi.org/10.1080/00401706.1970.10488634.

[8] Kibria, B. G., and Lukman, A. F. (2020). A new ridge‐type estimator for the linear regression model: simulations and applica-tions. Scientifica, 2020(1). https://doi.org/10.1155/2020/9758378.

[9] Kibria, B. M. G. (2003). Performance of some new ridge regression estimators. Communications in Statistics Simulation and Computation, 32(2), 419–435. https://doi.org/10.1081/SAC-120017499.

[10] Liu, K. (1993). A new class of biased estimate in linear regression. Communications in Statistics—Theory and Methods, 22(2), 393–402. https://doi.org/10.1080/03610929308831027.

[11] Luqman, A. F., Ayinde, K., and Dawoud, I. (2025). Data-driven biasing parameter selection strategies for shrinkage estimators under severe multi-collinearity. Statistical Papers, 66(2), 1123–1145.

[12] Luqman, M., Bhatti, S. H., Aydin, D., and Jamil, M. (2025). Addressing multicollinearity in general linear model: A novel approach for ridge param-eter with performance comparison. PLOS ONE, 20(10), e0335072. https://doi.org/10.1371/journal.pone.0335072.

[13] Ugwuowo, F. I., Oranye, H. E., and Arum, K. C. (2023). On the jackknife Kibria-Lukman estimator for the linear regression model. Communications in Statistics - Simulation and Computation, 52(12), 6116–6128. https://doi.org/10.1080/03610918.2021.2007401.