{"id":39294,"date":"2026-08-11T18:12:23","date_gmt":"2026-08-11T17:12:23","guid":{"rendered":"https:\/\/www.vtei.cz\/?p=39294"},"modified":"2026-08-11T18:12:23","modified_gmt":"2026-08-11T17:12:23","slug":"predicting-effluent-bod5-in-wastewater-treatment-using-machine-learning","status":"publish","type":"post","link":"https:\/\/www.vtei.cz\/en\/2026\/08\/predicting-effluent-bod5-in-wastewater-treatment-using-machine-learning\/","title":{"rendered":"Predicting effluent BOD5 in wastewater treatment using machine learning"},"content":{"rendered":"<h2>ABSTRACT<\/h2>\n<p>Accurate and quick prediction of effluent Biochemical Oxygen Demand (BOD\u2085) is important for maintaining regulatory standards and improv-ing the operation of wastewater treatment (WWT) plants. Traditional laboratory tests for BOD\u2085 take several days, causing delays in assessing water quality and adjusting plant performance. To overcome this issue, this study develops a machine learning model to estimate effluent BOD\u2085 using easily available plant data. A dataset from a large-scale wastewater treatment plant was used, comprising 12 variables from both influent and effluent data, such as pH, Chemical Oxygen Demand (COD), conductivity, Total Suspended Solids (TSS), BOD\u2085, and temperature, with effluent BOD\u2085 as the target variable. The data were cleaned and processed using the Interquartile Range (IQR) method to remove outli-ers, and new features were created to better reflect process behavior. Among several tested models, XGBoost showed the best performance, achieving an R\u00b2 of 0.9695 and a Mean Squared Error (MSE) of 0.05, based on a time-based train\/test split (80 % \/ 20 %). The study demon-strates that machine learning can serve as a powerful predictive tool for BOD\u2085 estimation, significantly reducing the five-day analytical lag of laboratory testing and supporting faster, data-driven decision-making in WWT management.<\/p>\n<h2 class=\"03NADPIS2\">INTRODUCTION<\/h2>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\">The\u00a0proper management and treatment of\u00a0municipal and industrial wastewater are essential to preserve water resources and protect human\u00a0and ecosystem health. Among the\u00a0various parameters used to evaluate effluent quality, the\u00a0Five-Day Biochemical Oxygen Demand (BOD\u2085) is one of\u00a0the\u00a0most widely adopted indicators. It reflects the\u00a0amount of\u00a0oxygen required by microorganisms to biologically decompose organic pollutants in\u00a0water over a\u00a0five-day period [1]. Because BOD\u2085 provides a\u00a0direct measure of\u00a0biodegradable organic load, it serves as a\u00a0critical benchmark for assessing treatment efficiency and environmental impact. Consequently, regulatory authorities across the\u00a0globe have established strict discharge limits to maintain\u00a0acceptable BOD\u2085 levels in\u00a0treated effluents [2].<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">Despite its importance, the\u00a0conventional BOD\u2085 test has significant limitations. The\u00a0standard method requires a\u00a0five-day incubation period, delaying feedback on treatment performance and limiting the\u00a0ability to make timely operational decisions [3]. Moreover, factors such as sample toxicity, nitrification effects, variability in\u00a0microbial seeding, and preservation challenges can\u00a0compromise measurement accuracy and reproducibility [4\u20137]. These drawbacks highlight the\u00a0need for rapid, reliable, and continuous BOD\u2085 estimation methods.<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\" style=\"letter-spacing: 0pt;\">Machine learning (ML) and soft-sensor modeling have emerged as effective tools for rapid BOD\u2085 estimation, enabling predictions from routinely measured operational parameters such as COD, pH, conductivity, temperature, and TSS\u00a0[8,\u00a09]. Unlike empirical or deterministic models [7, 8], which often fail under variable conditions due to the\u00a0complex biochemical and hydrodynamic processes in\u00a0full-scale wastewater treatment plants (WWTPs) [10, 11], ML approaches can\u00a0capture nonlinear interactions and system dynamics, supporting proactive process control, early anomaly detection, and more sustainable plant operation.<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\" style=\"letter-spacing: 0pt;\">Recent studies show that modern machine-learning methods can\u00a0achieve very strong predictive performance for BOD\u2085 and other water-quality parameters. For example, Dantas et al. [12] demonstrated that artificial neural networks (ANN), which are computational models inspired by the\u00a0way biological neurons process information, perform well when forecasting BOD\u2085 and COD in\u00a0full-scale treatment plants. Zhang et al. [13] used a\u00a0Random Forest model, an\u00a0ensemble of\u00a0many decision trees that collectively improve prediction accuracy, to estimate BOD<span class=\"01DOLNIINDEX\">5<\/span> in\u00a0river water, achieving an\u00a0R\u00b2 of\u00a0approximately 0.89. Jafar et al. (2022) [14] employed deep neural networks (DNN), a\u00a0more complex form of\u00a0neural-network architecture capable of\u00a0capturing highly non-linear relationships and reported R\u00b2 values between 0.89 and 0.93. Other studies have shown comparable results: XGBoost (short for Extreme Gradient Boosting) models, which use gradient-boosted decision trees, have reached R\u00b2 values around 0.92 in\u00a0ecological applications [15], while RVFL (Random Vector Functional Link) networks have achieved R\u00b2 \u2248 0.924 in\u00a0WWTP settings [16]. In\u00a0contrast, simpler models often provide lower accuracy. For instance, Ali et al. [17] reported an\u00a0R\u00b2 of\u00a0about 0.88 using CatBoost (Categorical Boosting), which efficiently handles categorical and numerical data with minimal preprocessing. It should be noted that linear regression models are not robust under varying wastewater treatment conditions, as they often fail to capture nonlinear interactions and complex system dynamics [18].<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">Accurate and timely estimation of\u00a0the\u00a0biochemical oxygen demand (BOD<span class=\"01DOLNIINDEX\">5<\/span>) remains a\u00a0persistent challenge in\u00a0WWT, as traditional laboratory analyses require a\u00a0five-day incubation period. This inherent delay limits the\u00a0speed of\u00a0process evaluation and prevents proactive operational control. To address this challenge, this study proposes a\u00a0predictive framework leveraging the\u00a0XGBoost algorithm, selected for its superior ability to handle high-dimensional data and non-linear relationships. By integrating a\u00a0robust data-driven methodology with domain-informed feature engineering, the\u00a0model seeks to accurately capture the\u00a0complex dynamics of\u00a0the\u00a0treatment process under realistic operating conditions. The\u00a0primary objective of\u00a0this work is to provide a\u00a0reliable alternative for rapidly assessing wastewater quality, thereby facilitating faster decision-making and enhancing wastewater management response. The\u00a0following sections detail the\u00a0methodology and model development process used to achieve these objectives.<\/span><\/p>\n<h2 class=\"03NADPIS2\">METHODOLOGY AND MODEL DEVELOPMENT<\/h2>\n<h3 class=\"03NADPIS3\" style=\"margin-top: 0cm;\">Presentation site and data collection, area characteristics<\/h3>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\">The\u00a0dataset used in\u00a0this study was collected from the\u00a0municipal wastewater treatment plant (WWTP) of\u00a0Skikda Province, located in\u00a0north-eastern Algeria (36.89\u00b0N, 7.01\u00b0E). The\u00a0facility operates a\u00a0medium-load activated sludge process designed to handle variable organic loads from domestic and municipal sources. The\u00a0plant has a\u00a0design capacity of\u00a046,000\u00a0m\u00b3\/day and serves approximately 279,700 population equivalents.<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">The\u00a0treatment process includes preliminary screening<span class=\"01BOLD\"><span style=\"font-weight: normal;\">, <\/span><\/span>grit removal<span class=\"01BOLD\"><span style=\"font-weight: normal;\">, <\/span><\/span>primary sedimentation<span class=\"01BOLD\"><span style=\"font-weight: normal;\">, <\/span><\/span>biological oxidation in\u00a0aeration tanks<span class=\"01BOLD\"><span style=\"font-weight: normal;\">, <\/span><\/span>secondary clarification, and final disinfection (see <em><span class=\"01ITALIC\">Fig. 2<\/span><\/em>). The\u00a0system is designed for an\u00a0average influent BOD\u2085 load of\u00a0about 280\u00a0mg\/L, corresponding to 14,950\u00a0kg of\u00a0BOD\u2085 per day, 34,383\u00a0kg of\u00a0COD per day<span class=\"01BOLD\"><span style=\"font-weight: normal;\">, <\/span><\/span>and 17,250\u00a0kg of\u00a0suspended solids (TSS) per day<span class=\"01BOLD\"><span style=\"font-weight: normal;\">.<\/span><\/span><\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">Treated effluent is discharged into the\u00a0Oued Safsaf River, a\u00a0sensitive water body that plays an\u00a0important ecological role in\u00a0the\u00a0region. The\u00a0produced excess sludge is dewatered and reused as an\u00a0agricultural fertilizer, which highlights the\u00a0need for reliable effluent quality control. Incomplete organic removal increases the\u00a0amount of\u00a0residual biodegradable matter, raising the\u00a0moisture content and hindering drying, thereby reducing the\u00a0quality of\u00a0the\u00a0final biosolid product.<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">Samples used in\u00a0this study (2015\u20132024) were digitized from the\u00a0WWTP\u2019s\u00a0historical archives. The\u00a0dataset represents continuous monitoring of\u00a0influent and effluent streams across all 12\u00a0months of\u00a0each year. The\u00a0monitored parameters included biochemical oxygen demand (BOD\u2085)<span class=\"01BOLD\"><span style=\"font-weight: normal;\">, <\/span><\/span>chemical oxygen demand (COD), total suspended solids (TSS), pH, temperature, Electrical Conductivity, and dissolved oxygen (DO). All analyses were performed in\u00a0accordance with the\u00a0Standard Methods for the\u00a0Examination of\u00a0Water and Wastewater (APHA,\u00a02017) [4].<\/span><\/p>\n<h3 class=\"03NADPIS3\">Workflow and modeling strategy<\/h3>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\">This section delineates the\u00a0development of\u00a0a\u00a0high-accuracy model for predicting effluent BOD<span class=\"01DOLNIINDEX\">5<\/span> in\u00a0a\u00a0WWTP. The\u00a0approach uses real operational dataset we applied the\u00a0XGBoost algorithm to optimize performance.<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">Domain-informed feature engineering was applied to enhance model accuracy, and a\u00a0time-based train\/test split (80\u00a0% \/ 20\u00a0%) ensured realistic evaluation. The\u00a0pipeline includes data preprocessing (handling missing values and removing outliers), feature creation, model training, and evaluation using R<span class=\"01HORNIINDEX\">2<\/span>, MSE, and MAE. <em><span class=\"01ITALIC\">Fig. 1<\/span><\/em> illustrates the\u00a0complete workflow, from raw data to final prediction, highlighting a\u00a0data-driven framework for effluent BOD\u2085 estimation.<\/span><\/p>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-1.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39558 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-1.jpg\" alt=\"\" width=\"800\" height=\"125\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-1.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-1-300x47.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-1-768x120.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/125;\" \/><\/a>\n<h6 class=\"05POPISKYobrazku\">Fig. 1. Flowchart of\u00a0the\u00a0modeling pipeline for predicting effluent BOD<span class=\"01DOLNIINDEX\">5<\/span><\/h6>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-2.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39559 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-2.jpg\" alt=\"\" width=\"800\" height=\"479\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-2.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-2-300x180.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-2-768x460.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/479;\" \/><\/a>\n<h6 class=\"05POPISKYobrazku\">Fig. 2. Time series of\u00a0BOD<span class=\"01DOLNIINDEX\">5<\/span> inflow and outflow showing consistent treatment performance despite significant fluctuations<\/h6>\n<h3>Data collection, cleaning, and preprocessing<\/h3>\n<p><span style=\"color: #19a112;\"><strong>Data collection and initial screening<\/strong><\/span><\/p>\n<p>The\u00a0dataset spans nine years (2015\u20132024) and includes 410 valid records, corresponding to an\u00a0average of\u00a0(3\u20134) measurements per month due to irregular sampling frequency. Each record contains 15 operational parameters measured at both the\u00a0influent (inlet) and effluent (outlet) of\u00a0the\u00a0treatment process. These variables were selected based on their relevance to biological treatment performance and the\u00a0availability of\u00a0consistent data.<\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">BOD<span class=\"01DOLNIINDEX\">5<\/span>_out was selected as the\u00a0target variable because it is the\u00a0primary measure of\u00a0wastewater quality. The\u00a0standard 5-day test is considered too slow for active management; however, rapid estimation is facilitated by predictive modeling compared to conventional measurement methods. This allows problems to be caught early and the\u00a0plant to be run more efficiently.<\/span><\/p>\n<p class=\"00TEXTenglish\"><em><span class=\"01ITALIC\"><span lang=\"EN-GB\">Fig. 2<\/span><\/span><\/em><span lang=\"EN-GB\"> illustrates the\u00a0temporal trends for influent and effluent BOD\u2085 concentrations between 2015 and 2025. While influent levels (BOD<span class=\"01DOLNIINDEX\">5<\/span>_int, blue) show significant volatility, peaking during 2017\u20132018 due to seasonal factors or hydraulic surges, the\u00a0effluent levels (BOD<span class=\"01DOLNIINDEX\">5<\/span>_out, green) remain\u00a0stable and low from 2019 onward. This stability confirms consistent treatment efficacy despite variable influent loads. Data gaps in\u00a02016, 2018, and 2022 resulting from sensor or logging issues were managed during the\u00a0preprocessing phase.<\/span><\/p>\n<p><em>Fig. 3<\/em> illustrates the relationship between influent and effluent BOD\u2085 values, highlighting the overall removal performance of the treatment system across a wide range of influent loads. A dense cluster of data points, where influent BOD\u2085 ranges from 80 to 180 mg\/L and effluent concentrations remain between 10 and 47 mg\/L, indicates consistent and effective organic load reduction across oper-ational cycles. This corresponds to an estimated removal efficiency of approximately 74\u201388 %, as illustrated in <em>Fig. 3<\/em> by the region bounded between the lines (y = 0,26x) and (y = 0,12x), which represent 74 % and 88 % removal, respectively. Notably, the majority of effluent BOD\u2085 values remain below the regulatory discharge limit of 35 mg\/L, in accordance with the discharge limits established by Algerian environmental regulations [19]. The observed non-linear distribution reflects the inherent complexity of biological treatment processes, thereby supporting the application of advanced machine learning approaches for performance modeling. As shown in Fig. 3, several isolated data points deviate from the main cluster, exhibiting relatively high effluent BOD\u2085 concentrations despite moderate influent loads. These deviations are not ran-dom but likely correspond to transient operational disturbances, such as temperature fluctuations or hydraulic shocks (sudden increases in flow rate or organic loading), that affect microbial activity and treatment efficiency. They may also reflect measurement errors or tempo-rary equipment malfunctions. Confirming the exact cause of these events is challenging, however, due to the absence of supplementary data such as detailed flow rates, rainfall records, ambient temperature, or on-site operational logs. Despite this limitation, the presence of such anomalies underscores the importance of rigorous data preprocessing, particularly for effluent BOD\u2085. It also highlights the need for adaptive modeling strategies capable of maintaining predictive accuracy under variable and non-ideal operating conditions.<\/p>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-3.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39560 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-3.jpg\" alt=\"\" width=\"800\" height=\"661\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-3.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-3-300x248.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-3-768x635.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/661;\" \/><\/a>\n<h6>Fig. 3. Scatter plot of effluent versus influent BOD\u2085<\/h6>\n<p>The BOD5_int and BOD5_out histograms in <em>Fig. 4<\/em> show right-skewed distributions, with most values clustered around central peaks (e.g. ~200 mg\/L for BOD5_int) and a few high values extending to the right (&gt; 400 mg\/L). These extremes are not necessarily errors; they may reflect real events like industrial discharges or stormwater inflows. However, they require careful handling during preprocessing stage, since they can distort model training and affect its accuracy and credibility. A balanced approach was used: extreme values linked to actual disturb-ances were retained; isolated outliers were removed to ensure data quality.<\/p>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-4.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39561 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-4.jpg\" alt=\"\" width=\"800\" height=\"349\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-4.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-4-300x131.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-4-768x335.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/349;\" \/><\/a>\n<h6>Fig. 4. Distribution of influent and effluent BOD\u2085 values between 2015 and 2024; there is a clustering around central values with right-skewed tails, which indicates the occurrence of rare, high-load events<\/h6>\n<p><span style=\"color: #19a112;\"><strong>Data preprocessing<\/strong><\/span><br \/>\nRaw data exhibited missing values, inconsistencies, and outliers due to measurement and logging issues. A structured preprocessing pipeline was applied to enhance data quality, covering missing data treatment, outlier detection, and temporal alignment.<\/p>\n<p><span style=\"color: #19a112;\"><strong>Handling missing data<\/strong><\/span><br \/>\nThe initial examination of the dataset revealed that missing values were represented by the symbol \u2018\/\u2019, particularly in the dissolved oxygen variables (O2_int and O2_out). More than 62 % of the observations in these columns were missing, likely due to sensor downtime, mainte-nance activities, or intermittent data recording issues, rendering them unsuitable for reliable modeling. Keeping such variables would intro-duce significant bias or require heavy imputation, potentially distorting the true process behavior. Therefore, to preserve data integrity, both variables were removed from the analysis.<br \/>\nAfter replacing all \u2018\/\u2019 entries with NaN and converting the remaining inlet (int) and outlet (out) measurements to numeric format, the other variables (pH, BOD\u2085, COD, TSS, temperature, and the recording date) were found to be complete with no missing values. It is evi-dent that the number of variables has been reduced from 15 (the sampling date, 7 influent variables, and 7 effluent variables, as listed in <em>Tab. 1<\/em>) to 13 following the exclusion of Dissolved Oxygen (DO) from\u00a0both the influent and effluent data. Moreover, no missing values were identified that required imputation. These predictors were therefore used directly in the modeling process.<br \/>\nIt should be emphasized that dissolved oxygen (DO) is a fundamental operational parameter in activated sludge systems, as it directly governs microbial activity and nitrification efficiency [11]. Its exclusion from the present analysis was driven by substantial data loss exceed-ing 60 %, rather than by any diminished relevance of (DO) to the underlying treatment processes.<\/p>\n<h5>Tab. 1. Water treatment process monitoring data<\/h5>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-1-1.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39567 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-1-1.jpg\" alt=\"\" width=\"800\" height=\"750\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-1-1.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-1-1-300x281.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-1-1-768x720.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/750;\" \/><\/a>\n<p><span style=\"color: #19a112;\"><strong>Outlier detection and removal<\/strong><\/span><br \/>\nOutliers in the target variable, effluent BOD5 (BOD5_out), were identified using the Inter-Quartile Range (IQR) method, a robust statistical approach commonly applied in environmental data analysis [20]. The IQR is defined as:<\/p>\n<h5><a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-1.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39620 size-medium lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-1-300x44.jpg\" alt=\"\" width=\"300\" height=\"44\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-1-300x44.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-1-768x111.jpg 768w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-1-780x116.jpg 780w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-1.jpg 800w\" data-sizes=\"(max-width: 300px) 100vw, 300px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 300px; --smush-placeholder-aspect-ratio: 300\/44;\" \/><\/a><\/h5>\n<p>where:<\/p>\n<p>Q\u2081 and\u00a0Q\u2083 represent the\u00a025th and 75th percentiles, respectively.<\/p>\n<p>Observations outside the\u00a0following bounds were considered outliers:<\/p>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-2.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39621 lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-2-300x18.jpg\" alt=\"\" width=\"500\" height=\"31\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-2-300x18.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-2-768x47.jpg 768w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-2-780x49.jpg 780w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-2.jpg 800w\" data-sizes=\"(max-width: 500px) 100vw, 500px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 500px; --smush-placeholder-aspect-ratio: 500\/31;\" \/><\/a>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-3.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39622 lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-3-300x18.jpg\" alt=\"\" width=\"500\" height=\"31\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-3-300x18.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-3-768x47.jpg 768w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-3-780x49.jpg 780w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-3.jpg 800w\" data-sizes=\"(max-width: 500px) 100vw, 500px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 500px; --smush-placeholder-aspect-ratio: 500\/31;\" \/><\/a>\n<p>For BOD5_out, Q<sub>1<\/sub> = 18.18 mg\/L and Q<sub>3<\/sub> = 28.63 mg\/L, yielding an IQR of\u00a010.45\u00a0mg\/L. Accordingly, the\u00a0lower and upper thresholds were calculated to be 2.50\u00a0mg\/L and 44.30\u00a0mg\/L, respectively. Applying this criterion resulted in\u00a0the\u00a0removal of\u00a053 extreme observations, including unusually high values (e.g., 59\u00a0mg\/L) that were likely attributable to measurement errors or occasional process disturbances. This reduced the\u00a0dataset from 410 to 357 samples.<\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">Next, rows containing missing values in\u00a0either the\u00a0target variable or key predictors were excluded. The\u00a0final curated dataset for modeling consisted of\u00a0N\u00a0=\u00a0357 complete samples. As shown in\u00a0<em><span class=\"01ITALIC\">Fig. 5<\/span><\/em>, these preprocessing steps improved the\u00a0symmetry and reduced the\u00a0spread of\u00a0the\u00a0distribution, thereby strengthening the\u00a0robustness and generalizability of\u00a0the\u00a0predictive models.<\/span><\/p>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-5.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39562 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-5.jpg\" alt=\"\" width=\"800\" height=\"475\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-5.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-5-300x178.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-5-768x456.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/475;\" \/><\/a>\n<h6>Fig. 5. Box plot comparison of\u00a0BOD5_out before and after outlier filtering<\/h6>\n<p class=\"03NADPIS4\"><strong><span style=\"text-transform: none; color: #19a112;\">Temporal Alignment of\u00a0Data<\/span><\/strong><\/p>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\">The\u00a0data column was parsed into a\u00a0datetime format and used to sort the\u00a0dataset chronologically. This approach was adopted to ensure the\u00a0preservation of\u00a0temporal dependencies, thereby facilitating the\u00a0validation of\u00a0the\u00a0model\u2019s\u00a0realism.<\/span><\/p>\n<p class=\"03NADPIS4\"><span style=\"color: #19a112;\"><strong><span style=\"text-transform: none;\">Feature expansion and target selection for improved prediction<\/span><\/strong><\/span><\/p>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\">When developing the\u00a0machine learning model, it became apparent that the\u00a0initial dataset contained insufficient features to achieve high predictive accuracy. To address this limitation, we expanded the\u00a0feature space through targeted engineering, carefully selecting the\u00a0most relevant input and output features. This section outlines the\u00a0process of\u00a0generating new features and defining the\u00a0target variable to improve model performance.<\/span><\/p>\n<p class=\"03NADPIS4\"><span style=\"color: #19a112;\"><strong><span style=\"text-transform: none;\">Selection of\u00a0feature and target variable<\/span><\/strong><\/span><\/p>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\">In\u00a0predictive modelling, all available system parameters may be considered as\u00a0candidate inputs, but not all contribute equally to performance. Selecting and engineering informative features is therefore essential, as irrelevant variables can\u00a0reduce accuracy [21]. To improve predictive capacity, the\u00a0effluent outlet parameters were also integrated as inputs, a\u00a0strategy shown to enhance accuracy in\u00a0similar applications.<\/span><\/p>\n<p><em>Fig. 6<\/em> displays the\u00a0Pearson correlation matrix, showing the\u00a0linear relationships among operational variables in\u00a0the\u00a0dataset. Strong positive correlations (red) indicate variables that increase together (e.g., COD_int and BOD5_int), while negative correlations (blue) show opposite trends. As expected, influent BOD\u2085 and COD display a\u00a0strong positive correlation (r = 0.84), and effluent BOD\u2085 and COD are also highly correlated (r = 0.76). Temperature and TSS_int show weak correlations with other variables, suggesting they have little impact on short-term BOD\u2085 changes. Temperature (Temp_int, Temp_out) and suspended solids (TSS_int) exhibit weak correlations with the\u00a0remaining features (<em>Fig. 6<\/em>). However, these variables were retained to capture subtle influences on biological oxygen demand (BOD5) and enhance model precision. This brings the\u00a0total feature count to 13 variables for the\u00a0final model architecture.<\/p>\n<h6><a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-6.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39563 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-6.jpg\" alt=\"\" width=\"800\" height=\"684\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-6.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-6-300x257.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-6-768x657.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/684;\" \/><\/a><\/h6>\n<h6 class=\"05POPISKYobrazku\">Fig. 6. Pearson correlation matrix<\/h6>\n<p><span style=\"color: #19a112;\"><strong>Feature engineering: creating new predictive input<\/strong><\/span><br \/>\nFeature engineering transforms raw data into informative variables that enhance model interpretability, predictive performance, and generalization [20, 21]. Guided by WWT principles, domain-informed features were derived to better capture treatment dynamics. Six additional features were engineered: removal efficiencies, ratios, and rate-of-change indicators. This expanded the initial 13-variable set to a total of 18 variables, excluding the sampling date. Preliminary evaluation confirmed that these domain-specific variables improved model accuracy, particularly for process control tasks. All engineered features are summarized in <em>Tab. 2<\/em>, including their mathematical definitions and physical interpretations. Each feature was validated for physical plausibility and analyzed for its correlation with the target variable, BOD5_out.<\/p>\n<h5 class=\"04TABULKApopisek\"><span class=\"01ITALIC\">Tab.\u00a02. Engineered features and their mathematical descriptions<\/span><\/h5>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-2-1.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39568 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-2-1.jpg\" alt=\"\" width=\"800\" height=\"1031\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-2-1.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-2-1-233x300.jpg 233w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-2-1-795x1024.jpg 795w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-2-1-768x990.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/1031;\" \/><\/a>\n<p class=\"03NADPIS4\"><span style=\"color: #19a112;\"><strong><span style=\"text-transform: none;\">Train\/test split strategy<\/span><\/strong><\/span><\/p>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\">Because the\u00a0dataset spans nearly a\u00a0decade (2015\u20132024) and contains sequential time-stamped measurements, a\u00a0time-based train\/test split was applied to preserve temporal order and prevent data leakage. Unlike random splitting, which risks introducing temporal bias by allowing future data to influence training, this approach chronologically allocated the\u00a0first 80\u00a0% of\u00a0samples (286 observations, January 2015\u2013December 2022) for training\/validation and the\u00a0remaining 20\u00a0% (71 observations, January 2023-May 2024) for testing, thereby simulating real-world deployment where models are trained on historical data evaluated on feature periods [22, 23].<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">This approach evaluates the\u00a0model under realistic forecasting conditions that capture both seasonal variation and long-term operational trends. The\u00a0performance metrics used in\u00a0this study, including R\u00b2 (which indicates how much of\u00a0the\u00a0variability in\u00a0effluent BOD\u2085 the\u00a0model can\u00a0explain), MSE (which highlights the\u00a0presence of\u00a0large prediction errors), and MAE (which reflects the\u00a0average deviation between predicted and measured values), therefore provide a\u00a0meaningful picture of\u00a0the\u00a0model\u2019s\u00a0true predictive ability. By strictly separating the\u00a0test data from the\u00a0training and tuning phases, the\u00a0evaluation remains unbiased and representative of\u00a0real-world deployment.<\/span><\/p>\n<h3 class=\"03NADPIS3\">Selected model: XGBoost architecture and\u00a0hyperparameters<\/h3>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\">XGBoost (Extreme Gradient Boosting) was selected for its high computational efficiency and superior performance in\u00a0handling non-linear biological data. It is a\u00a0highly optimized implementation of\u00a0gradient boosting that combines decision trees in\u00a0an\u00a0additive manner, using gradient descent to minimize prediction error. It enhances training efficiency and model performance through:<\/span><\/p>\n<ul>\n<li class=\"01TEXT-ODRAZKY\">Second-order gradients (Hessian-based optimization), enabling faster convergence,<\/li>\n<li class=\"01TEXT-ODRAZKY\">Built-in\u00a0L1 (Lasso) and L2 (Ridge) regularization to control complexity and reduce overfitting,<\/li>\n<li class=\"01TEXT-ODRAZKY\">Column and row subsampling to improve generalization.<\/li>\n<\/ul>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">Due to its robustness, speed, and high predictive accuracy, XGBoost has become a\u00a0leading algorithm in\u00a0machine learning competitions (e.g., Kaggle) and real-world applications across engineering, finance, and environmental sciences [24\u201325].<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">\u00a0<\/span><span lang=\"EN-GB\">In\u00a0this study, the\u00a0XGBoost model was trained using the\u00a0following hyperparameters:<\/span><\/p>\n<ul>\n<li class=\"01TEXT-ODRAZKY\">n_estimators = 200: number of\u00a0boosting rounds (trees),<\/li>\n<li class=\"01TEXT-ODRAZKY\">learning_rate = 0.05: step size shrinkage to prevent overfitting,<\/li>\n<li class=\"01TEXT-ODRAZKY\">max_depth = 5: limits tree depth to balance model flexibility and interpretability,<\/li>\n<li class=\"01TEXT-ODRAZKY\">subsample = 0.8: uses 80\u00a0% of\u00a0samples per tree to introduce randomness and\u00a0improve robustness,<\/li>\n<li class=\"01TEXT-ODRAZKY\">colsample_bytree = 0.8: randomly selects 80\u00a0% of\u00a0features at each split, enhancing diversity,<\/li>\n<li class=\"01TEXT-ODRAZKY\">reg_alpha = 0.1, reg_lambda = 0.1: applies mild L1 and L2 regularization to\u00a0penalize large coefficients.<\/li>\n<\/ul>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">These parameters were selected based on prior experimentation and align with best practices for regression tasks. The\u00a0model was implemented using the\u00a0xgboost Python library (version 1.7.6).<\/span><\/p>\n<h3>Performance evaluation metrics<\/h3>\n<p>The predictive performance of the models was assessed using three standard regression metrics: Mean Absolute Error (MAE), Mean Squared Error (MSE), and the Coefficient of Determination (R<sup>2<\/sup>). The mathematical definitions are as follows:<\/p>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-10.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39626 lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-10-300x29.jpg\" alt=\"\" width=\"500\" height=\"49\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-10-300x29.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-10-768x75.jpg 768w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-10-780x78.jpg 780w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-10.jpg 800w\" data-sizes=\"(max-width: 500px) 100vw, 500px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 500px; --smush-placeholder-aspect-ratio: 500\/49;\" \/><\/a>\n<p>&nbsp;<\/p>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-11.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39627 lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-11.jpg\" alt=\"\" width=\"500\" height=\"49\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-11.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-11-300x29.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-11-768x75.jpg 768w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-11-780x78.jpg 780w\" data-sizes=\"(max-width: 500px) 100vw, 500px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 500px; --smush-placeholder-aspect-ratio: 500\/49;\" \/><\/a>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-12.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39628 lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-12.jpg\" alt=\"\" width=\"500\" height=\"74\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-12.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-12-300x44.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-vzorec-12-768x113.jpg 768w\" data-sizes=\"(max-width: 500px) 100vw, 500px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 500px; --smush-placeholder-aspect-ratio: 500\/74;\" \/><\/a>\n<p>where:<\/p>\n<p>y<sub>i<\/sub>\u00a0\u00a0\u00a0 is\u00a0\u00a0 the\u00a0observed BOD5_out value<\/p>\n<p>\u0177<sub>i<\/sub>\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 the\u00a0predicted value<\/p>\n<p>\u0233<sub>i<\/sub>\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 the\u00a0mean\u00a0of\u00a0the\u00a0observed values<\/p>\n<p>n\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 the\u00a0number of\u00a0samples in\u00a0the\u00a0test set<\/p>\n<p>MAE measures the\u00a0average magnitude of\u00a0prediction errors, treating all deviations equally and providing a\u00a0clear and intuitive indicator of\u00a0overall accuracy. It is widely used for its interpretability, particularly in\u00a0environmental and geospatial modeling. MSE, by squaring each error before averaging, assigns greater weight to large deviations, making it effective for detecting models prone to producing substantial outliers [26]. R\u00b2 complements these metrics by indicating the\u00a0proportion of\u00a0variance in\u00a0the\u00a0target variable explained by the\u00a0model, a\u00a0value close to 1 reflects a\u00a0strong fit, whereas lower values indicate weaker explanatory capability. All performance metrics were calculated on the\u00a0held-out test set (January 2023\u2013May 2024), ensuring an\u00a0unbiased assessment of\u00a0the\u00a0model\u2019s\u00a0generalization ability.<\/p>\n<h2>RESULTS AND DISCUSSION<\/h2>\n<p>This section presents the\u00a0performance of\u00a0the\u00a0individual base models and the\u00a0stacked ensemble in\u00a0predicting effluent BOD5 (BOD5_out). The\u00a0analysis includes quantitative results, visualizations, and contextual interpretation to highlight the\u00a0model\u2019s\u00a0accuracy, robustness, and practical relevance.<\/p>\n<h3>Model performance analysis<\/h3>\n<p>The\u00a0predictive performance of\u00a0the\u00a0XGBoost model was evaluated after rigorous data cleaning and feature engineering. From the\u00a0initial dataset, 357 high-quality samples remained for analysis, which were partitioned into a\u00a0training set (n\u00a0=\u00a0286) and a\u00a0test set (n = 71).<\/p>\n<p>The\u00a0model was constructed using 18 features, comprising both raw sensor measurements and engineered indicators such as removal efficiencies and COD\/BOD\u2085 ratios. Of\u00a0these, 17 features served as input variables, including 11 original influent and effluent parameters and 6 engineered variables. The\u00a0remaining variable (BOD\u2085_out) was designated as the\u00a0target output. Performance metrics for the test dataset are summarized in <em>Tab. 3<\/em>.<\/p>\n<p>The results demonstrate a high degree of predictive accuracy. An\u00a0R<sup>2<\/sup> of\u00a096.95\u00a0% suggests that the\u00a0model effectively captures the\u00a0non-linear dynamics of\u00a0the\u00a0treatment process. The\u00a0low MAE (0.146) is particularly significant, as it indicates that the\u00a0model\u2019s\u00a0average prediction error is negligible relative to the\u00a0regulatory limits of\u00a0BOD5.<\/p>\n<h5>Tab. 3. XGBoost performance metrics for BOD5\u200b prediction<\/h5>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-3-1.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39569 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-3-1.jpg\" alt=\"\" width=\"800\" height=\"282\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-3-1.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-3-1-300x106.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-3-1-768x271.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/282;\" \/><\/a>\n<p>&nbsp;<\/p>\n<p><em>Fig. 7<\/em> shows the close alignment between actual and predicted BOD\u2085 levels. The model accurately tracks both stable periods and fluctua-tions, maintaining high precision even during high-load events. This performance exceeds results in previous studies [13\u201316]. Additionally, our use of time-based validation ensures the model is reliable for real-world operations and free from data leakage.<\/p>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-7.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39564 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-7.jpg\" alt=\"\" width=\"800\" height=\"567\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-7.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-7-300x213.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-7-768x544.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/567;\" \/><\/a>\n<h6>Fig. 7. Actual vs. predicted BOD5_out values; the red dashed line represents perfect prediction (y = x)<\/h6>\n<h3 class=\"03NADPIS3\">Feature importance and process insights<\/h3>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\">The inclusion of a total of 17 features was pivotal to the model\u2019s success. Specifically, the integration of domain-informed features, such as COD_removal_eff, COD_out_BOD5_int_ratio, BOD5_int_x_pH and the\u00a0net change in\u00a0conductivity Cond_ratio, (equation 4, 7, 8 and 9 in\u00a0<em><span class=\"01ITALIC\">Tab. 2<\/span><\/em>), provided the\u00a0model with deeper context regarding the\u00a0biological health of\u00a0the\u00a0system. The\u00a0low MSE (0.05) further confirms the\u00a0model\u2019s\u00a0stability, showing that it does not produce large, erratic errors (outliers).<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">This robustness is critical for dependable monitoring under operational conditions, as it allows plant operators to trust the\u00a0predicted effluent values for proactive process adjustments.<\/span><\/p>\n<h3 class=\"03NADPIS3\">Temporal prediction accuracy<\/h3>\n<p class=\"00TEXTbezodsazenienglish\"><em><span class=\"01ITALIC\"><span lang=\"EN-GB\">Fig. 8<\/span><\/span><\/em><span lang=\"EN-GB\"> illustrates the\u00a0temporal evolution of\u00a0actual versus predicted BOD\u2085 over the\u00a0test period. The\u00a0model accurately tracks the\u00a0dynamic behavior of\u00a0the\u00a0treatment process, including short-term spikes and gradual trends. Despite minor fluctuations, the\u00a0model swiftly recovers and continues to provide precise tracking. Residual analysis further confirms stable performance. As shown in\u00a0<em><span class=\"01ITALIC\">Fig. 9<\/span><\/em>, residuals are randomly distributed around zero, with no visible trend or increasing variance over time. This indicates homoscedastic errors and the\u00a0absence of\u00a0structural bias, both of\u00a0which are essential properties for reliable deployment in\u00a0real-world monitoring systems.<\/span><\/p>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\"> <a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-8.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39565 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-8.jpg\" alt=\"\" width=\"800\" height=\"411\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-8.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-8-300x154.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-8-768x395.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/411;\" \/><\/a><\/span><\/p>\n<h6 class=\"05POPISKYobrazku\">Fig. 8. Temporal comparison of\u00a0actual and predicted effluent BOD\u2085 over the\u00a0test period<\/h6>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-9.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39566 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-9.jpg\" alt=\"\" width=\"800\" height=\"411\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-9.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-9-300x154.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-fig-9-768x395.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/411;\" \/><\/a>\n<h6 class=\"05POPISKYobrazku\">Fig. 9. Residuals over the\u00a0test period; errors are centered around zero, with no systematic drift<\/h6>\n<h3 class=\"03NADPIS3\">Comparison with literature<\/h3>\n<p class=\"00TEXTbezodsazenienglish\"><em><span class=\"01ITALIC\"><span lang=\"EN-GB\">Tab. 4<\/span><\/span><\/em><span lang=\"EN-GB\"> compares the\u00a0performance of\u00a0this study with recent works on BOD<span class=\"01DOLNIINDEX\">5<\/span>\/COD prediction in\u00a0WWTPs. Our model outperforms several existing approaches in\u00a0terms of\u00a0R\u00b2. Notably, many previous studies used random train\/test splits, which may lead to inflated performance estimates due to temporal leakage. In\u00a0contrast, this study employs a\u00a0strict chronological split (80\u00a0% \/ 20\u00a0%), preserving the\u00a0time-series nature of\u00a0the\u00a0data and ensuring its generalizability under real deployment scenarios. In\u00a0contrast, our model provides high accuracy while using strict validation procedures and ensuring full reproducibility.<\/span><\/p>\n<h5 class=\"04TABULKApopisek\"><span class=\"01ITALIC\">\u00a0<\/span><span class=\"01ITALIC\">Tab.\u00a04. Comparative performance of\u00a0BOD<\/span><span class=\"01DOLNIINDEX\">5<\/span><span class=\"01ITALIC\"> prediction models in\u00a0recent literature<\/span><\/h5>\n<a href=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-4-1.jpg\" rel=\"shadowbox[sbpost-39294];player=img;\"><img decoding=\"async\" class=\"alignnone wp-image-39570 size-full lazyload\" data-src=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-4-1.jpg\" alt=\"\" width=\"800\" height=\"439\" data-srcset=\"https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-4-1.jpg 800w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-4-1-300x165.jpg 300w, https:\/\/www.vtei.cz\/wp-content\/uploads\/2026\/08\/Lekouaghet-tab-4-1-768x421.jpg 768w\" data-sizes=\"(max-width: 800px) 100vw, 800px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 800px; --smush-placeholder-aspect-ratio: 800\/439;\" \/><\/a>\n<h3 class=\"03NADPIS3\">Generalizability and transferability considerations<\/h3>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\">Despite the\u00a0strong performance achieved at the\u00a0study site (R\u00b2 = 0.9695), the\u00a0model\u2019s\u00a0applicability beyond this context remains constrained by its reliance on site-specific training data. Predictive performance is strongly influenced by local influent characteristics, operational conditions, and environmental factors embedded in\u00a0the\u00a0training dataset. Consequently, reliable performance is expected primarily for facilities with similar process configurations and operating regimes.<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">Factors limiting the\u00a0model\u2019s\u00a0transferability include variations in\u00a0influent composition, hydraulic loading, and process design (e.g., aeration strategy, sludge age, and nutrient removal stages), as well as regional conditions such as climate variability and industrial discharge profiles. Applying the\u00a0model to plants with substantially different characteristics without recalibration may result in\u00a0reduced accuracy or increased uncertainty. Additionally, the\u00a0model is primarily suitable for medium-load operations and may not be directly applicable to low-load treatment plants.<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">Future work should focus on multi-site validation and the\u00a0integration of\u00a0domain\u00a0adaptation or transfer learning techniques to enhance model robustness and enable broader deployment across diverse wastewater treatment systems.<\/span><\/p>\n<h2 class=\"03NADPIS2\">CONCLUSION<\/h2>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\">This study demonstrates that a\u00a0standalone XGBoost model can\u00a0achieve highly accurate prediction of\u00a0effluent BOD\u2085 (BOD<span class=\"01DOLNIINDEX\">5<\/span>_out) using nearly a\u00a0decade of\u00a0operational data (2015\u20132024) from a\u00a0full-scale WWTP. With an\u00a0R\u00b2 of\u00a096.95\u00a0%, MSE\u00a0of\u00a00.05, and MAE of\u00a00.146\u00a0mg\/L on a\u00a0chronologically held-out test set, the\u00a0model delivers exceptional performance while avoiding the\u00a0complexity of\u00a0ensemble fusion strategies.<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">Crucially, this level of\u00a0accuracy was achieved through a\u00a0combination of\u00a0domain-informed feature engineering, including COD removal efficiency, pH\u00a0variation, and COD\/BOD\u2085 ratios, together with rigorous time-based validation that ensures realistic evaluation under forecasting conditions.<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\" style=\"letter-spacing: 0pt;\">Despite the\u00a0absence of\u00a0dissolved oxygen from the\u00a0dataset, the\u00a0model maintained strong predictive performance, indicating that routinely monitored variables such as effluent COD, conductivity, and influent pH captured sufficient information about process conditions and biological activity. Consequently, while the\u00a0fundamental importance of\u00a0DO in\u00a0treatment kinetics remains undisputed, these findings indicate that highly correlated operational parameters can\u00a0partially mitigate the\u00a0absence of\u00a0direct DO measurements in\u00a0predictive modeling.<\/span><\/p>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">More importantly, the\u00a0model acts as a\u00a0practical sensor, providing nearly instantaneous estimates of\u00a0BOD\u2085 from direct measurements. By replacing the\u00a05-day laboratory delay with instant predictions, it turns BOD\u2085 from a\u00a0static compliance value into a\u00a0useful control variable. This allows:<\/span><\/p>\n<ul>\n<li class=\"01TEXT-ODRAZKY\">Early detection of\u00a0process upsets: Continuous BOD\u2085 prediction enables early detection of\u00a0abnormal biological performance caused by load variations or hydraulic disturbances, allowing timely operational interventions before effluent quality deteriorates.<\/li>\n<li class=\"01TEXT-ODRAZKY\">Optimization of\u00a0aeration energy use: Predicted BOD<span class=\"01DOLNIINDEX\">5\u200b <\/span>allows for precise aeration control, ensuring oxygen supply matches organic load. This eliminates energy waste from over-aeration while maintaining high process performance.<\/li>\n<li class=\"01TEXT-ODRAZKY\">Proactive adjustments to maintain\u00a0regulatory compliance: Anticipatory BOD<span class=\"01DOLNIINDEX\">5\u200b<\/span> forecasting enables preventive operational adjustments when predicted values approach discharge limits. This ensures consistent regulatory compliance, thereby avoiding heavy fines as well as environmental and economic damage.<\/li>\n<\/ul>\n<p class=\"00TEXTenglish\"><span lang=\"EN-GB\">In\u00a0summary, this study shows that a\u00a0straightforward model can\u00a0still achieve excellent accuracy. With proper data handling and domain\u00a0knowledge, a\u00a0tuned XGBoost model can\u00a0equal or surpass more complex methods while remaining faster, easier to interpret, and simpler to deploy. This makes it a\u00a0strong option for scalable wastewater management applications.<\/span><\/p>\n<p class=\"00TEXTbezodsazenienglish\"><span lang=\"EN-GB\">This paper has been peer-reviewed.<\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Accurate and quick prediction of effluent Biochemical Oxygen Demand (BOD5) is important for maintaining regulatory standards and improving the operation of wastewater treatment (WWT) plants. Traditional laboratory tests for BOD5 take several days, causing delays in assessing water quality and adjusting plant performance. To overcome this issue, this study develops a machine learning model to estimate effluent BOD5 using easily available plant data. A dataset from a large-scale wastewater treatment plant was used, comprising 12 variables from both influent and effluent data, such as pH, Chemical Oxygen Demand (COD), conductivity, Total Suspended Solids (TSS), BOD5, and temperature, with effluent BOD5 as the target variable.<\/p>\n","protected":false},"author":8,"featured_media":39460,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":"","_members_access_role":[],"_members_access_error":""},"categories":[94,90,93],"tags":[4194,4197,4195,4198,4196,3355,4179],"coauthors":[4182,4183,4184],"class_list":["post-39294","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-current-issue","category-waste-management","category-two-articles","tag-biochemical-oxygen-demand-bod5","tag-bod5-prediction","tag-machine-learning","tag-process-optimization","tag-soft-sensor","tag-wastewater-treatment","tag-xgboost"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.vtei.cz\/en\/wp-json\/wp\/v2\/posts\/39294","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.vtei.cz\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.vtei.cz\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.vtei.cz\/en\/wp-json\/wp\/v2\/users\/8"}],"replies":[{"embeddable":true,"href":"https:\/\/www.vtei.cz\/en\/wp-json\/wp\/v2\/comments?post=39294"}],"version-history":[{"count":2,"href":"https:\/\/www.vtei.cz\/en\/wp-json\/wp\/v2\/posts\/39294\/revisions"}],"predecessor-version":[{"id":39635,"href":"https:\/\/www.vtei.cz\/en\/wp-json\/wp\/v2\/posts\/39294\/revisions\/39635"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.vtei.cz\/en\/wp-json\/wp\/v2\/media\/39460"}],"wp:attachment":[{"href":"https:\/\/www.vtei.cz\/en\/wp-json\/wp\/v2\/media?parent=39294"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.vtei.cz\/en\/wp-json\/wp\/v2\/categories?post=39294"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.vtei.cz\/en\/wp-json\/wp\/v2\/tags?post=39294"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/www.vtei.cz\/en\/wp-json\/wp\/v2\/coauthors?post=39294"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}