Constarium
← Search

Data · dataset · 2024

Conjunto de datos de: Water-quality data imputation with a high percentage of missing values: a machine learning approach

Listed in Redata

The monitoring of surface-water quality followed by water-quality modeling and analysis is essential for generating effective strategies in water resource management.

Description

However, water-quality studies are limited by the lack of complete and reliable data sets on surface-water-quality variables. These deficiencies are particularly noticeable in developing countries.

This work focuses on surface-water-quality data from Santa Lucía Chico river (Uruguay), a mixed lotic and lentic river system. Data collected at six monitoring stations are publicly available at dinama.gub.uy/oan/datos-abiertos/calidad-agua/. The high temporal and spatial variability that characterizes water-quality variables and the high rate of missing values (between 50% and 70%) raises significant challenges.

Read the rest (3 more)

To deal with missing values, we applied several statistical and machine-learning imputation methods. The competing algorithms implemented belonged to both univariate and multivariate imputation methods (inverse distance weighting (IDW), Random Forest Regressor (RFR), Ridge (R), Bayesian Ridge (BR), AdaBoost (AB), Huber Regressor (HR), Support Vector Regressor (SVR), and K-nearest neighbors Regressor (KNNR)). IDW outperformed the others, achieving a very good performance (NSE greater than 0.8) in most cases.

In this dataset, we include the original and imputed values for the following variables: - Water temperature (Tw) - Dissolved oxygen (DO) - Electrical conductivity (EC) - pH - Turbidity (Turb) - Nitrite (NO2-) - Nitrate (NO3-) - Total Nitrogen (TN) Each variable is identified as [STATION] VARIABLE FULL NAME (VARIABLE SHORT NAME) [UNIT METRIC]. More details about the study area, the original datasets, and the methodology adopted can be found in our paper mdpi.com/2071-1050/13/11/6318.

If you use this dataset in your work, please cite our paper: Rodríguez, R.; Pastorini, M.; Etcheverry, L.; Chreties, C.; Fossati, M.; Castro, A.; Gorgoglione, A. Water-Quality Data Imputation with a High Percentage of Missing Values: A Machine Learning Approach. Sustainability 2021, 13, 6318. doi.org/10.3390/su13116318

Links

Where it is published

Catalogue records · 1

Topics

Provenance · 1 source records, 11 field assertions
SourceKeyLast seenRaw
Redatadoi:10.60895/redata/TNRT8Q8 d agoJSON v1
FieldAssertionExtractorEvidence
concepts[field].anzsrc:field:410504enrichment · redata anii org uytaxonomy-embedding@1.1.0title+keywords+description (76%)
concepts[field].dataverse_subject:ciencias-de-la-informaci-n-y-computaci-nsource · redata anii org uyconnector:redata_anii_org_uy@1.0.0/subjects
concepts[field].dataverse_subject:ciencias-de-la-tierra-y-el-medioambientesource · redata anii org uyconnector:redata_anii_org_uy@1.0.0/subjects
concepts[field].dataverse_subject:ingenier-asource · redata anii org uyconnector:redata_anii_org_uy@1.0.0/subjects
concepts[field].local:field:earth-environmentalmapping · redata anii org uyconnector:redata_anii_org_uy@1.0.0/subjects
created_datesource · redata anii org uyconnector:redata_anii_org_uy@1.0.0
descriptionsource · redata anii org uyconnector:redata_anii_org_uy@1.0.0/description
publication_datesource · redata anii org uyconnector:redata_anii_org_uy@1.0.0
titlesource · redata anii org uyconnector:redata_anii_org_uy@1.0.0/name
updated_datesource · redata anii org uyconnector:redata_anii_org_uy@1.0.0
version_labelsource · redata anii org uyconnector:redata_anii_org_uy@1.0.0