{"doi":"10.5815/ijitcs.2017.09.06","title":"A New Dynamic Data Cleaning Technique for Improving Incomplete Dataset Consistency","abstract":null,"journal":"International Journal of Information Technology and Computer Science","year":2017,"id":681420,"datarank":0.10397207708399181,"base_score":0.6931471805599453,"endowment":0.6931471805599453,"self_citation_contribution":0.10397207708399181,"citation_network_contribution":0.0,"self_endowment_contribution":0.10397207708399181,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":1,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1780360,"name":"Sreedhar Kumar S","orcid":null,"position":1,"is_corresponding":false},{"id":1780363,"name":"Meenakshi Sundaram S","orcid":null,"position":2,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"A New Dynamic Data Cleaning Technique for Improving Incomplete Dataset Consistency","abstract":"This paper presents a new approach named Dynamic Data Cleaning (DDC) aims to improve incomplete dataset consistency by identifying, reconstructing and removing inconsistent data objects for future data analysis process. The proposed DDC approach consists of three methods: Identify Normal Object (INO), Reconstruct Normal Object (RNO) and Dataset Quality Measure (DQM). The first method INO divides the incomplete dataset into normal objects and abnormal objects (outliers) based on degree of missing attributes values in each individual object. Second, the (RNO) method reconstructs missed attributes values in the normal objects by the closest object based on a distance metric and removes inconsistent data objects (outliers) with higher missed data. Finally, the DQM method measures the consistency and inconsistency among the objects in improved dataset with and without outlier. Experimental results show that the proposed DDC approach is suitable to identify and reconstruct the incomplete data objects for improving dataset consistency from lower to higher level without user knowledge.","is_dataset_classified":null,"base_score":0.6931471805599453,"endowment":0.6931471805599453,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"21097893","pmcid":null,"openalex_id":"https://openalex.org/W2749515417","authors":[],"funders":[],"total_grants":0,"fwci":0.2274,"citation_percentile":0.68324223,"influential_citations":0,"citation_trend":[{"year":2018,"count":1}],"oa_status":"gold","license":null,"oa_locations":[{"url":"http://www.mecs-press.org/ijitcs/ijitcs-v9-n9/IJITCS-V9-N9-6.pdf","host_type":"journal"},{"url":"http://www.mecs-press.org/ijitcs/ijitcs-v9-n9/IJITCS-V9-N9-6.pdf","host_type":"publisher"},{"url":"https://doi.org/10.5815/ijitcs.2017.09.06","host_type":"journal"}],"fields_of_study":["Machine Learning and Data Classification","Data Quality and Management","Data Mining Algorithms and Applications"],"mesh_terms":[],"keywords":["Consistency (knowledge bases)","Outlier","Computer science","Object (grammar)","Data consistency","Metric (unit)","Data mining","Process (computing)","Missing data","Measure (data warehouse)","Quality (philosophy)","Artificial intelligence","Anomaly detection","Pattern recognition (psychology)","Machine learning","Database"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Industry, innovation and infrastructure"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-17T17:47:08.799514Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}