{"doi":"10.21608/jaet.2020.40032.1035","title":"Pre-processing Steps for Genome-wide High-density NARAC Dataset Facilitates its Haplotype Block Partitioning","abstract":"The pre-processing ‎ ‎ phase‎ is a crucial step to prepare any data for deep considerable ‎ analysis. ‎Genome-wide data ‎is considered ‎ big data; dealing with such data is not an easy task and still poses ‎a significant challenge. The ‎genome-wide association study (GWAS) ‎ is based on enormous high-‎density data with high throughput. This paper has illustrated the main pre-processing ‎ steps on data ‎from North American Rheumatoid Arthritis Consortium ‎‎(NARAC) for preparing it for haplotype ‎block partitioning using different methods and with different platforms. This paper’s main ‎objective is to summarize the steps of pre-processing the raw genotyped dataset to prepare it for ‎haplotype block partitioning and further analyses. Besides, we present each practical step by clear ‎tables for better visualizing, elucidation, and workflow interpretation. Besides, we aimed to ‎overcome the missing data and normalize the output in a standardized format. Eventually, this will ‎improve the understanding of such data formats and build the foundation stone of critical genome-wide experiments and studies. Thus, this work could a guide for other researchers who use similar ‎data. The pre-processed data will be applied to imputation, BigLD block partitioning under R and ‎Haploview methods. Our sequence of ‎pre-processing steps includes preparing the characters to be ‎in a form that is suitable for imputation. The next step is ‎recording data in 0,1,2 format to be ‎proper for the BigLD. We were finally preparing data for Haploview to ‎provide clear haplotype ‎block partitioning, association analysis, and furthermore.‎","journal":"Journal of Advanced Engineering Trends","year":2020,"id":114197,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.8938,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2020-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":537621,"name":"Mohamed N. Saad","orcid":"0000-0001-8229-0280","position":1,"is_corresponding":false},{"id":538165,"name":"Ashraf Said lrm","orcid":null,"position":2,"is_corresponding":false},{"id":537622,"name":"Hesham F. A. Hamed","orcid":"0000-0002-4208-1190","position":3,"is_corresponding":false},{"id":538164,"name":"F. Ibrahim","orcid":null,"position":0,"is_corresponding":true}],"reference_count":49,"raw_metadata":null,"created_at":"2026-07-18T23:13:29.674622Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}