{"doi":"10.1101/2021.07.06.451263","title":"Efficient gradient boosting for prognostic biomarker discovery","abstract":"Abstract Motivation Gradient boosting decision tree (GBDT) is a powerful ensemble machine learning method that has the potential to accelerate biomarker discovery from high-dimensional molecular data. Recent algorithmic advances, such as Extreme Gradient Boosting (XGB) and Light Gradient Boosting (LGB), have rendered the GBDT training more efficient, scalable and accurate. These modern techniques, however, have not yet been widely adopted in biomarkers discovery based on patient survival data, which are key clinical outcomes or endpoints in cancer studies. Results In this paper, we present a new R package Xsurv as an integrated solution which applies two modern GBDT training framework namely, XGB and LGB, for the modeling of censored survival outcomes. Based on a comprehensive set of simulations, we benchmark the new approaches against traditional methods including the stepwise Cox regression model and the original gradient boosting function implemented in the package gbm . We also demonstrate the application of Xsurv in analyzing a melanoma methylation dataset. Together, these results suggest that Xsurv is a useful and computationally viable tool for screening a large number of prognostic candidate biomarkers, which may facilitate cancer translational and clinical research. Availability Xsurv is freely available as an R package at: https://github.com/topycyao/Xsurv Contact xuefeng.wang@moffitt.org Supplementary information Supplementary data are available at Bioinformatics online.","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2021,"id":217631,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":3,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9507,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":650251,"name":"Sijie Yao","orcid":"0000-0003-4228-0832","position":1,"is_corresponding":false},{"id":326458,"name":"Zhenyu Zhang","orcid":"0000-0001-5570-090X","position":2,"is_corresponding":false},{"id":650252,"name":"Biwei Cao","orcid":"0009-0000-4255-8369","position":3,"is_corresponding":false},{"id":814555,"name":"C. M. Wilson","orcid":"0000-0003-2480-9229","position":4,"is_corresponding":false},{"id":264437,"name":"Pei Fen Kuan","orcid":"0000-0001-7861-916X","position":5,"is_corresponding":false},{"id":541037,"name":"Ruoqing Zhu","orcid":"0000-0002-0753-5716","position":6,"is_corresponding":false},{"id":377318,"name":"Xuefeng Wang","orcid":"0000-0001-5775-408X","position":7,"is_corresponding":false},{"id":650250,"name":"Kaiqiao Li","orcid":"0000-0001-8069-5639","position":0,"is_corresponding":true}],"reference_count":27,"raw_metadata":null,"created_at":"2026-07-18T23:53:20.374789Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}