{"doi":"10.1162/tacl_a_00669","title":"Source-Free Domain Adaptation for Question Answering with Masked Self-training","abstract":"<jats:title>Abstract</jats:title>\n               <jats:p>Previous unsupervised domain adaptation (UDA) methods for question answering (QA) require access to source domain data while fine-tuning the model for the target domain. Source domain data may, however, contain sensitive information and should be protected. In this study, we investigate a more challenging setting, source-free UDA, in which we have only the pretrained source model and target domain data, without access to source domain data. We propose a novel self-training approach to QA models that integrates a specially designed mask module for domain adaptation. The mask is auto-adjusted to extract key domain knowledge when trained on the source domain. To maintain previously learned domain knowledge, certain mask weights are frozen during adaptation, while other weights are adjusted to mitigate domain shifts with pseudo-labeled samples generated in the target domain. Our empirical results on four benchmark datasets suggest that our approach significantly enhances the performance of pretrained QA models on the target domain, and even outperforms models that have access to the source data during adaptation.</jats:p>","journal":"Transactions of the Association for Computational Linguistics","year":2024,"id":646084,"datarank":0.24141568686511508,"base_score":1.6094379124341003,"endowment":1.6094379124341003,"self_citation_contribution":0.24141568686511508,"citation_network_contribution":0.0,"self_endowment_contribution":0.24141568686511508,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":4,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1290705,"name":"Boyu Wang","orcid":"0009-0001-5829-3366","position":1,"is_corresponding":false},{"id":363463,"name":"Yue Dong","orcid":"0000-0002-1737-6536","position":2,"is_corresponding":false},{"id":1682617,"name":"Charles Ling","orcid":null,"position":3,"is_corresponding":false},{"id":1682615,"name":"Maxwell J. Yin","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Source-Free Domain Adaptation for Question Answering with Masked Self-training","abstract":"<jats:title>Abstract</jats:title>\n               <jats:p>Previous unsupervised domain adaptation (UDA) methods for question answering (QA) require access to source domain data while fine-tuning the model for the target domain. Source domain data may, however, contain sensitive information and should be protected. In this study, we investigate a more challenging setting, source-free UDA, in which we have only the pretrained source model and target domain data, without access to source domain data. We propose a novel self-training approach to QA models that integrates a specially designed mask module for domain adaptation. The mask is auto-adjusted to extract key domain knowledge when trained on the source domain. To maintain previously learned domain knowledge, certain mask weights are frozen during adaptation, while other weights are adjusted to mitigate domain shifts with pseudo-labeled samples generated in the target domain. Our empirical results on four benchmark datasets suggest that our approach significantly enhances the performance of pretrained QA models on the target domain, and even outperforms models that have access to the source data during adaptation.</jats:p>","is_dataset_classified":null,"base_score":1.3862943611198906,"endowment":1.3862943611198906,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19767382","pmcid":null,"openalex_id":"https://openalex.org/W4399426318","authors":[],"funders":[],"total_grants":0,"fwci":0.7303,"citation_percentile":0.75047497,"influential_citations":0,"citation_trend":[{"year":2026,"count":3}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"http://dx.doi.org/10.1162/tacl_a_00669","host_type":"journal"},{"url":"http://dx.doi.org/10.1162/tacl_a_00669","host_type":"publisher"},{"url":"https://direct.mit.edu/tacl/article-pdf/doi/10.1162/tacl_a_00669/2377802/tacl_a_00669.pdf","host_type":"publisher"},{"url":"https://doaj.org/article/7e465f499fb94d969a5086bbe3c342b1","host_type":"repository"}],"fields_of_study":["Topic Modeling","Domain Adaptation and Few-Shot Learning","Multimodal Machine Learning Applications"],"mesh_terms":[],"keywords":["Computer science","Domain adaptation","Adaptation (eye)","Domain (mathematical analysis)","Question answering","Training (meteorology)","Training set","Natural language processing","Artificial intelligence","Information retrieval","Speech recognition","Psychology"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-09T11:33:06.988183Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}