{"doi":"10.17760/d20467253","title":"Self-repetition in abstractive neural summarizers","abstract":"We provide a quantitative and qualitative analysis of self-repetition in the output of neural summarizers. We measure self-repetition as the number of n-grams of length four or longer that appear in multiple outputs of the same system. We analyze the behavior of three popular summarization architectures (BART, T5, and Pegasus), fine-tuned on five representative datasets. In a regression analysis, we find that the three architectures have different propensities for repeating con-tent in their summaries for different inputs, with BART being the most and considerably prone to self-repeat. Fine-tuning on more abstractive data is associated with a higher rate of self-repetition. In the qualitative analysis, we find systems produce artefacts such as ads and disclaimers unrelated to the content being summarized, as well as formulaic phrases common in the fine-tuning domain, which often are hallucinations not supported by the original input. While some hallucinations are suitable in the context of the summary and its input, often times they are incongruous and completely unrelated to the input. Our approach to corpus-level analysis of self-repetition may help practitioners clean up training data for summarizers, and ultimately support methods for minimizing the amount of hallucinations in abstractive neural summaries.--Author's abstract","journal":"PubMed","year":2022,"id":300734,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":3,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9564,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2022-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":992081,"name":"Nikita Salkar","orcid":null,"position":0,"is_corresponding":true}],"reference_count":22,"raw_metadata":null,"created_at":"2026-07-19T00:31:57.812980Z","pmid":"37484061","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}