{"doi":"10.18653/v1/w18-3001","title":"Corpus Specificity in LSA and Word2vec: The Role of Out-of-Domain Documents","abstract":null,"journal":"Proceedings of The Third Workshop on Representation Learning for NLP","year":2018,"id":619123,"datarank":0.3596842909197557,"base_score":2.3978952727983707,"endowment":2.3978952727983707,"self_citation_contribution":0.3596842909197557,"citation_network_contribution":0.0,"self_endowment_contribution":0.3596842909197557,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":10,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":771348,"name":"Mariano Sigman","orcid":"0000-0003-0589-0033","position":1,"is_corresponding":false},{"id":988464,"name":"Diego Fernández Slezak","orcid":"0000-0001-6348-1559","position":2,"is_corresponding":false},{"id":1597589,"name":"Edgar Altszyler","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Corpus Specificity in LSA and Word2vec: The Role of Out-of-Domain Documents","abstract":"Despite the popularity of word embeddings, the precise way by which they acquire semantic relations between words remain unclear. In the present article, we investigate whether LSA and word2vec capacity to identify relevant semantic relations increases with corpus size. One intuitive hypothesis is that the capacity to identify relevant associations should increase as the amount of data increases. However, if corpus size grows in topics which are not specific to the domain of interest, signal to noise ratio may weaken.","is_dataset_classified":null,"base_score":1.9459101490553132,"endowment":1.9459101490553132,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19767382","pmcid":null,"openalex_id":"https://openalex.org/W2778440942","authors":[],"funders":[],"total_grants":0,"fwci":0.5659,"citation_percentile":0.69766683,"influential_citations":0,"citation_trend":[{"year":2020,"count":3},{"year":2021,"count":2},{"year":2022,"count":1}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://www.aclweb.org/anthology/W18-3001.pdf","host_type":""},{"url":"https://www.aclweb.org/anthology/W18-3001.pdf","host_type":""},{"url":"https://doi.org/10.18653/v1/w18-3001","host_type":""},{"url":"http://arxiv.org/abs/1712.10054","host_type":"repository"},{"url":"https://arxiv.org/pdf/1712.10054.pdf","host_type":"repository"},{"url":"https://doi.org/10.48550/arxiv.1712.10054","host_type":"repository"},{"url":"https://doi.org/10.60692/jshkc-mt746","host_type":"repository"},{"url":"https://doi.org/10.60692/v75aj-c1630","host_type":"repository"},{"url":"https://arxiv.org/pdf/1712.10054","host_type":"repository"}],"fields_of_study":["Topic Modeling","Natural Language Processing Techniques","Domain Adaptation and Few-Shot Learning"],"mesh_terms":[],"keywords":["Word2vec","Computer science","Latent semantic analysis","Natural language processing","Curse of dimensionality","Artificial intelligence","Word (group theory)","Point (geometry)","Domain (mathematical analysis)","Text corpus","Popularity","Set (abstract data type)","Information retrieval","Linguistics","Mathematics","Psychology"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-03T06:04:43.557184Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}