{"doi":"10.1109/scis-isis.2012.6505411","title":"Smoothing of ngram language models of human chats","abstract":null,"journal":"The 6th International Conference on Soft Computing and Intelligent Systems, and The 13th International Symposium on Advanced Intelligence Systems","year":2012,"id":652237,"datarank":0.29188652235829704,"base_score":1.9459101490553132,"endowment":1.9459101490553132,"self_citation_contribution":0.29188652235829704,"citation_network_contribution":0.0,"self_endowment_contribution":0.29188652235829704,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":6,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1701331,"name":"Joe Dumoulin","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Smoothing of ngram language models of human chats","abstract":"Ngram language models are ubiquitous in speech applications and many other natural language systems. One issue with n-gram language models is that the language is not completely represented in the model. When words appear that are not in the model, we may need to provide a smoothing method to distribute the model probabilities over the unknown values. Many techniques exist for language model smoothing with many different performance characteristics. Often the performance of smoothing algorithms may depend on the application of the language model (so, for example, unigram models with interpolation smoothing may perform better with information retrieval applications, but trigram models with backoff smoothing might perform better for speech). This paper examines the relative performance of some selected smoothing methods with bigram language models created using chat data. The language models are used for machine translation of chat data and for creating text classification models.","is_dataset_classified":null,"base_score":1.9459101490553132,"endowment":1.9459101490553132,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19162232","pmcid":null,"openalex_id":"https://openalex.org/W2028371633","authors":[],"funders":[],"total_grants":0,"fwci":0.7805,"citation_percentile":0.6905356,"influential_citations":0,"citation_trend":[{"year":2014,"count":1},{"year":2015,"count":2},{"year":2016,"count":1},{"year":2019,"count":1},{"year":2023,"count":1}],"oa_status":"closed","license":"https://doi.org/10.15223/policy-029","oa_locations":[{"url":"http://xplorestaging.ieee.org/ielx7/6495582/6504994/06505411.pdf?arnumber=6505411","host_type":"publisher"},{"url":"https://doi.org/10.1109/scis-isis.2012.6505411","host_type":""}],"fields_of_study":["Speech Recognition and Synthesis","Speech and dialogue systems","Topic Modeling"],"mesh_terms":[],"keywords":["Bigram","Language model","Trigram","Smoothing","Computer science","Natural language processing","Artificial intelligence","Data modeling","Machine learning","Natural language","Machine translation","Cache language model","Speech recognition","Interpolation (computer graphics)","Universal Networking Language","Database"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Quality Education"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-10T12:32:40.780963Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}