{"doi":"10.1093/bioinformatics/18.2.362","title":"Structure motif discovery and mining the PDB","abstract":"<jats:title>Abstract</jats:title>\n               <jats:p>Motivation: Many of the most interesting functional and evolutionary relationships among proteins are so ancient that they cannot be reliably detected through sequence analysis and are apparent only through a comparison of the tertiary structures. The conserved features can often be described as structural motifs consisting of a few single residues or Secondary Structure (SS) elements. Confidence in such motifs is greatly boosted when they are found in more than a pair of proteins.</jats:p>\n               <jats:p>Results: We describe an algorithm for the automatic discovery of recurring patterns in protein structures. The patterns consist of individual residues having a defined order along the protein’s backbone that come close together in the structure and whose spatial conformations are similar. The residues in a pattern need not be close in the protein’s sequence. The work described in this paper builds on an earlier reported algorithm for motif discovery. This paper describes a significant improvement of the algorithm which makes it very efficient. The improved efficiency allows us to use it for doing unsupervised learning of patterns occurring in small subsets in a large set of structures, a non-redundant subset of the Protein Data Bank (PDB) database of all known protein structures.</jats:p>\n               <jats:p>Availability: The program is freely available to academia, requests can be sent to Inge.Jonassen@ii.uib.no.</jats:p>\n               <jats:p>Contact: Inge.Jonassen@ii.uib.no</jats:p>\n               <jats:p>* To whom correspondence should be addressed.</jats:p>","journal":"Bioinformatics","year":2002,"id":14947,"datarank":2.380657851270929,"base_score":3.8918202981106265,"endowment":3.8918202981106265,"self_citation_contribution":0.5837730447165941,"citation_network_contribution":1.796884806554335,"self_endowment_contribution":0.5837730447165941,"citer_contribution":1.796884806554335,"corpus_percentile":null,"corpus_rank":null,"citation_count":48,"citer_count":47,"citers_with_citation_signal":34,"citers_with_endowment":34,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":115904,"name":"Ingvar Eidhammer","orcid":null,"position":1,"is_corresponding":false},{"id":115905,"name":"Darrell Conklin","orcid":null,"position":2,"is_corresponding":false},{"id":115906,"name":"William R. Taylor","orcid":null,"position":3,"is_corresponding":false},{"id":87086,"name":"Inge Jonassen","orcid":"0000-0003-4110-0748","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"base_score":3.8918202981106265,"endowment":3.8918202981106265,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"11847094","pmcid":null,"openalex_id":"https://openalex.org/W2133452849","authors":[],"funders":[],"total_grants":0,"fwci":1.9552,"citation_percentile":0.86206189,"influential_citations":2,"citation_trend":[{"year":2012,"count":1},{"year":2013,"count":1},{"year":2014,"count":3},{"year":2015,"count":2},{"year":2017,"count":2},{"year":2020,"count":2},{"year":2022,"count":1}],"oa_status":"bronze","license":null,"oa_locations":[{"url":"https://academic.oup.com/bioinformatics/article-pdf/18/2/362/606303/180362.pdf","host_type":"journal"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/18/2/362/606303/180362.pdf","host_type":"BRONZE"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/18/2/362/606303/180362.pdf","host_type":"publisher"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/18/2/362/48850516/bioinformatics_18_2_362.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/bioinformatics/18.2.362","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/11847094","host_type":"repository"}],"fields_of_study":["Biochemical and Structural Characterization","Glycosylation and Glycoproteins Research","Chemical Synthesis and Analysis","Computer Science","Medicine","Biology","Algorithms","Amino Acid Motifs","Computational Biology","Cystine","Databases, Protein","Molecular Structure","Protein Structure, Secondary","Proteins","Software"],"mesh_terms":["Algorithms","Cystine","Proteins","Software","Molecular Structure","Protein Structure, Secondary","Computational Biology","Amino Acid Motifs","Databases, Protein"],"keywords":["Motif (music)","Protein Data Bank (RCSB PDB)","Computer science","Computational biology","Structural motif","Data mining","Bioinformatics","Biology","Biochemistry"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-06-01T15:41:29.066138Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}