{"doi":"10.1101/2024.07.22.604620","title":"FuncFetch: An LLM-assisted workflow enables mining thousands of enzyme-substrate interactions from published manuscripts","abstract":"<jats:title>Abstract</jats:title>\n                <jats:sec>\n                  <jats:title>Motivation</jats:title>\n                  <jats:p>Thousands of genomes are publicly available, however, most genes in those genomes have poorly defined functions. This is partly due to a gap between previously published, experimentally-characterized protein activities and activities deposited in databases. This activity deposition is bottlenecked by the time-consuming biocuration process. The emergence of large language models (LLMs) presents an opportunity to speed up text-mining of protein activities for biocuration.</jats:p>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>We developed FuncFetch — a workflow that integrates NCBI E-Utilities, OpenAI’s GPT-4 and Zotero — to screen thousands of manuscripts and extract enzyme activities. Extensive validation revealed high precision and recall of GPT-4 in determining whether the abstract of a given paper indicates presence of a characterized enzyme activity in that paper. Provided the manuscript, FuncFetch extracted data such as species information, enzyme names, sequence identifiers, substrates and products, which were subjected to extensive quality analyses. Comparison of this workflow against a manually curated dataset of BAHD acyltransferase activities demonstrated a precision/recall of 0.86/0.64 in extracting substrates. We further deployed FuncFetch on nine large plant enzyme families. Screening 27,120 papers, FuncFetch retrieved 32,605 entries from 5547 selected papers. We also identified multiple extraction errors including incorrect associations, non-target enzymes, and hallucinations, which highlight the need for further manual curation. The BAHD activities were verified, resulting in a comprehensive functional fingerprint of this family and revealing that ∼70% of the experimentally characterized enzymes are uncurated in the public domain. FuncFetch represents an advance in biocuration and lays the groundwork for predicting functions of uncharacterized enzymes.</jats:p>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Availability and Implementation</jats:title>\n                  <jats:p>\n                    Code and minimally-curated activities available at:\n                    <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://github.com/moghelab/funcfetch\">https://github.com/moghelab/funcfetch</jats:ext-link>\n                    and\n                    <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://tools.moghelab.org/funczymedb\">https://tools.moghelab.org/funczymedb</jats:ext-link>\n                  </jats:p>\n                </jats:sec>","journal":null,"year":null,"id":604235,"datarank":0.26876392038420827,"base_score":1.791759469228055,"endowment":1.791759469228055,"self_citation_contribution":0.26876392038420827,"citation_network_contribution":0.0,"self_endowment_contribution":0.26876392038420827,"citer_contribution":0.0,"corpus_percentile":41.5,"corpus_rank":7610,"citation_count":5,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1550259,"name":"Xinyu Yuan","orcid":"0009-0006-9199-7652","position":1,"is_corresponding":false},{"id":1550261,"name":"Chesney Melissinos","orcid":"0009-0005-4443-8797","position":2,"is_corresponding":false},{"id":497005,"name":"Gaurav D. Moghe","orcid":"0000-0002-8761-064X","position":3,"is_corresponding":false},{"id":886132,"name":"Nathaniel Smith","orcid":"0000-0002-9215-8109","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"FuncFetch: An LLM-assisted workflow enables mining thousands of enzyme-substrate interactions from published manuscripts","abstract":"<jats:title>Abstract</jats:title>\n                <jats:sec>\n                  <jats:title>Motivation</jats:title>\n                  <jats:p>Thousands of genomes are publicly available, however, most genes in those genomes have poorly defined functions. This is partly due to a gap between previously published, experimentally-characterized protein activities and activities deposited in databases. This activity deposition is bottlenecked by the time-consuming biocuration process. The emergence of large language models (LLMs) presents an opportunity to speed up text-mining of protein activities for biocuration.</jats:p>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>We developed FuncFetch — a workflow that integrates NCBI E-Utilities, OpenAI’s GPT-4 and Zotero — to screen thousands of manuscripts and extract enzyme activities. Extensive validation revealed high precision and recall of GPT-4 in determining whether the abstract of a given paper indicates presence of a characterized enzyme activity in that paper. Provided the manuscript, FuncFetch extracted data such as species information, enzyme names, sequence identifiers, substrates and products, which were subjected to extensive quality analyses. Comparison of this workflow against a manually curated dataset of BAHD acyltransferase activities demonstrated a precision/recall of 0.86/0.64 in extracting substrates. We further deployed FuncFetch on nine large plant enzyme families. Screening 27,120 papers, FuncFetch retrieved 32,605 entries from 5547 selected papers. We also identified multiple extraction errors including incorrect associations, non-target enzymes, and hallucinations, which highlight the need for further manual curation. The BAHD activities were verified, resulting in a comprehensive functional fingerprint of this family and revealing that ∼70% of the experimentally characterized enzymes are uncurated in the public domain. FuncFetch represents an advance in biocuration and lays the groundwork for predicting functions of uncharacterized enzymes.</jats:p>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Availability and Implementation</jats:title>\n                  <jats:p>\n                    Code and minimally-curated activities available at:\n                    <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://github.com/moghelab/funcfetch\">https://github.com/moghelab/funcfetch</jats:ext-link>\n                    and\n                    <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://tools.moghelab.org/funczymedb\">https://tools.moghelab.org/funczymedb</jats:ext-link>\n                  </jats:p>\n                </jats:sec>","is_dataset_classified":null,"base_score":1.791759469228055,"endowment":1.791759469228055,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"23304386","pmcid":null,"openalex_id":"https://openalex.org/W4400966753","authors":[],"funders":[{"funder_name":"National Science Foundation","grant_id":"2310395","title":"Collaborative Research: TRTech-PGR: PlantSynBio: FuncZyme: Building a pipeline for rapid prediction and functional validation of plant enzyme activities"}],"total_grants":1,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2024,"count":3},{"year":2025,"count":2}],"oa_status":"green","license":"https://www.biorxiv.org/about/FAQ#license","oa_locations":[{"url":"https://www.biorxiv.org/content/biorxiv/early/2024/07/23/2024.07.22.604620.full.pdf","host_type":"repository"},{"url":"https://www.biorxiv.org/content/biorxiv/early/2024/07/23/2024.07.22.604620.full.pdf","host_type":"repository"},{"url":"https://syndication.highwire.org/content/doi/10.1101/2024.07.22.604620","host_type":"publisher"},{"url":"https://doi.org/10.1101/2024.07.22.604620","host_type":"repository"},{"url":"https://doi.org/10.1093/bioinformatics/btae756","host_type":""},{"url":"https://pubmed.ncbi.nlm.nih.gov/39718779","host_type":""},{"url":"http://dx.doi.org/10.1093/bioinformatics/btae756","host_type":""}],"fields_of_study":["Biomedical Text Mining and Ontologies","Semantic Web and Ontologies","Bioinformatics and Genomic Networks","0301 basic medicine","03 medical and health sciences","0206 medical engineering","02 engineering and technology"],"mesh_terms":[],"keywords":["Workflow","Substrate (aquarium)","Computer science","Substrate specificity","World Wide Web","Data science","Chemistry","Database","Enzyme","Biology","Biochemistry","Ecology","Original Paper","Data Mining","Computational Biology","Databases, Protein","Software","Enzymes"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Quality Education"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-29T23:36:47.770485Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}