{"doi":"10.1186/1471-2164-9-501","title":"Matching curated genome databases: a non trivial task","abstract":"<jats:title>Abstract</jats:title>\n          <jats:sec>\n            <jats:title>Background</jats:title>\n            <jats:p>Curated databases of completely sequenced genomes have been designed independently at the NCBI (RefSeq) and EBI (Genome Reviews) to cope with non-standard annotation found in the version of the sequenced genome that has been published by databanks GenBank/EMBL/DDBJ. These curation attempts were expected to review the annotations and to improve their pertinence when using them to annotate newly released genome sequences by homology to previously annotated genomes. However, we observed that such an uncoordinated effort has two unwanted consequences. First, it is not trivial to map the protein identifiers of the same sequence in both databases. Secondly, the two reannotated versions of the same genome differ at the level of their structural annotation.</jats:p>\n          </jats:sec>\n          <jats:sec>\n            <jats:title>Results</jats:title>\n            <jats:p>Here, we propose CorBank, a program devised to provide cross-referencing protein identifiers no matter what the level of identity is found between their matching sequences. Approximately 98% of the 1,983,258 amino acid sequences are matching, allowing instantaneous retrieval of their respective cross-references. CorBank further allows detecting any differences between the independently curated versions of the same genome. We found that the RefSeq and Genome Reviews versions are perfectly matching for only 50 of the 641 complete genomes we have analyzed. In all other cases there are differences occurring at the level of the coding sequence (CDS), and/or in the total number of CDS in the respective version of the same genome.</jats:p>\n            <jats:p>CorBank is freely accessible at <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" xlink:href=\"http://www.corbank.u-psud.fr\" ext-link-type=\"uri\">http://www.corbank.u-psud.fr</jats:ext-link>. The CorBank site contains also updated publication of the exhaustive results obtained by comparing RefSeq and Genome Reviews versions of each genome. Accordingly, this web site allows easy search of cross-references between RefSeq, Genome Reviews, and UniProt, for either a single CDS or a whole replicon.</jats:p>\n          </jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion</jats:title>\n            <jats:p>CorBank is very efficient in rapid detection of the numerous differences existing between RefSeq and Genome Reviews versions of the same curated genome. Although such differences are acceptable as reflecting different views, we suggest that curators of both genome databases could help reducing further divergence by agreeing on a minimal dialogue and attempting to publish the point of view of the other database whenever it is technically possible.</jats:p>\n          </jats:sec>","journal":"BMC Genomics","year":2008,"id":17079,"datarank":0.27159558671876644,"base_score":1.3862943611198906,"endowment":1.3862943611198906,"self_citation_contribution":0.20794415416798362,"citation_network_contribution":0.0636514325507828,"self_endowment_contribution":0.20794415416798362,"citer_contribution":0.0636514325507828,"corpus_percentile":42.33774270905856,"corpus_rank":7455,"citation_count":3,"citer_count":3,"citers_with_citation_signal":2,"citers_with_endowment":2,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":42.7083,"fair_percentile":30.242825607064017,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":123465,"name":"Matthieu Barba","orcid":null,"position":1,"is_corresponding":false},{"id":123467,"name":"Bernard Labedan","orcid":null,"position":2,"is_corresponding":false},{"id":123464,"name":"Stéphane Descorps-Declère","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"base_score":1.3862943611198906,"endowment":1.3862943611198906,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"18950477","pmcid":"PMC2596144","openalex_id":"https://openalex.org/W2143236684","authors":[],"funders":[{"funder_name":"French National Research Agency (ANR)","grant_id":"ANR-05-MMSA-0009","title":"Extraction optimisée d'informations pertinentes à partir de données complexes et hétérogènes issues de comparaisons génomiques"}],"total_grants":1,"fwci":0.1404,"citation_percentile":0.56640062,"influential_citations":0,"citation_trend":[{"year":2014,"count":1},{"year":2018,"count":1}],"oa_status":"gold","license":"CC BY","oa_locations":[{"url":"https://bmcgenomics.biomedcentral.com/counter/pdf/10.1186/1471-2164-9-501","host_type":"journal"},{"url":"https://bmcgenomics.biomedcentral.com/counter/pdf/10.1186/1471-2164-9-501","host_type":"GOLD"},{"url":"https://bmcgenomics.biomedcentral.com/counter/pdf/10.1186/1471-2164-9-501","host_type":"publisher"},{"url":"https://link.springer.com/content/pdf/10.1186/1471-2164-9-501.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1186/1471-2164-9-501","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/18950477","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/2596144","host_type":"repository"},{"url":"https://hal.science/hal-00353392","host_type":"repository"},{"url":"https://doaj.org/article/44b7fbe41ba04f8f8309ab7b21cfa305","host_type":"repository"},{"url":"http://www.biomedcentral.com/content/pdf/1471-2164-9-501.pdf","host_type":"BioMedCentral"},{"url":"http://www.biomedcentral.com/1471-2164/9/501","host_type":"BioMedCentral"},{"url":"http://www.biomedcentral.com/1471-2164/9/501/abstract","host_type":"BioMedCentral"},{"url":"https://europepmc.org/articles/PMC2596144","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC2596144?pdf=render","host_type":"Europe_PMC"},{"url":"https://bmcgenomics.biomedcentral.com/track/pdf/10.1186/1471-2164-9-501","host_type":""},{"url":"http://dx.doi.org/10.1186/1471-2164-9-501","host_type":""},{"url":"https://hal.science/hal-00353392v1","host_type":""},{"url":"https://dx.doi.org/10.1186/1471-2164-9-501","host_type":""}],"fields_of_study":["Genomics and Phylogenetic Studies","Genomics and Rare Diseases","Machine Learning in Bioinformatics","Biology","Medicine","Computer Science","0301 basic medicine","0303 health sciences","03 medical and health sciences","Computational Biology","Database Management Systems","Databases, Nucleic Acid","Databases, Protein","Genomics","Sequence Alignment"],"mesh_terms":["Database Management Systems","Sequence Alignment","Computational Biology","Genomics","Databases, Nucleic Acid","Databases, Protein"],"keywords":["RefSeq","Genome","GenBank","Annotation","Genome project","Identifier","Biology","Ensembl","Computational biology","Reference genome","Computer science","Database","Information retrieval","Genetics","Genomics","Gene","[SDV.BIBS] Life Sciences [q-bio]/Quantitative Methods [q-bio.QM]","QH426-470","[SDV.BBM] Life Sciences [q-bio]/Biochemistry, Molecular Biology","Database Management Systems","Databases, Nucleic Acid","Databases, Protein","Sequence Alignment","TP248.13-248.65","Biotechnology","Research Article"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-06-02T16:36:09.225580Z","pmid":"18950477","pmcid":"PMC2596144","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":52.5,"fair_a":60.0,"fair_i":25.0,"fair_r":33.3333,"fair_zscore":-0.5292,"fair_rationale":{"fair_score":42.71,"has_llm":true,"dimensions":{"F":{"name":"Findable","score":52.5,"criteria":[{"key":"f_has_doi","label":"Has a persistent DOI","kind":"deterministic","weight":1.0,"fraction":1.0,"signal":"DOI present","rationale":null},{"key":"f_repository_presence","label":"Indexed in repositories / literature DBs","kind":"deterministic","weight":1.0,"fraction":1.0,"signal":"datacite=0, pmcid=True, pmid=True","rationale":null},{"key":"f_persistent_ids","label":"Resolvable scholarly identifiers (OpenAlex)","kind":"deterministic","weight":0.5,"fraction":0.0,"signal":"no OpenAlex id","rationale":null},{"key":"f_metadata_richness","label":"Rich, machine-readable metadata","kind":"llm","weight":1.0,"fraction":0.25,"signal":null,"rationale":"The paper provides human-readable metadata (title, authors, abstract, etc.) but lacks explicit machine-readable metadata (e.g., structured data in a standard format like JSON-LD, RDF, or XML with FAIR-compliant vocabularies)."}]},"A":{"name":"Accessible","score":60.0,"criteria":[{"key":"a_open_access","label":"Open Access / files deposited","kind":"deterministic","weight":1.5,"fraction":0.5,"signal":"files/OA location present but not flagged OA","rationale":null},{"key":"a_retrievable","label":"Free full text retrievable","kind":"deterministic","weight":1.0,"fraction":1.0,"signal":"18 OA location(s)","rationale":null},{"key":"a_access_protocol","label":"Clear data/code access protocol","kind":"llm","weight":1.0,"fraction":0.5,"signal":null,"rationale":"The paper states that CorBank is freely accessible at a website and includes FTP URLs for downloading curated databases, but it does not provide a clear, step-by-step access protocol (e.g., authentication, download method) or a persistent identifier for the software."}]},"I":{"name":"Interoperable","score":25.0,"criteria":[{"key":"i_linked_data","label":"Linked datasets / DataCite relations","kind":"deterministic","weight":1.0,"fraction":0.0,"signal":"linked_datasets=0, datacite=0","rationale":null},{"key":"i_standard_ids","label":"References data via standard accessions","kind":"deterministic","weight":1.0,"fraction":0.0,"signal":"accessions=0, trials=0","rationale":null},{"key":"i_standards","label":"Standard formats, vocabularies & identifiers","kind":"llm","weight":1.0,"fraction":0.5,"signal":null,"rationale":"The work uses common identifiers (INSD, NCBI tax_id, UniProt, RefSeq, Genome Reviews) and standard formats (amino acid sequences), but it does not explicitly mention adherence to community-endorsed ontologies, terminologies, or file format standards for interoperability beyond basic sequence comparison."}]},"R":{"name":"Reusable","score":33.33,"criteria":[{"key":"r_license","label":"Clear, open reuse license","kind":"deterministic","weight":1.5,"fraction":0.0,"signal":"no license","rationale":null},{"key":"r_downloads","label":"Demonstrated reuse (downloads)","kind":"deterministic","weight":0.5,"fraction":0.0,"signal":"downloads=0","rationale":null},{"key":"r_version","label":"Versioned / maintained","kind":"deterministic","weight":0.5,"fraction":0.0,"signal":"no version chain","rationale":null},{"key":"r_dataset","label":"Classified as a data resource","kind":"deterministic","weight":0.5,"fraction":1.0,"signal":"is_dataset","rationale":null},{"key":"r_reusability","label":"Data-availability statement, license & reproducibility","kind":"llm","weight":2.0,"fraction":0.5,"signal":null,"rationale":"The paper is published under a CC BY license, which enables reuse, and provides a web site and FTP links for data/code, but it lacks a formal data-availability statement, explicit software citation, and details on reproducibility (e.g., full versioned code, dependencies, or containerized environment)."}]}},"suggestions":["Provide metadata in a machine-readable format (e.g., JSON-LD or XML with schema.org/Dublin Core) alongside the paper.","Document a clear access protocol for CorBank and the downloaded datasets, including expected input/output formats and any API endpoints.","Use community-standard vocabularies (e.g., EDAM for tools, OBO for ontologies) and specify file formats (e.g., FASTA, GFF3) to improve interoperability.","Add a formal data-availability statement specifying that all data and code are available under the CC BY license, with links to version-controlled repositories (e.g., GitHub, Zenodo).","Provide a software citation (e.g., with a DOI from Zenodo) and include a step-by-step reproducibility guide (e.g., conda environment, workflow script)."],"model":"deepseek/deepseek-v4-flash","agent_version":"fair_agent_v1","fulltext_source":"epmc_xml"},"fair_model":"deepseek/deepseek-v4-flash","fair_agent_version":"fair_agent_v1","fair_fulltext_source":"epmc_xml","fair_has_llm":true,"fair_computed_at":"2026-06-17T22:59:57.351716Z","clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}