{"doi":"10.1093/nar/gkq1157","title":"MIPS: curated databases and comprehensive secondary data resources in 2010","abstract":"The Munich Information Center for Protein Sequences (MIPS at the Helmholtz Center for Environmental Health, Neuherberg, Germany) has many years of experience in providing annotated collections of biological data. Selected data sets of high relevance, such as model genomes, are subjected to careful manual curation, while the bulk of high-throughput data is annotated by automatic means. High-quality reference resources developed in the past and still actively maintained include Saccharomyces cerevisiae, Neurospora crassa and Arabidopsis thaliana genome databases as well as several protein interaction data sets (MPACT, MPPI and CORUM). More recent projects are PhenomiR, the database on microRNA-related phenotypes, and MIPS PlantsDB for integrative and comparative plant genome research. The interlinked resources SIMAP and PEDANT provide homology relationships as well as up-to-date and consistent annotation for 38,000,000 protein sequences. PPLIPS and CCancer are versatile tools for proteomics and functional genomics interfacing to a database of compilations from gene lists extracted from literature. A novel literature-mining tool, EXCERBT, gives access to structured information on classified relations between genes, proteins, phenotypes and diseases extracted from Medline abstracts by semantic analysis. All databases described here, as well as the detailed descriptions of our projects can be accessed through the MIPS WWW server (http://mips.helmholtz-muenchen.de).","journal":"Nucleic Acids Research","year":2010,"id":8797,"datarank":0.6966586348712059,"base_score":4.6443908991413725,"endowment":4.6443908991413725,"self_citation_contribution":0.6966586348712059,"citation_network_contribution":0.0,"self_endowment_contribution":0.6966586348712059,"citer_contribution":0.0,"corpus_percentile":71.31585054537015,"corpus_rank":3709,"citation_count":103,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.954,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2010-11-24","fair_score":38.3333,"fair_percentile":19.426048565121413,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":5457,"name":"Andreas Ruepp","orcid":"0000-0003-1705-3515","position":1,"is_corresponding":false},{"id":75597,"name":"Thomas Rattei","orcid":"0000-0002-0592-7791","position":3,"is_corresponding":false},{"id":75598,"name":"Mathias Walter","orcid":null,"position":4,"is_corresponding":false},{"id":65055,"name":"Dmitrij Frishman","orcid":"0000-0002-9006-4707","position":5,"is_corresponding":false},{"id":2350,"name":"Karsten Suhre","orcid":"0000-0001-9638-3912","position":6,"is_corresponding":false},{"id":75599,"name":"Manuel Spannagl","orcid":null,"position":7,"is_corresponding":false},{"id":75600,"name":"Klaus F.X. Mayer","orcid":null,"position":8,"is_corresponding":false},{"id":71445,"name":"Volker Stümpflen","orcid":null,"position":9,"is_corresponding":false},{"id":75601,"name":"Alexey Antonov","orcid":null,"position":10,"is_corresponding":false},{"id":75602,"name":"Hans‐Werner Mewes","orcid":"0000-0002-9713-6398","position":11,"is_corresponding":false},{"id":42,"name":"Fabian Joachim Theis","orcid":"0000-0002-2419-1943","position":12,"is_corresponding":false},{"id":75603,"name":"Mathias C. Walter","orcid":"0000-0003-3012-2626","position":13,"is_corresponding":false},{"id":75604,"name":"M. Spannagl","orcid":"0000-0003-0701-7035","position":14,"is_corresponding":false},{"id":45544,"name":"Klaus Mayer","orcid":"0000-0001-6484-1077","position":15,"is_corresponding":false},{"id":75605,"name":"Alexey V. Antonov","orcid":"0000-0002-6909-8902","position":16,"is_corresponding":false},{"id":75596,"name":"H. Werner Mewes","orcid":null,"position":0,"is_corresponding":true}],"reference_count":22,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-03-01T18:20:47.508186Z","pmid":"21109531","pmcid":"PMC3013725","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":40.0,"fair_a":42.5,"fair_i":37.5,"fair_r":33.3333,"fair_zscore":-0.7998,"fair_rationale":{"fair_score":38.33,"has_llm":true,"dimensions":{"F":{"name":"Findable","score":40.0,"criteria":[{"key":"f_has_doi","label":"Has a persistent DOI","kind":"deterministic","weight":1.0,"fraction":1.0,"signal":"DOI present","rationale":null},{"key":"f_repository_presence","label":"Indexed in repositories / literature DBs","kind":"deterministic","weight":1.0,"fraction":1.0,"signal":"datacite=0, pmcid=True, pmid=True","rationale":null},{"key":"f_persistent_ids","label":"Resolvable scholarly identifiers (OpenAlex)","kind":"deterministic","weight":0.5,"fraction":0.0,"signal":"no OpenAlex id","rationale":null},{"key":"f_metadata_richness","label":"Rich, machine-readable metadata","kind":"llm","weight":1.0,"fraction":0.0,"signal":null,"rationale":"The paper does not provide any machine-readable metadata (e.g., structured metadata in standard schemas) either in the text or through reference to a metadata record, and describes databases accessible via web pages without FAIR metadata distribution."}]},"A":{"name":"Accessible","score":42.5,"criteria":[{"key":"a_open_access","label":"Open Access / files deposited","kind":"deterministic","weight":1.5,"fraction":1.0,"signal":"Open Access","rationale":null},{"key":"a_retrievable","label":"Free full text retrievable","kind":"deterministic","weight":1.0,"fraction":0.0,"signal":"0 OA location(s)","rationale":null},{"key":"a_access_protocol","label":"Clear data/code access protocol","kind":"llm","weight":1.0,"fraction":0.25,"signal":null,"rationale":"The paper lists URLs for each database and mentions web services (e.g., BioMOBY, SIMAP Web-Services) but does not specify standard access protocols, authentication requirements, or provide a clear, unambiguous protocol for programmatic access to data/code."}]},"I":{"name":"Interoperable","score":37.5,"criteria":[{"key":"i_linked_data","label":"Linked datasets / DataCite relations","kind":"deterministic","weight":1.0,"fraction":0.0,"signal":"linked_datasets=0, datacite=0","rationale":null},{"key":"i_standard_ids","label":"References data via standard accessions","kind":"deterministic","weight":1.0,"fraction":0.0,"signal":"accessions=0, trials=0","rationale":null},{"key":"i_standards","label":"Standard formats, vocabularies & identifiers","kind":"llm","weight":1.0,"fraction":0.75,"signal":null,"rationale":"The paper uses standard identifiers (e.g., Gene Ontology, EC numbers, RefSeq), standard formats are implied (e.g., sequence similarity, relations), and references standards like GO, KEGG, IntAct, but does not explicitly state adherence to machine-readable standard formats (e.g., RDF, JSON-LD) or vocabularies for all resources."}]},"R":{"name":"Reusable","score":33.33,"criteria":[{"key":"r_license","label":"Clear, open reuse license","kind":"deterministic","weight":1.5,"fraction":0.0,"signal":"no license","rationale":null},{"key":"r_downloads","label":"Demonstrated reuse (downloads)","kind":"deterministic","weight":0.5,"fraction":0.0,"signal":"downloads=0","rationale":null},{"key":"r_version","label":"Versioned / maintained","kind":"deterministic","weight":0.5,"fraction":0.0,"signal":"no version chain","rationale":null},{"key":"r_dataset","label":"Classified as a data resource","kind":"deterministic","weight":0.5,"fraction":1.0,"signal":"is_dataset","rationale":null},{"key":"r_reusability","label":"Data-availability statement, license & reproducibility","kind":"llm","weight":2.0,"fraction":0.5,"signal":null,"rationale":"The paper states a Creative Commons Attribution Non-Commercial License (CC BY-NC 2.5) for the article, but does not provide a clear data-availability statement for the databases themselves, lacks explicit code or data licensing, and does not describe reproducibility procedures beyond mentioning web access and service availability."}]}},"suggestions":["Provide the databases' metadata as machine-readable JSON-LD or RDF using a standard schema (e.g., DCAT, schema.org) and reference a DOI for each dataset.","Specify a standard machine-access protocol (e.g., REST/SPARQL endpoints) for each resource, including authentication methods and rate limits if any, in a dedicated 'Access' section.","Declare explicit adherence to community-standard formats (e.g., FASTA, GFF3, SBML) and vocabularies (e.g., OBO Foundry ontologies) for all data types, and use persistent identifiers (DOI, ORCID) for authors and resources.","Add a data-availability statement that clearly licenses all underlying data and code under a standard open license (e.g., CC0 or MIT) and describes how to reproduce the databases, including versioning and archival (e.g., in a repository like Zenodo)."],"model":"deepseek/deepseek-v4-flash","agent_version":"fair_agent_v2","fulltext_source":"epmc_xml"},"fair_model":"deepseek/deepseek-v4-flash","fair_agent_version":"fair_agent_v2","fair_fulltext_source":"epmc_xml","fair_has_llm":true,"fair_computed_at":"2026-06-18T00:39:49.227518Z","clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}