{"doi":"10.64898/2026.01.05.697809","title":"Can AI Conduct Autonomous Scientific Research? Case Studies on Two Real-World Tasks","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>Recent advances in artificial intelligence (AI) have prompted claims about autonomous “AI scientists,” yet systematic evaluations of these capabilities remain scarce. This exploratory study investigates whether current AI frameworks can execute scientific research tasks beyond isolated demonstrations. We tested eight open-source AI frameworks (Agent Laboratory, AutoGen, BabyAGI, GPT Researcher, MOOSE-Chem2, SciAgents, SciMON, and Virtual Lab) on two tasks that aimed to reproduce research on algorithm development from recent papers in uncertainty quantification and protein interaction discovery. In our evaluation, no framework completed a full research cycle from literature understanding through computational execution to validated results and scientific paper writing. While all systems showed competence in conceptual tasks such as planning and summarization, they consistently failed at robust implementation. Every framework produced sophisticated hallucinations. Deployment proved demanding, requiring substantial debugging and technical expertise, which undermines common claims about the democratization of science with AI. Despite these limitations, the frameworks showed promise as research assistants for methodological planning and ideation under careful human supervision. Our findings suggest that the explored AI systems cannot yet autonomously conduct scientific research, but may provide real value for specific subtasks within the research workflow. We offer preliminary observations to help researchers and developers better understand the gap between advertised and actual capabilities of AI in science.</jats:p>","journal":null,"year":null,"id":619445,"datarank":0.16479184330021646,"base_score":1.0986122886681096,"endowment":1.0986122886681096,"self_citation_contribution":0.16479184330021646,"citation_network_contribution":0.0,"self_endowment_contribution":0.16479184330021646,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1598540,"name":"Harsh B. Anadkat","orcid":null,"position":1,"is_corresponding":false},{"id":1598541,"name":"Kiran K. Athimoolam","orcid":null,"position":2,"is_corresponding":false},{"id":1598542,"name":"Harsh Bhardwaj","orcid":null,"position":3,"is_corresponding":false},{"id":1598543,"name":"Trishul Chowdhury","orcid":null,"position":4,"is_corresponding":false},{"id":1598544,"name":"Shengtao Gao","orcid":null,"position":5,"is_corresponding":false},{"id":1598545,"name":"Purva K. Kamat","orcid":null,"position":6,"is_corresponding":false},{"id":1598546,"name":"Vishwadeepsinh Makwana","orcid":null,"position":7,"is_corresponding":false},{"id":1598547,"name":"Mohammed H. Shariff","orcid":null,"position":8,"is_corresponding":false},{"id":89571,"name":"Amitesh Badkul","orcid":"0009-0000-9207-1954","position":9,"is_corresponding":false},{"id":1235972,"name":"Lei Xie","orcid":"0000-0002-7669-1886","position":10,"is_corresponding":false},{"id":834096,"name":"Anton V. Sinitskiy","orcid":"0000-0002-8610-0162","position":11,"is_corresponding":false},{"id":1598539,"name":"Shreyansh Agrawal","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Can AI Conduct Autonomous Scientific Research? Case Studies on Two Real-World Tasks","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>Recent advances in artificial intelligence (AI) have prompted claims about autonomous “AI scientists,” yet systematic evaluations of these capabilities remain scarce. This exploratory study investigates whether current AI frameworks can execute scientific research tasks beyond isolated demonstrations. We tested eight open-source AI frameworks (Agent Laboratory, AutoGen, BabyAGI, GPT Researcher, MOOSE-Chem2, SciAgents, SciMON, and Virtual Lab) on two tasks that aimed to reproduce research on algorithm development from recent papers in uncertainty quantification and protein interaction discovery. In our evaluation, no framework completed a full research cycle from literature understanding through computational execution to validated results and scientific paper writing. While all systems showed competence in conceptual tasks such as planning and summarization, they consistently failed at robust implementation. Every framework produced sophisticated hallucinations. Deployment proved demanding, requiring substantial debugging and technical expertise, which undermines common claims about the democratization of science with AI. Despite these limitations, the frameworks showed promise as research assistants for methodological planning and ideation under careful human supervision. Our findings suggest that the explored AI systems cannot yet autonomously conduct scientific research, but may provide real value for specific subtasks within the research workflow. We offer preliminary observations to help researchers and developers better understand the gap between advertised and actual capabilities of AI in science.</jats:p>","is_dataset_classified":null,"base_score":0.0,"endowment":0.0,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19767382","pmcid":null,"openalex_id":null,"authors":[],"funders":[],"total_grants":0,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[],"oa_status":"green","license":"cc-by","oa_locations":[{"url":"https://www.biorxiv.org/content/biorxiv/early/2026/01/06/2026.01.05.697809.full.pdf","host_type":"repository"},{"url":"https://syndication.highwire.org/content/doi/10.64898/2026.01.05.697809","host_type":"publisher"}],"fields_of_study":[],"mesh_terms":[],"keywords":[],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-03T07:12:22.544692Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}