{"doi":"10.1016/j.patter.2020.100064","title":"Bridging Domain and Data","abstract":"Dr. Anne Carpenter addresses her career path from cell biology toward computation. Why would a researcher move outside their comfort zone into a different field, from a domain into data science? What is the best way to bridge domain and data? What is challenging about moving from domain toward data? What is amazing about bridging domain and data? Dr. Anne Carpenter addresses her career path from cell biology toward computation. Why would a researcher move outside their comfort zone into a different field, from a domain into data science? What is the best way to bridge domain and data? What is challenging about moving from domain toward data? What is amazing about bridging domain and data? I am the daughter of an engineer. The sister of an engineer. I aced math classes and assessments. But in undergraduate and graduate school, I studied biology because engineering seemed too dry, too constrained. And frankly, no one ever suggested it to me, a young girl growing up on a farm in Indiana, more interested in reading the classics than tinkering with electronics. I’m now listed on a top-100 list of AI Leaders in Drug Discovery & Healthcare, and I lead the new National Institutes of Health-funded Center for Open Bioimage Analysis (COBA) together with Kevin Eliceiri of University of Wisconsin-Madison. I was just inducted into the AIMBE, the American Institute for Medical and Biological Engineering, representing the top 2% of engineers in the United States. What happened? And is it a great tragedy that I was steered away from engineering in my early years, or a blessing in disguise? Biology is messy. Not just in its physical methods—dissection, pulverizing living things into bits, working with chemicals. It’s also messy in its manner of approaching problems, and I mean that in a good way. Biology is about making connections between disparate pieces of data, about hypothesizing mechanisms when only a small fraction of a system is known. Like all sciences, it is a logic puzzle, but it’s a puzzle where defining which puzzle to solve is a major part of the work. Where it is quite likely there are many possible solutions. Where some of the input data are unknown, some are unknown by you (unless you have encyclopedic knowledge of the literature), and some of what is “known” is wrong. On top of all that, the data you have are almost certainly insufficient to solve the problem conclusively. While earning my PhD in cell biology, my hidden engineer emerged: after observing cell samples by eye for hours on the microscope, I had an intense desire to automate and quantify biology for my project. It became clear I was more enamored with how to answer a question rather than learning the answer. So I dove headfirst into coding in Pascal, then spent my postdoc almost entirely in MATLAB rather than at the bench. When I launched my own laboratory at the Broad Institute in 2007, I had no microscopes, no incubators, no pipettes. I had gone full computational. This year, the software project I started, CellProfiler, will hit its 10,000th citation, and aside from accelerating the research of thousands of biologists, my laboratory’s work has contributed to at least five clinical trials. For me, it’s been a fruitful transition from domain to data. What is the best way to bridge domain and data? I’m in no position to definitively answer that—probably no one is. I can, however, report my experience. I wrote CellProfiler in collaboration with Thouis Jones, a graduate student in computer science at MIT. What was so productive about this interaction is that each of us had the time and motivation to learn, say, 20% of the others’ field. I’ve come to believe that it’s impossible for one person to be sufficiently immersed in both biology and data science to be at the cutting edge of both, much less remain there over time. Being a deep and up-to-date expert in two fields is too much for one human brain, so one solution is to go with two brains. This requires a particula","journal":"Patterns","year":2020,"id":115729,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9467,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2020-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":11498,"name":"Anne E. Carpenter","orcid":"0000-0003-1555-8261","position":0,"is_corresponding":true}],"reference_count":0,"raw_metadata":null,"created_at":"2026-07-18T23:13:40.596604Z","pmid":"33205116","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}