Please note: This master’s thesis presentation will take place online.
Han Zhou, Master’s candidate
David R. Cheriton School of Computer Science
Supervisor: Professor Yang Lu
As biological analysis becomes increasingly mediated by foundation models and automated multi-step workflows, computational methods are used not only to process data, but also to generate outputs that shape downstream biological interpretation. These outputs may include integrated representations, inferred data alignments, similarity scores, retrieval rankings, and other intermediate products of analysis. They are often used to support biological candidates, such as predicted protein functions, putative homologous relationships, candidate cell-state correspondences, proposed gene-regulatory programs, annotations, or hypotheses about molecular interactions and biological mechanisms. However, if the intermediate computational outputs are ambiguous, unstable, or difficult to interpret, then the biological candidates derived from them may not be reliable evidence. This thesis argues that diagnostics are needed to test the meaning, reliability, and failure modes of computational outputs before they are used to support biological interpretation.
This diagnostic perspective is developed across two critical layers of artificial AI-driven biological analysis: the data layer and the model layer. First, at the data layer, this thesis presents SONATA, a diagnostic framework designed for diagonal multimodal single-cell data integration. In the absence of shared cells or features, multiple cross-modality alignments can appear computationally coherent yet remain biologically ambiguous. SONATA exposes these alternative integration solutions and quantifies mapping ambiguity, preventing users from treating unstable data alignments as definitive biological facts.
Second, at the model layer, the thesis introduces PLM-GUARD, a diagnostic suite that evaluates protein language models (PLMs) used in similarity search. PLM-GUARD scrutinizes model-derived similarity scores across biological fidelity, semantic validity, and manipulation safety. Its evaluations demonstrate that a model’s retrieval utility does not inherently imply evidential reliability, highlighting the need for diagnostic caution before interpreting embedding-space scores as true biological meaning.
Finally, this thesis points to a future paradigm at the agent layer. As autonomous AI agents begin to chain together complex analytical workflows, errors and ambiguities from early stages risk propagating silently. Agent-level diagnostics are therefore proposed as an indispensable requirement to ensure that intermediate computational candidates are robust enough to support downstream reasoning. Ultimately, the frameworks developed in this work shift the focus of computational biology from merely accelerating candidate generation to systematically validating outputs as trustworthy scientific evidence.