01
Patient trajectory modeling
Continuous-time latent representations of how a patient's state moves through years of sparse, irregularly sampled EHR data — built so that the trajectory itself, not just the final score, is inspectable.
I build AI systems that reason over a patient's history the way a clinician does — across years of messy, irregular records — and that can be trusted with the answer.
I am a postdoctoral research fellow in the Department of Biomedical Informatics at Harvard Medical School, working with Tianxi Cai. My research develops machine learning methods for longitudinal electronic health records — representing how a patient's state evolves over years of sparse, irregularly sampled data, and turning that representation into predictions clinicians can interrogate rather than merely accept — and computational drug repurposing, where the goal is to find new indications for existing drugs by reasoning across multi-modal molecular, clinical, and real-world evidence.
I completed my PhD in Biomedical Informatics at the University of Washington in 2026, advised by Ruth Etzioni and Meliha Yetisgen, with Matthew Thompson and Hoifung Poon on my committee. My dissertation on trustworthy modeling of patient trajectories ran along three threads: continuous-time latent trajectory models for survival and risk prediction, multi-agent LLM frameworks that make temporal reasoning over raw clinical notes explicit and auditable, and benchmarks that test whether biomedical AI holds up under the conditions it will actually meet.
Earlier, I studied Electronic Engineering with a minor in Statistics at Tsinghua University. There I worked with Sheng Yu on biomedical term representation and large-scale knowledge graph construction, and with Bowen Zhou on language models for scientific and biomedical reasoning. That grounding in biomedical language still shapes how I think about clinical text as evidence.
I stay involved with both lines of work as a remote collaborator with the Tsinghua C3I lab and Frontis AI, on multi-agent LLM systems, reinforcement learning for reasoning, and benchmarks for AI-driven science.
Five connected problems: representing a patient over time, acting on that representation, and making both hold up across the messy, heterogeneous data health systems actually produce.
01
Continuous-time latent representations of how a patient's state moves through years of sparse, irregularly sampled EHR data — built so that the trajectory itself, not just the final score, is inspectable.
02
Finding new indications for existing drugs by reasoning jointly over molecular structure, target biology, and the real-world clinical record, rather than over any one of them alone.
03
LLM agent systems that can solve complex biomedical tasks or understand a patient's EHR before making any decisions — training-free where possible, self-evolving where it helps.
04
Pretrained models over longitudinal patient records that transfer across tasks, sites, and populations, and whose behavior can be probed rather than taken on faith.
05
Making heterogeneous vocabularies, coding systems, and health systems comparable — the representation-learning problem that has to be solved before anything above generalizes.
PheMART connects missense variants to 4,179 clinical phenotypes in a shared metric space, combining protein language models, interaction networks, medical knowledge graphs, and EHR-derived signal to support rare disease diagnosis.
Pairs an experience pool of retrieved similar patients with RL-optimized agent coordination, beating nine baselines over five years of patient history.
Trains review generation with verifiable rewards so that automated reviews stay grounded in the paper rather than drifting into fluent generality.
Learns a continuous-time latent trajectory per patient, so survival predictions come with an inspectable account of how risk accumulated.
A chain of LLM agents reads a patient's notes in order and passes forward a compressed timeline, making zero-shot lung cancer risk prediction traceable to source evidence.
A large biomedical instruction dataset and model suite that reaches domain specialization without giving up general capability.
Joined Harvard Medical School as a Postdoctoral Research Fellow Position
Working with Tianxi Cai in the Department of Biomedical Informatics on patient trajectory modeling and computational drug repurposing from multi-modal data.
STRATOS-P published in JCO Precision Oncology Paper
A validated genomic classifier for overall survival in metastatic prostate cancer, developed across 7,201 veterans.
Completed PhD at the University of Washington Milestone
Dissertation: “Towards Trustworthy Modeling of Patient Trajectory with Longitudinal Electronic Health Records.”
Traj-Evolve preprint released Paper
A self-evolving multi-agent system for lung cancer early detection that outperforms nine baselines over five years of patient history.
CONTEXTOR accepted at ICML 2026 Paper
Contextualized high-order contrastive learning for asymmetric relation inference.
PheMART published in Nature Biomedical Engineering Paper
Deep contrastive learning linking 5.1 million missense variants to 4,179 clinical phenotypes.
TrajOnco preprint released Paper
A training-free multi-agent LLM framework for multi-cancer early detection across 15 cancer types.
MARTI accepted at ICLR 2026 Paper
An open framework for reinforced training and inference in multi-agent LLM systems.