Sihang Zeng

I build AI systems that reason over a patient's history the way a clinician does — across years of messy, irregular records — and that can be trusted with the answer.

Boston, Massachusetts

About

I am a postdoctoral research fellow in the Department of Biomedical Informatics at Harvard Medical School, working with Tianxi Cai. My research develops machine learning methods for longitudinal electronic health records — representing how a patient's state evolves over years of sparse, irregularly sampled data, and turning that representation into predictions clinicians can interrogate rather than merely accept — and computational drug repurposing, where the goal is to find new indications for existing drugs by reasoning across multi-modal molecular, clinical, and real-world evidence.

I completed my PhD in Biomedical Informatics at the University of Washington in 2026, advised by Ruth Etzioni and Meliha Yetisgen, with Matthew Thompson and Hoifung Poon on my committee. My dissertation on trustworthy modeling of patient trajectories ran along three threads: continuous-time latent trajectory models for survival and risk prediction, multi-agent LLM frameworks that make temporal reasoning over raw clinical notes explicit and auditable, and benchmarks that test whether biomedical AI holds up under the conditions it will actually meet.

Earlier, I studied Electronic Engineering with a minor in Statistics at Tsinghua University. There I worked with Sheng Yu on biomedical term representation and large-scale knowledge graph construction, and with Bowen Zhou on language models for scientific and biomedical reasoning. That grounding in biomedical language still shapes how I think about clinical text as evidence.

I stay involved with both lines of work as a remote collaborator with the Tsinghua C3I lab and Frontis AI, on multi-agent LLM systems, reinforcement learning for reasoning, and benchmarks for AI-driven science.

Research interests

Five connected problems: representing a patient over time, acting on that representation, and making both hold up across the messy, heterogeneous data health systems actually produce.

01

Patient trajectory modeling

Continuous-time latent representations of how a patient's state moves through years of sparse, irregularly sampled EHR data — built so that the trajectory itself, not just the final score, is inspectable.

Longitudinal EHRSurvival analysisEarly detection

02

Computational drug repurposing

Finding new indications for existing drugs by reasoning jointly over molecular structure, target biology, and the real-world clinical record, rather than over any one of them alone.

Multi-modalTarget biologyReal-world evidence

03

Healthcare agents

LLM agent systems that can solve complex biomedical tasks or understand a patient's EHR before making any decisions — training-free where possible, self-evolving where it helps.

LLM agentsSelf-evolvingTemporal reasoningClinical decision support

04

EHR foundation models

Pretrained models over longitudinal patient records that transfer across tasks, sites, and populations, and whose behavior can be probed rather than taken on faith.

PretrainingTransferTrustworthy ML

05

Medical data harmonization

Making heterogeneous vocabularies, coding systems, and health systems comparable — the representation-learning problem that has to be solved before anything above generalizes.

TerminologiesKnowledge graphsInteroperability

Selected work

All publications

Nature Biomedical Engineering 2026

Phenotypic prediction of missense variants via deep contrastive learning

Jun Wen, Sihang Zeng, Clara-Lea Bonzel, Shilpa Nadimpalli Kobren, Jiangchuan Du, Yi Chai, Hao Wang, Meng Zhu, Siwei Chen, Fangwei Leng, Harrison G. Zhang, Katherine P. Liao, Kelly Cho, Isaac S. Kohane, Marinka Zitnik, Alexandre C. Pereira, Jun S. Liu, Tianxi Cai

PheMART connects missense variants to 4,179 clinical phenotypes in a shared metric space, combining protein language models, interaction networks, medical knowledge graphs, and EHR-derived signal to support rare disease diagnosis.

Contrastive learningVariant effectRare disease

Conference on Empirical Methods in Natural Language Processing (EMNLP) 2025

ReviewRL: towards automated scientific review with reinforcement learning

Sihang Zeng, Kai Tian, Kaiyan Zhang, Junqi Gao, Runze Liu, Sa Yang, Jingxuan Li, Xinwei Long, Jiaheng Ma, Biqing Qi, Bowen Zhou

Trains review generation with verifiable rewards so that automated reviews stay grounded in the paper rather than drifting into fluent generality.

Scientific AIRLVREvaluation

NeurIPS GenAI4Health Workshop 2025

Traj-CoA: patient trajectory modeling via chain-of-agents for lung cancer risk prediction

Sihang Zeng, Yujuan Fu, Sitong Zhou, Zixuan Yu, Lucas Jing Liu, Jun Wen, Matthew Thompson, Ruth Etzioni, Meliha Yetisgen

A chain of LLM agents reads a patient's notes in order and passes forward a compressed timeline, making zero-shot lung cancer risk prediction traceable to source evidence.

LLM agentsCancer screeningZero-shot

NeurIPS Datasets and Benchmarks Track 2024 Spotlight

UltraMedical: building specialized generalists in biomedicine

Kaiyan Zhang, Sihang Zeng, Ermo Hua, Ning Ding, Zhang-Ren Chen, Zhiyuan Ma, Haoxin Li, Ganqu Cui, Biqing Qi, Xuekai Zhu, et al.

A large biomedical instruction dataset and model suite that reaches domain specialization without giving up general capability.

Biomedical LLMsDatasetsInstruction tuning
  1. Joined Harvard Medical School as a Postdoctoral Research Fellow Position

    Working with Tianxi Cai in the Department of Biomedical Informatics on patient trajectory modeling and computational drug repurposing from multi-modal data.

  2. STRATOS-P published in JCO Precision Oncology Paper

    A validated genomic classifier for overall survival in metastatic prostate cancer, developed across 7,201 veterans.

  3. Completed PhD at the University of Washington Milestone

    Dissertation: “Towards Trustworthy Modeling of Patient Trajectory with Longitudinal Electronic Health Records.”

  4. Traj-Evolve preprint released Paper

    A self-evolving multi-agent system for lung cancer early detection that outperforms nine baselines over five years of patient history.

  5. CONTEXTOR accepted at ICML 2026 Paper

    Contextualized high-order contrastive learning for asymmetric relation inference.

  6. PheMART published in Nature Biomedical Engineering Paper

    Deep contrastive learning linking 5.1 million missense variants to 4,179 clinical phenotypes.

  7. TrajOnco preprint released Paper

    A training-free multi-agent LLM framework for multi-cancer early detection across 15 cancer types.

  8. MARTI accepted at ICLR 2026 Paper

    An open framework for reinforced training and inference in multi-agent LLM systems.