Sihang Zeng
Postdoctoral Research Fellow, Department of Biomedical Informatics, Harvard Medical School
Boston, Massachusetts · sihangzeng[at]gmail[dot]com
- Peer-reviewed
- 16
- Citations
- 653
- h-index
- 10
- i10-index
- 10
Citation counts from Google Scholar, August 2026. “Download PDF” opens your browser’s print dialog — choose “Save as PDF”.
Education
-
2023 – 2026
PhD, Biomedical Informatics and Medical Education
University of Washington
- Dissertation: “Towards Trustworthy Modeling of Patient Trajectory with Longitudinal Electronic Health Records” — interpretable and generalizable frameworks for patient trajectory modeling
- Committee: Ruth Etzioni (co-chair), Meliha Yetisgen (co-chair), Matthew Thompson, Hoifung Poon, Noemi Kreif
- Degree conferred June 2026
-
2019 – 2023
BEng, Electronic Engineering — minor in Statistics
Tsinghua University
- Research on biomedical term representation and algorithmic knowledge graph construction
Research positions
-
2026 – present
Postdoctoral Research Fellow
Department of Biomedical Informatics, Harvard Medical School
Working with Tianxi Cai on patient trajectory modeling over longitudinal EHR and on computational drug repurposing from multi-modal clinical and biological data.
-
2025 – 2026
Machine Learning Intern
Patient-journey foundation models and multi-agent risk prediction over large-scale de-identified clinical data, with results presented at ASCO and ISPOR.
-
2023 – 2026
Graduate Research Assistant
University of Washington and Fred Hutchinson Cancer Center
Trajectory modeling for cancer early detection and survival prediction in the Etzioni Lab and UW BioNLP.
-
2020 – 2023
Undergraduate Researcher
Tsinghua University
Biomedical term representation learning and construction of the BIOS knowledge graph with Sheng Yu; language models for scientific and biomedical reasoning with Bowen Zhou.
Awards and honors
-
2025
1st place, ChemoTimelines 2025 shared task (subtask 2)
Clinical NLP Workshop, ACL — best official score of 0.678
-
2024
Spotlight, NeurIPS Datasets and Benchmarks Track
UltraMedical: building specialized generalists in biomedicine
Publications
Annotated listJournal articles
5- 2026
Phenotypic prediction of missense variants via deep contrastive learning
Nature Biomedical Engineering
- 2026
- 2025
The role of Whole Health in enhancing tobacco cessation outcomes for veterans: a retrospective cohort study
Journal of General Internal Medicine
- 2024
CoRTEx: contrastive learning for representing terms via explanations with applications on constructing biomedical knowledge graphs
Journal of the American Medical Informatics Association 31(9), 1912–1920
- 2021
A feasibility study of 2-D microwave thorax imaging based on the supervised descent method
Electronics 10(3), 352
Conference and workshop papers
11- 2026
CONTEXTOR: contextualized high-order contrastive learning
International Conference on Machine Learning (ICML)
- 2026
MARTI: a framework for multi-agent LLM systems reinforced training and inference
International Conference on Learning Representations (ICLR)
- 2025
ReviewRL: towards automated scientific review with reinforcement learning
Conference on Empirical Methods in Natural Language Processing (EMNLP) Main conference
- 2025
TrajSurv: learning continuous latent trajectories from electronic health records for trustworthy survival prediction
Machine Learning for Healthcare Conference (MLHC) PMLR 298
- 2025
Traj-CoA: patient trajectory modeling via chain-of-agents for lung cancer risk prediction
NeurIPS GenAI4Health Workshop
- 2025
UW-BioNLP at ChemoTimelines 2025: thinking, fine-tuning, and dictionary-enhanced LLM systems for chemotherapy timeline extraction
Clinical NLP Workshop (ACL) pp. 40–56 1st place, subtask 2
- 2025
Scalability of LLM-based multi-agent systems for scientific code generation: a preliminary study
MathNLP Workshop (EMNLP)
- 2024
UltraMedical: building specialized generalists in biomedicine
NeurIPS Datasets and Benchmarks Track Spotlight
- 2024
Large language models as biomedical hypothesis generators: a comprehensive evaluation
Conference on Language Modeling (COLM)
- 2023
- 2022
Automatic biomedical term clustering by learning fine-grained term representations
BioNLP Workshop (ACL) pp. 91–96
Preprints and working papers
9- 2026
- 2026
- 2026
- 2026
- 2026
- 2025
- 2025
- 2023
- 2022
Conference abstracts
9- 2026
Predicting multi-cancer risk from EHR data using multi-agent LLMs
Journal of Clinical Oncology ASCO Annual Meeting
- 2026
STRATOS-P clinical model for prognosis and selection of patients with metastatic hormone-sensitive prostate cancer for intermittent therapy
Journal of Clinical Oncology ASCO Annual Meeting
- 2026
Development of an oncology generative AI foundation model trained on more than a million longitudinal patient journeys across the United States
Journal of Clinical Oncology ASCO Annual Meeting
- 2026
MSR67 — Zero-shot lung cancer risk prediction from longitudinal electronic health records with a chain-of-agents framework
Value in Health ISPOR
- 2026
MSR68 — Characterizing oncology patient journeys and health state transitions using a data-driven Markov transition matrix in large-scale electronic health records
Value in Health ISPOR
- 2026
MSR172 — Can a generative patient journey foundation model alleviate the burden of cancer screening?
Value in Health ISPOR
- 2026
RWD140 — Patient journey foundational model for scalable imputation of missing units of measurement in electronic health records data
Value in Health ISPOR
- 2026
Adapting the Global Burden of Disease Healthcare Access and Quality Index for the Veterans Health Administration: a feasibility study
AcademyHealth Annual Research Meeting
- 2025
Population-level tobacco cessation outcomes associated with implementing Whole Health at the Veterans Health Administration
AcademyHealth Annual Research Meeting
Research areas and affiliations
Research interests
- Patient trajectory modeling
- Computational drug repurposing
- Healthcare agents
- EHR foundation models
- Medical data harmonization
Methods
- Longitudinal and survival modeling
- LLM agents and multi-agent systems
- Contrastive and representation learning
- Reinforcement learning with verifiable rewards
Research groups
- Cai Lab / CELEHS, Harvard Medical School
- Etzioni Lab, Fred Hutchinson Cancer Center
- UW BioNLP, University of Washington
Remote collaborations
- Tsinghua C3I
- Frontis AI