Sihang Zeng

Postdoctoral Research Fellow, Department of Biomedical Informatics, Harvard Medical School
Boston, Massachusetts · sihangzeng[at]gmail[dot]com

Peer-reviewed
16
Citations
653
h-index
10
i10-index
10

Citation counts from Google Scholar, August 2026. “Download PDF” opens your browser’s print dialog — choose “Save as PDF”.

Education

  1. 2023 – 2026

    PhD, Biomedical Informatics and Medical Education

    University of Washington

  2. 2019 – 2023

    BEng, Electronic Engineering — minor in Statistics

    Tsinghua University

    • Research on biomedical term representation and algorithmic knowledge graph construction

Research positions

  1. 2026 – present

    Postdoctoral Research Fellow

    Department of Biomedical Informatics, Harvard Medical School

    Working with Tianxi Cai on patient trajectory modeling over longitudinal EHR and on computational drug repurposing from multi-modal clinical and biological data.

  2. 2025 – 2026

    Machine Learning Intern

    Truveta

    Patient-journey foundation models and multi-agent risk prediction over large-scale de-identified clinical data, with results presented at ASCO and ISPOR.

  3. 2023 – 2026

    Graduate Research Assistant

    University of Washington and Fred Hutchinson Cancer Center

    Trajectory modeling for cancer early detection and survival prediction in the Etzioni Lab and UW BioNLP.

  4. 2020 – 2023

    Undergraduate Researcher

    Tsinghua University

    Biomedical term representation learning and construction of the BIOS knowledge graph with Sheng Yu; language models for scientific and biomedical reasoning with Bowen Zhou.

Awards and honors

  1. 2025

    1st place, ChemoTimelines 2025 shared task (subtask 2)

    Clinical NLP Workshop, ACL — best official score of 0.678

  2. 2024

    Spotlight, NeurIPS Datasets and Benchmarks Track

    UltraMedical: building specialized generalists in biomedicine

Publications

Annotated list

Journal articles

5
  1. 2026

    Phenotypic prediction of missense variants via deep contrastive learning

    Jun Wen, Sihang Zeng, Clara-Lea Bonzel, Shilpa Nadimpalli Kobren, Jiangchuan Du, Yi Chai, Hao Wang, Meng Zhu, Siwei Chen, Fangwei Leng, Harrison G. Zhang, Katherine P. Liao, Kelly Cho, Isaac S. Kohane, Marinka Zitnik, Alexandre C. Pereira, Jun S. Liu, Tianxi Cai

    Nature Biomedical Engineering

  2. 2026

    Genomic classification to predict survival in metastatic prostate cancer: development of Somatic Tumor Risk Assessment for Overall Survival–Prostate (STRATOS-P)

    Martin W. Schoen, Jiannong Li, Sihang Zeng, Heena Desai, Ryan Hausler, Candace L. Haroldsen, Lukas Owens, Luca F. Valle, Ruth B. Etzioni, Timothy R. Rebbeck, Brent S. Rose, Michael J. Kelley, R. Bruce Montgomery, Nicholas G. Nickols, Matthew B. Rettig, Kosj Yamoah, Kara N. Maxwell, Isla P. Garraway

    JCO Precision Oncology

  3. 2025

    The role of Whole Health in enhancing tobacco cessation outcomes for veterans: a retrospective cohort study

    Sihang Zeng, Scott S. Coggeshall, Ethan W. Rosser, Stephanie L. Taylor, Diana J. Burgess, Gang Luo, Steven B. Zeliadt

    Journal of General Internal Medicine

  4. 2024

    CoRTEx: contrastive learning for representing terms via explanations with applications on constructing biomedical knowledge graphs

    Huaiyuan Ying, Zhengyun Zhao, Yang Zhao, Sihang Zeng, Sheng Yu

    Journal of the American Medical Informatics Association 31(9), 1912–1920

  5. 2021

    A feasibility study of 2-D microwave thorax imaging based on the supervised descent method

    Haolin Zhang, Maokun Li, Fan Yang, Shenheng Xu, Yan Yin, Hongyu Zhou, Yubo Yang, Sihang Zeng, Jianchong Shao

    Electronics 10(3), 352

Conference and workshop papers

11
  1. 2026

    CONTEXTOR: contextualized high-order contrastive learning

    Ze Cai, Hanzhe Liang, Sihang Zeng, Binbin Zhou, Jun Wen

    International Conference on Machine Learning (ICML)

  2. 2026

    MARTI: a framework for multi-agent LLM systems reinforced training and inference

    Kaiyan Zhang, Kai Tian, Runze Liu, Sihang Zeng, Xuekai Zhu, Guoli Jia, Yuchen Fan, Xingtai Lv, Yuxin Zuo, Che Jiang, et al.

    International Conference on Learning Representations (ICLR)

  3. 2025

    ReviewRL: towards automated scientific review with reinforcement learning

    Sihang Zeng, Kai Tian, Kaiyan Zhang, Junqi Gao, Runze Liu, Sa Yang, Jingxuan Li, Xinwei Long, Jiaheng Ma, Biqing Qi, Bowen Zhou

    Conference on Empirical Methods in Natural Language Processing (EMNLP) Main conference

  4. 2025

    TrajSurv: learning continuous latent trajectories from electronic health records for trustworthy survival prediction

    Sihang Zeng, Lucas Jing Liu, Jun Wen, Meliha Yetisgen, Ruth Etzioni, Gang Luo

    Machine Learning for Healthcare Conference (MLHC) PMLR 298

  5. 2025

    Traj-CoA: patient trajectory modeling via chain-of-agents for lung cancer risk prediction

    Sihang Zeng, Yujuan Fu, Sitong Zhou, Zixuan Yu, Lucas Jing Liu, Jun Wen, Matthew Thompson, Ruth Etzioni, Meliha Yetisgen

    NeurIPS GenAI4Health Workshop

  6. 2025

    UW-BioNLP at ChemoTimelines 2025: thinking, fine-tuning, and dictionary-enhanced LLM systems for chemotherapy timeline extraction

    Tianmai M. Zhang, Zhaoyi Sun, Sihang Zeng, Chenxi Li, Neil F. Abernethy, Barbara D. Lam, Fei Xia, Meliha Yetisgen

    Clinical NLP Workshop (ACL) pp. 40–56 1st place, subtask 2

  7. 2025

    Scalability of LLM-based multi-agent systems for scientific code generation: a preliminary study

    Yuru Wang, Kaiyan Zhang, Kai Tian, Sihang Zeng, Xingtai Lv, Ning Ding, Biqing Qi, Bowen Zhou

    MathNLP Workshop (EMNLP)

  8. 2024

    UltraMedical: building specialized generalists in biomedicine

    Kaiyan Zhang, Sihang Zeng, Ermo Hua, Ning Ding, Zhang-Ren Chen, Zhiyuan Ma, Haoxin Li, Ganqu Cui, Biqing Qi, Xuekai Zhu, et al.

    NeurIPS Datasets and Benchmarks Track Spotlight

  9. 2024

    Large language models as biomedical hypothesis generators: a comprehensive evaluation

    Biqing Qi, Kaiyan Zhang, Kai Tian, Haoxiang Li, Zhang-Ren Chen, Sihang Zeng, Ermo Hua, Jinfang Hu, Bowen Zhou

    Conference on Language Modeling (COLM)

  10. 2023

    Large language models are zero shot hypothesis proposers

    Biqing Qi, Kaiyan Zhang, Haoxiang Li, Kai Tian, Sihang Zeng, Zhang-Ren Chen, Bowen Zhou

    NeurIPS Instruction Tuning Workshop

  11. 2022

    Automatic biomedical term clustering by learning fine-grained term representations

    Sihang Zeng, Zheng Yuan, Sheng Yu

    BioNLP Workshop (ACL) pp. 91–96

Preprints and working papers

9
  1. 2026
  2. 2026

    NatureBench: can coding agents match the published SOTA of Nature-family papers?

    Yuru Wang, Lejun Cheng, Yuxin Zuo, Sihang Zeng, Bingxiang He, Che Jiang, Junlin Yang, Yuchong Wang, Kaikai Zhao, et al.

    arXiv:2606.24530

  3. 2026

    EpiEvolve: self-evolving agents for streaming pandemic forecasting under regime shifts

    Yiming Lu, Sihang Zeng, Zhengxu Tang, Max Lau, Fei Liu, Wei Jin

    arXiv:2606.05513

  4. 2026

    TrajOnco: a multi-agent framework for temporal reasoning over longitudinal EHR for multi-cancer early detection

    Sihang Zeng, Young Won Kim, Wilson Lau, Ehsan Alipour, Ruth Etzioni, Meliha Yetisgen, Anand Oka

    arXiv:2604.10386

  5. 2026

    Self-improving agents in the era of experience: a survey of self- to meta-evolution

    Che Jiang, Jincheng Zhong, Yu Fu, Kai Tian, Junlin Yang, Kaikai Zhao, Yuchong Wang, Tianwei Luo, Weizhi Wang, Yuxin Zuo, Guoli Jia, Xingtai Lv, Dianqiao Lei, Sihang Zeng, et al.

    Preprint

  6. 2025

    A survey of reinforcement learning for large reasoning models

    Kaiyan Zhang, Yuxin Zuo, Bingxiang He, Youbang Sun, Runze Liu, Che Jiang, Yuchen Fan, Kai Tian, Guoli Jia, Pengfei Li, Yu Fu, Xingtai Lv, Yuchen Zhang, Sihang Zeng, et al.

    arXiv:2509.08827

  7. 2025

    CaSBRE: causality-inspired semi-supervised biomedical relation extraction

    Sihang Zeng, Jun Wen, Jiangchuan Du, Jing Qian, Tianxi Cai, Hao Wang

    Under review

  8. 2023

    Hierarchical pretraining for biomedical term embeddings

    Bryan Cai, Sihang Zeng, Yucong Lin, Zheng Yuan, Doudou Zhou, Lu Tian

    arXiv:2307.00266

  9. 2022

    BIOS: an algorithmically generated biomedical knowledge graph

    Sheng Yu, Zheng Yuan, Jun Xia, Shengxuan Luo, Huaiyuan Ying, Sihang Zeng, Jingyi Ren, Hongyi Yuan, Zhengyun Zhao, Yucong Lin, et al.

    arXiv:2203.09975

Conference abstracts

9
  1. 2026

    Predicting multi-cancer risk from EHR data using multi-agent LLMs

    Sihang Zeng, Young Won Kim, Wilson Lau, Ehsan Alipour, Ruth B. Etzioni, Meliha Yetisgen, Anand Oka, Jayashree Nanduri

    Journal of Clinical Oncology ASCO Annual Meeting

  2. 2026

    STRATOS-P clinical model for prognosis and selection of patients with metastatic hormone-sensitive prostate cancer for intermittent therapy

    Martin W. Schoen, Joshua Gruber, Jason M. Doherty, David B. Eaton Jr., Sihang Zeng, Lukas Owens, et al.

    Journal of Clinical Oncology ASCO Annual Meeting

  3. 2026

    Development of an oncology generative AI foundation model trained on more than a million longitudinal patient journeys across the United States

    Wilson Lau, Ehsan Alipour, Young Won Kim, Sihang Zeng, Anand Oka, Jayashree Nanduri

    Journal of Clinical Oncology ASCO Annual Meeting

  4. 2026

    MSR67 — Zero-shot lung cancer risk prediction from longitudinal electronic health records with a chain-of-agents framework

    Sihang Zeng, Young Won Kim, Wilson Lau, Ehsan Alipour, Ruth Etzioni, Meliha Yetisgen, Anand Oka, Jayashree Nanduri

    Value in Health ISPOR

  5. 2026

    MSR68 — Characterizing oncology patient journeys and health state transitions using a data-driven Markov transition matrix in large-scale electronic health records

    Young Won Kim, Wilson Lau, Ehsan Alipour, Sihang Zeng, Anand Oka, Jayashree Nanduri

    Value in Health ISPOR

  6. 2026

    MSR172 — Can a generative patient journey foundation model alleviate the burden of cancer screening?

    Wilson Lau, Ehsan Alipour, Young Won Kim, Sihang Zeng, Anand Oka, Jayashree Nanduri

    Value in Health ISPOR

  7. 2026

    RWD140 — Patient journey foundational model for scalable imputation of missing units of measurement in electronic health records data

    Ehsan Alipour, Wilson Lau, Young Won Kim, Sihang Zeng, Anand Oka, Jayashree Nanduri

    Value in Health ISPOR

  8. 2026

    Adapting the Global Burden of Disease Healthcare Access and Quality Index for the Veterans Health Administration: a feasibility study

    Sihang Zeng, Christine Wilson, Gang Luo, Steven Zeliadt

    AcademyHealth Annual Research Meeting

  9. 2025

    Population-level tobacco cessation outcomes associated with implementing Whole Health at the Veterans Health Administration

    Sihang Zeng, Scott Coggeshall, Ethan Rosser, Stephanie Taylor, Diana Burgess, Gang Luo, Steven Zeliadt

    AcademyHealth Annual Research Meeting

Research areas and affiliations

Research interests

  • Patient trajectory modeling
  • Computational drug repurposing
  • Healthcare agents
  • EHR foundation models
  • Medical data harmonization

Methods

  • Longitudinal and survival modeling
  • LLM agents and multi-agent systems
  • Contrastive and representation learning
  • Reinforcement learning with verifiable rewards

Research groups

  • Cai Lab / CELEHS, Harvard Medical School
  • Etzioni Lab, Fred Hutchinson Cancer Center
  • UW BioNLP, University of Washington

Remote collaborations

  • Tsinghua C3I
  • Frontis AI