Yu Bao (鲍宇)

About Me

I am a Principal Research Scientist at ByteDance Seed, working on large language models, multilingual and multimodal intelligence, and AI for Science. My current research focuses on reliable LLM reasoning and alignment, machine translation, and generative modeling for scientific discovery.

I received my Ph.D. from Nanjing University in 2022, advised by Prof. Shujian Huang and Prof. Jiajun Chen. During my doctoral studies, I interned at ByteDance AI Lab under the mentorship of Prof. Hao Zhou and Prof. Lei Li, where I worked on non-autoregressive generation and latent-variable models.

Recent Highlights

  1. "Improving Long-Context Translation via Self-Supervised Dual Learning" is published in ACL 2026 (paper).

  2. DuPO is published at ICLR 2026 (poster) (official page).

  3. Seed-X released: a 7B-parameter multilingual translation LLM with open-sourced models and demos (arXiv, Hugging Face, Demo).

  4. Seed LiveInterpret 2.0 released: an end-to-end simultaneous speech-to-speech translation system with 3-second latency (70% reduction from prior solutions) and voice cloning (Technical Report, Homepage, Demo).

Publications/Preprints

[Full list] [*: equal contributions] [interns/students I mentored]

Find a publication

Use the filters or search below, or show all 20 publications.

20 publications + 4 technical reports.

2026

  1. Original figureFigure 1, page 3
    DuPO: Enabling Reliable Self-Verification via Dual Preference Optimization
    Shuaijie She, Yu Bao, Yu Lu, Lu Xu, Tao Li, Wenhao Zhu, Jianbing Zhang, Shujian Huang, Shanbo Cheng, Lu Lu, Yuxuan Wang
    ICLR 2026

    Turns a primal task into known and unknown components, reconstructs the unknown part as a dual task, and aligns the model with generalized duality rewards, with no verifiable labels.

  2. Original figureFigure 1 (page 3)
    Improving Long-Context Translation via Self-Supervised Dual Learning
    Shanbo Cheng, Shuaijie She, Yu Bao, Jianbing Zhang, Jiajun Chen, Shujian Huang
    ACL 2026

    Round-trip self-supervised dual learning: translate x to y and back to x-hat, reward reconstructability, and add Stick/Carrot constraints against reward hacking.

2025

  1. Original figureFigure 1, page 4
    Decomposed Direct Preference Optimization for Structure-Based Drug Design
    Xiwei Cheng, Xiangxin Zhou, Yuwei Yang, Yu Bao, Quanquan Gu
    TMLR 2025

    Aligns a structure-based diffusion model to pharmaceutical preference with multi-granularity pairs, global QED and SA plus decomposable arm and scaffold rewards, under a physics-informed energy constraint.

  2. Adapted schematicFigure 1 (page 4), simplified
    TRANS-ZERO: Self-Play Incentivizes Large Language Models for Multilingual Translation Without Parallel Data
    Wei Zou, Sen Yang, Yu Bao, Shujian Huang, Jiajun Chen, Shanbo Cheng
    Findings of ACL 2025

    Self-play without parallel data: one LLM expands a search tree by merging and mutating candidates, simulates to score them, and harvests preference pairs.

  3. Original figureFigure 2 (page 3)
    EnAnchored-X2X: English-Anchored Optimization for Many-to-Many Translation
    Sen Yang, Yu Bao, Yu Lu, Jiajun Chen, Shujian Huang, Shanbo Cheng
    EMNLP 2025

    English-anchored many-to-many translation: compare direct, pivot and English-anchored decoding, then build an English-anchored reward model and cross-lingual preference pairs.

2024

  1. Original figureFigure 2, page 4
    DecompOpt: Controllable and Decomposed Diffusion Models for Structure-based Molecular Optimization
    Xiangxin Zhou*, Xiwei Cheng*, Yuwei Yang, Yu Bao, Liang Wang, Quanquan Gu
    ICLR 2024

    Iteratively optimizes a ligand molecule: sample a reference arm per subpocket, generate and decompose molecules with a controllable diffusion model, then replace poor arms.

  2. Original figureFigure 2, page 3
    Binding-Adaptive Diffusion Models for Structure-Based Drug Design
    Zhilin Huang, Ling Yang, Zaixi Zhang, Xiangxin Zhou, Yu Bao, Xiawu Zheng, Yuwei Yang, Yu Wang, Wenming Yang
    AAAI 2024

    Alternates complex-to-subcomplex and subcomplex-to-complex interaction blocks with a gated transmission module, letting the diffusion model adapt ligand generation to each binding site.

  3. Original figure图 1(Figure 1, 模型总览)— 全文 https://www.jos.org.cn/html/2024/9/6956.htm ,图 https://www.jos.org.cn/html/2024/9/PIC/6956-1.jpg
    Pobe: 一种基于生成式模型的分布外文本检测方法
    欧阳亚文, 高源, 宗石, 鲍宇, 戴新宇
    软件学报 (Journal of Software) 2024

    Generation-based OOD text detection: combine language-model prediction with KNN retrieval, take the higher score, then divide by a pretrained GPT-2 probability to calibrate bias.

  4. Original figureFigure 2, page 4 (right column)
    EDT: Improving Large Language Models' Generation by Entropy-based Dynamic Temperature Sampling
    Shimao Zhang, Yu Bao, Shujian Huang
    Preprint 2024

    Chooses the sampling temperature at each decoding step from the entropy of the predicted token distribution, raising temperature when the model is uncertain.

2023

  1. Original figureFigure 2, page 4
    DecompDiff: Diffusion Models with Decomposed Priors for Structure-Based Drug Design
    Jiaqi Guan*, Xiangxin Zhou*, Yuwei Yang, Yu Bao, Jian Peng, Jianzhu Ma, Qiang Liu, Liang Wang, Quanquan Gu
    ICML 2023

    Generates 3D ligands from decomposed priors, separate arm and scaffold centers, denoised by an equivariant heterogeneous-graph network with validity guidance for clash-free, connected molecules.

  2. Original figureFigure 1 (page 1)
    Selective Knowledge Distillation for Non-Autoregressive Neural Machine Translation
    Min Liu, Yu Bao, Chengqi Zhao, Shujian Huang
    AAAI 2023

    Standard KD distils from all raw data; selectively distilling only high-quality, low-complexity examples via data selection yields the selected dataset used for distillation.

  3. Original figureFigure 2 (p. 6)
    DINOISER: Diffused Conditional Sequence Learning by Manipulating Noises
    Jiasheng Ye, Zaixiang Zheng, Yu Bao, Lihua Qian, Mingxuan Wang
    TACL (accepted); presented at ACL 2024 · Preprint 2023

    DINOISER clips the diffusion noise scale so corrupted token embeddings stay far enough apart, making discrete diffused conditional sequence learning trainable.

  4. Original figureFigure 1 (p. 2)
    Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-Finetuning
    Jiasheng Ye, Zaixiang Zheng, Yu Bao, Lihua Qian, Quanquan Gu
    Preprint 2023

    Large pre-trained masked language models are reprogrammed into diffusion LMs by generative surgery, so scaling plus instruction fine-tuning turns them into multitask generators.

2022

  1. Original figureFigure 1, PDF page 3 (ACL 2022 long paper, pp.8398-8409); DOI 10.18653/v1/2022.acl-long.575
    latent-GLAT: Glancing at Latent Variables for Parallel Text Generation
    Yu Bao, Hao Zhou, Shujian Huang, Dongqi Wang, Lihua Qian, Xinyu Dai, Jiajun Chen, Lei Li
    ACL 2022

    Add discrete latent variables to a glancing non-autoregressive model, learning target categorical codes via vector quantization to ease the multi-modality problem in parallel generation.

  2. Abstract-based schematicAbstract (full text and figures not retrievable — the publisher page is a JavaScript SPA and its PDF endpoints return an empty shell)
    基于句法模板采样的无监督复述生成方法 (Unsupervised Paraphrasing via Syntactic Template Sampling)
    鲍宇, 黄书剑, 周浩, 李磊, 戴新宇, 陈家骏
    中国科学: 信息科学 (SCIENTIA SINICA Informationis) 2022

    Unsupervised paraphrase with a syntactic-template latent variable linking meaning and syntax; a two-step sampling: prior over templates for diversity, posterior over syntactic representations for semantic match.

2021

  1. Original figureFigure 2, PDF page 3 (ACL-IJCNLP 2021 long paper, pp.1993-2003); DOI 10.18653/v1/2021.acl-long.155
    Glancing Transformer for Non-Autoregressive Neural Machine Translation
    Lihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang, Lin Qiu, Weinan Zhang, Yong Yu, Lei Li
    ACL-IJCNLP 2021

    Train a non-autoregressive Transformer with adaptive glancing sampling, letting it learn word dependencies so it can decode a whole sentence in a single parallel pass.

  2. Original figureFigure 1, PDF page 2 (NAACL-HLT 2021 long paper, pp.5749-5759); DOI 10.18653/v1/2021.naacl-main.458
    Non-Autoregressive Translation by Learning Target Categorical Codes
    Yu Bao, Shujian Huang, Tong Xiao, Dongqi Wang, Xinyu Dai, Jiajun Chen
    NAACL-HLT 2021

    Learn target categorical codes as latent variables inside non-autoregressive translation, so the decoder captures word-level structure and translates more accurately in parallel.

2020

  1. Original figureFigure 1, PDF page 3 (ACL 2020 long paper, pp.708-717); DOI 10.18653/v1/2020.acl-main.65
    Explicit Semantic Decomposition for Definition Generation
    Jiahuan Li*, Yu Bao*, Shujian Huang, Xinyu Dai, Jiajun Chen
    ACL 2020

    Decompose a word's meaning into explicit semantic components modelled with discrete latent variables, giving interpretable components for generating a dictionary definition.

2019

  1. Original figureFigure 2, PDF page 4 (ACL 2019 long paper, pp.6008-6019); DOI 10.18653/v1/P19-1602
    Generating Sentences from Disentangled Syntactic and Semantic Spaces
    Yu Bao*, Hao Zhou*, Shujian Huang, Lei Li, Lili Mou, Olga Vechtomova, Xinyu Dai, Jiajun Chen
    ACL 2019

    Disentangle sentence generation into separate syntactic and semantic latent spaces, modelling syntax as a linearized parse-tree sequence so each factor can be controlled independently.

  2. Original figureFigure 1, PDF page 3 (arXiv:1911.10677, 2019 preprint; no formal venue pages)
    PNAT: Non-Autoregressive Transformer by Position Learning
    Yu Bao, Hao Zhou, Jiangtao Feng, Mingxuan Wang, Shujian Huang, Jiajun Chen, Lei Li
    Preprint 2019

    Treat positions as a latent variable in the text generation process, so a non-autoregressive Transformer learns where to place each output token.

Technical Reports

  1. Original figureFigure 2 (p. 2)
    Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
    Shanbo Cheng, Yu Bao, et al. (see the full author list on arXiv)
    Technical Report 2025

    An end-to-end duplex speech-to-speech model performs simultaneous interpretation, keeps each speaker's own voice, and cuts cloned-speech latency to about three seconds.

  2. Adapted schematicAbstract; §2 Pre-training (pp. 2–3); §3 Post-training incl. §3.1 SFT and the RL/PPO paragraph (pp. 4–5)
    Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters
    Shanbo Cheng, Yu Bao, et al. (see the full author list on arXiv)
    Technical Report 2025

    A 7B translation LLM built in stages — multilingual pre-training, Chain-of-Thought fine-tuning, then PPO — with a revise-and-filter loop that keeps improving parallel data.

  3. Adapted schematic§1 Introduction (p. 1) and the evaluation-framework paragraph (pp. 2 and 7)
    Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
    Bytedance Seed (collective author)
    Model Card 2026

    Seed2.0 targets real-world complexity, tracked by four evaluation dimensions — Science Discovery, Vibe Coding, Context Learning, Real-World Tasks — across three model sizes.

  4. Adapted schematic§1 Introduction (p. 1) and the §2 evaluation overview (p. 2)
    Seed1.8 Model Card: Towards Generalized Real-World Agency
    Bytedance Seed (collective author)
    Model Card 2026

    Seed1.8 pursues generalized real-world agency by integrating perception, reasoning and action in one model rather than task-specific agent pipelines.

Awards

  • 2022, Excellent Doctoral Paper Award, Jiangsu Association of Artificial Intelligence (JSAI)
  • 2020, Outstanding Ph.D. Candidate, Nanjing University
  • 2019, Outstanding Graduate Student, Nanjing University