I am a Principal Research Scientist at ByteDance Seed, working on large language models, multilingual and multimodal intelligence, and AI for Science. My current research focuses on reliable LLM reasoning and alignment, machine translation, and generative modeling for scientific discovery.
I received my Ph.D. from Nanjing University in 2022, advised by Prof. Shujian Huang and Prof. Jiajun Chen. During my doctoral studies, I interned at ByteDance AI Lab under the mentorship of Prof. Hao Zhou and Prof. Lei Li, where I worked on non-autoregressive generation and latent-variable models.
Recent Highlights
"Improving Long-Context Translation via Self-Supervised Dual Learning" is published in ACL 2026 (paper).
DuPO is published at ICLR 2026 (poster) (official page).
Seed-X released: a 7B-parameter multilingual translation LLM with open-sourced models and demos (arXiv, Hugging Face, Demo).
Seed LiveInterpret 2.0 released: an end-to-end simultaneous speech-to-speech translation system with 3-second latency (70% reduction from prior solutions) and voice cloning (Technical Report, Homepage, Demo).
Turns a primal task into known and unknown components, reconstructs the unknown part as a dual task, and aligns the model with generalized duality rewards, with no verifiable labels.
Round-trip self-supervised dual learning: translate x to y and back to x-hat, reward reconstructability, and add Stick/Carrot constraints against reward hacking.
Aligns a structure-based diffusion model to pharmaceutical preference with multi-granularity pairs, global QED and SA plus decomposable arm and scaffold rewards, under a physics-informed energy constraint.
Self-play without parallel data: one LLM expands a search tree by merging and mutating candidates, simulates to score them, and harvests preference pairs.
English-anchored many-to-many translation: compare direct, pivot and English-anchored decoding, then build an English-anchored reward model and cross-lingual preference pairs.
Iteratively optimizes a ligand molecule: sample a reference arm per subpocket, generate and decompose molecules with a controllable diffusion model, then replace poor arms.
Alternates complex-to-subcomplex and subcomplex-to-complex interaction blocks with a gated transmission module, letting the diffusion model adapt ligand generation to each binding site.
Original figure图 1(Figure 1, 模型总览)— 全文 https://www.jos.org.cn/html/2024/9/6956.htm ,图 https://www.jos.org.cn/html/2024/9/PIC/6956-1.jpgPobe: 一种基于生成式模型的分布外文本检测方法
Generation-based OOD text detection: combine language-model prediction with KNN retrieval, take the higher score, then divide by a pretrained GPT-2 probability to calibrate bias.
Chooses the sampling temperature at each decoding step from the entropy of the predicted token distribution, raising temperature when the model is uncertain.
Generates 3D ligands from decomposed priors, separate arm and scaffold centers, denoised by an equivariant heterogeneous-graph network with validity guidance for clash-free, connected molecules.
Standard KD distils from all raw data; selectively distilling only high-quality, low-complexity examples via data selection yields the selected dataset used for distillation.
Large pre-trained masked language models are reprogrammed into diffusion LMs by generative surgery, so scaling plus instruction fine-tuning turns them into multitask generators.
Add discrete latent variables to a glancing non-autoregressive model, learning target categorical codes via vector quantization to ease the multi-modality problem in parallel generation.
Unsupervised paraphrase with a syntactic-template latent variable linking meaning and syntax; a two-step sampling: prior over templates for diversity, posterior over syntactic representations for semantic match.
Train a non-autoregressive Transformer with adaptive glancing sampling, letting it learn word dependencies so it can decode a whole sentence in a single parallel pass.
Learn target categorical codes as latent variables inside non-autoregressive translation, so the decoder captures word-level structure and translates more accurately in parallel.
Decompose a word's meaning into explicit semantic components modelled with discrete latent variables, giving interpretable components for generating a dictionary definition.
Disentangle sentence generation into separate syntactic and semantic latent spaces, modelling syntax as a linearized parse-tree sequence so each factor can be controlled independently.
An end-to-end duplex speech-to-speech model performs simultaneous interpretation, keeps each speaker's own voice, and cuts cloned-speech latency to about three seconds.
A 7B translation LLM built in stages — multilingual pre-training, Chain-of-Thought fine-tuning, then PPO — with a revise-and-filter loop that keeps improving parallel data.
Seed2.0 targets real-world complexity, tracked by four evaluation dimensions — Science Discovery, Vibe Coding, Context Learning, Real-World Tasks — across three model sizes.