Benyou Wang王本友

Benyou Wang(王本友)

Benyou Wang王本友

Assistant Professor, PhD Supervisor and Presidential Young Professor, School of Data Science, CUHK-Shenzhen 香港中文大学(深圳)数据科学学院助理教授、博士生导师、校长青年教授

Benyou Wang leads FreedomAI and has hands-on experience in large-scale model pre-training, post-training, and deployment across agents, world models, healthcare, multimodal AI, and related areas. His publications have received more than 12,000 Google Scholar citations and five best-paper awards or equivalent distinctions, including a SIGIR 2017 Best Paper Award Honorable Mention and the NAACL 2019 Best Explainable Paper award (recognized alongside the authors of four other award-winning papers, including BERT). 王本友带领 FreedomAI 团队,在智能体、世界模型、医疗、多模态等方向具有大规模模型预训练、后训练和部署经验。论文在 Google Scholar 上被引用超过 1.2 万次,曾五次获得最佳论文或同等级荣誉,包括 SIGIR 2017 最佳论文荣誉提名和 NAACL 2019 Best Explainable Paper(与 BERT 等其他四篇获奖论文的作者一同领奖)。

他获天津大学硕士(2017)和帕多瓦大学博士(2022)学位,曾访问哥本哈根大学、蒙特利尔大学,曾在腾讯从事研发,并于 2020–2022 年在华为实习两年零两个月,导师为尚利锋、蒋欣和刘群。研究获欧盟玛丽·居里奖学金及华为、腾讯等项目支持。He earned his master’s degree from Tianjin University (2017) and PhD from the University of Padua (2022), visited the University of Copenhagen and Université de Montréal, and worked in R&D at Tencent. From 2020 to 2022, he interned at Huawei for two years and two months, mentored by Lifeng Shang, Xin Jiang, and Qun Liu. His research has been supported by the Marie Skłodowska-Curie Fellowship and programmes from Huawei and Tencent.

王本友的工作涵盖研究方向规划、大规模模型训练、开源工程、实际部署与科技创业。他组织跨学科团队开展研究,并将研究中积累的方法和工具用于后续项目。FreedomAI 开放模型与数据在 Hugging Face 的下载量已超过百万,GitHub 收藏星标超过 10K;华佗 GPT 已部署到十几家医院和数百家社康,累计访问量近 500 万次,为实际医疗服务提供支持。团队也支持学生探索技术创业,已有近十位校友(alumni)成为创业公司 CTO 或技术创始人,其中,博士生陈志鸿担任创始人兼 CTO 的医疗影像公司以约 8,000 万美元被收购。Benyou Wang works on research planning, large-scale model training, open-source engineering, deployment, and technology ventures. He organizes interdisciplinary teams and carries methods and tools developed in one project into subsequent research. FreedomAI's open models and datasets have surpassed one million downloads on Hugging Face and 10,000 GitHub stars. HuatuoGPT has been deployed in more than a dozen hospitals and hundreds of community health centers, with nearly five million visits, bringing medical AI research into practical service. The group also supports students in technology entrepreneurship; nearly ten alumni have become startup CTOs or technical founders. Among them, doctoral student Zhihong Chen served as founder and CTO of a medical-imaging company that was acquired for approximately US$80 million.

从基础模型到真实世界智能From foundation models to real-world intelligence

团队长期关注如何让人工智能持续学习,并解决实际问题。研究以基础模型为基础,重点探索持续学习的智能体,同时从有实际价值的应用场景中寻找研究问题。The group's long-term agenda asks how AI can keep learning and create value in the real world. It connects foundation models, continuously learning agents, and high-value applications as mutually reinforcing research directions.

  • 基础模型与训练方法。Foundation models and training. 研究预训练、后训练与评测方法,在多语言、医疗和多模态方向开展大规模训练与开源实践,开发可复用的模型与工具。Connect pre-training, post-training, and evaluation across multilingual, medical, and multimodal AI, turning methodological advances and large-scale training experience into reusable models and tools.
  • 持续学习的智能体。Continuously learning agents. 研究推理、工具使用、环境反馈与递归自提升,构建可重复的任务环境和评测体系,检验模型能力的变化,并据此改进方法。Study reasoning, tool use, environmental feedback, and recursive self-improvement, with reproducible task environments and evaluations that make progress measurable and cumulative.
  • 真实场景与技术转化。Real-world applications and translation. 从医疗、运筹优化、金融、教育与机器人等应用中提炼研究问题,通过系统部署和产业合作检验方法的实际效果。Turn needs in healthcare, optimization, finance, education, and robotics into research questions, and assess their value through deployed systems and industry collaboration.

在研究推进中,团队重视方向选择、跨学科合作,以及从模型开发到系统部署的具体工作。长期目标会被拆解为可检验的阶段任务,模型、数据和评测资源则尽可能开放,供团队及合作伙伴继续研究和使用。The team selects research directions, organizes interdisciplinary work, and follows projects from model development through deployment. Long-term goals are broken into testable milestones, with models, data, and evaluations shared to support further work by the team and its collaborators.

Selected recognition代表性荣誉

  • SIGIR Best Paper Award Honorable MentionSIGIR 最佳论文荣誉提名
  • NAACL 2019 Best Explainable Paper (recognized alongside the authors of four other award-winning papers, including BERT)NAACL 2019 Best Explainable Paper(与 BERT 等其他四篇获奖论文的作者一同领奖)
  • NLPCC Best PaperNLPCC 最佳论文
  • ICLR Financial AI Workshop Best PaperICLR 金融人工智能研讨会最佳论文
  • NeurIPS ResponsibleFM Workshop Outstanding PaperNeurIPS ResponsibleFM 研讨会杰出论文

Competition results竞赛成绩

  • 与智子芯元合作,以 AI 自动编写华为 Ascend 算子的全自动方案,曾取得 CannBench 榜单第一名。
  • Led a CUHK-Shenzhen team to a Gold Medal in AIMO Progress Prize 2, placing in the top 0.4% among more than 2,000 teams.带领团队与华为联合参加第二届人工智能数学奥林匹克竞赛(AIMO2),在 2,000 余支队伍中进入前 0.4%,获得金牌。
  • With Professor Bingyi Jing, co-supervised master's student Shuqi Guo, a core member of team sZs, which won the IEEE ICRA LeHome Challenge real-robot final with 895 points.与荆炳义教授共同指导硕士生郭书杞;郭书杞作为 sZs 队核心成员,在 IEEE ICRA LeHome Challenge 真机赛决赛中以 895 分获得冠军。

Academic service.学术服务。 Publicity Chair of NLPCC 2023, Website Chair of EMNLP 2023, and Area Chair or Senior Area Chair for ICLR, NeurIPS, COLM, ACL, and EMNLP.曾任 NLPCC 2023 宣传主席、EMNLP 2023 网站主席,并多次担任 ICLR、NeurIPS、COLM、ACL、EMNLP 等会议的领域主席或高级领域主席。

Models, systems, and applications模型、系统与应用

The group works on large language models, agents, and their applications in healthcare, multilingual communication, operations research, finance, education, and robotics. The links below connect representative systems with their papers, code, models, data, and application pages.团队围绕大语言模型、智能体及其在医疗、多语言交流、运筹优化、金融、教育和机器人等场景中的应用开展研究。以下介绍代表性项目,并附上相关论文、代码、模型、数据和应用链接。

华佗GPT:面向全球的医疗AIHuatuoGPT · Medical AI for Global Access

In collaboration with Haizhou Li, Xiang Wan, Guangjun Yu, Ruoyu Sun, and clinical partners, he participated in the HuatuoGPT series. The work has expanded from medical dialogue to one-stage domain adaptation, verifiable medical reasoning, and medical vision-language understanding. As of September 2026, public releases across the series had reached more than one million downloads and over 10,000 GitHub stars; related systems have been deployed in more than a dozen hospitals and hundreds of community health centers, with nearly five million visits.与李海洲、万翔、于广军、孙若愚等学者及医疗团队合作,参与华佗GPT系列研发。相关工作已从医学对话扩展到一阶段领域适配、可验证医学推理和医学视觉语言理解。截至 2026 年 9 月,系列公开模型累计下载量超过百万,GitHub Stars 超过一万;相关系统已部署到十几家医院和数百家社康,累计访问量近 500 万次。

从医学对话到专业知识适配。HuatuoGPT 探索面向医疗场景的语言交互;HuatuoGPT-II 将不同来源、格式的医学知识与指令数据统一为输入输出对,以一阶段训练完成领域适配。这条研究路线关注如何把通用模型转化为具备医学知识、能够理解专业问题的医疗助手。From medical dialogue to domain expertise. HuatuoGPT studies language interaction in healthcare. HuatuoGPT-II brings heterogeneous medical knowledge and instruction data into a shared input–output format for one-stage domain adaptation, studying how general models can become assistants that understand medical questions.

从多模态理解到可验证推理。HuatuoGPT-Vision 将医学视觉知识引入多模态模型,连接医学图像与语言理解;HuatuoGPT-o1 则通过可验证医学问题、推理轨迹搜索与强化学习,研究如何提升复杂医学问题的求解能力。两者分别推进医学信息的感知与推理,为更完整的医疗AI系统提供能力基础。From multimodal understanding to verifiable reasoning. HuatuoGPT-Vision connects medical images with language understanding. HuatuoGPT-o1 uses verifiable medical problems, reasoning-trajectory search, and reinforcement learning to improve complex medical problem solving. Together, these research directions develop perception and reasoning capabilities for broader medical AI systems.

跨越语言与资源门槛。Apollo 与 ApolloMoE 在这一医疗AI方向下推进多语言普惠:Apollo 开放轻量医学模型、医学语料与评测资源,ApolloMoE 按语系组织专家模块,将支持范围扩展到 50 种语言,探索训练效率与跨语言医学能力的平衡。通过开放模型、数据和工具,相关工作旨在让更多语言社区和资源有限的研究团队能够参与医疗AI的开发与评测。Across languages and resource constraints. Apollo and ApolloMoE extend this medical AI research toward multilingual access. Apollo releases lightweight medical models, corpora, and evaluation resources; ApolloMoE organizes experts by language family to support 50 languages while studying efficiency and cross-language medical capabilities. Open models, data, and tools aim to help more language communities and resource-constrained research groups develop and evaluate medical AI.

Phoenix

Phoenix is an open multilingual dialogue model released in 2023. It made training data, model weights, code, and evaluation resources available for Chinese, English, and several lower-resource languages. Its results were competitive with contemporary open chat models, and it entered the leading group in third-party Chinese LLM evaluations at release.Phoenix 是团队于 2023 年发布的开放多语言对话模型,面向中文、英文和多种低资源语言开放训练数据、模型权重、代码与评测资源。其效果在同期开放对话模型中具有竞争力,发布初期进入第三方中文大模型评测前列。

AceGPT

With Professor Jinchao Xu's group and support from Peng Cheng Laboratory, the team developed Arabic language models adapted to local language, culture, and values. AceGPT was among the leading open Arabic large language models at release; follow-up work explored native alignment during pre-training and progressive vocabulary expansion.在许进超教授团队带领和鹏城实验室支持下,团队开展面向阿拉伯语言、文化与价值观适配的大模型研究。AceGPT 发布时是表现领先的开源阿拉伯语大模型;后续工作进一步探索预训练原生对齐(native alignment)与渐进式词表扩展(progressive vocabulary expansion),并使用大规模华为昇腾 910A 集群训练更大规模模型。

ORLM

Together with Professor Zizhuo Wang and Cardinal Operations, the team developed ORLM for translating natural-language business problems into executable mathematical optimization models and solver code. The work introduced OR-Instruct and IndustryOR and connects with Cardinal Operations' COLORMind decision-intelligence product line. Mamo extends this direction with a mathematical-modeling benchmark and solvers that connect natural-language problems to executable formal models. CALM Before the STORM further studies native reasoning for optimization modeling.与王子卓教授及杉数科技共同开发 ORLM,用于将自然语言业务问题转化为可执行的数学优化模型与求解代码,并提出 OR-Instruct 和 IndustryOR。相关研究与杉数科技 COLORMind 智能决策产品线相结合。Mamo 则通过数学建模基准与求解器,探索如何把自然语言问题转化为可执行的形式化模型。CALM Before the STORM 进一步研究优化建模中的原生推理能力。

TinyDeepSeek

从零开始训练高效的小规模模型。TinyDeepSeek 从零开始预训练,构建了超过 3T tokens 的数据,工作覆盖数据处理、模型架构、预训练与后训练。根据团队实验结果,模型在约 3B 参数规模的对比中达到 SOTA 水平,团队完成了从数据准备到训练系统的完整研发。Training from scratch for capable small-scale models. TinyDeepSeek builds a pre-training corpus of more than 3 trillion tokens and connects data processing, model architecture, pre-training, and post-training. In the team's evaluations, it achieved state-of-the-art performance among models at approximately the 3B-parameter scale, demonstrating end-to-end development from data to training systems.

团队自 2025 年 1 月左右开始探索 Loop Transformer 的效果,研究通过循环计算与参数复用提升模型能力。项目开放了训练代码、0.5B 与 3.3B 基础模型及中间检查点,便于开展低成本的架构实验和持续训练研究。The team began exploring Loop Transformers around January 2025, studying how recurrent computation and parameter reuse can improve model capabilities. The project releases training code, 0.5B and 3.3B base models, and intermediate checkpoints to support lower-cost architecture experiments and continued training research.

Code · HF · 3.3B · HF · 0.5B · 训练检查点Training checkpoints

多模态与长上下文:ALLaVA、LongLLaVAMultimodal & Long-context Models · ALLaVA & LongLLaVA

从高质量图文数据到长上下文理解。ALLaVA 使用 GPT-4V 生成细粒度图像描述和视觉问答数据,支持轻量视觉语言模型的图文对齐与指令微调,探索在有限计算资源下提升多模态能力。LongLLaVA 进一步研究如何高效理解大量图像,MileBench 则提供长上下文评测,形成数据、模型与评测相互支撑的研究路线。From high-quality visual data to long-context understanding. ALLaVA uses GPT-4V to synthesize detailed image descriptions and visual question-answer pairs for alignment and instruction tuning of lightweight vision-language models. LongLLaVA studies efficient understanding of many images, while MileBench provides long-context evaluation, connecting data, models, and benchmarks.

LongLLaVA 结合 Mamba 与 Transformer 架构,配合时空数据构建和渐进训练,在单张 A100 80GB 上处理近千张图像。MileBench 从真实任务与诊断性测试两个角度,评估多模态大模型的长上下文理解能力。LongLLaVA combines Mamba and Transformer blocks with temporal and spatial data construction and progressive training to process nearly 1,000 images on a single A100 80GB GPU. MileBench evaluates long-context multimodal understanding through realistic tasks and diagnostic tests.

PhoneBuddy

PhoneBuddy 系列研究手机智能体的训练,将真实应用与 PhoneWorld 重建的模拟应用纳入统一流程:先通过轨迹数据进行监督微调,再在真实或混合环境中进行强化学习。系列工作还包括隐私评测 MyPhoneBench、安全评测、PhoneWorld 环境构建,以及结合 GUI、CLI 和工具动作的 PhoneHarness。The PhoneBuddy series studies training for phone agents using both real apps and reconstructed apps in PhoneWorld. Agents first learn from trajectories through supervised fine-tuning, then undergo reinforcement learning in real or mixed environments. The series also covers privacy evaluation with MyPhoneBench, safety evaluation, PhoneWorld environment construction, and PhoneHarness for mixed GUI, CLI, and tool actions.

语音大模型Speech models

语音方向涵盖模型训练、可控生成与交互评测:Soundwave 研究语音与文本的高效对齐;TTS-Hub 通过模块化 LoRA 及其组合实现可控语音合成;S2S-Arena 评估语音到语音模型对副语言指令的遵循能力,EchoMind 则关注语音交互中的共情表现。Human or Machine? 通过初步图灵测试,研究语音到语音交互中人类与模型的可区分性。The speech research covers model training, controllable generation, and interaction evaluation. Soundwave studies efficient speech–text alignment; TTS-Hub uses modular LoRAs and their composition for controllable speech synthesis. S2S-Arena evaluates paralinguistic instruction following in speech-to-speech models, while EchoMind examines empathy in spoken interaction. Human or Machine? uses a preliminary Turing test to study whether people can distinguish human from model speech-to-speech interaction.

智能体学习与创造Agent Learning & Creation

从环境交互到可执行成果。我们研究如何构建可交互、可反馈、可检验的环境,让智能体通过工具使用与强化学习提升解决问题的能力,并将人的需求转化为可运行的代码、三维模型和可玩的游戏。环境工程提供学习和验证的条件,Agentic RL 研究如何利用反馈改进行为,智能体创作则把这些能力连接到具体成果。From interaction to executable artifacts. We study environments that support interaction, feedback, and verification, enabling agents to develop problem-solving skills through tool use and reinforcement learning and to turn human intent into runnable code, 3D models, and playable games. Environment engineering provides the conditions for learning and validation; agentic RL studies how feedback improves behavior; agentic creation connects these capabilities to concrete outputs.

这一方向围绕智能体可交互的环境、工具和反馈机制展开。SepsisAgent 将 Clinical World Model 用作患者动态模拟器与训练环境,研究脓毒症场景中的状态预测和决策学习,并在回顾性患者轨迹上评估。This direction studies interactive environments, tools, and feedback for agents. SepsisAgent uses a Clinical World Model as a patient-dynamics simulator and training environment for state prediction and decision learning in sepsis, evaluated on retrospective patient trajectories.

CoRT 将代码执行融入推理过程,结合监督微调与强化学习训练模型使用代码解释器。TwinMarket 则用大语言模型驱动的多智能体模拟金融市场,研究个体行为如何形成市场层面的现象。CoRT integrates code execution into reasoning, using supervised fine-tuning and reinforcement learning to teach code-interpreter use. TwinMarket uses LLM-driven agents to simulate financial markets and study how individual behavior produces market-level phenomena.

面向专业任务的工具调用与证据推理。OpenClaw-Medical-Skills 汇集可供智能体调用的医学技能与工具,支持医学信息检索和相关工作流程;Legal-R1 系列通过多轮检索与推理,在获取证据、分析问题和继续检索之间迭代,围绕证据完成法律推理。这些工作将智能体的工具使用与推理能力连接到医疗和法律领域的具体任务。Tool use and evidence-based reasoning for professional tasks. OpenClaw-Medical-Skills brings together medical skills and tools that agents can call for medical information retrieval and related workflows. Legal-R1 alternates between retrieving evidence, analyzing questions, and further retrieval to support legal reasoning. These projects connect agent tool use and reasoning to concrete tasks in healthcare and law.

这一方向研究模型如何生成可检验的创作结果。MicroVerse 面向微观生物过程生成科学可视化视频,并建立相应数据与评测;BlenderLLM 将文本需求转化为可在 Blender 中运行的建模脚本,通过自我改进训练提升三维建模能力;GameCraft-Bench 在真实 Godot 引擎中评测智能体端到端制作可玩游戏的能力,通过运行项目、重放操作和评分规则检验结果。This direction studies verifiable creative outputs. MicroVerse generates scientific videos of microscopic biological processes and provides supporting data and evaluation. BlenderLLM translates text instructions into modeling scripts that run in Blender, using self-improvement training to develop 3D modeling capabilities. GameCraft-Bench evaluates end-to-end creation of playable games in the Godot engine by running projects, replaying actions, and checking scoring rubrics.

这些工作的共同目标,是让AI更好地服务研究、教育与创作:帮助研究者探索复杂系统,让创作者表达并实现设计意图,也让学习者获得可交互的内容。我们关注成果是否能够运行、检验和迭代,让智能体能力与实际任务的完成质量建立联系。The shared goal is to make AI more useful for research, education, and creation: helping researchers explore complex systems, creators realize design ideas, and learners engage with interactive content. We focus on whether outputs can be run, evaluated, and iteratively improved, connecting agent capabilities to the quality of completed tasks.

Student ventures and technology transfer学生创业与技术转化

The group supports students and research staff in translating research into practical products and services. Nearly ten alumni have gone on to serve as startup CTOs or technical founders.团队支持学生和研究人员将科研成果转化为实际产品与服务,已有近十位校友(alumni)成为创业公司 CTO 或技术创始人。

  • Zhihong Chen陈志鸿PhD graduate · Founder and CTO. Cognita Imaging, medical-imaging AI; acquired by Mosaic Clinical Technologies in 2025 in a transaction valued at approximately US$80 million.博士毕业生 · 创始人兼 CTO。Cognita Imaging,医疗影像 AI;公司于 2025 年被 Mosaic Clinical Technologies 收购,交易规模约 8,000 万美元。
  • Jianye Hou侯建业PhD student · CTO. Metastone, high-performance computing. The company has revenue approaching RMB 1 billion.博士生 · CTO。是石科技,高性能计算;收入接近十亿。
  • Shunlin Lu路舜麟PhD student · CTO. NeoteAI, embodied intelligence; financing totals close to RMB 100 million.博士生 · CTO。新智具身(NeoteAI),具身智能;融资规模近人民币 1 亿元。
  • Jiajun You游佳君PhD student · CEO. Yuniu Technology (驭牛科技), currently raising funds.博士生 · CEO。驭牛科技,目前正在融资。
  • Xiao Yang肖杨PhD student · CEO. Shenzhen Sikoo Intelligent Information Services, developing AI education products including Shuta AI.博士生 · CEO。深圳思酷智能信息服务有限公司,开展 AI 教育业务,包括薯塔 AI。
  • Zhiqi Gao高治淇MPhil student · Founder. Heryin Yuanzhu (Shenzhen) Intelligent Technology, working on spatial intelligence, 3D reconstruction, and digital twins.MPhil 学生 · 创始人。和瑛元筑(深圳)智能科技有限公司,主要开展空间智能、三维重建与数字孪生相关业务。
  • Huiteng Xiao肖徽腾Research assistant · Founder. An early-stage venture combining AI and hardware.研究助理 · 创始人。开展 AI 与硬件结合的早期创业项目。
  • Shijun Chu褚士钧Research assistant · COO. Chaotic Pendulum Technology (Shenzhen), working on financial AI.研究助理 · COO。混沌摆科技(深圳)有限公司,主要开展 AI 金融相关业务。

Operating and financing figures above reflect information provided by the group as of September 2026; where public disclosures are available, the latest disclosure takes precedence.以上经营与融资信息根据团队提供的截至 2026 年 9 月资料整理;如有公开披露,以最新披露为准。

More about the startup ecosystem了解更多创业与转化项目

Natural language processing自然语言处理课程

Benyou Wang teaches Natural Language Processing each year to approximately 500 students. The course covers NLP foundations, large language models, agents, and practical system development, with course materials available online.王本友每年讲授自然语言处理课程,约有 500 名学生修读。课程涵盖自然语言处理基础、大语言模型、智能体与系统实践,并持续开放课程资料。

Course website ↗课程网站 ↗

自由互助,以学生为中心Freedom, mutual support, and student-led innovation

每个人追求自己的自由就是给整个世界追求自由Each person's pursuit of their own freedom is a pursuit of freedom for the whole world.

实验室重视自由探索和相互帮助,以学生的研究兴趣与成长为中心。我们尊重每个人的研究兴趣和发展选择,鼓励独立探索、开放交流,也鼓励同学之间相互支持。The laboratory has developed a research culture of freedom and mutual support, with an innovation network centered on students. We respect individual research interests and career choices, and encourage independent exploration, open exchange, and collaboration.

学生在 HY、Kimi、MiniMax、华为、大疆、阶跃、百川、蚂蚁、Qwen 等团队与企业实习,也有近十位同学选择 AI 创业。这些经历让实验室与学术界、产业界保持了持续的交流与合作。Students undertake internships at HY, Kimi, MiniMax, DJI, StepFun, Baichuan, Ant Group, Qwen, and other teams and companies. Nearly ten students have also embarked on AI ventures, contributing to sustained collaboration between research and industry.

He also co-supervises doctoral students with Professors Haizhou Li, Bingyi Jing, Ming Yan, Yongtao Guan, Hongyuan Zha, and Yilun Chen in areas including language intelligence, speech, multimodal learning, optimization, statistical learning, and robotics.他还与李海洲、荆炳义、严明、官永涛、查宏远、陈逸伦等教授联合指导博士生,研究方向涵盖语言智能、语音、多模态学习、运筹优化、统计学习和机器人。

Selected publications代表性论文

  1. Jiaxi Bi, Tongxu Luo, Wenyu Du, Zhengyang Tang, Benyou Wang. Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning. ACL, 2026. arXiv
  2. Bingquan Zhang, Xiaoxiao Liu, Yuchi Wang, Lei Zhou, Qianqian Xie, Benyou Wang. Human or LLM as Standardized Patients? A Comparative Study in Medical Education. ACL, 2026. arXiv
  3. Feng Jiang, Zhiyu Lin, Yiyang Liu, Liumeng Xue, Fan Bu, Yuhao Du, Xiangying Chen, Benyou Wang, Haizhou Li. S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models. ACL, 2026. arXiv · Code · HF
  4. Shunian Chen, Xinyuan Xie, Zheshu Chen, Liyan Zhao, Owen Lee, Zhan Su, Qilin Sun, Benyou Wang. Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion. ACL, 2026. arXiv · Code · HF
  5. Zhiqi Gao, Guo Zhu, Huarui Luo, Dongyijie Primo Pan, Haoming Tang, Bingquan Zhang, Jiahuan Pei, Jie Li, Benyou Wang. It Talks Like a Patient, But Feels Different: Co-Designing AI Standardized Patients. CHI, 2026. Paper · arXiv
  6. Shunian Chen, Rui Yu, Zhengyang Tang, Yuhao Du, Zejian Xie, Songxin Zhang, Benyou Wang. Agentic Rubrics: Interfacing Agents with Heterogeneous Environments using Executable Rubrics. EMNLP, Main, 2026. Paper
  7. Zichen Xu, Geng Zhao, Aoxiang Qin, Li Zhou, Benyou Wang. Steer on Rubrics: Auditable Steering of Schwartz Values in Language Models. EMNLP, Main, 2026. Paper
  8. Xiang Li, Shiqi Zhang, Zichen Xu, Wenyuan Gu, Hongru Xiao, Bo Cheng, Li Zhou, Jiale Han, Benyou Wang. TTS-Hub: Leveraging Modular LoRAs and Arithmetic Composition for Controllable Text-to-Speech. EMNLP, Main, 2026. Paper
  9. Xu Wang, Yan Hu, Benyou Wang, Difan Zou. Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse Autoencoders. ICLR, 2026. arXiv
  10. Li Zhou, Lutong Yu, You Lyu, Yihang Lin, Zefeng Zhao, Junyi Ao, Yuhao Zhang, Benyou Wang, Haizhou Li. EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models. ICLR, 2026. arXiv · Code · Website · HF
  11. Xiang Li, Jiabao Gao, Sipei Lin, Xuan Zhou, Chi Zhang, Bo Cheng, Jiale Han, Benyou Wang. Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction. ICLR, 2026. arXiv
  12. Xidong Wang, Shuqi Guo, Yue Shen, Junying Chen, Jian Wang, Jinjie Gu, Ping Zhang, Lei Liu, Benyou Wang. LiveClin: A Live Clinical Benchmark without Leakage. ICLR, 2026. arXiv · Code · HF
  13. Rongsheng Wang, Minghao Wu, Hongru Zhou, Zhihan Yu, Zhenyang Cai, Junying Chen, Benyou Wang. MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation. ICLR, 2026. arXiv · Code · HF
  14. Zhengyang Tang, Zihan Ye, Chenyu Huang, Xuhan Huang, Chengpeng Li, Sihang Li, Guanhua Chen, Ming Yan, Zizhuo Wang, Hongyuan Zha, Dayiheng Liu, Benyou Wang. CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling. ICML, 2026. arXiv · Code · HF
  15. Yihong Tang, Kehai Chen, Liang Yue, Benyou Wang, Min Zhang. CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning. ICML, 2026. arXiv
  16. Junying Chen, Xinyuan Xie, Ziniu Li, Benyou Wang. OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation. ICML, 2026.
  17. Hongru Xiao, Xiang Li, Duyi Pan, Longfei Zhang, Zhixue Song, Jiale Han, Songning Lai, Wenshuo Chen, Jing Tang, Benyou Wang. Can Audio Language Models Listen Between the Lines? A Study on Metaphorical Reasoning via Unspoken. ACM MM, 2025. Paper · Code · HF
  18. Zhenyang Cai, Junying Chen, Rongsheng Wang, Weihong Wang, Yonglin Deng, Dingjie Song, Yize Chen, Zixu Zhang, Benyou Wang. Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging. ACL, 2025. arXiv · Code · HF
  19. Jianqing Zhu, Huang Huang, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An, Juncai He, Xiangbo Wu, Fei Yu, Junying Chen, Ma Zhuoheng, Yuhao Du, He Zhang, Saied Alshahrani, Emad A. Alghamdi, Lian Zhang, Ruoyu Sun, Haizhou Li, Benyou Wang, Jinchao Xu. Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion. ACL, 2025. Paper · arXiv · Code · HF
  20. Yuhao Zhang, Zhiheng Liu, Fan Bu, Ruiyu Zhang, Benyou Wang, Haizhou Li. Soundwave: Less is More for Speech-Text Alignment in LLMs. ACL, 2025. arXiv · Code · HF
  21. Xunlian Dai, Li Zhou, Benyou Wang, Haizhou Li. From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test. EMNLP, Main, 2025. arXiv
  22. Xu Wang, Zihao Li, Benyou Wang, Yan Hu, Difan Zou. Model Unlearning via Sparse Autoencoder Subspace Guided Projections. EMNLP, Main, 2025. arXiv
  23. Wanlong Liu, Junying Chen, Ke Ji, Li Zhou, Wenyu Chen, Benyou Wang. RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions. EMNLP Oral, Main, 2025. arXiv
  24. Guorui Zheng, Xidong Wang, Juhao Liang, Nuo Chen, Yuping Zheng, Benyou Wang. Efficiently Democratizing Medical LLMs for 50 Languages via Mixture of Language Family Experts. ICLR, 2025. arXiv · Code · HF
  25. Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, Baobao Chang. Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language Models. ICLR, 2025. arXiv · Code · HF
  26. Xu Wang, Yan Hu, Wenyu Du, Reynold Cheng, Benyou Wang, Difan Zou. Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis. ICML, 2025. arXiv
  27. Ling-Hao Chen, Shunlin Lu, Ailing Zeng, Hao Zhang, Benyou Wang, Ruimao Zhang, Lei Zhang. MotionLLM: Understanding Human Behaviors from Human Motions and Videos. IEEE TPAMI, 2025. Paper · arXiv · Code · Website · HF
  28. Chengpeng Li, Zhengyang Tang, Ziniu Li, Mingfeng Xue, Keqin Bao, Tian Ding, Ruoyu Sun, Benyou Wang, Xiang Wang, Junyang Lin, Dayiheng Liu. CoRT: Code-integrated Reasoning within Thinking. NeurIPS, 2025. arXiv · Code · HF
  29. Ke Ji, Jiahao Xu, Tian Liang, Qiuzhi Liu, Zhiwei He, Xingyu Chen, Xiaoyuan Liu, Zhijie Wang, Junying Chen, Benyou Wang, Zhaopeng Tu, Haitao Mi, Dong Yu. The First Few Tokens Are All You Need: An Efficient Unsupervised Prefix Fine-Tuning Method. NeurIPS, 2025. arXiv
  30. Yuzhe Yang, Yifei Zhang, Minghao Wu, Kaidi Zhang, Yunmiao Zhang, Honghai Yu, Yan Hu, Benyou Wang. TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets. NeurIPS, 2025. arXiv
  31. Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, Xiangyu Yue. Video-R1: Reinforcing Video Reasoning in MLLMs. NeurIPS, 2025. arXiv · Code · HF
  32. Wanlong Liu, Junxiao Xu, Fei Yu, Yukang Lin, Ke Ji, Wenyu Chen, Yan Xu, Yasheng Wang, Lifeng Shang, Benyou Wang. Question-Free Fine-Tuning: Towards Efficient and Adaptive Reasoning in LLMs. NeurIPS Spotlight, 2025. Paper · arXiv · Code · HF
  33. Junying Chen, Zhenyang Cai, Ke Ji, Xidong Wang, Wanlong Liu, Rongsheng Wang, Jianye Hou, Benyou Wang. Towards Medical Complex Reasoning with LLMs through Medical Verifiable Problems. ACL 2025 Findings. Paper · arXiv · Code · HF
  34. Chenyu Huang, Zhengyang Tang, Dongdong Ge, Shixi Hu, Ruoqing Jiang, Benyou Wang, Zizhuo Wang, Xin Zheng. ORLM: A Customizable Framework in Training Large Models for Automated Optimization Modeling. Operations Research, 2025. Paper · arXiv · Code · HF
  35. Xuhan Huang, Qingning Shen, Yan Hu, Anningzhe Gao, Benyou Wang. LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages. NAACL Findings, 2025. arXiv · Code · HF
  36. Chuyi Kong, Yaxin Fan, Xiang Wan, Feng Jiang, Benyou Wang. Large Language Model as a User Simulator. ACL, 2024. arXiv · Code
  37. Chen Zhang, Benyou Wang, Dawei Song. On Elastic Language Models. ACM TOIS, 2024. Paper · arXiv
  38. Junying Chen, Chi Gui, Ruyi Ouyang, Anningzhe Gao, Shunian Chen, Guiming Hardy Chen, Xidong Wang, Ruifei Zhang, Zhenyang Cai, Ke Ji, Guangjun Yu, Xiang Wan, Benyou Wang. HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale. EMNLP, Main, 2024. arXiv · Code · HF
  39. Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, Benyou Wang. Humans or LLMs as the Judge? A Study on Judgement Biases. EMNLP, Main, 2024. Paper · arXiv
  40. Lei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu. VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment. EMNLP, Main, 2024. arXiv · Code · Website · HF
  41. Xianghong Fang, Jian Li, Qiang Sun, Benyou Wang. Rethinking the Uniformity Metric in Self-Supervised Learning. ICLR, 2024. arXiv · Paper
  42. Zhengyang Tang, Xingxing Zhang, Benyou Wang, Furu Wei. MathScale: Scaling Instruction Tuning for Mathematical Reasoning. ICML, 2024. arXiv
  43. Juhao Liang, Zhenyang Cai, Jianqing Zhu, Huang Huang, Kewei Zong, Bang An, Mosen Alharthi, Juncai He, Lian Zhang, Haizhou Li, Benyou Wang, Jinchao Xu. Alignment at Pre-training! Towards Native Alignment for Arabic LLMs. NeurIPS, 2024. Paper · arXiv · Code · HF
  44. Qianqian Xie, Weiguang Han, Zhengyu Chen, Ruoyu Xiang, Xiao Zhang, Yueru He, Mengxi Xiao, Dong Li, Yongfu Dai, Duanyu Feng, Yijing Xu, Haoqiang Kang, Ziyan Kuang, Chenhan Yuan, Kailai Yang, Zheheng Luo, Tianlin Zhang, Zhiwei Liu, Guojun Xiong, Zhiyang Deng, Yuechen Jiang, Zhiyuan Yao, Haohang Li, Yangyang Yu, Gang Hu, Jiajia Huang, Xiao-Yang Liu, Alejandro Lopez-Lira, Benyou Wang, Yanzhao Lai, Hao Wang, Min Peng, Sophia Ananiadou, Jimin Huang. FinBen: An Holistic Financial Benchmark for Large Language Models. NeurIPS D&B, 2024. Paper · arXiv · Code
  45. Pengcheng Chen, Jin Ye, Guoan Wang, Yanjun Li, Zhongying Deng, Wei Li, Tianbin Li, Haodong Duan, Ziyan Huang, Yanzhou Su, Benyou Wang, Shaoting Zhang, Bin Fu, Jianfei Cai, Bohan Zhuang, Eric J Seibel, Junjun He, Yu Qiao. GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI. NeurIPS D&B, 2024. arXiv · Code · Website · HF
  46. Junying Chen, Xidong Wang, Ke Ji, Anningzhe Gao, Feng Jiang, Shunian Chen, Hongbo Zhang, Dingjie Song, Wenya Xie, Chuyi Kong, Jianquan Li, Xiang Wan, Haizhou Li, Benyou Wang. HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs. COLM 2024. Paper · arXiv · Code · HF
  47. Huang Huang, Fei Yu, Jianqing Zhu, Xuening Sun, Hao Cheng, Song Dingjie, Zhihong Chen, Mosen Alharthi, Bang An, Juncai He, Ziche Liu, Junying Chen, Jianquan Li, Benyou Wang, Lian Zhang, Ruoyu Sun, Xiang Wan, Haizhou Li, Jinchao Xu. AceGPT, Localizing Large Language Models in Arabic. NAACL 2024. Paper · arXiv · Code · HF
  48. Benyou Wang, Xiangbo Wu, Xiaokang Liu, Jianquan Li, Prayag Tiwari, Qianqian Xie. Can Language Models Make Fun? A Case Study in Chinese Comical Crosstalk. ACL, 2023. arXiv · Code
  49. Chen Zhang, Yang Yang, Jiahao Liu, Jingang Wang, Yunsen Xian, Benyou Wang, Dawei Song. Lifting the Curse of Capacity Gap in Distilling Language Models. ACL, 2023. arXiv
  50. Yajiao Liu, Xin Jiang, Yichun Yin, Yasheng Wang, Fei Mi, Qun Liu, Xiang Wan, Benyou Wang. One Cannot Stand for Everyone! Leveraging Multiple User Simulators to Train Task-oriented Dialogue Systems. ACL, 2023. Paper
  51. Benyou Wang, Qianqian Xie, Jiahuan Pei, Prayag Tiwari, Zhao Li, Fu Jie. Pre-trained Language Models in Biomedical Domain: A Survey from Multiscale Perspective. ACM Computing Surveys, 2023. Paper · arXiv
  52. Zhongwei Wan, Xin Liu, Benyou Wang, Jiezhong Qiu, Boyu Li, Ting Guo, Guangyong Chen, Yang Wang. Spatio-Temporal Contrastive Learning Enhanced GNNs for Session-based Recommendation. ACM TOIS, 2023. arXiv
  53. Zhihong Chen, Shizhe Diao, Benyou Wang, Guanbin Li, Xiang Wan. Towards Unifying Medical Vision-and-Language Pre-training via Soft Prompts. ICCV, 2023. arXiv · Code
  54. Zhongwei Wan, Che Liu, Mi Zhang, Jie Fu, Benyou Wang, Sibo Cheng, Lei Ma, César Quilodrán-Casas, Rossella Arcucci. Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias. NeurIPS, 2023. arXiv
  55. Yazhou Zhang, Yang Yu, Qing Guo, Benyou Wang, Dongming Zhao, Sagar Uprety, Dawei Song, Qiuchi Li, Jing Qin. All In One: A Chinese Multi-Modal Dataset for Multi-Affection Detection in Conversations. NeurIPS D&B, 2023. Paper
  56. Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu, Zhihong Chen, Jianquan Li, Guiming Chen, Xiangbo Wu, Zhiyi Zhang, Qingying Xiao, Xiang Wan, Benyou Wang, Haizhou Li. HuatuoGPT, Towards Taming Language Model to Be a Doctor. EMNLP 2023 Findings. Paper · arXiv · Code · HF
  57. Zhihong Chen, Feng Jiang, Junying Chen, Tiannan Wang, Fei Yu, Guiming Chen, Hongbo Zhang, Juhao Liang, Chen Zhang, Zhiyi Zhang, Jianquan Li, Xiang Wan, Benyou Wang, Haizhou Li. Phoenix: Democratizing ChatGPT across Languages. 2023. arXiv · Code · HF
  58. Xiaokang Liu, Jianquan Li, Jingjing Mu, Min Yang, Ruifeng Xu, Benyou Wang. Effective Open Intent Classification with K-center Contrastive Learning and Adjustable Decision Boundary. AAAI, 2023. arXiv
  59. Peng Zhang, Wenjie Hui, Benyou Wang, Donghao Zhao, Dawei Song, Christina Lioma, Jakob Grue Simonsen. Complex-valued Neural Network-based Quantum Language Models. ACM TOIS, 2022. Paper
  60. Sunzhu Li, Peng Zhang, Guobing Gan, Xiuqing Lv, Benyou Wang, Junqiu Wei, Xin Jiang. Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation. EMNLP, Main, 2022. Paper
  61. Benyou Wang, Yuxin Ren, Lifeng Shang, Xin Jiang, Qun Liu. Exploring Extreme Parameter Compression for Pre-trained Language Models. ICLR, 2022. Paper · arXiv
  62. Guobing Gan, Peng Zhang, Sunzhu Li, Xiuqing Lu, Benyou Wang. MorphTE: Injecting Morphology in Tensorized Embeddings. NeurIPS, 2022. arXiv · Code
  63. Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang, Qun Liu, Jakob Grue Simonsen. On Position Embeddings in BERT. ICLR, 2021. Paper
  64. Benyou Wang, Emanuele Di Buiccio, Massimo Melucci. Word2Fun: Modeling Words as Functions for Dynamic Word Embeddings. NeurIPS, 2021. Paper
  65. Benyou Wang, Donghao Zhao, Christina Lioma, Qiuchi Li, Peng Zhang, Jakob Grue Simonsen. Encoding Word Order in Complex Embeddings. ICLR Spotlight, 2020. Paper
  66. Benyou Wang, Qiuchi Li, Massimo Melucci, Dawei Song. Semantic Hilbert Space for Text Representation Learning. WWW, 2019. arXiv · Paper
  67. Wei Zhao, Benyou Wang, Jianbo Ye, Min Yang, Zhou Zhao, Ruotian Luo, Yu Qiao. A Multi-task Learning Approach for Image Captioning. IJCAI, 2018. Paper
  68. Wei Zhao, Benyou Wang, Jianbo Ye, Yongqiang Gao, Min Yang, Xiaojun Chen. PLASTIC: Prioritize Long and Short-term Information in Top-n Recommendation using Adversarial Training. IJCAI, 2018. Paper · Code
  69. Peng Zhang, Jiabin Niu, Zhan Su, Benyou Wang, Liqun Ma, Dawei Song. End-to-End Quantum-like Language Models with Application to Question Answering. AAAI, 2018. Paper
  70. Jun Wang, Lantao Yu, Weinan Zhang, Yu Gong, Yinghui Xu, Benyou Wang, Peng Zhang, Dell Zhang. IRGAN: A Minimax Game for Unifying Generative and Discriminative Information Retrieval Models. SIGIR, 2017. arXiv

View all publications查看完整论文列表