Benyou Wang王本友
Benyou Wang王本友
Assistant Professor, PhD Supervisor and Presidential Young Professor, School of Data Science, CUHK-Shenzhen 香港中文大学(深圳)数据科学学院助理教授、博士生导师、校长青年教授
Benyou Wang leads FreedomAI and has hands-on experience in large-scale model pre-training, post-training, and deployment across agents, world models, healthcare, multimodal AI, and related areas. His publications have received more than 12,000 Google Scholar citations and five best-paper awards or equivalent distinctions, including a SIGIR 2017 Best Paper Award Honorable Mention and the NAACL 2019 Best Explainable Paper award (recognized alongside the authors of four other award-winning papers, including BERT). 王本友带领 FreedomAI 团队,在智能体、世界模型、医疗、多模态等方向具有大规模模型预训练、后训练和部署经验。论文在 Google Scholar 上被引用超过 1.2 万次,曾五次获得最佳论文或同等级荣誉,包括 SIGIR 2017 最佳论文荣誉提名和 NAACL 2019 Best Explainable Paper(与 BERT 等其他四篇获奖论文的作者一同领奖)。
他获天津大学硕士(2017)和帕多瓦大学博士(2022)学位,曾访问哥本哈根大学、蒙特利尔大学,曾在腾讯从事研发,并于 2020–2022 年在华为实习两年零两个月,导师为尚利锋、蒋欣和刘群。研究获欧盟玛丽·居里奖学金及华为、腾讯等项目支持。He earned his master’s degree from Tianjin University (2017) and PhD from the University of Padua (2022), visited the University of Copenhagen and Université de Montréal, and worked in R&D at Tencent. From 2020 to 2022, he interned at Huawei for two years and two months, mentored by Lifeng Shang, Xin Jiang, and Qun Liu. His research has been supported by the Marie Skłodowska-Curie Fellowship and programmes from Huawei and Tencent.
王本友的工作涵盖研究方向规划、大规模模型训练、开源工程、实际部署与科技创业。他组织跨学科团队开展研究,并将研究中积累的方法和工具用于后续项目。FreedomAI 开放模型与数据在 Hugging Face 的下载量已超过百万,GitHub 收藏星标超过 10K;华佗 GPT 已部署到十几家医院和数百家社康,累计访问量近 500 万次,为实际医疗服务提供支持。团队也支持学生探索技术创业,已有近十位校友(alumni)成为创业公司 CTO 或技术创始人,其中,博士生陈志鸿担任创始人兼 CTO 的医疗影像公司以约 8,000 万美元被收购。Benyou Wang works on research planning, large-scale model training, open-source engineering, deployment, and technology ventures. He organizes interdisciplinary teams and carries methods and tools developed in one project into subsequent research. FreedomAI's open models and datasets have surpassed one million downloads on Hugging Face and 10,000 GitHub stars. HuatuoGPT has been deployed in more than a dozen hospitals and hundreds of community health centers, with nearly five million visits, bringing medical AI research into practical service. The group also supports students in technology entrepreneurship; nearly ten alumni have become startup CTOs or technical founders. Among them, doctoral student Zhihong Chen served as founder and CTO of a medical-imaging company that was acquired for approximately US$80 million.
Research vision and leadership
从基础模型到真实世界智能From foundation models to real-world intelligence
团队长期关注如何让人工智能持续学习,并解决实际问题。研究以基础模型为基础,重点探索持续学习的智能体,同时从有实际价值的应用场景中寻找研究问题。The group's long-term agenda asks how AI can keep learning and create value in the real world. It connects foundation models, continuously learning agents, and high-value applications as mutually reinforcing research directions.
- 基础模型与训练方法。Foundation models and training. 研究预训练、后训练与评测方法,在多语言、医疗和多模态方向开展大规模训练与开源实践,开发可复用的模型与工具。Connect pre-training, post-training, and evaluation across multilingual, medical, and multimodal AI, turning methodological advances and large-scale training experience into reusable models and tools.
- 持续学习的智能体。Continuously learning agents. 研究推理、工具使用、环境反馈与递归自提升,构建可重复的任务环境和评测体系,检验模型能力的变化,并据此改进方法。Study reasoning, tool use, environmental feedback, and recursive self-improvement, with reproducible task environments and evaluations that make progress measurable and cumulative.
- 真实场景与技术转化。Real-world applications and translation. 从医疗、运筹优化、金融、教育与机器人等应用中提炼研究问题,通过系统部署和产业合作检验方法的实际效果。Turn needs in healthcare, optimization, finance, education, and robotics into research questions, and assess their value through deployed systems and industry collaboration.
在研究推进中,团队重视方向选择、跨学科合作,以及从模型开发到系统部署的具体工作。长期目标会被拆解为可检验的阶段任务,模型、数据和评测资源则尽可能开放,供团队及合作伙伴继续研究和使用。The team selects research directions, organizes interdisciplinary work, and follows projects from model development through deployment. Long-term goals are broken into testable milestones, with models, data, and evaluations shared to support further work by the team and its collaborators.
Honors and service
Selected recognition代表性荣誉
- SIGIR Best Paper Award Honorable MentionSIGIR 最佳论文荣誉提名
- NAACL 2019 Best Explainable Paper (recognized alongside the authors of four other award-winning papers, including BERT)NAACL 2019 Best Explainable Paper(与 BERT 等其他四篇获奖论文的作者一同领奖)
- NLPCC Best PaperNLPCC 最佳论文
- ICLR Financial AI Workshop Best PaperICLR 金融人工智能研讨会最佳论文
- NeurIPS ResponsibleFM Workshop Outstanding PaperNeurIPS ResponsibleFM 研讨会杰出论文
Competition results竞赛成绩
- 与智子芯元合作,以 AI 自动编写华为 Ascend 算子的全自动方案,曾取得 CannBench 榜单第一名。
- Led a CUHK-Shenzhen team to a Gold Medal in AIMO Progress Prize 2, placing in the top 0.4% among more than 2,000 teams.带领团队与华为联合参加第二届人工智能数学奥林匹克竞赛(AIMO2),在 2,000 余支队伍中进入前 0.4%,获得金牌。
- With Professor Bingyi Jing, co-supervised master's student Shuqi Guo, a core member of team sZs, which won the IEEE ICRA LeHome Challenge real-robot final with 895 points.与荆炳义教授共同指导硕士生郭书杞;郭书杞作为 sZs 队核心成员,在 IEEE ICRA LeHome Challenge 真机赛决赛中以 895 分获得冠军。
Academic service.学术服务。 Publicity Chair of NLPCC 2023, Website Chair of EMNLP 2023, and Area Chair or Senior Area Chair for ICLR, NeurIPS, COLM, ACL, and EMNLP.曾任 NLPCC 2023 宣传主席、EMNLP 2023 网站主席,并多次担任 ICLR、NeurIPS、COLM、ACL、EMNLP 等会议的领域主席或高级领域主席。
Selected projects
Models, systems, and applications模型、系统与应用
The group works on large language models, agents, and their applications in healthcare, multilingual communication, operations research, finance, education, and robotics. The links below connect representative systems with their papers, code, models, data, and application pages.团队围绕大语言模型、智能体及其在医疗、多语言交流、运筹优化、金融、教育和机器人等场景中的应用开展研究。以下介绍代表性项目,并附上相关论文、代码、模型、数据和应用链接。
华佗GPT:面向全球的医疗AIHuatuoGPT · Medical AI for Global Access
In collaboration with Haizhou Li, Xiang Wan, Guangjun Yu, Ruoyu Sun, and clinical partners, he participated in the HuatuoGPT series. The work has expanded from medical dialogue to one-stage domain adaptation, verifiable medical reasoning, and medical vision-language understanding. As of September 2026, public releases across the series had reached more than one million downloads and over 10,000 GitHub stars; related systems have been deployed in more than a dozen hospitals and hundreds of community health centers, with nearly five million visits.与李海洲、万翔、于广军、孙若愚等学者及医疗团队合作,参与华佗GPT系列研发。相关工作已从医学对话扩展到一阶段领域适配、可验证医学推理和医学视觉语言理解。截至 2026 年 9 月,系列公开模型累计下载量超过百万,GitHub Stars 超过一万;相关系统已部署到十几家医院和数百家社康,累计访问量近 500 万次。
从医学对话到专业知识适配。HuatuoGPT 探索面向医疗场景的语言交互;HuatuoGPT-II 将不同来源、格式的医学知识与指令数据统一为输入输出对,以一阶段训练完成领域适配。这条研究路线关注如何把通用模型转化为具备医学知识、能够理解专业问题的医疗助手。From medical dialogue to domain expertise. HuatuoGPT studies language interaction in healthcare. HuatuoGPT-II brings heterogeneous medical knowledge and instruction data into a shared input–output format for one-stage domain adaptation, studying how general models can become assistants that understand medical questions.
从多模态理解到可验证推理。HuatuoGPT-Vision 将医学视觉知识引入多模态模型,连接医学图像与语言理解;HuatuoGPT-o1 则通过可验证医学问题、推理轨迹搜索与强化学习,研究如何提升复杂医学问题的求解能力。两者分别推进医学信息的感知与推理,为更完整的医疗AI系统提供能力基础。From multimodal understanding to verifiable reasoning. HuatuoGPT-Vision connects medical images with language understanding. HuatuoGPT-o1 uses verifiable medical problems, reasoning-trajectory search, and reinforcement learning to improve complex medical problem solving. Together, these research directions develop perception and reasoning capabilities for broader medical AI systems.
跨越语言与资源门槛。Apollo 与 ApolloMoE 在这一医疗AI方向下推进多语言普惠:Apollo 开放轻量医学模型、医学语料与评测资源,ApolloMoE 按语系组织专家模块,将支持范围扩展到 50 种语言,探索训练效率与跨语言医学能力的平衡。通过开放模型、数据和工具,相关工作旨在让更多语言社区和资源有限的研究团队能够参与医疗AI的开发与评测。Across languages and resource constraints. Apollo and ApolloMoE extend this medical AI research toward multilingual access. Apollo releases lightweight medical models, corpora, and evaluation resources; ApolloMoE organizes experts by language family to support 50 languages while studying efficiency and cross-language medical capabilities. Open models, data, and tools aim to help more language communities and resource-constrained research groups develop and evaluate medical AI.
- HuatuoGPT, Towards Taming Language Model to Be a Doctor. EMNLP 2023 Findings. Paper · arXiv · Code ★ 1,326 · HF
- HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs. COLM 2024. Paper · arXiv · Code ★ 414 · HF
- HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale. EMNLP 2024. Paper · arXiv · Code ★ 401 · HF
- Towards Medical Complex Reasoning with LLMs through Medical Verifiable Problems. ACL 2025 Findings. Paper · arXiv · Code ★ 1,357 · HF
- Apollo: A Lightweight Multilingual Medical LLM towards Democratizing Medical AI to 6B People. arXiv, 2024. arXiv · Code · HF · Data
- Efficiently Democratizing Medical LLMs for 50 Languages via Mixture of Language Family Experts. ICLR, 2025. arXiv · Code · HF
Phoenix
Phoenix is an open multilingual dialogue model released in 2023. It made training data, model weights, code, and evaluation resources available for Chinese, English, and several lower-resource languages. Its results were competitive with contemporary open chat models, and it entered the leading group in third-party Chinese LLM evaluations at release.Phoenix 是团队于 2023 年发布的开放多语言对话模型,面向中文、英文和多种低资源语言开放训练数据、模型权重、代码与评测资源。其效果在同期开放对话模型中具有竞争力,发布初期进入第三方中文大模型评测前列。
AceGPT
With Professor Jinchao Xu's group and support from Peng Cheng Laboratory, the team developed Arabic language models adapted to local language, culture, and values. AceGPT was among the leading open Arabic large language models at release; follow-up work explored native alignment during pre-training and progressive vocabulary expansion.在许进超教授团队带领和鹏城实验室支持下,团队开展面向阿拉伯语言、文化与价值观适配的大模型研究。AceGPT 发布时是表现领先的开源阿拉伯语大模型;后续工作进一步探索预训练原生对齐(native alignment)与渐进式词表扩展(progressive vocabulary expansion),并使用大规模华为昇腾 910A 集群训练更大规模模型。
- AceGPT, Localizing Large Language Models in Arabic. NAACL 2024. Paper · arXiv · Code · HF
- Alignment at Pre-training! Towards Native Alignment for Arabic LLMs. NeurIPS, 2024. Paper · arXiv · Code · HF
- Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion. ACL, 2025. Paper · arXiv · Code · HF
ORLM
Together with Professor Zizhuo Wang and Cardinal Operations, the team developed ORLM for translating natural-language business problems into executable mathematical optimization models and solver code. The work introduced OR-Instruct and IndustryOR and connects with Cardinal Operations' COLORMind decision-intelligence product line. Mamo extends this direction with a mathematical-modeling benchmark and solvers that connect natural-language problems to executable formal models. CALM Before the STORM further studies native reasoning for optimization modeling.与王子卓教授及杉数科技共同开发 ORLM,用于将自然语言业务问题转化为可执行的数学优化模型与求解代码,并提出 OR-Instruct 和 IndustryOR。相关研究与杉数科技 COLORMind 智能决策产品线相结合。Mamo 则通过数学建模基准与求解器,探索如何把自然语言问题转化为可执行的形式化模型。CALM Before the STORM 进一步研究优化建模中的原生推理能力。
- ORLM: A Customizable Framework in Training Large Models for Automated Optimization Modeling. Operations Research, 2025. Paper · arXiv · Code ★ 275 · HF
- LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages. NAACL Findings, 2025. arXiv · Code ★ 15 · HF
- CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling. ICML, 2026. arXiv · Code · HF
TinyDeepSeek
从零开始训练高效的小规模模型。TinyDeepSeek 从零开始预训练,构建了超过 3T tokens 的数据,工作覆盖数据处理、模型架构、预训练与后训练。根据团队实验结果,模型在约 3B 参数规模的对比中达到 SOTA 水平,团队完成了从数据准备到训练系统的完整研发。Training from scratch for capable small-scale models. TinyDeepSeek builds a pre-training corpus of more than 3 trillion tokens and connects data processing, model architecture, pre-training, and post-training. In the team's evaluations, it achieved state-of-the-art performance among models at approximately the 3B-parameter scale, demonstrating end-to-end development from data to training systems.
团队自 2025 年 1 月左右开始探索 Loop Transformer 的效果,研究通过循环计算与参数复用提升模型能力。项目开放了训练代码、0.5B 与 3.3B 基础模型及中间检查点,便于开展低成本的架构实验和持续训练研究。The team began exploring Loop Transformers around January 2025, studying how recurrent computation and parameter reuse can improve model capabilities. The project releases training code, 0.5B and 3.3B base models, and intermediate checkpoints to support lower-cost architecture experiments and continued training research.
多模态与长上下文:ALLaVA、LongLLaVAMultimodal & Long-context Models · ALLaVA & LongLLaVA
从高质量图文数据到长上下文理解。ALLaVA 使用 GPT-4V 生成细粒度图像描述和视觉问答数据,支持轻量视觉语言模型的图文对齐与指令微调,探索在有限计算资源下提升多模态能力。LongLLaVA 进一步研究如何高效理解大量图像,MileBench 则提供长上下文评测,形成数据、模型与评测相互支撑的研究路线。From high-quality visual data to long-context understanding. ALLaVA uses GPT-4V to synthesize detailed image descriptions and visual question-answer pairs for alignment and instruction tuning of lightweight vision-language models. LongLLaVA studies efficient understanding of many images, while MileBench provides long-context evaluation, connecting data, models, and benchmarks.
LongLLaVA 结合 Mamba 与 Transformer 架构,配合时空数据构建和渐进训练,在单张 A100 80GB 上处理近千张图像。MileBench 从真实任务与诊断性测试两个角度,评估多模态大模型的长上下文理解能力。LongLLaVA combines Mamba and Transformer blocks with temporal and spatial data construction and progressive training to process nearly 1,000 images on a single A100 80GB GPU. MileBench evaluates long-context multimodal understanding through realistic tasks and diagnostic tests.
- ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models. arXiv, 2024. arXiv · Code
- LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture. EMNLP Findings, 2025. arXiv · Code
- MileBench: Benchmarking MLLMs in Long Context. COLM, 2024. arXiv · Code · HF · Website
PhoneBuddy
PhoneBuddy 系列研究手机智能体的训练,将真实应用与 PhoneWorld 重建的模拟应用纳入统一流程:先通过轨迹数据进行监督微调,再在真实或混合环境中进行强化学习。系列工作还包括隐私评测 MyPhoneBench、安全评测、PhoneWorld 环境构建,以及结合 GUI、CLI 和工具动作的 PhoneHarness。The PhoneBuddy series studies training for phone agents using both real apps and reconstructed apps in PhoneWorld. Agents first learn from trajectories through supervised fine-tuning, then undergo reinforcement learning in real or mixed environments. The series also covers privacy evaluation with MyPhoneBench, safety evaluation, PhoneWorld environment construction, and PhoneHarness for mixed GUI, CLI, and tool actions.
- Do Phone-Use Agents Respect Your Privacy?. arXiv, 2026. arXiv · Code
- Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents. arXiv, 2026. arXiv
- PhoneWorld: Scaling Phone-Use Agent Environments. arXiv, 2026. arXiv
- PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions. arXiv, 2026. arXiv · Code · HF
- PhoneBuddy: Training Open Models for Agentic Phone Use. arXiv, 2026. arXiv
语音大模型Speech models
语音方向涵盖模型训练、可控生成与交互评测:Soundwave 研究语音与文本的高效对齐;TTS-Hub 通过模块化 LoRA 及其组合实现可控语音合成;S2S-Arena 评估语音到语音模型对副语言指令的遵循能力,EchoMind 则关注语音交互中的共情表现。Human or Machine? 通过初步图灵测试,研究语音到语音交互中人类与模型的可区分性。The speech research covers model training, controllable generation, and interaction evaluation. Soundwave studies efficient speech–text alignment; TTS-Hub uses modular LoRAs and their composition for controllable speech synthesis. S2S-Arena evaluates paralinguistic instruction following in speech-to-speech models, while EchoMind examines empathy in spoken interaction. Human or Machine? uses a preliminary Turing test to study whether people can distinguish human from model speech-to-speech interaction.
- Soundwave: Less is More for Speech-Text Alignment in LLMs. ACL, 2025. arXiv · Code · HF
- TTS-Hub: Leveraging Modular LoRAs and Arithmetic Composition for Controllable Text-to-Speech. EMNLP, Main, 2026. Paper
- S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models. ACL, 2026. arXiv · Code · HF
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models. ICLR, 2026. arXiv · Code · Website · HF
- Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction. ICLR, 2026. arXiv
智能体学习与创造Agent Learning & Creation
从环境交互到可执行成果。我们研究如何构建可交互、可反馈、可检验的环境,让智能体通过工具使用与强化学习提升解决问题的能力,并将人的需求转化为可运行的代码、三维模型和可玩的游戏。环境工程提供学习和验证的条件,Agentic RL 研究如何利用反馈改进行为,智能体创作则把这些能力连接到具体成果。From interaction to executable artifacts. We study environments that support interaction, feedback, and verification, enabling agents to develop problem-solving skills through tool use and reinforcement learning and to turn human intent into runnable code, 3D models, and playable games. Environment engineering provides the conditions for learning and validation; agentic RL studies how feedback improves behavior; agentic creation connects these capabilities to concrete outputs.
这一方向围绕智能体可交互的环境、工具和反馈机制展开。SepsisAgent 将 Clinical World Model 用作患者动态模拟器与训练环境,研究脓毒症场景中的状态预测和决策学习,并在回顾性患者轨迹上评估。This direction studies interactive environments, tools, and feedback for agents. SepsisAgent uses a Clinical World Model as a patient-dynamics simulator and training environment for state prediction and decision learning in sepsis, evaluated on retrospective patient trajectories.
CoRT 将代码执行融入推理过程,结合监督微调与强化学习训练模型使用代码解释器。TwinMarket 则用大语言模型驱动的多智能体模拟金融市场,研究个体行为如何形成市场层面的现象。CoRT integrates code execution into reasoning, using supervised fine-tuning and reinforcement learning to teach code-interpreter use. TwinMarket uses LLM-driven agents to simulate financial markets and study how individual behavior produces market-level phenomena.
面向专业任务的工具调用与证据推理。OpenClaw-Medical-Skills 汇集可供智能体调用的医学技能与工具,支持医学信息检索和相关工作流程;Legal-R1 系列通过多轮检索与推理,在获取证据、分析问题和继续检索之间迭代,围绕证据完成法律推理。这些工作将智能体的工具使用与推理能力连接到医疗和法律领域的具体任务。Tool use and evidence-based reasoning for professional tasks. OpenClaw-Medical-Skills brings together medical skills and tools that agents can call for medical information retrieval and related workflows. Legal-R1 alternates between retrieving evidence, analyzing questions, and further retrieval to support legal reasoning. These projects connect agent tool use and reasoning to concrete tasks in healthcare and law.
这一方向研究模型如何生成可检验的创作结果。MicroVerse 面向微观生物过程生成科学可视化视频,并建立相应数据与评测;BlenderLLM 将文本需求转化为可在 Blender 中运行的建模脚本,通过自我改进训练提升三维建模能力;GameCraft-Bench 在真实 Godot 引擎中评测智能体端到端制作可玩游戏的能力,通过运行项目、重放操作和评分规则检验结果。This direction studies verifiable creative outputs. MicroVerse generates scientific videos of microscopic biological processes and provides supporting data and evaluation. BlenderLLM translates text instructions into modeling scripts that run in Blender, using self-improvement training to develop 3D modeling capabilities. GameCraft-Bench evaluates end-to-end creation of playable games in the Godot engine by running projects, replaying actions, and checking scoring rubrics.
这些工作的共同目标,是让AI更好地服务研究、教育与创作:帮助研究者探索复杂系统,让创作者表达并实现设计意图,也让学习者获得可交互的内容。我们关注成果是否能够运行、检验和迭代,让智能体能力与实际任务的完成质量建立联系。The shared goal is to make AI more useful for research, education, and creation: helping researchers explore complex systems, creators realize design ideas, and learners engage with interactive content. We focus on whether outputs can be run, evaluated, and iteratively improved, connecting agent capabilities to the quality of completed tasks.
- Agentifying Patient Dynamics within LLMs through Interacting with Clinical World Model. arXiv, 2026. arXiv · Code · HF
- CoRT: Code-integrated Reasoning within Thinking. NeurIPS, 2025. arXiv · Code · HF
- TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets. NeurIPS, 2025. arXiv
- MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation. ICLR, 2026. arXiv · Code · HF
- BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement. arXiv, 2024. arXiv · Code ★ 302 · HF
- GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?. arXiv, 2026. arXiv · Code · Website
- OpenClaw-Medical-Skills. Code ★ 3,004
- Legal-R1: Agentic Retrieval for Evidence-Based Legal Reasoning. Paper
Innovation and mentorship
Student ventures and technology transfer学生创业与技术转化
The group supports students and research staff in translating research into practical products and services. Nearly ten alumni have gone on to serve as startup CTOs or technical founders.团队支持学生和研究人员将科研成果转化为实际产品与服务,已有近十位校友(alumni)成为创业公司 CTO 或技术创始人。
- Zhihong Chen陈志鸿PhD graduate · Founder and CTO. Cognita Imaging, medical-imaging AI; acquired by Mosaic Clinical Technologies in 2025 in a transaction valued at approximately US$80 million.博士毕业生 · 创始人兼 CTO。Cognita Imaging,医疗影像 AI;公司于 2025 年被 Mosaic Clinical Technologies 收购,交易规模约 8,000 万美元。
- Jianye Hou侯建业PhD student · CTO. Metastone, high-performance computing. The company has revenue approaching RMB 1 billion.博士生 · CTO。是石科技,高性能计算;收入接近十亿。
- Shunlin Lu路舜麟PhD student · CTO. NeoteAI, embodied intelligence; financing totals close to RMB 100 million.博士生 · CTO。新智具身(NeoteAI),具身智能;融资规模近人民币 1 亿元。
- Jiajun You游佳君PhD student · CEO. Yuniu Technology (驭牛科技), currently raising funds.博士生 · CEO。驭牛科技,目前正在融资。
- Xiao Yang肖杨PhD student · CEO. Shenzhen Sikoo Intelligent Information Services, developing AI education products including Shuta AI.博士生 · CEO。深圳思酷智能信息服务有限公司,开展 AI 教育业务,包括薯塔 AI。
- Zhiqi Gao高治淇MPhil student · Founder. Heryin Yuanzhu (Shenzhen) Intelligent Technology, working on spatial intelligence, 3D reconstruction, and digital twins.MPhil 学生 · 创始人。和瑛元筑(深圳)智能科技有限公司,主要开展空间智能、三维重建与数字孪生相关业务。
- Huiteng Xiao肖徽腾Research assistant · Founder. An early-stage venture combining AI and hardware.研究助理 · 创始人。开展 AI 与硬件结合的早期创业项目。
- Shijun Chu褚士钧Research assistant · COO. Chaotic Pendulum Technology (Shenzhen), working on financial AI.研究助理 · COO。混沌摆科技(深圳)有限公司,主要开展 AI 金融相关业务。
Operating and financing figures above reflect information provided by the group as of September 2026; where public disclosures are available, the latest disclosure takes precedence.以上经营与融资信息根据团队提供的截至 2026 年 9 月资料整理;如有公开披露,以最新披露为准。
Teaching and supervision
Natural language processing自然语言处理课程
Benyou Wang teaches Natural Language Processing each year to approximately 500 students. The course covers NLP foundations, large language models, agents, and practical system development, with course materials available online.王本友每年讲授自然语言处理课程,约有 500 名学生修读。课程涵盖自然语言处理基础、大语言模型、智能体与系统实践,并持续开放课程资料。
Culture and mentorship
自由互助,以学生为中心Freedom, mutual support, and student-led innovation
每个人追求自己的自由就是给整个世界追求自由Each person's pursuit of their own freedom is a pursuit of freedom for the whole world.
实验室重视自由探索和相互帮助,以学生的研究兴趣与成长为中心。我们尊重每个人的研究兴趣和发展选择,鼓励独立探索、开放交流,也鼓励同学之间相互支持。The laboratory has developed a research culture of freedom and mutual support, with an innovation network centered on students. We respect individual research interests and career choices, and encourage independent exploration, open exchange, and collaboration.
学生在 HY、Kimi、MiniMax、华为、大疆、阶跃、百川、蚂蚁、Qwen 等团队与企业实习,也有近十位同学选择 AI 创业。这些经历让实验室与学术界、产业界保持了持续的交流与合作。Students undertake internships at HY, Kimi, MiniMax, DJI, StepFun, Baichuan, Ant Group, Qwen, and other teams and companies. Nearly ten students have also embarked on AI ventures, contributing to sustained collaboration between research and industry.
He also co-supervises doctoral students with Professors Haizhou Li, Bingyi Jing, Ming Yan, Yongtao Guan, Hongyuan Zha, and Yilun Chen in areas including language intelligence, speech, multimodal learning, optimization, statistical learning, and robotics.他还与李海洲、荆炳义、严明、官永涛、查宏远、陈逸伦等教授联合指导博士生,研究方向涵盖语言智能、语音、多模态学习、运筹优化、统计学习和机器人。
Publications
Selected publications代表性论文
显示 篇论文Results:
没有符合当前条件的论文,请调整筛选条件。No publications match these filters. Try a different selection.
- Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning. ACL, 2026. arXiv
- Human or LLM as Standardized Patients? A Comparative Study in Medical Education. ACL, 2026. arXiv
- S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models. ACL, 2026. arXiv · Code · HF
- Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion. ACL, 2026. arXiv · Code · HF
- It Talks Like a Patient, But Feels Different: Co-Designing AI Standardized Patients. CHI, 2026. Paper · arXiv
- Agentic Rubrics: Interfacing Agents with Heterogeneous Environments using Executable Rubrics. EMNLP, Main, 2026. Paper
- Steer on Rubrics: Auditable Steering of Schwartz Values in Language Models. EMNLP, Main, 2026. Paper
- TTS-Hub: Leveraging Modular LoRAs and Arithmetic Composition for Controllable Text-to-Speech. EMNLP, Main, 2026. Paper
- Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse Autoencoders. ICLR, 2026. arXiv
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models. ICLR, 2026. arXiv · Code · Website · HF
- Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction. ICLR, 2026. arXiv
- LiveClin: A Live Clinical Benchmark without Leakage. ICLR, 2026. arXiv · Code · HF
- MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation. ICLR, 2026. arXiv · Code · HF
- CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling. ICML, 2026. arXiv · Code · HF
- CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning. ICML, 2026. arXiv
- OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation. ICML, 2026.
- Can Audio Language Models Listen Between the Lines? A Study on Metaphorical Reasoning via Unspoken. ACM MM, 2025. Paper · Code · HF
- Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging. ACL, 2025. arXiv · Code · HF
- Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion. ACL, 2025. Paper · arXiv · Code · HF
- Soundwave: Less is More for Speech-Text Alignment in LLMs. ACL, 2025. arXiv · Code · HF
- From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test. EMNLP, Main, 2025. arXiv
- Model Unlearning via Sparse Autoencoder Subspace Guided Projections. EMNLP, Main, 2025. arXiv
- RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions. EMNLP Oral, Main, 2025. arXiv
- Efficiently Democratizing Medical LLMs for 50 Languages via Mixture of Language Family Experts. ICLR, 2025. arXiv · Code · HF
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language Models. ICLR, 2025. arXiv · Code · HF
- Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis. ICML, 2025. arXiv
- MotionLLM: Understanding Human Behaviors from Human Motions and Videos. IEEE TPAMI, 2025. Paper · arXiv · Code · Website · HF
- CoRT: Code-integrated Reasoning within Thinking. NeurIPS, 2025. arXiv · Code · HF
- The First Few Tokens Are All You Need: An Efficient Unsupervised Prefix Fine-Tuning Method. NeurIPS, 2025. arXiv
- TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets. NeurIPS, 2025. arXiv
- Video-R1: Reinforcing Video Reasoning in MLLMs. NeurIPS, 2025. arXiv · Code · HF
- Question-Free Fine-Tuning: Towards Efficient and Adaptive Reasoning in LLMs. NeurIPS Spotlight, 2025. Paper · arXiv · Code · HF
- Towards Medical Complex Reasoning with LLMs through Medical Verifiable Problems. ACL 2025 Findings. Paper · arXiv · Code · HF
- ORLM: A Customizable Framework in Training Large Models for Automated Optimization Modeling. Operations Research, 2025. Paper · arXiv · Code · HF
- LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages. NAACL Findings, 2025. arXiv · Code · HF
- Large Language Model as a User Simulator. ACL, 2024. arXiv · Code
- On Elastic Language Models. ACM TOIS, 2024. Paper · arXiv
- HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale. EMNLP, Main, 2024. arXiv · Code · HF
- Humans or LLMs as the Judge? A Study on Judgement Biases. EMNLP, Main, 2024. Paper · arXiv
- VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment. EMNLP, Main, 2024. arXiv · Code · Website · HF
- Rethinking the Uniformity Metric in Self-Supervised Learning. ICLR, 2024. arXiv · Paper
- MathScale: Scaling Instruction Tuning for Mathematical Reasoning. ICML, 2024. arXiv
- Alignment at Pre-training! Towards Native Alignment for Arabic LLMs. NeurIPS, 2024. Paper · arXiv · Code · HF
- FinBen: An Holistic Financial Benchmark for Large Language Models. NeurIPS D&B, 2024. Paper · arXiv · Code
- GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI. NeurIPS D&B, 2024. arXiv · Code · Website · HF
- HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs. COLM 2024. Paper · arXiv · Code · HF
- AceGPT, Localizing Large Language Models in Arabic. NAACL 2024. Paper · arXiv · Code · HF
- Can Language Models Make Fun? A Case Study in Chinese Comical Crosstalk. ACL, 2023. arXiv · Code
- Lifting the Curse of Capacity Gap in Distilling Language Models. ACL, 2023. arXiv
- One Cannot Stand for Everyone! Leveraging Multiple User Simulators to Train Task-oriented Dialogue Systems. ACL, 2023. Paper
- Pre-trained Language Models in Biomedical Domain: A Survey from Multiscale Perspective. ACM Computing Surveys, 2023. Paper · arXiv
- Spatio-Temporal Contrastive Learning Enhanced GNNs for Session-based Recommendation. ACM TOIS, 2023. arXiv
- Towards Unifying Medical Vision-and-Language Pre-training via Soft Prompts. ICCV, 2023. arXiv · Code
- Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias. NeurIPS, 2023. arXiv
- All In One: A Chinese Multi-Modal Dataset for Multi-Affection Detection in Conversations. NeurIPS D&B, 2023. Paper
- HuatuoGPT, Towards Taming Language Model to Be a Doctor. EMNLP 2023 Findings. Paper · arXiv · Code · HF
- Phoenix: Democratizing ChatGPT across Languages. 2023. arXiv · Code · HF
- Effective Open Intent Classification with K-center Contrastive Learning and Adjustable Decision Boundary. AAAI, 2023. arXiv
- Complex-valued Neural Network-based Quantum Language Models. ACM TOIS, 2022. Paper
- Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation. EMNLP, Main, 2022. Paper
- Exploring Extreme Parameter Compression for Pre-trained Language Models. ICLR, 2022. Paper · arXiv
- MorphTE: Injecting Morphology in Tensorized Embeddings. NeurIPS, 2022. arXiv · Code
- On Position Embeddings in BERT. ICLR, 2021. Paper
- Word2Fun: Modeling Words as Functions for Dynamic Word Embeddings. NeurIPS, 2021. Paper
- Encoding Word Order in Complex Embeddings. ICLR Spotlight, 2020. Paper
- Semantic Hilbert Space for Text Representation Learning. WWW, 2019. arXiv · Paper
- A Multi-task Learning Approach for Image Captioning. IJCAI, 2018. Paper
- PLASTIC: Prioritize Long and Short-term Information in Top-n Recommendation using Adversarial Training. IJCAI, 2018. Paper · Code
- End-to-End Quantum-like Language Models with Application to Question Answering. AAAI, 2018. Paper
- IRGAN: A Minimax Game for Unifying Generative and Discriminative Information Retrieval Models. SIGIR, 2017. arXiv