I am Bo-Wen Zhang (张博闻), a Ph.D. student at the School of Intelligence Science and Technology, Nanjing University.

我是张博闻,南京大学智能科学与技术学院博士生。

Research Interests研究方向
Long-term Interest长期研究兴趣

My long-term research goal is recursive self-improvement (RSI): agents that learn from experience to improve their capabilities and their ability to learn and improve. I am interested in the loop between task design, environments, evaluation, and updates to model weights, memory, tools, and harnesses. One question within this goal is learning from extremely delayed real-world rewards: how can agents learn and provide credible evidence of improvement when true outcomes are unavailable in time and historical supervision risks leaking future information?

我的长期研究目标是实现递归式自我改进(RSI):让智能体从经验中提升能力,并进一步提升自身学习与改进的能力。我关注任务设计、环境、评价,以及模型参数、记忆、工具与 harness 的更新如何形成闭环。其中一个问题是真实世界超级延迟奖励下的学习:当真实标签来不及用于训练、历史监督又容易泄漏未来信息时,如何学习并提供可信的改进证据?

One possible route is learning in accelerated simulations that preserve the mechanisms relevant to our decisions, then testing whether the resulting improvements transfer to the real world.

一条可能的路径是在可加速的模拟环境中学习:保留影响决策的关键机制,再检验其中学到的改进能否迁移到现实。

Toward RSI: experience and delayed rewards走向 RSI:经验闭环与延迟奖励

I am pursuing my Ph.D. under the supervision of Prof. Lan-Zhe Guo, a member of the LAMDA Group led by Prof. Zhi-Hua Zhou.

目前,我在郭兰哲教授的指导下攻读博士学位;郭老师是周志华教授领导的 LAMDA 团队成员。

News动态

  • 2026.01 Our paper ChinaTravel was accepted to ICLR 2026.论文 ChinaTravel 被 ICLR 2026 接收。
  • 2025.12 Mind the Gap to Trustworthy LLM Agents received the Best Student Paper Award at the AAAI 2026 TrustAgent Workshop.Mind the Gap to Trustworthy LLM Agents 获 AAAI 2026 TrustAgent Workshop 最佳学生论文奖。

Publications论文

* denotes equal contribution.

* 表示同等贡献。

Journal Papers期刊论文

Conference & Workshop Papers会议与研讨会论文

Preprints预印本

Projects项目

Illustration of an agent planning and validating a multi-day journey through a Chinese city

ChinaTravel

An open-ended benchmark and sandbox for realistic multi-day, multi-POI travel planning with compositional constraint validation for language agents.

一个面向语言智能体的开放式基准与沙盒,用于真实的多日、多景点旅行规划,并支持组合式约束校验。

Neuro-Symbolic AI神经符号 AILLM Agents大模型智能体

Experience经历

ByteDance

2026.04 - Present2026.04 - 至今

LLM Algorithm Intern大模型算法实习生

Douyin E-commerce抖音电商

Training and exploratory work in LLM reinforcement learning and agentic RL.参与大模型强化学习与 Agentic RL 的训练及探索性工作。

Education教育

Engineering工学

Ph.D. Student博士生

2026.09 - Present2026.09 - 至今

School of Intelligence Science and Technology, Nanjing University南京大学智能科学与技术学院

Advisor: Prof. Lan-Zhe Guo导师:郭兰哲教授

Science理学

Bachelor of Science理学学士

2022.09 - 2026.06

School of Intelligence Science and Technology, Nanjing University南京大学智能科学与技术学院

Advisor: Prof. Lan-Zhe Guo导师:郭兰哲教授

Personal个人

Deutschland football德国足球 Bayern München