Welcome

About me

I am an M.S. student in Artificial Intelligence at the School of Data Science, Fudan University, advised by Prof. Yanwei Fu. I am also a research intern at Shanghai Artificial Intelligence Laboratory, working with Dr. Dongrui Liu.

My research focuses on safe and trustworthy AI agents and data-centric visual learning. I work on evaluating agent behavior, building diagnostic guardrails, and learning across domains with limited supervision.

Before joining Fudan, I received my B.S. in Software Engineering from Dalian University of Technology. I am open to collaboration and discussions — feel free to get in touch.

Current researchExplore research →

Agent safety & trustworthiness

Trajectory evaluation, step-level guardrails, and conflicts in shared environments.

Data-centric visual learning

Retrieval-guided generation and cross-domain few-shot object detection.

Recent updatesView publications →

  • Sep 2026ClashBench was released on arXiv, benchmarking agent conflicts over shared resources.
  • Aug 2026StepGuard was accepted to EMNLP 2026.
  • Jul 2026ATBench was accepted to COLM 2026.
Earlier updates
  • Apr 2026ATBench was released on arXiv for trajectory-level agent safety evaluation and diagnosis.
  • Jan 2026AgentDoG was released as a diagnostic guardrail framework for AI agent safety and security.
  • 2025Domain-RAG was accepted to NeurIPS 2025.

Research

What I work on

Trustworthy agents and visual learning across domains.

01 / AGENT SAFETY

Agent safety & trustworthiness

I study how to evaluate and diagnose safety failures in AI agent trajectories, and how guardrails can balance safety with task utility. This includes step-level supervision, realistic evaluation benchmarks, and conflicts between agents sharing resources.

02 / VISUAL LEARNING

Data-centric visual learning

I explore retrieval-guided compositional image generation for cross-domain few-shot object detection, with an interest in improving visual learning when target-domain annotations are limited.

Also interested in

Trustworthy multimodal foundation models, jailbreak and defense, alignment, interpretability, efficient reasoning, computer-use and mobile agents, and embodied agent safety.

Publications

Selected publications

* denotes equal contribution. Full list on Google Scholar ↗

Peer-reviewed publications3 papers

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
Zhijie Zheng*, Yu Li*, Chen Qian, Yuqian Fu, Yanwei Fu, Lu Sheng, Jing Shao, Dongrui Liu
* Equal contribution.
EMNLP 2026
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
Yu Li, Haoyu Luo, Yuejin Xie, Yuqian Fu, Zhonghao Yang, Shuai Shao, Qihan Ren, Wanying Qu, Yanwei Fu, Yujiu Yang, Jing Shao, Xia Hu, Dongrui Liu
COLM 2026
Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection
Yu Li, Xingyu Qiu, Yuqian Fu, Jie Chen, Tianwen Qian, Xu Zheng, Danda Pani Paudel, Yanwei Fu, Xuanjing Huang, Luc Van Gool, Yu-Gang Jiang
NeurIPS 2025

Preprints3 papers

ClashBench: Conflicts Leading Agents to Seize and Harm
Yuejin Xie*, Yu Li*, Dadi Guo, Qingyu Liu, Yuqian Fu, Yanwei Fu, Yujiu Yang, Xia Hu, Dongrui Liu
* Equal contribution.
arXiv 2026
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-CodeX
Zhonghao Yang, Yu Li, Yanxu Zhu, Tianyi Zhou, Yuejin Xie, Haoyu Luo, Jing Shao, Xia Hu, Dongrui Liu
arXiv 2026
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
Dongrui Liu, Qihan Ren, Chen Qian, Shuai Shao, Yuejin Xie, Yu Li, Zhonghao Yang, Haoyu Luo, et al.
arXiv 2026

Experience

Research experience

Nov 2025 – Present

Shanghai Artificial Intelligence Laboratory

Research Intern · Center for Safety and Trustworthiness

Working with Dr. Dongrui Liu.

Worked on data pipelines and benchmark construction for trajectory-level agent safety diagnosis, including scalable synthesis, filtering, and evaluation data design.

Education

Academic background

2024 – 2027

Fudan University

M.S. in Artificial Intelligence · School of Data Science

Advisor: Prof. Yanwei Fu

2020 – 2024

Dalian University of Technology

B.S. in Software Engineering

Contact

Get in touch

I welcome conversations and collaborations on agent safety, trustworthy AI, and visual learning. Email is the best way to reach me.