Research & Projects

Selected research topics and systems developed in the lab (safe RL, robot learning, LLM agents, and beyond).

Research

Concretely, the lab builds algorithms, theory, benchmarks, and open-source systems for safe decision-making under uncertainty — spanning safe reinforcement learning, robot learning, and LLM-based agents, and evaluates them in robotics, autonomous driving, smart grid, and semiconductor-manufacturing scenarios.

His work has been featured in leading publications, including top-tier journals and conferences such as Science Advances, the Journal of Artificial Intelligence, IEEE Transactions on Pattern Analysis and Machine Intelligence, NeurIPS, and other prestigious venues.

Safe Reinforcement Learning Robust Reinforcement Learning Robot Learning Large Language Models & Agents Multi-Agent Reinforcement Learning AI Safety & Trustworthy AI Semiconductor Manufacturing

More on our publications and open-source softwares.

Projects

2026

CheetahClaws demo

CheetahClaws: An Agent Harness Infrastructure for Long-Horizon, Multi-Model, and Tool-Using AI Systems

CheetahClaws is a fast and easy-to-use agent harness infrastructure designed for long-horizon, multi-model, and tool-using AI systems. It provides a flexible and extensible framework supporting frontier and local models, tool use, memory, multi-agent orchestration, and autonomous task execution for building and studying advanced AI agents.

2025

Robust Gymnasium overview

Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning

This benchmark aims to advance robust reinforcement learning for real-world applications and domain adaptation. The benchmark provides a comprehensive set of tasks that cover various robustness requirements in the face of uncertainty on state, action, reward, and environmental dynamics, and spans diverse applications including control, robot manipulation, dexterous hands, and so on.

2024

TeaMs-RL framework

TeaMs-RL: Teaching LLMs to Teach Themselves Better Instructions via Reinforcement Learning

This work presents TeaMs-RL, a method that leverages reinforcement learning to directly generate instruction datasets for fine-tuning LLMs. By minimizing reliance on human feedback and external queries, TeaMs-RL enhances data quality, privacy, and model capabilities, making it a valuable approach for foundation-model post-training.

2024

Safe MARL roundabout scenario

Safe Multi-Agent Reinforcement Learning with Bilevel Optimization in Autonomous Driving

This study introduces a safe MARL method based on a Stackelberg model with bi-level optimization, featuring two algorithms, CSQ and CS-MADDPG, for autonomous driving applications. Experimental results demonstrate its superior safety and performance compared with strong MARL baselines in challenging driving scenarios.

2023

Safe multi-robot control

Safe Multi-Agent Reinforcement Learning for Multi-Robot Control

We investigate safe MARL for multi-robot control on cooperative tasks, in which each individual robot has to not only meet its own safety constraints while maximizing its reward, but also consider those of others to guarantee safe team behaviors.

More topics we work on — including safe RL for semiconductor manufacturing and smart-grid operation, robust RL, offline RL, and agentic web research — are summarized on the homepage and in our publications.