🎯 I'm on the 2026 Job Market — actively seeking research positions in both industry and academia. 🤝 I'm also always excited to explore new research collaborations and exchange ideas with fellow researchers. 📩 Please feel free to reach out via ruhwang@iu.edu or LinkedIn — I'd love to chat!
About
I am a Ph.D. candidate in Computer Engineering at Indiana University, advised by Prof. Dongruo Zhou. My research focuses on LLM agents, agent harnesses, reinforcement learning for LLMs, and scalable post-training. I build self-evolving, tool-using agents that interact with environments, complete long-horizon tasks, and improve through execution feedback.
My recent work spans agent harness design and interpretability, long-horizon agent evaluation, efficient RLVR, uncertainty-aware federated reasoning, and privacy-preserving LLM collaboration. I also work on multimodal agentic recommender systems and offline and safe reinforcement learning.
I am currently a Ph.D. Research Intern at Tencent AI Lab (Hunyuan Frontier Lab), where I work on agentic reinforcement learning and self-evolving agent infrastructure. Previously, I was a Ph.D. Research Intern at Mitsubishi Electric Research Laboratories, where I developed generative quantum models for few-shot learning.
Selected Publications
View All →More Memory, Worse Agents: Error Reproduction and Anti-Persistence in LLM Agents
Ruhan Wang, Kishan Panaganti, Dongruo Zhou
The Fortieth Conference on Neural Information Processing Systems (NeurIPS)
Identifies a structural error-reproduction failure mode in persistent abstraction memory for self-improving LLM agents, and proposes ANTI-PERSISTENCE: a memory module that stores only factual interaction records and synthesizes task-specific abstractions on demand via an adaptive contextual UCB bandit over abstraction modes.
Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable
Ruhan Wang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Yue Yu, Junyao Yang, Kishan Panaganti, Haitao Mi, Dongruo Zhou, Leowei Liang
arXiv preprint arXiv:2607.13285
A behavior-centric representation that maps high-level agent-system behaviors to concrete implementation sites, making evolving harnesses easier to navigate, audit, and modify.
FERA: Uncertainty-Aware Federated Reasoning for Large Language Models
Ruhan Wang, Chengkai Huang, Zhiyong Wang, Junda Wu, Rui Wang, Tong Yu, Julian McAuley, Lina Yao, Dongruo Zhou
Conference on Language Modeling (COLM)
A parameter-free federated reasoning framework (FERA) using uncertainty quantification and a dual-pipeline aggregation mechanism for cross-client knowledge integration.
Federated In-Context Learning: Iterative Refinement for Improved Answer Quality
Ruhan Wang, Zhiyong Wang, Chengkai Huang, Rui Wang, Tong Yu, Lina Yao, John C.S. Lui, Dongruo Zhou
Forty-Second International Conference on Machine Learning (ICML)
Privacy-preserving framework (Fed-ICL) that combines federated learning with in-context learning to collaboratively train diverse LLM agents, with theoretical equivalence to established FL algorithms.
Safe Decision Transformer with Learning-based Constraints
Ruhan Wang, Dongruo Zhou
7th Annual Learning for Dynamics and Control Conference (L4DC)
Constrained Q-learning Decision Transformer (CQDT) for safe offline RL, addressing stitching limitations of CDT while strictly adhering to safety constraints.
Return Augmented Decision Transformer for Off-Dynamics Reinforcement Learning
Ruhan Wang, Yu Yang, Zhishuai Liu, Dongruo Zhou, Pan Xu
Transactions on Machine Learning Research (TMLR)
Return Augmented Decision Transformer (RADT) for offline off-dynamics RL with rigorous suboptimality analysis and D4RL evaluation across off-dynamics shifts.
News
📄 Released Recursive Synthesis for Long-Horizon Terminal Tasks, our Tencent AI Lab work on recursively constructing verified long-horizon terminal-agent tasks at scale.
📄 Released Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning, our Tencent AI Lab work on stabilizing asynchronous RL with staleness-aware trust regions.
📄 Released Harness Handbook, our Tencent AI Lab work introducing a behavior-centric representation that makes evolving agent harnesses readable, navigable, and editable.
📄 Released Long-Horizon-Terminal-Bench, our Tencent AI Lab benchmark of 46 tasks across nine categories with dense reward-based grading for long-horizon terminal agents.
🎉 FERA: Uncertainty-Aware Federated Reasoning for Large Language Models was accepted to COLM 2026, held October 6–9 at Hilton Union Square in San Francisco. See you in San Francisco!
🏅 Recognized as a Gold Reviewer for ICML 2026, placing among the top reviewers for this year's conference (with complimentary registration).
💼 Started my Ph.D. Research Internship at Tencent AI Lab (Hunyuan Frontier Lab) in Bellevue, WA. Working on Agentic Reinforcement Learning for large language model agents.
🎉 Our paper Instance-Dependent Continuous-Time RL via Maximum Likelihood Estimation — joint work with Runze Zhao, Yue Yu, Chunfeng Huang, and Prof. Dongruo Zhou — has been accepted to ICML 2026 (Seoul)!
📝 Submitted FERA: Uncertainty-Aware Federated Reasoning with Large Language Models to COLM 2026 — a parameter-free federated reasoning framework that uses uncertainty quantification for cross-client knowledge integration. With Chengkai Huang, Zhiyong Wang, Rui Wang, Tong Yu, Lina Yao, and Prof. Dongruo Zhou.
🏅 Recognized as a Top 25% Reviewer for ICLR 2026.
✈️ Attended ICML 2025 in Vancouver to present Federated In-Context Learning: Iterative Refinement for Improved Answer Quality — a privacy-preserving Fed-ICL framework with theoretical equivalence to established FL algorithms. With Zhiyong Wang, Chengkai Huang, Rui Wang, Tong Yu, Lina Yao, John C.S. Lui, and Prof. Dongruo Zhou.
✈️ Attended L4DC 2025 in Ann Arbor, MI to present Safe Decision Transformer with Learning-based Constraints — our Constrained Q-learning Decision Transformer (CQDT) for safe offline RL. Joint work with Prof. Dongruo Zhou; previously presented at the NeurIPS 2024 Safe Generative AI Workshop.
🎓 Advanced to Ph.D. Candidacy in Computer Engineering at Indiana University after passing the qualifying examination.
🎉 Our work Towards Agentic Recommender Systems in the Era of Multimodal LLMs has been accepted to ACM TIST — a formal LLM-ARS framework spanning user profiling, memory, planning, and action selection, identifying seven open challenges. Collaboration led by Chengkai Huang with the team at UNSW, Adobe Research, UCSD, and Indiana University.
🎓 Completed my M.S. in Computer Engineering at Indiana University Bloomington.
💼 Wrapped up my Ph.D. Research Internship at Mitsubishi Electric Research Laboratories (MERL) in Cambridge, MA, hosted by Dr. Toshiaki Koike-Akino. Worked on quantum machine learning and generative models — resulted in Quantum Diffusion Models for Few-Shot Learning, accepted to ICAD 2025 and the AAAI 2024 Quantum Computing & AI Workshop.
🎉 LLMCarbon: Modeling the End-to-End Carbon Footprint of Large Language Models selected for Oral Presentation at ICLR 2024 (Vienna) — projecting LLM carbon footprints across training, inference, experimentation, and storage. Joint work with Ahmad Faiz, Sotaro Kaneda, Rita Osi, Parteek Sharma, Fan Chen, and Lei Jiang.
🤝 Joined the Machine Learning Lab at Indiana University, advised by Prof. Dongruo Zhou — shifting research focus toward reinforcement learning, foundation models, and agentic AI.
🎓 Started my Ph.D. journey in Computer Engineering at Indiana University Bloomington, joining the Quantum Computing Lab under Prof. Fan Chen.
