About Me
Hi there, this is Junyao. I am a first-year graduate student at the School of Computing, National University of Singapore (NUS), where I am pursuing a specialization in Artificial Intelligence. I am currently working as a Research Intern at Tencent Hy Frontier Lab, Singapore, working on Agentic RL, Long-Horizon Terminus Agent and RL Stability. My research interests lie in Agentic AI, Large Language Models, Reinforcement Learning and Explainable Artificial Intelligence.
My Chinese name is 杨竣尧 (/jɑːŋ dʒuːn jaʊ/).
CVs: EN, 中文
My Chinese name is 杨竣尧 (/jɑːŋ dʒuːn jaʊ/).
CVs: EN, 中文
News
- 2026.08 Conducting my Dissertation under the guidance of Professor Shuicheng Yan.
- 2026.04 Joined Tencent Hy LLM Team, working on Agentic RL Stability and Self-Evolving Long-Horizon Terminus Agent.
- 2026.04 First-Author paper ReasonAny accepted to ACL 2026 Main.
- 2026.01 Tech report: AgentDoG! State-of-the-art diagnostic guardrail framework with an Agentic XAI attribution module.
- 2025.11 First-Author paper RCP-Merging accepted to AAAI 2026 Main Track.
- 2025.08 RewardDS accepted to EMNLP 2025 Main.
- 2025.08 Joined Shanghai AI Lab as a Research Intern, working with Dongrui Liu.
- 2025.05 Co-First-Author paper PrivacyRestore accepted to ACL 2025 Main.
- 2024.07 Joined ZeroNLP as a Research Assistant, advised by Prof. Ziqian Zeng.
- 2024.07 Contextless CS reached 20,000 DAU.
- 2024.04 Joined Tencent as a machine learning intern.
- 2024.03 Completed internship at ShenZhen Stock Exchange as a machine learning intern.
Selected Publications
ACL 2026
ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging
Junyao Yang, Chen Qian, Dongrui Liu†, Wen Shen, Yong Liu†, Jing Shao†
TL;DR: Merging robust chain-of-thought capabilities into domain-specific models (Safety, Biomedicine) using Contrastive Gradient Identification.
AAAI 2026
RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior
Junyao Yang, Jianwei Wang, Huiping Zhuang, Cen Chen, Ziqian Zeng*†
TL;DR: Enhancing domain performance while preserving chain-of-thought reasoning abilities by treating reasoning as a prior.
ACL 2025
PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration
Ziqian Zeng*†, Jianwei Wang*, Junyao Yang*, Zhengdong Lu, Haoran Li, Huiping Zhuang, Cen Chen
TL;DR: Protecting privacy via activation steering using a protected meta-vector without retraining.
Tech Report
Shanghai Artificial Intelligence Laboratory (Contributor)
TL;DR: A state-of-the-art diagnostic guardrail framework utilizing a unified three-dimensional taxonomy to provide fine-grained monitoring and root-cause analysis of AI agent safety risks.
arXiv Preprint
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Zongxia Li, Zhongzhi Li, Yucheng Shi, Ruhan Wang, Junyao Yang, Zhichao Liu, Xiyang Wu, Anhao Li, Yue Yu, Ninghao Liu, Lichao Sun, Haotao Mi, Leowei Liang
TL;DR: A dense-reward terminal benchmark for measuring how far agents can progress on long, complex terminal workflows.
arXiv Preprint
RST: Recursive Synthesis for Long-Horizon Terminal Tasks
Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang
TL;DR: A recursive verified synthesis framework that scales long-horizon terminal-agent tasks, yielding 37,484 tasks at ~$0.05 each and lifting Qwen3.5 by up to 10 points across terminal benchmarks.
arXiv Preprint
Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable
Ruhan Wang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Yue Yu, Junyao Yang, Kishan Panaganti, Haitao Mi, Dongruo Zhou, Leoweiliang
TL;DR: A behavior-centric map and BGPD framework that help agents find code and plan harness edits.
EMNLP 2025
RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis
Jianwei Wang, Chengming Shi, Junyao Yang, Haoran Li, Qianli Ma, Huiping Zhuang, Cen Chen, Ziqian Zeng†
TL;DR: Using client-side reward models to filter synthetic data, mitigating noise while protecting privacy.
ACL 2026
ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging
Junyao Yang, Chen Qian, Dongrui Liu†, Wen Shen, Yong Liu†, Jing Shao†
TL;DR: Merging robust chain-of-thought capabilities into domain-specific models (Safety, Biomedicine) using Contrastive Gradient Identification.
AAAI 2026
RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior
Junyao Yang, Jianwei Wang, Huiping Zhuang, Cen Chen, Ziqian Zeng*†
TL;DR: Enhancing domain performance while preserving chain-of-thought reasoning abilities by treating reasoning as a prior.
arXiv Preprint
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang
TL;DR: Stabilizing asynchronous RL by adapting trust regions to rollout staleness, tightening high-mismatch updates while preserving ordinary-token behavior.
ACL 2025
PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration
Ziqian Zeng*†, Jianwei Wang*, Junyao Yang*, Zhengdong Lu, Haoran Li, Huiping Zhuang, Cen Chen
TL;DR: Protecting privacy via activation steering using a protected meta-vector without retraining.
arXiv Preprint
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Zongxia Li, Zhongzhi Li, Yucheng Shi, Ruhan Wang, Junyao Yang, Zhichao Liu, Xiyang Wu, Anhao Li, Yue Yu, Ninghao Liu, Lichao Sun, Haotao Mi, Leowei Liang
TL;DR: A dense-reward terminal benchmark for measuring how far agents can progress on long, complex terminal workflows.
arXiv Preprint
RST: Recursive Synthesis for Long-Horizon Terminal Tasks
Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang
TL;DR: A recursive verified synthesis framework that scales long-horizon terminal-agent tasks, yielding 37,484 tasks at ~$0.05 each and lifting Qwen3.5 by up to 10 points across terminal benchmarks.
arXiv Preprint
Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable
Ruhan Wang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Yue Yu, Junyao Yang, Kishan Panaganti, Haitao Mi, Dongruo Zhou, Leoweiliang
TL;DR: A behavior-centric map and BGPD framework that help agents find code and plan harness edits.
Tech Report
Shanghai Artificial Intelligence Laboratory (Contributor)
TL;DR: A state-of-the-art diagnostic guardrail framework utilizing a unified three-dimensional taxonomy to provide fine-grained monitoring and root-cause analysis of AI agent safety risks.
arXiv Preprint
Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models
Junyao Yang, Chen Qian, Kun Wang, Linfeng Zhang, Quanshi Zhang, Yong Liu, Dongrui Liu†
TL;DR: Identifying Entropy-Gradient Inversion—a robust negative correlation between token entropy and logit gradients—as a geometric fingerprint of LRM reasoning capability, and proposing CorR-PO that embeds this signature into RL reward regularization to stabilize reasoning optimization without costly external verifiers.
arXiv Preprint
The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution
Chen Qian, Peng Wang, Dongrui Liu†, Junyao Yang, Dadi Guo, Ling Tang, Jilin Mei, Qihan Ren, Shuai Shao, Yong Liu, Jie Fu, Jing Shao, Xia Hu
TL;DR: A hierarchical framework for agentic attribution, using temporal likelihood and perturbation-based analysis to unveil internal factors driving LLM-based agent actions.
EMNLP 2025
RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis
Jianwei Wang, Chengming Shi, Junyao Yang, Haoran Li, Qianli Ma, Huiping Zhuang, Cen Chen, Ziqian Zeng†
TL;DR: Using client-side reward models to filter synthetic data, mitigating noise while protecting privacy.
Blogs
Long-Horizon Terminal-Bench: Measuring the Progress Agents Can Sustain, Not Just What They Can Finish
TL;DR: LHTB evaluates agent progress on long terminal tasks in Docker with hidden dense-reward grading, not just final completion.
Harness Handbook: Making Agent Harnesses Understandable, Auditable, and Editable
TL;DR: Harness Handbook maps behavior to implementation, making agent harnesses easier to inspect, navigate, and edit.
The Entropy-Gradient Inversion: A New Perspective on LLM Reasoning Capabilities
TL;DR: We discover that reasoning models exhibit a unique "fingerprint": a significant negative correlation between gradient strength and token entropy, which contradicts traditional base models. This capability emerges rapidly within the first 200 steps of SFT.
Education
MComp in AI
National University of Singapore
2025 - 2027 (Expected)
B.S. in CS (with honor)
South China University of Technology
2021 - 2025
Internship and Experience
Tencent Hy
青云计划 - Research Intern | 2026.04 - Present
Shanghai AI Lab
Research Intern | 2025.06 - 2026.04
South China University of Technology
Research Intern | 2024.07 - 2025.06
Tencent
Machine Learning Intern | 2024.04 - 2024.07
SZSE
Machine Learning Intern | 2024.01 - 2024.04
Honor & Awards
- Excellent Graduation Thesis (2025.06)
- Outstanding Student Leader (2022-2024)
- Second-Class Scholarship of SCUT (2024.10)
- Second-Class Award in CUMCM at Guangdong Province (2022.09)
- Second-Prize of Olympic Mathematics Competition (2020.05)
- Second-Prize of Olympic Physics Competition (2020.02)