Publications
Preprint 2026
JevOut: Natural Context Can Flip Decision Models
Preprint 2026
AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds
Findings of EMNLP 2026
SocialSim @ COLM 2025 · Spotlight
SocialMaze: A Benchmark for Evaluating and Enhancing Social Reasoning in Large Language Models in Complex Social Environments
ACL 2025
Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models
NeurIPS 2025
Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search
NeurIPS 2025
DyFlow: Dynamic Workflow Framework for Agentic Reasoning
Preprint 2026
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
Preprint 2026
GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs
ICML 2026
Do LLMs “Feel”? Emotion Circuits Discovery and Control
ACL 2026
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
AAAI 2026
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models
TRB 2025
Oral
Temporal-IRL: Modeling Port Congestion and Berth Scheduling with Inverse Reinforcement Learning
Findings of EMNLP 2025
Under the Shadow of Babel: How Language Shapes Reasoning in LLMs
ACM MM 2024
GenUDC: High-Quality 3D Mesh Generation with Unsigned Dual Contouring Representation
Preprint 2025
Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies
Preprint 2025
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
PRCV 2024
Oral
MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
CCBR 2023
Oral