Learning from interaction
How can an agent learn which decisions mattered over a long task? I study credit assignment, agentic reinforcement learning, and learning from experimental feedback.
Hi! I am Changjian Liu (刘昌健), a master’s student in Spatio-Temporal Big Data at Peking University, where I study at the Institute of Remote Sensing and Geographic Information Systems, under the supervision of Prof. Yong Gao. I am currently a Research Intern at Galbot, working on agentic systems for embodied brains and long-horizon tasks.
Previously, I was an Intern at Alibaba · Taobao & Tmall Group · Alimama, where I worked on agentic reinforcement learning, causal decision systems, and large-scale decision optimization, and an Intern at DiDi, focusing on advertising algorithms and decision making. I received my B.Eng. in Computer Science and Technology from China University of Geosciences, Beijing.
My research interests lie broadly in agentic reinforcement learning, LLM post-training, and sequential decision making. I am particularly interested in how intelligent agents can learn from interaction, assign credit over long horizons, and reliably translate reasoning into actions in both digital and embodied environments.
From learning signals to real-world decisions.
How can an agent learn which decisions mattered over a long task? I study credit assignment, agentic reinforcement learning, and learning from experimental feedback.
How can learned policies stay useful under constraints and distribution shift? My work connects causal learning with decision objectives, from resource allocation to LLM agents.
How can robots reason about a task and carry it out? At Galbot, I am exploring agentic approaches to embodied reasoning and task execution.
Feedback informs the next observation and decision
Current projects, submissions, and published work.
Execution-aware credit assignment for long-horizon agents. The project uses local verification signals to connect feedback to relevant decisions while preserving the original reward objective, with a training pipeline built on verl.
Helping research agents learn from experiments through attribution checks and hierarchical memory, so that verified insights can inform later tasks.
ReAlloc formulates fixed-budget multi-channel marketing as simplex-constrained uplift policy learning, combining an orthogonal teacher, explanation-guided student, and support-aware local reallocation for stable production decisions.
A causal response learning framework for continuous decisions under temporal drift, using anchored link-scale contrasts, orthogonal pilots, and profiled morphology selection to learn deployable action geometry.
Research on anchor-regularized ROI–uplift frontier learning for advertising decisions.
My earlier work on structured environments, mobility, and generalization.
A mechanism-constrained framework for robust origin-destination flow prediction that learns row-centered choice potentials and separates transferable allocation laws from origin demand scale.
A trajectory semantic modeling framework that combines movement traces with demographic structure for travel-flow prediction and interaction analysis.
A dynamic graph forecasting model for time-varying photovoltaic power prediction under volatile renewable-energy generation patterns.
A generative trajectory forecasting study that uses recurrent visit patterns to improve long-horizon individual mobility prediction.
Explainable spatial regression combining global homogeneity and local heterogeneity through meta-learning.
Tourist attraction recommendation with signed feedback and a spatial-sentiment knowledge graph.
Learning and decision systems in practice.
Sep 2026 – Present
Research internship · Embodied intelligence brain
Working on agentic approaches to robot reasoning and task execution, exploring how robots can think through tasks and translate decisions into actions.
Nov 2025 – Jun 2026
LLM Algorithm Intern
Worked on agentic reinforcement learning, self-evolving research agents, and causal decision systems. Projects covered execution-aware credit assignment, experimental attribution, pricing, and constrained multi-channel allocation.
Jun 2025 – Oct 2025
Advertising Algorithm Intern
Developed models for personalized coupon allocation, combining purchase, redemption, and short-term value prediction with uncertainty-aware treatment modeling and budget-constrained optimization.
Research Notes
A research note on why observational response models become brittle when logs, policies, and environments move together.
Engineering Notes
A short technical outline for connecting spatial flow modeling with downstream allocation, pricing, and intervention decisions.
Our work on deployable causal action geometry under temporal non-stationarity has been accepted at NeurIPS 2026.
I recently received the Academic Scholarship at Peking University.
Since September 2026, I have been working on agentic approaches to robot reasoning and task execution at Galbot.
Visitor Analytics
A view of visits over the last 30 days.
Loading visitor data…
Loading…
Earth imagery: NASA Earth Observatory