Offline goal-conditioned RL
Learning to reach any goal from a fixed dataset. I study value learning that stays correct over long horizons and under stochastic dynamics, and which representations actually help a goal-reaching agent act better.
Research Assistant · AVIS Lab, University of Dhaka
I work on embodied agents that learn long-horizon skills from offline data and from language feedback.
Applying for PhD positions · Fall 2027I’m a Research Assistant at the AVIS Lab, University of Dhaka, working on offline goal-conditioned reinforcement learning and long-horizon robot manipulation. Most recently, I grounded divide-and-conquer value learning with temporal differences so it holds up under stochastic dynamics (GTRL), and turned a vision–language model’s reflections on failed episodes into reward for robot manipulation (LAGEA, ICML 2026).
Before this, I was a Research Assistant at the MAIM Lab, building a wearable fetal-movement monitor for stillbirth prevention, a Wellcome Leap In Utero project with Dr. Abhishek Kumar Ghosh and Dr. Niamh Nowlan. I hold a BSc in Robotics and Mechatronics Engineering from the University of Dhaka (2024), where Dr. Md Mehedi Hasan supervised my thesis on UAV-based human action recognition.
I want agents that can carry out long, multi-step tasks in the real world. That takes value estimates that stay accurate over long horizons, feedback that tells an agent where an attempt went wrong and not just that it failed, and representations that bring together what an agent sees, reads and measures. My work so far follows three threads.
Learning to reach any goal from a fixed dataset. I study value learning that stays correct over long horizons and under stochastic dynamics, and which representations actually help a goal-reaching agent act better.
Vision–language models as critics for robot learning: structured reflections on failed episodes, grounded in time and turned into dense reward shaping for manipulation.
Fusing complementary views of one signal, such as time, frequency and language, or the visible and hidden parts of a scene, into representations that hold up with little data.
Fall 2027
I’m applying to PhD programs starting in Fall 2027. I want to keep building embodied agents that learn long-horizon skills, where reinforcement learning, vision–language models and robot manipulation meet. My first-author work has appeared at ICML 2026 and AAAI 2026.
If you work on robot learning, embodied AI or reinforcement learning and think we’d be a good fit, I’d love to hear from you.
ICLR '27: GTRL: Grounding Divide-and-Conquer Value Learning with Temporal Differences and Do Better Goal Representations Improve Goal-Conditioned Reinforcement Learning? LAGEA: Language Guided Embodied Agents for Robotic Manipulation got accepted for a Poster at ICML '26 ECCV '26, and the paper is published at ArXiv T3Time: Tri-Modal Time Series Forecasting via Adaptive Multi-Head Alignment and Residual Fusion got accepted for a Poster at AAAI '26 U-ActionNet: Dual-Pathway Fourier Networks with Region-of-Interest Module for Efficient Action Recognition in UAV Surveillance, based on my fourth year thesis work, has been published in IEEE Access University of Dhaka, Bangladesh with a BSc. in Robotics & Mechatronics Engineering. [Report] FFT-UAVNet: FFT Based Human Action Recognition for Drone Surveillance System got accepted to the 5th IEEE International Conference on Sustainable Technologies for Industry 5.0 (STI) conference arXiv 2026
Grounding divide-and-conquer value learning with temporal differences, so offline goal-conditioned RL stays correct when the dynamics are stochastic.
Preprint 2026
Handed a perfect goal representation, a goal-conditioned agent barely acts better. The headroom is in how it sees its own position.
Project page Paper & code soon
ICML 2026
A vision–language model reflects on each failed episode in structured language, and the reflection becomes time-localised reward shaping for robot manipulation.
AAAI 2026
T3Time reads a series three ways (in time, in frequency, and as a language prompt) and lets the forecast horizon decide how to weigh them.
A live snapshot of where visitors have been dropping in from. Click anywhere on the map to open the more detailed view.