Vision–language models can score how close a robot is to its goal, but that score is noisy and says nothing about why an attempt failed. LAGEA asks a small VLM to reflect on each episode in a fixed schema (what went wrong, at which frames, how to fix it), aligns that reflection with the robot's visual state, and turns it into a bounded, step-wise shaping reward that is strongest early in training and fades as the policy improves.
Figure 1. The LAGEA loop. (a) After each rollout, key frames mark the decisive moments and a frozen VLM writes a schema-constrained reflection, encoded as a feedback vector \(f\). (b) Small trainable projectors map images, the goal image, the instruction and \(f\) into one space, trained so that feedback agrees with the states it describes. (c) Progress toward the goal and toward the feedback become two delta rewards, fused with the sparse task reward to train the policy.
A score is not a diagnosis
Hand-designing dense rewards for every manipulation task does not scale, so a common shortcut is to use a pretrained VLM as a zero-shot reward model: embed the camera image and the instruction, and reward their similarity. FuRL showed that these “fuzzy” rewards help once they are aligned to the task. But a similarity score summarises the outcome, not the cause. It drifts with viewpoint, can be hallucinated, and on long, sparse-reward tasks an early misalignment compounds into misdirected exploration.
Self-reflection works for language agents: Reflexion improves decisions with textual self-critique, but in text worlds where observation, action and feedback share one symbolic interface. A robot sees pixels and acts in joint space. LAGEA closes that gap in three steps: make the reflection structured so a small VLM cannot ramble, anchor it in time so it credits the right steps, and align it with the visual state so it can be turned into a per-step reward.
What a reflection looks like
At the end of an episode, Qwen2.5-VL-3B sees frames from the rollout, the task instruction, a five-code error taxonomy, one success and one failure exemplar, and the last few attempts. It must answer in this JSON schema; malformed answers are rejected and retried.
- wrong_object
- bad_approach_direction
- failed_grasp
- insufficient_force
- drift_from_goal
{
"task": "button-press-topdown-v2-goal-observable",
"outcome": "failure",
"primary_error": {
"code": "bad_approach_direction",
"explanation": "The gripper came from the side, sliding off the button instead of a vertical press."
},
"secondary_factors": [],
"key_frame_indices": [18, 22],
"suggested_fix": "Approach from directly above the button; align gripper normal to the button surface, then press straight down.",
"confidence": 0.85,
"summary": "The robot failed to press the button correctly because it approached from the side instead of a vertical press."
} {
"task": "button-press-topdown-v2-goal-observable",
"outcome": "success",
"primary_error": {
"code": "good_grasp",
"explanation": "The gripper successfully grasped the button."
},
"secondary_factors": [],
"key_frame_indices": [12, 18],
"suggested_fix": "(n/a)",
"confidence": 0.9,
"summary": "The agent succeeded because it grasped the button securely and pressed it straight down, achieving the goal."
} The reflection is embedded with a lightweight GPT-2 into a 768-dimensional feedback vector \(f\). Because the schema fixes the vocabulary, near-duplicate episodes give near-duplicate vectors, which is what makes the signal usable for learning.
From one reflection to a reward at every step
A reflection arrives once per episode. The agent needs a signal at every transition, concentrated where the outcome was decided, and small enough not to drown out the task reward. LAGEA builds it in four stages.
1 · Find the moments that mattered
Broadcasting one feedback vector across a whole episode assigns credit to steps that had nothing to do with the outcome. LAGEA picks key frames from the goal-similarity trajectory itself, which keeps the selector deterministic and model-agnostic. With \(x_t\) the image embedding at step \(t\) and \(g\) the goal embedding,
where \(z(\cdot)\) z-scores within the episode, so frames score highly when they are close to the goal, changing fast, or turning sharply. Up to \(M\) of the most salient frames are kept, a minimum distance apart, together with both endpoints. Each key frame \(k \in \mathcal{K}\) then spreads weight over its neighbours with a triangular kernel of half-width \(h\) and a floor \(\beta\), and dividing by the episode mean \(\bar w\) keeps the average weight at one:
The episode is synthetic (an approach, a slip away from the goal, then contact), and the saliency weights are fixed at \(\omega = (0.5, 0.3, 0.2)\) for the demo. Everything else follows the formulas above.
2 · Put feedback and pixels in one space
Small MLP projectors \(E_i, E_f\) map image and feedback embeddings to unit vectors \(z_t\) and \(z_f\), with \(\psi_t = \langle z_t, z_f\rangle\). Two losses shape the space, each step weighted by \(u_t\) (key-frame saliency times goal proximity):
The first pulls successful steps toward their feedback and pushes failed ones away, using the success label \(y_t\). The second is an InfoNCE term over the batch, \(S_{ij} = \langle z_f^{(i)}, z^{(j)}\rangle / \tau_{\text{nce}}\), so each reflection prefers its own images to everyone else's. A symmetric, weighted contrastive step with alignment and uniformity regularisers then polishes the geometry.
Does the space actually separate success from failure? Fig. 5
Early on, successful and unsuccessful states score alike and success stays at zero. Around 0.5M steps a discrimination gap opens, and success rises right after it. The alignment losses then increase: a better policy produces harder negatives, which forces finer distinctions.
3 · Reward progress, not position
Two potentials are read off the aligned space: \(\phi_t\), the state's agreement with the instruction \(z_y\) and the goal image \(z_g\), and \(\psi_t\), its agreement with the feedback. As in potential-based shaping, only their change is rewarded, so a state that merely sits near the goal earns nothing. Everything is squashed with \(\tanh\), so each term stays in \([-1, 1]\).
The feedback term is gated by the key-frame weights \(\hat w_t\), and its share \(\alpha = \mathrm{clip}\big(\alpha_{\text{base}} \cdot \tfrac12(1 + \langle z_y, z_f\rangle),\, [\alpha_{\min}, \alpha_{\max}]\big)\) grows when the reflection agrees with the instruction, so feedback that has wandered off-task counts for less.
Figure 2. The two delta rewards. (a) The goal potential \(\phi_t\) aligns the current state with the goal image and the instruction. (b) The feedback potential \(\psi_t\) aligns it with the reflection, weighted by the key frames. Their temporal differences form the fused reward.
4 · Let the guidance fade
A dense signal can overpower a sparse one. LAGEA shapes only failing steps, \(m_t = \mathbb{1}[r^{\text{task}}_t \lt 0]\), and scales the shaping by an estimate of competence \(P \in [0,1]\): the larger of a success-rate moving average \(\bar s\) and the squared fraction of steps that moved toward the goal.
Early on, when \(P \approx 0\), language guidance is at full strength; as the policy starts succeeding it recedes and the task reward takes over. The critic of a standard SAC agent trains on \(r_t\).
Higher success on two manipulation suites
We compare against reward-learning baselines on sparse-reward tasks: SAC on the task reward alone; LIV, a robotics reward model pretrained on large-scale data, and LIV-Proj, with random fixed projection heads; Relay, which adds relay exploration to LIV; and FuRL, which aligns fuzzy VLM rewards and uses relay RL. Visual observations are embedded with LIV throughout.
Ten Sawyer-arm tasks with a sparse task reward and a goal that stays fixed and unobserved. Baselines range from plain SAC to FuRL, which already aligns VLM rewards and adds relay exploration.
80.0% against 76.0% for FuRL: +4.0 points, a 5.3% relative gain. LAGEA solves eight of the ten tasks in every seed, including push, where FuRL reaches 80%.
Peg-insert-side and pick-place stay at zero for every method, LAGEA included.
Per-task results · MT10 · fixed goals 10 tasks × 7 methods
| Task | SAC | LIV | LIV-Proj | Relay | FuRL w/o goal image | FuRL | LAGEA |
|---|---|---|---|---|---|---|---|
| button-press-topdown | 0 | 0 | 0 | 60 | 80 | 100 | 100 |
| door-open | 50 | 0 | 0 | 80 | 100 | 100 | 100 |
| drawer-close | 100 | 100 | 100 | 100 | 100 | 100 | 100 |
| drawer-open | 20 | 0 | 0 | 40 | 80 | 80 | 100 |
| peg-insert-side | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| pick-place | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| push | 0 | 0 | 0 | 0 | 40 | 80 | 100 |
| reach | 60 | 80 | 80 | 100 | 100 | 100 | 100 |
| window-close | 60 | 60 | 40 | 80 | 100 | 100 | 100 |
| window-open | 80 | 40 | 20 | 80 | 100 | 100 | 100 |
| Average | 37.0 | 28.0 | 24.0 | 54.0 | 70.0 | 76.0 | 80.0 |
Success rate (%) averaged over five seeds; the paper reports no spread for this table. Bold marks the best in each row and underline the runner-up.
The same ten tasks, but the goal moves every episode and is visible to the agent, so the policy has to generalise across goal positions.
70.4% against 64.6% for FuRL: +5.8 points, a 9.0% relative gain, with the smallest spread across seeds (±1.85 against ±5.0).
Contact-rich tasks remain hard. Pick-place reaches 4% and push 12%, and peg-insert-side stays at zero.
Per-task results · MT10 · random goals 10 tasks × 4 methods
| Task | SAC | Relay | FuRL | LAGEA |
|---|---|---|---|---|
| button-press-topdown | 16.0±32.0 | 56.0±38.3 | 64.0±32.6 | 96±8 |
| door-open | 78.0±39.2 | 80.0±30.3 | 96.0±8.0 | 100±0 |
| drawer-close | 100±0 | 100±0 | 100±0 | 100±0 |
| drawer-open | 40.0±49.0 | 50.0±42.0 | 84.0±27.3 | 92±9.8 |
| pick-place | 0±0 | 0±0 | 0±0 | 4±4.9 |
| peg-insert-side | 0±0 | 0±0 | 0±0 | 0±0 |
| push | 0±0 | 0±0 | 6.0±8.0 | 12±4 |
| reach | 100±0 | 100±0 | 100±0 | 100±0 |
| window-close | 86.0±28.0 | 96.0±4.9 | 100±0 | 100±0 |
| window-open | 78.0±39.2 | 92.0±7.5 | 96.0±4.9 | 100±0 |
| Average | 49.8 | 57.4 | 64.6 | 70.4 |
Success rate (%) ± standard deviation over five seeds. Bold marks the best in each row and underline the runner-up.
A 7-DoF arm with a two-finger gripper, 50-step episodes and the standard sparse binary reward: reach, push, pick-and-place and slide.
51.7% against 44.2% for FuRL: +7.5 points, a 17% relative gain. Reach is solved by everyone; on push, pick-and-place and slide LAGEA is best.
Slide, where the puck has to be struck and left to glide, stays hard at 10%.
Per-task results · Fetch 4 tasks × 4 methods
| Task | SAC | Relay | FuRL | LAGEA |
|---|---|---|---|---|
| Reach | 100±0 | 100±0 | 100±0 | 100±0 |
| Push | 26.67±4.71 | 30±8.16 | 40±8.16 | 53.33±4.71 |
| PickAndPlace | 10±8.16 | 20±0 | 33.33±9.43 | 43.33±4.71 |
| Slide | 0±0 | 0±0 | 3.33±4.71 | 10±8.16 |
| Average | 34.17 | 37.5 | 44.17 | 51.67 |
Success rate (%) ± standard deviation over three seeds. Bold marks the best in each row and underline the runner-up.
Faster to converge, not just higher at the end
Figure 3. Success over 1M environment steps on eight MT10 tasks. LAGEA reaches high success well before FuRL on most tasks, and SAC with the sparse reward alone mostly stalls. The extra VLM call per episode is paid back in samples: averaged over nine tasks, LAGEA converges in 92.4 minutes of wall-clock time against 94.9 for FuRL, and is faster on six of the nine.
The benchmarks Meta-World MT10 and Fetch
Meta-World MT10. Ten tasks with a Sawyer arm, a 4-D action (end-effector motion and gripper) and 500-step episodes, from reaching to peg insertion. Each task is paired with a one-line instruction, such as “Press a button from the top.”
Gymnasium-Robotics Fetch. A 7-DoF arm with a two-finger gripper on reach, push, slide and pick-and-place, with 50-step episodes and a sparse binary reward.
Which parts carry the gain
Each ablation keeps the encoders, the SAC learner and the goal image fixed and changes one thing. The components turn out to be complementary: removing any one of them costs about as much as removing another.
Dropping the goal-delta reward, the feedback-delta reward or the adaptive schedule each leaves about 80%. Only all three together reach 99%: long-term progress, short-term correction, and knowing when to step back.
Random or evenly spaced frames lose 12 to 13 points. The gap is widest on button-press, 96% against 37% and 50%, where the decisive moment is a brief contact.
Drawer-open, with and without key frames. With key frames the shaped reward turns upward as the agent learns the approach, and success follows. Without them the feedback is smeared over the whole episode and the agent stays stuck.
Letting the VLM answer in free text costs 20 points. The biggest single drop is button-press-topdown, 93% to 10%: verbose, ambiguous explanations make noisy feedback vectors, and noisy vectors make misleading rewards.
The reflecting VLM matters more than the text encoder. All three alternative encoders stay above 70%, but a weaker VLM costs 12 to 24 points, with SmolVLM2 falling to 56%. Qwen2.5-VL-3B with GPT-2 is the best pairing, which is why it is the default.
Trained from one diagonal front camera and evaluated zero-shot from three others, LAGEA stays within three points of its training-view score (77–79% against 80%).
Heavy self-occlusion and cluttered scenes were not tested and remain open.
Limitations
LAGEA inherits occasional hallucinations from its VLM; the schema and the alignment reduce them but cannot remove them. All experiments are in simulation, and contact-rich insertion tasks remain unsolved. Moving to real robots, closing the sim-to-real gap, and handling long horizons under partial observability are the natural next steps.
Citation
If this work is useful to you, please cite it as:
@inproceedings{chowdhury2026lagea,
title = {{LAGEA}: Language Guided Embodied Agents for Robotic Manipulation},
author = {Chowdhury, Abdul Monaf and Mazumder, Akm Moshiur Rahman and
Arib, Safaeid Hossain and Akter, Rabeya},
booktitle = {Proceedings of the 43rd International Conference on Machine Learning},
series = {Proceedings of Machine Learning Research},
volume = {306},
year = {2026}
}