ICML 2026Embodied AI · reinforcement learning

LAGEA: Language Guided Embodied Agents for Robotic Manipulation

  • Abdul Monaf Chowdhury1
  • Akm Moshiur Rahman Mazumder2
  • Safaeid Hossain Arib1
  • Rabeya Akter1
  • 1University of Dhaka
  • 2Independent University, Bangladesh
TL;DR

Vision–language models can score how close a robot is to its goal, but that score is noisy and says nothing about why an attempt failed. LAGEA asks a small VLM to reflect on each episode in a fixed schema (what went wrong, at which frames, how to fix it), aligns that reflection with the robot's visual state, and turns it into a bounded, step-wise shaping reward that is strongest early in training and fades as the policy improves.

80.0%
MT10 with hidden, fixed goals
+4.0 over FuRL
70.4%
MT10 with observable, random goals
+5.8 over FuRL
51.7%
Gymnasium-Robotics Fetch, 4 tasks
+7.5 over FuRL
92min
Mean wall-clock to converge, against 95 for FuRL, VLM calls included
LAGEA overview: trajectories go to a replay buffer; key frames and a VLM produce structured feedback; projectors align images, goal, instruction and feedback; goal-delta and feedback-delta rewards are fused with the task reward for the policy.

Figure 1. The LAGEA loop. (a) After each rollout, key frames mark the decisive moments and a frozen VLM writes a schema-constrained reflection, encoded as a feedback vector \(f\). (b) Small trainable projectors map images, the goal image, the instruction and \(f\) into one space, trained so that feedback agrees with the states it describes. (c) Progress toward the goal and toward the feedback become two delta rewards, fused with the sparse task reward to train the policy.

Motivation

A score is not a diagnosis

Hand-designing dense rewards for every manipulation task does not scale, so a common shortcut is to use a pretrained VLM as a zero-shot reward model: embed the camera image and the instruction, and reward their similarity. FuRL showed that these “fuzzy” rewards help once they are aligned to the task. But a similarity score summarises the outcome, not the cause. It drifts with viewpoint, can be hallucinated, and on long, sparse-reward tasks an early misalignment compounds into misdirected exploration.

How can an embodied policy get reliable, time-localised explanations of its own failures, straight from visual trajectories, when exploration is expensive?

Self-reflection works for language agents: Reflexion improves decisions with textual self-critique, but in text worlds where observation, action and feedback share one symbolic interface. A robot sees pixels and acts in joint space. LAGEA closes that gap in three steps: make the reflection structured so a small VLM cannot ramble, anchor it in time so it credits the right steps, and align it with the visual state so it can be turned into a per-step reward.

What a reflection looks like

At the end of an episode, Qwen2.5-VL-3B sees frames from the rollout, the task instruction, a five-code error taxonomy, one success and one failure exemplar, and the last few attempts. It must answer in this JSON schema; malformed answers are rejected and retried.

  • wrong_object
  • bad_approach_direction
  • failed_grasp
  • insufficient_force
  • drift_from_goal
button-press-topdown · examples from the paper's appendix
{
  "task": "button-press-topdown-v2-goal-observable",
  "outcome": "failure",
  "primary_error": {
    "code": "bad_approach_direction",
    "explanation": "The gripper came from the side, sliding off the button instead of a vertical press."
  },
  "secondary_factors": [],
  "key_frame_indices": [18, 22],
  "suggested_fix": "Approach from directly above the button; align gripper normal to the button surface, then press straight down.",
  "confidence": 0.85,
  "summary": "The robot failed to press the button correctly because it approached from the side instead of a vertical press."
}
{
  "task": "button-press-topdown-v2-goal-observable",
  "outcome": "success",
  "primary_error": {
    "code": "good_grasp",
    "explanation": "The gripper successfully grasped the button."
  },
  "secondary_factors": [],
  "key_frame_indices": [12, 18],
  "suggested_fix": "(n/a)",
  "confidence": 0.9,
  "summary": "The agent succeeded because it grasped the button securely and pressed it straight down, achieving the goal."
}

The reflection is embedded with a lightweight GPT-2 into a 768-dimensional feedback vector \(f\). Because the schema fixes the vocabulary, near-duplicate episodes give near-duplicate vectors, which is what makes the signal usable for learning.

Method

From one reflection to a reward at every step

A reflection arrives once per episode. The agent needs a signal at every transition, concentrated where the outcome was decided, and small enough not to drown out the task reward. LAGEA builds it in four stages.

1 · Find the moments that mattered

Broadcasting one feedback vector across a whole episode assigns credit to steps that had nothing to do with the outcome. LAGEA picks key frames from the goal-similarity trajectory itself, which keeps the selector deterministic and model-agnostic. With \(x_t\) the image embedding at step \(t\) and \(g\) the goal embedding,

\[ \begin{aligned} s_t &= \cos(x_t, g), \qquad v_t = s_t - s_{t-1}, \qquad a_t = v_t - v_{t-1}, \\ p_t &= \omega_s\,[z(s_t)]_+ + \omega_v\, z(|v_t|) + \omega_a\, z(|a_t|), \end{aligned} \]
saliency

where \(z(\cdot)\) z-scores within the episode, so frames score highly when they are close to the goal, changing fast, or turning sharply. Up to \(M\) of the most salient frames are kept, a minimum distance apart, together with both endpoints. Each key frame \(k \in \mathcal{K}\) then spreads weight over its neighbours with a triangular kernel of half-width \(h\) and a floor \(\beta\), and dividing by the episode mean \(\bar w\) keeps the average weight at one:

\[ \begin{aligned} \tilde w_t &= \max_{k \in \mathcal{K}} \Big(1 - \tfrac{|t-k|}{h+1}\Big)_+, \\ w_t &= \beta + (1-\beta)\,\tilde w_t, \qquad \hat w_t = w_t / \bar w, \end{aligned} \]
credit weights
InteractiveWhere does one episode's feedback land?
The equations above, run on an illustrative episode
Key frames and credit weights Top: the goal similarity of one episode over 64 steps, with the selected key frames marked. Bottom: the per-step credit weight, compared with a uniform broadcast.

The busiest quarter of the episode carries most of the credit; a uniform broadcast would give it exactly a quarter.

The episode is synthetic (an approach, a slip away from the goal, then contact), and the saliency weights are fixed at \(\omega = (0.5, 0.3, 0.2)\) for the demo. Everything else follows the formulas above.

2 · Put feedback and pixels in one space

Small MLP projectors \(E_i, E_f\) map image and feedback embeddings to unit vectors \(z_t\) and \(z_f\), with \(\psi_t = \langle z_t, z_f\rangle\). Two losses shape the space, each step weighted by \(u_t\) (key-frame saliency times goal proximity):

\[ \begin{aligned} \mathcal{L}_{\text{bce}} &= \tfrac{1}{\sum_t u_t} \textstyle\sum_t u_t\, \mathrm{BCE}\big(\sigma(\psi_t/\tau_{\text{bce}}),\, y_t\big), \\ \mathcal{L}_{\text{nce}} &= \tfrac{1}{\sum_{i:\,y_i=1} u_i} \textstyle\sum_{i:\,y_i=1} u_i\, \mathrm{CE}\big(\mathrm{softmax}(S_{i:}),\, i\big), \\ \mathcal{L}_{\text{align}} &= \lambda_{\text{bce}}\, \mathcal{L}_{\text{bce}} + \lambda_{\text{nce}}\, \mathcal{L}_{\text{nce}}. \end{aligned} \]
alignment

The first pulls successful steps toward their feedback and pushes failed ones away, using the success label \(y_t\). The second is an InfoNCE term over the batch, \(S_{ij} = \langle z_f^{(i)}, z^{(j)}\rangle / \tau_{\text{nce}}\), so each reflection prefers its own images to everyone else's. A symmetric, weighted contrastive step with alignment and uniformity regularisers then polishes the geometry.

Does the space actually separate success from failure? Fig. 5
Three plots over training: the gap between positive and negative logits opens around 0.5M steps; success rises after it; BCE and NCE losses rise as negatives get harder.

Early on, successful and unsuccessful states score alike and success stays at zero. Around 0.5M steps a discrimination gap opens, and success rises right after it. The alignment losses then increase: a better policy produces harder negatives, which forces finer distinctions.

3 · Reward progress, not position

Two potentials are read off the aligned space: \(\phi_t\), the state's agreement with the instruction \(z_y\) and the goal image \(z_g\), and \(\psi_t\), its agreement with the feedback. As in potential-based shaping, only their change is rewarded, so a state that merely sits near the goal earns nothing. Everything is squashed with \(\tanh\), so each term stays in \([-1, 1]\).

\[ \begin{aligned} r^{\text{goal}}_t &= \tanh\!\Big(\tfrac{\gamma\,\phi_{t+1} - \phi_t}{\tau_{\text{goal}}}\Big), \qquad r^{\text{fb}}_t = \hat w_t \tanh\!\Big(\tfrac{\gamma\,\psi_{t+1} - \psi_t}{\tau_f}\Big), \\ \tilde r_t &= (1-\alpha)\, r^{\text{goal}}_t + \alpha\, r^{\text{fb}}_t. \end{aligned} \]
delta rewards

The feedback term is gated by the key-frame weights \(\hat w_t\), and its share \(\alpha = \mathrm{clip}\big(\alpha_{\text{base}} \cdot \tfrac12(1 + \langle z_y, z_f\rangle),\, [\alpha_{\min}, \alpha_{\max}]\big)\) grows when the reflection agrees with the instruction, so feedback that has wandered off-task counts for less.

Computation of the delta rewards: goal potential from the goal image and instruction, feedback potential from the VLM feedback, and their temporal differences.

Figure 2. The two delta rewards. (a) The goal potential \(\phi_t\) aligns the current state with the goal image and the instruction. (b) The feedback potential \(\psi_t\) aligns it with the reflection, weighted by the key frames. Their temporal differences form the fused reward.

4 · Let the guidance fade

A dense signal can overpower a sparse one. LAGEA shapes only failing steps, \(m_t = \mathbb{1}[r^{\text{task}}_t \lt 0]\), and scales the shaping by an estimate of competence \(P \in [0,1]\): the larger of a success-rate moving average \(\bar s\) and the squared fraction of steps that moved toward the goal.

\[ \begin{aligned} P &= \max\Big(\bar s,\ \big(\tfrac{1}{B}\textstyle\sum_t \mathbb{1}[r^{\text{goal}}_t \gt 0]\big)^2\Big), \\ \rho_t &= \rho_{\min} + (\rho_{\max} - \rho_{\min})(1 - P), \\ r_t &=r^{\text{task}}_t + m_t\, \rho_t\, \tilde r_t. \end{aligned} \]
adaptive shaping

Early on, when \(P \approx 0\), language guidance is at full strength; as the policy starts succeeding it recedes and the task reward takes over. The critic of a standard SAC agent trains on \(r_t\).

Results

Higher success on two manipulation suites

We compare against reward-learning baselines on sparse-reward tasks: SAC on the task reward alone; LIV, a robotics reward model pretrained on large-scale data, and LIV-Proj, with random fixed projection heads; Relay, which adds relay exploration to LIV; and FuRL, which aligns fuzzy VLM rewards and uses relay RL. Visual observations are embedded with LIV throughout.

Hover or focus a bar for its value

Ten Sawyer-arm tasks with a sparse task reward and a goal that stays fixed and unobserved. Baselines range from plain SAC to FuRL, which already aligns VLM rewards and adds relay exploration.

Meta-World MT10 · hidden, fixed goals Average success rate (%) over 10 tasks and five seeds
SAC 37.0
LIV 28.0
LIV-Proj 24.0
Relay 54.0
FuRL w/o goal image 70.0
FuRL 76.0
LAGEA 80.0

80.0% against 76.0% for FuRL: +4.0 points, a 5.3% relative gain. LAGEA solves eight of the ten tasks in every seed, including push, where FuRL reaches 80%.

Peg-insert-side and pick-place stay at zero for every method, LAGEA included.

Per-task results · MT10 · fixed goals 10 tasks × 7 methods
Task SAC LIV LIV-Proj Relay FuRL w/o goal image FuRL LAGEA
button-press-topdown 0 0 0 60 80 100 100
door-open 50 0 0 80 100 100 100
drawer-close 100 100 100 100 100 100 100
drawer-open 20 0 0 40 80 80 100
peg-insert-side 0 0 0 0 0 0 0
pick-place 0 0 0 0 0 0 0
push 0 0 0 0 40 80 100
reach 60 80 80 100 100 100 100
window-close 60 60 40 80 100 100 100
window-open 80 40 20 80 100 100 100
Average 37.0 28.0 24.0 54.0 70.0 76.0 80.0

Success rate (%) averaged over five seeds; the paper reports no spread for this table. Bold marks the best in each row and underline the runner-up.

The same ten tasks, but the goal moves every episode and is visible to the agent, so the policy has to generalise across goal positions.

Meta-World MT10 · observable, random goals Average success rate (%) over 10 tasks and five seeds
SAC 49.8
Relay 57.4
FuRL 64.6
LAGEA 70.4

70.4% against 64.6% for FuRL: +5.8 points, a 9.0% relative gain, with the smallest spread across seeds (±1.85 against ±5.0).

Contact-rich tasks remain hard. Pick-place reaches 4% and push 12%, and peg-insert-side stays at zero.

Per-task results · MT10 · random goals 10 tasks × 4 methods
Task SAC Relay FuRL LAGEA
button-press-topdown 16.0±32.0 56.0±38.3 64.0±32.6 96±8
door-open 78.0±39.2 80.0±30.3 96.0±8.0 100±0
drawer-close 100±0 100±0 100±0 100±0
drawer-open 40.0±49.0 50.0±42.0 84.0±27.3 92±9.8
pick-place 0±0 0±0 0±0 4±4.9
peg-insert-side 0±0 0±0 0±0 0±0
push 0±0 0±0 6.0±8.0 12±4
reach 100±0 100±0 100±0 100±0
window-close 86.0±28.0 96.0±4.9 100±0 100±0
window-open 78.0±39.2 92.0±7.5 96.0±4.9 100±0
Average 49.8 57.4 64.6 70.4

Success rate (%) ± standard deviation over five seeds. Bold marks the best in each row and underline the runner-up.

A 7-DoF arm with a two-finger gripper, 50-step episodes and the standard sparse binary reward: reach, push, pick-and-place and slide.

Gymnasium-Robotics Fetch · 4 tasks Average success rate (%) over three seeds
SAC 34.17
Relay 37.5
FuRL 44.17
LAGEA 51.67

51.7% against 44.2% for FuRL: +7.5 points, a 17% relative gain. Reach is solved by everyone; on push, pick-and-place and slide LAGEA is best.

Slide, where the puck has to be struck and left to glide, stays hard at 10%.

Per-task results · Fetch 4 tasks × 4 methods
Task SAC Relay FuRL LAGEA
Reach 100±0 100±0 100±0 100±0
Push 26.67±4.71 30±8.16 40±8.16 53.33±4.71
PickAndPlace 10±8.16 20±0 33.33±9.43 43.33±4.71
Slide 0±0 0±0 3.33±4.71 10±8.16
Average 34.17 37.5 44.17 51.67

Success rate (%) ± standard deviation over three seeds. Bold marks the best in each row and underline the runner-up.

Faster to converge, not just higher at the end

Learning curves on eight Meta-World tasks: LAGEA reaches high success earlier than FuRL and SAC.

Figure 3. Success over 1M environment steps on eight MT10 tasks. LAGEA reaches high success well before FuRL on most tasks, and SAC with the sparse reward alone mostly stalls. The extra VLM call per episode is paid back in samples: averaged over nine tasks, LAGEA converges in 92.4 minutes of wall-clock time against 94.9 for FuRL, and is faster on six of the nine.

The benchmarks Meta-World MT10 and Fetch
The ten Meta-World MT10 tasks.

Meta-World MT10. Ten tasks with a Sawyer arm, a 4-D action (end-effector motion and gripper) and 500-step episodes, from reaching to peg insertion. Each task is paired with a one-line instruction, such as “Press a button from the top.”

The four Gymnasium-Robotics Fetch tasks: reach, push, slide and pick-and-place.

Gymnasium-Robotics Fetch. A 7-DoF arm with a two-finger gripper on reach, push, slide and pick-and-place, with 50-step episodes and a sparse binary reward.

Ablations

Which parts carry the gain

Each ablation keeps the encoders, the SAC learner and the goal image fixed and changes one thing. The components turn out to be complementary: removing any one of them costs about as much as removing another.

Remove one reward term at a time Average success rate (%), as reported in Fig. 4b
w/o goal-delta reward 79
w/o feedback-delta reward 80
w/o adaptive ρ 80
Full LAGEA 99

Dropping the goal-delta reward, the feedback-delta reward or the adaptive schedule each leaves about 80%. Only all three together reach 99%: long-term progress, short-term correction, and knowing when to step back.

Which frames carry the feedback Average success rate (%) on five observable MT10 tasks
Random keyframes 68.0
Uniform keyframes 67.3
LAGEA keyframes 80.0

Random or evenly spaced frames lose 12 to 13 points. The gap is widest on button-press, 96% against 37% and 50%, where the decisive moment is a brief contact.

Drawer-open: with key frames, the VLM reward rises and success climbs to about 78% by 0.6M steps; without them, success stays near zero.

Drawer-open, with and without key frames. With key frames the shaped reward turns upward as the agent learns the approach, and success follows. Without them the feedback is smeared over the whole episode and the agent stays stuck.

Structured versus free-form reflections Average success rate (%) on six MT10 tasks
Free-form text 78.9
Schema-constrained JSON 98.9

Letting the VLM answer in free text costs 20 points. The biggest single drop is button-press-topdown, 93% to 10%: verbose, ambiguous explanations make noisy feedback vectors, and noisy vectors make misleading rewards.

Swap the reflecting VLM Average success rate (%) on five observable MT10 tasks
SmolVLM2 56.0
InternVL2 68.0
OpenQwen2VL 66.7
Qwen2.5-VL-3B 80.0
Swap the text encoder Average success rate (%) on the same five tasks
MPNet 70.7
BGE 71.3
LIV 72.7
GPT-2 80.0

The reflecting VLM matters more than the text encoder. All three alternative encoders stay above 70%, but a weaker VLM costs 12 to 24 points, with SmolVLM2 falling to 56%. Qwen2.5-VL-3B with GPT-2 is the best pairing, which is why it is the default.

Evaluate from cameras never seen in training Average success rate (%) on five observable MT10 tasks
Training view 80.0
Directly overhead 79.3
Front-left diagonal 77.3
Behind-left diagonal 79.3

Trained from one diagonal front camera and evaluated zero-shot from three others, LAGEA stays within three points of its training-view score (77–79% against 80%).

Heavy self-occlusion and cluttered scenes were not tested and remain open.

Limitations

LAGEA inherits occasional hallucinations from its VLM; the schema and the alignment reduce them but cannot remove them. All experiments are in simulation, and contact-rich insertion tasks remain unsolved. Moving to real robots, closing the sim-to-real gap, and handling long horizons under partial observability are the natural next steps.

Cite

Citation

If this work is useful to you, please cite it as:

@inproceedings{chowdhury2026lagea,
  title     = {{LAGEA}: Language Guided Embodied Agents for Robotic Manipulation},
  author    = {Chowdhury, Abdul Monaf and Mazumder, Akm Moshiur Rahman and
               Arib, Safaeid Hossain and Akter, Rabeya},
  booktitle = {Proceedings of the 43rd International Conference on Machine Learning},
  series    = {Proceedings of Machine Learning Research},
  volume    = {306},
  year      = {2026}
}