Preprint2026Goal-conditioned RL

Do Better Goal Representations Improve Goal-Conditioned Reinforcement Learning?

  • Syed Nazmus Sakib1
  • Abdul Monaf Chowdhury1
  • Nafiul Haque1
  • Shifat E Arman1,2
  • Md Mehedi Hasan1
  • 1University of Dhaka
  • 2University of Oxford
TL;DR

Goal-representation research assumes that encoding reachability more accurately makes goal-conditioned agents act better. We test this directly by handing agents an exactly computed goal representation. It barely helps, and neither corrupting it nor replacing it with random noise hurts much. The same intervention applied to the agent's own state more than doubles success, and a map-free version, random Fourier features of the agent's \((x, y)\), raises GCIVL on antmaze-large from 32 to 84.

6of 9
Exact-vs-learned goal comparisons that are statistically indistinguishable
3 split both ways
7.6pts
Largest shift across 16 corrupted goal representations
every CI includes zero
+39.6pts
From moving one random code from the goal to the state
32 → 84
GCIVL on antmaze-large with map-free position features
Left: a goal described by its temporal distances to landmark states, and nested distance radii. Right: raw observations lifted into a higher-dimensional state encoding from which the policy acts.

Figure 1. Two sides of the same interface. Left, top: goal-side representations, the axis this literature optimises, describe a goal by its temporal distances to other states. Left, bottom: the property that motivates them: such distances compose, so a policy accurate within one radius extends to two and four. Right: our proposal encodes the agent's own position instead, lifting observations that are close together in raw coordinates into a representation in which they are far apart, and from which the policy acts.

The question

The premise under test

Offline goal-conditioned RL learns one policy \(\pi(a \mid s, g)\) that reaches diverse target states, from a fixed dataset whose goals are relabelled in hindsight. A central design question is how the goal should be represented. Recent work characterises goals through future occupancy, temporal distance or controllability: contrastive representations, value-implicit pre-training, quasimetric and temporal-distance embeddings, and dual goal representations. The objectives differ, but the premise is shared: a representation that captures reachability more accurately should enable better control.

Existing evaluations don't isolate that premise. Comparisons change the representation-learning procedure and the resulting embedding at the same time, so it is hard to tell whether a gain comes from the information the representation carries or from how it was learned. We ask the question directly:

If a goal-conditioned agent were handed a perfect goal representation, how much better would it act?

In deterministic mazes, the question has an exact answer, because the perfect representation can be computed. The answer turns out to be barely, and it points somewhere else: the other argument of the policy, the agent's own state.

Setup

Measuring with a perfect representation

The goal interface

Any goal-conditioned network combines its two arguments before producing an output. Writing that combination explicitly as the goal interface \(I\),

\[ \pi(a \mid s, g) = h\big(I(s, \varphi(g))\big), \qquad I_{\text{concat}}(s, \varphi) = \mathrm{MLP}\big([\,s\,;\,\varphi\,]\big). \]
goal interface

Every goal-representation method we know of fixes the interface to concatenation. The representation \(\varphi\), the interface \(I\) and the state pathway are separate design axes, and this work measures all three.

An exact dual representation

Dual goal representations (DGR; Park et al., 2026) describe a goal by its temporal distances from other states, which is provably sufficient for optimal control. In practice it is evaluated against \(K\) landmark states \(a_1, \dots, a_K\) sampled from the data:

\[ \varphi(g) = \big[\, d(a_1, g),\ \dots,\ d(a_K, g) \,\big] \in \mathbb{R}^K. \]
dual representation

In continuous domains DGR has to learn these distances, with a bilinear value parameterisation. In the deterministic OGBench mazes, though, temporal distance is a graph geodesic, so a breadth-first search from every landmark gives the exact table (\(K = 256\) for AntMaze and HumanoidMaze, \(64\) for PointMaze). We call this the ideal representation and swap it in for the learned one, leaving the downstream algorithm, objective and every other component unchanged. The two downstream algorithms are GCIVL (expectile value learning with advantage-weighted regression) and CRL (a contrastive critic with a behaviour-regularised policy).

Measuring quality without a probe

To score how much distance structure a representation keeps, we decode distances with a min-plus rule and take the Spearman correlation \(\rho\) with the true geodesic, overall and within distance quartiles:

\[ \hat d(s, g) = \min_i\, \big[\, d(a_i, s) + \varphi_i(g) \,\big]. \]
min-plus decode

A learned probe is uninformative here: in a maze, the goal state determines its own position, so any representation that preserves goal identity contains complete distance information. An MLP probe scores 0.997 to 0.999 on a raw 2-D goal, a frozen random projection and a learned representation alike. The decode reads the representation directly and fits nothing.

A corruption ladder

Starting from the exact table, we remove distance information in two ways: Gaussian noise on every entry, and far-field noise on only the furthest quartile, which keeps nearby relationships intact:

\[ \tilde\varphi_i(g) = \max\big(0,\ \varphi_i(g) + \sigma\,\epsilon_{i,g}\big), \qquad \epsilon_{i,g} \sim \mathcal{N}(0, 1). \]
corruption

The noise is drawn once and fixed, so each setting is a deterministic representation rather than observation noise. The same constructions can be applied to the state instead, by appending a fixed per-cell code, \(\tilde s = [\,s\,;\, c(s)\,]\), of the same dimensionality and indexing. State codes need the maze layout, so they are diagnostics, not a method.

Rank correlation between decoded and true distance within each distance quartile, as Gaussian and far-field corruption grow.

What the two corruptions do to the geometry. Spearman correlation between decoded and true distance, within quartiles of true distance. Gaussian corruption degrades every range at once. Far-field corruption keeps the nearest quartile exact while driving the furthest negative, so at \(\sigma = 8\) distant goals are ranked in reverse.

Finding 1

Goal representations are saturated

Replacing the learned representation with the exact one, with everything else held fixed, does not reliably improve control.

Environment Algo DGR (published) Learned (ours) Ideal (ours) Ideal − learned, 95% CI
pointmaze-medium GCIVL 76 ± 7 70.9 ± 6.1 (5) 62.7 ± 3.0 (5) −8.2 [−15.6, −0.7]
pointmaze-large GCIVL 46 ± 6 46.4 ± 6.5 (5) 46.3 ± 7.9 (5) −0.1 [−10.7, +10.5]
antmaze-medium GCIVL 75 ± 4 76.5 ± 7.4 (5) 68.4 ± 4.0 (5) −8.0 [−17.2, +1.1]
antmaze-large GCIVL 28 ± 11 32.2 ± 8.9 (8) 30.0 ± 8.0 (8) −2.2 [−11.3, +6.9]
humanoidmaze-medium GCIVL 29 ± 3 27.8 ± 3.6 (5) 34.2 ± 3.1 (5) +6.4 [+1.5, +11.4]
pointmaze-medium CRL 33 ± 1 37.8 ± 3.3 (5) 57.4 ± 11.6 (5) +19.5 [+5.4, +33.7]
pointmaze-large CRL 39 ± 12 35.4 ± 7.9 (5) 43.0 ± 17.6 (4) +7.6 [−18.8, +34.1]
antmaze-medium CRL 93 ± 3 94.6 ± 1.6 (2) 92.5 ± 0.4 (2) −2.1 [−13.8, +9.7]
antmaze-large CRL 87 ± 2 82.2 ± 3.4 (8) 80.8 ± 3.4 (8) −1.4 [−5.1, +2.2]

Success rate (%), mean ± std with seed counts in parentheses; published DGR numbers are over 8 seeds. Intervals are 95% Welch intervals, drawn on a shared −20 to +35 scale with the vertical line at zero. Highlighted intervals exclude zero.

Six of the nine comparisons are statistically indistinguishable. Of the other three, the exact representation wins by 6.4 points (humanoidmaze-medium, GCIVL) and 19.5 points (pointmaze-medium, CRL), and loses by 8.2 (pointmaze-medium, GCIVL). The differences have no consistent direction, and across the five GCIVL tasks the exact representation averages 48% against 51% for the learned one. In the tabular Lights-Out setting of the original DGR paper the exact representation does help; in these continuous-control mazes it does not help consistently.

Learning curves: exact vs. learned six settings
Evaluation curves over one million steps for learned and exact goal representations in six settings.

Mean ± 1 SE over 2 to 8 seeds per panel; violet is learned, gold exact. Learning speed and asymptotic success are indistinguishable.

Degrading the representation barely matters

The learned representation might simply be good enough already. So we walk the exact one down the corruption ladder, from exact geodesics to essentially no distance information.

Success rate against representation quality for Gaussian and far-field corruption, for GCIVL and CRL, staying close to the uncorrupted line.

Control is insensitive to goal quality. antmaze-large, mean ± 1 SD over 5 to 8 seeds. Across corruption ladders from exact geodesics to zero distance information, no performance difference is distinguishable from zero.

Far-field corruption at \(\sigma = 8\) removes essentially all rank information about distant goals: the far-quartile correlation falls from 1.00 to −0.41 while the near quartile stays at 0.997. Yet success drops by only 6.4 points for GCIVL and 5.1 for CRL. Across all sixteen settings, every point estimate is within 7.6 points of the uncorrupted representation, every interval includes zero, and the correlation between measured quality and success stays between +0.21 and +0.27, none distinguishable from zero.

This is not a claim of exact flatness: the pre-stated tolerance was five points, and several intervals extend past it. The supported conclusion is narrower. Over nearly the full range of representation quality, no difference is distinguishable from zero, while the state-side interventions below move success by 30 to 50 points.

A code with no geometry does just as well

Finally, we replace the goal representation with a table of the same shape and indexing whose entries are i.i.d. \(\mathcal{N}(0, 1)\). It identifies the goal cell and contains no geometry at all (rank correlation with true distance: 0.004). It scores 31.9 ± 2.4 on antmaze-large, against 30.0 ± 8.0 for exact geodesics (+1.9 [−4.9, +8.8]) and 32.2 ± 8.9 for DGR's learned representation (−0.2 [−7.9, +7.4]). Whatever these agents extract from the goal, it is little more than its identity.

Finding 2

The state pathway is the bottleneck

If the goal side is saturated, the headroom lies elsewhere. So we move the identical intervention to the other argument of the policy.

The same random code, on either side of the interface antmaze-large · GCIVL · success (%)
Learned goal rep. (DGR) 32.2
Exact goal rep. (BFS) 30.0
Random code as goal 31.9
Random code on state 71.5
Code on the goal Code on the state

In the last two rows the construction, dimensionality, cell indexing and information content are identical; only the argument the code is applied to differs. That alone is worth 39.6 points. The same holds for CRL, where an exact state code raises success from 80.8 ± 3.4 to 89.9 ± 1.5 (+9.2 [+6.1, +12.3]) on a task where the goal-side intervention was worth −1.4 [−5.1, +2.2].

What must the state code contain?

We run the state code through the same corruption ladder used for the goal, alongside the random table.

What the state code needs to contain antmaze-large · GCIVL · success (%)
None 30.0
Exact geodesics · ρ = 1.00 64.7
Gaussian σ = 8 · ρ = 0.39 84.5
Far-field σ = 8 · ρ = 0.00 77.6
Random N(0, 1) · ρ = 0.00 71.5
No state code Exact geodesic code Corrupted or random code

Every code helps, and the exact geodesics are the weakest of them. Corrupting the table improves it, by 19.8 points for Gaussian \(\sigma = 8\) and 12.9 for far-field \(\sigma = 8\), and a table with no geometry at all is worth 6.8 points more than the exact one (+6.8 [−4.8, +18.3]). The effect is positional encoding: what the network gains is a high-dimensional, distinguishable representation of where it currently is. The exact geodesic table is a comparatively poor one, because it varies smoothly and nearly linearly across neighbouring cells.

Method

Position features: a map-free state encoding

State codes need the maze layout. Position features need only the agent's coordinates: fixed random Fourier features of \((x, y)\), appended to the observation.

\[ \tilde s = \big[\, s\ ;\ \sin(2\pi B\,[x, y]^\top)\ ;\ \cos(2\pi B\,[x, y]^\top) \,\big] \]
position features

The frequency matrix \(B \in \mathbb{R}^{F \times 2}\) has entries drawn from \(\mathcal{N}(0, \varsigma^2)\), where \(\varsigma\) sets the frequency scale in cycles per coordinate unit. With \(F = 128\), that is 256 fixed features. \(B\) is drawn once and kept fixed across training and seeds. The objective, optimiser, network widths, goal representation and hyperparameters are all unchanged, and the only requirement is knowing which two observation dimensions hold the agent's position: no map, no transition graph.

InteractiveWhat the policy sees of its own position
Click or drag in the maze to move the agent

Similarity 1 cell away –
Similarity 3 cells away –

Colour is the kernel the features induce, \(k(x, x') = \tfrac{1}{F}\sum_i \cos\big(2\pi\, b_i^\top (x - x')\big)\), between each location and the agent, with \(F = 128\) frequencies drawn once. Wavelengths are the paper's three settings in maze-cell units. Fourier features are Euclidean, so they ignore walls; the exact geodesic code respects walls but changes too slowly to tell neighbours apart.

Position features, and the controls around them antmaze-large · GCIVL · success (%)
Raw goal 15.7
Raw goal + position features 71.0
Learned goal rep. (DGR) 32.2
DGR + position features 83.9
DGR + standardised x, y 7.1
DGR + 1024-wide actor 31.8
Without position features With position features

On antmaze-large, position features raise GCIVL from 32.2 to 83.9 with DGR's learned representation; published goal-representation methods span 9 to 28 on this task. On humanoidmaze-medium they roughly double success under both representations. They work without any goal representation at all, lifting a raw-goal baseline from 15.7 to 71.0, and DGR adds +12.9 [+9.0, +16.9] on top: goal representations are not useless, but their contribution is second-order. Across settings, the longest wavelength works best and the shortest worst.

Success rate against feature wavelength for antmaze-large and humanoidmaze-medium, with and without position features.

Success against feature wavelength. Dashed lines mark the same agent without position features. Longer wavelengths keep nearby states similar while separating distant ones; short ones behave like random codes. The sweep stops at 8.5 cells, so where the trend peaks is open.

Ruling out other explanations

Capacity

Doubling the actor to 1024 units changes nothing (−0.4 [−8.8, +8.0]). Across interfaces, parameter count does not track performance: the best uses 635k parameters against 676k for concatenation, and the largest, at 1.47M, is the worst.

Input scale

Standardising \((x, y)\) in place, the smallest change that aligns position with the proprioceptive dimensions, drops success from 32.2 to 7.1. The features help by separating positions, not by rescaling them.

Goal input width

Passing \(\varphi(g)\) through a bottleneck matched to the attention context width changes nothing: −2.8 [−13.5, +7.8] with the exact representation, −0.1 [−6.5, +6.3] with raw goals.

Shortest-path routing

Cross-attention over landmark tokens could implement the min-plus decode, but doesn't. The landmark it would select gets 1.59× and 1.43× chance attention in two seeds; a state-independent "nearest to the goal" heuristic gets more, 2.52× and 1.85×.

The interface is a design axis too

Swapping concatenation for cross-attention, with queries formed from the state, is worth +23.2 [+15.6, +30.9] on antmaze-large with the exact representation, and +33.2 [+25.2, +41.1] when both actor and critic use it. The same actor-and-critic change with the learned representation gives only +3.9 [−3.9, +11.7]. Because the query comes from the state, cross-attention carries an implicit positional encoding, and it becomes largely redundant once an explicit state code is present (+8.9 [−3.4, +21.2]): two routes to the same quantity. Across fifteen controlled comparisons the swap moves success by −8.4 to +26.9 points. Architectural defaults fixed by convention shift performance more than optimising the goal representation does.

Full interface sweep 15 comparisons
Cross-attention minus concatenation with 95% intervals for fifteen settings.

Cross-attention minus concatenation, everything else fixed within a row; whiskers are 95% Welch intervals. Violet marks intervals that exclude zero in favour of cross-attention, gold in favour of concatenation, grey a tie.

Scope

Where it holds, and where it doesn't

PointMaze

None of the interventions help. Observations are four-dimensional and the coordinates are most of them, so there is little for a position encoding to disambiguate. Raising the landmark count from 64 to 256 does not change this.

Manipulation

The interface result does not transfer: cross-attention is worse than concatenation on cube-single-play (−3.6 [−6.5, −0.7]) and no better on scene-play (−6.5 [−13.2, +0.2]). These tasks have no maze geometry and no single pair of coordinates that locates the agent.

We rule out capacity, input scale and internal shortest-path routing, but the precise mechanism by which spatial distinguishability aids policy learning remains open. The usual explanation for Fourier features, spectral bias, is not supported by the frequency sweep. The evaluation is restricted to state-based navigation, position features are untested in pixel-based domains, and we do not claim state-of-the-art results over hierarchical methods.

Cite

Citation

If this work is useful to you, please cite it as:

@article{sakib2026goalreps,
  title   = {Do Better Goal Representations Improve Goal-Conditioned Reinforcement Learning?},
  author  = {Sakib, Syed Nazmus and Chowdhury, Abdul Monaf and Haque, Nafiul and Arman, Shifat E and Hasan, Md Mehedi},
  journal = {Preprint},
  year    = {2026}
}