Agents & ControlintermediateGPU optional~45 minColab

Reward shaping, and how it goes wrong

تشكيل المكافأة — وكيف ينحرف

The agent optimises what you wrote, not what you meant

Sparse rewards are hard to learn from: an that only scores on success spends most of training seeing nothing at all. So you help it. You add a term for progress — distance closed, fuel saved, altitude held — and learning speeds up immediately. That term is a shaping reward, and it is the most common intervention in applied RL.

It is also where the failures live. The agent does not read your intent; it reads the sum. If some behaviour you never considered scores higher than the behaviour you wanted, it will find that behaviour, and it will look like success on your own metric while doing the opposite of the job.

The goal

Train the same agent on three reward functions — the environment's own, a correct potential-based shaping, and one plausible mis-specification — then show that the mis-specified agent scores best on the shaped reward and worst on the thing you actually wanted.

Colab opens a read-only copy. Save a copy to Drive to keep your edits.

The notebook needs a keyboard — best opened on a desktop.

The papers behind this

The exploit you are about to see is not the one most people predict. Ask yourself, before running anything, what an agent does when the bonus you added is negative at every point in the state space.

LunarLander gives you a clean case to break. Its own reward already balances several things — distance to the pad, fuel spent, a large bonus for landing and a large penalty for crashing — and an agent trained on it learns to land.

We will keep that agent as the control, then add two shaping terms to it. One is provably safe. The other is the kind of thing a reasonable engineer writes on a Tuesday afternoon.

Preflight. No GPU, no downloads — the is generated, not fetched.

import arabic_reshaper
import azimuth_nb as azimuth
from bidi.algorithm import get_display

env = azimuth.setup(SLUG, lang=LANG, profile=PROFILE)
Workshop code

Read the three functions before you run anything. Decide now which one you think will land most often — the point is partly to be wrong.

# THE THREE REWARDS. Read them before running anything.
#
# LunarLander's observation is
#   [x, y, vx, vy, angle, angular_velocity, leg_left_contact, leg_right_contact]
# so `-(x**2 + y**2) ** 0.5` is distance to the pad, which sits at the origin.

GAMMA = env.cfg["gamma"]


def potential(obs):
    """Φ(s): closer to the pad is higher. Used by BOTH shaped rewards, so the
    difference between them is purely HOW it is applied, not what it measures."""
    return -((obs[0] ** 2 + obs[1] ** 2) ** 0.5)


def reward_honest(reward, obs, next_obs, done):
    """The environment's own reward, untouched. The control."""
    return reward


def reward_shaped(reward, obs, next_obs, done):
    """POTENTIAL-BASED shaping (Ng, Harada & Russell, 1999).

    Rewards the CHANGE in potential, γ·Φ(s′) − Φ(s). Because the added terms
    telescope over an episode, they cannot change which policy is optimal —
    they only make the gradient less sparse. This is the safe form, and it is
    the only shaping with a proof attached.
    """
    return reward + 10.0 * (GAMMA * potential(next_obs) - potential(obs))


def reward_hacked(reward, obs, next_obs, done):
    """The mis-specification. Rewards the STATE, not the change.

    This is not a strawman. "Give it points for being close to the target" is
    the single most natural thing to write, it reads as obviously helpful, and
    it is wrong for a reason that is invisible until you watch the agent: a
    bonus paid every step for BEING somewhere is a bonus for STAYING there.
    Landing ends the episode, and ending the episode ends the income.
    """
    return reward + 10.0 * potential(next_obs)


REWARDS = {
    "honest": reward_honest,
    "shaped": reward_shaped,
    "hacked": reward_hacked,
}

if env.lang == "ar":
    print("ثلاث دوال مكافأة:", " · ".join(REWARDS))
    print("أيّها تظنّه سيهبط أكثر؟ قرّر قبل التشغيل.")
else:
    print("three reward functions:", " · ".join(REWARDS))
    print("which do you think lands most often? decide before you run.")
Workshop code
three reward functions: honest · shaped · hacked
which do you think lands most often? decide before you run.

Three agents, one architecture, one seed. Only the reward differs.

import gymnasium as gym
import numpy as np
import torch
import torch.nn as nn


class Policy(nn.Module):
    """A small categorical policy with a value baseline."""

    def __init__(self, n_obs, n_actions, hidden):
        super().__init__()
        self.body = nn.Sequential(
            nn.Linear(n_obs, hidden), nn.Tanh(), nn.Linear(hidden, hidden), nn.Tanh()
        )
        self.actor = nn.Linear(hidden, n_actions)
        self.critic = nn.Linear(hidden, 1)

    def forward(self, obs):
        h = self.body(obs)
        return self.actor(h), self.critic(h).squeeze(-1)


def train(reward_fn, seed):
    """One agent, one reward function. Same seed for all three, so the only
    thing that differs between the runs is the function itself."""
    torch.manual_seed(seed)
    environment = gym.make("LunarLander-v3")
    net = Policy(
        environment.observation_space.shape[0], environment.action_space.n, env.cfg["hidden"]
    )
    optimizer = torch.optim.Adam(net.parameters(), lr=env.cfg["learningRate"])

    for episode in range(env.cfg["episodes"]):
        obs, _ = environment.reset(seed=seed + episode)
        log_probs, values, rewards = [], [], []
        done = False
        while not done:
            obs_t = torch.as_tensor(obs, dtype=torch.float32)
            logits, value = net(obs_t)
            dist = torch.distributions.Categorical(logits=logits)
            action = dist.sample()
            next_obs, raw_reward, terminated, truncated, _ = environment.step(action.item())
            done = terminated or truncated

            log_probs.append(dist.log_prob(action))
            values.append(value)
            # THE ONLY LINE THAT DIFFERS between the three agents.
            rewards.append(reward_fn(raw_reward, obs, next_obs, done))
            obs = next_obs

        returns, running = [], 0.0
        for r in reversed(rewards):
            running = r + GAMMA * running
            returns.append(running)
        returns = torch.tensor(list(reversed(returns)), dtype=torch.float32)
        returns = (returns - returns.mean()) / (returns.std() + 1e-8)

        values_t = torch.stack(values)
        advantage = returns - values_t.detach()
        loss = -(torch.stack(log_probs) * advantage).sum() + nn.functional.mse_loss(
            values_t, returns, reduction="sum"
        )
        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

        if episode % 100 == 0:
            print(f"  {episode:5d}  return {sum(rewards):8.1f}")

    environment.close()
    return net


agents = {}
for name, fn in REWARDS.items():
    print(f"\n{name}:")
    agents[name] = train(fn, env.cfg["seed"])

episodes_run = env.cfg["episodes"] * len(REWARDS)
Workshop code
honest:
      0  return   -264.6
    100  return   -132.8
    200  return   -102.0
    300  return    -65.5
    400  return     -7.1
    500  return    -22.2
    600  return    -99.2
    700  return    -83.4
    800  return      8.3
    900  return     93.6
   1000  return     -2.7
   1100  return    -40.8
   1200  return     52.2
   1300  return   -100.0
   1400  return    277.0

shaped:
      0  return   -248.9
    100  return   -159.0
    200  return    -98.8
    300  return    -63.4
    400  return     -0.6
    500  return     25.3
    600  return    -37.4
    700  return   -334.1
    800  return     -9.0
    900  return    -67.7
   1000  return   -116.3
   1100  return     50.5
   1200  return     -1.7
   1300  return    237.0
   1400  return    166.9

hacked:
      0  return  -1887.8
    100  return  -1665.9
    200  return  -1426.3
    300  return  -1058.9
    400  return  -1399.0
    500  return   -683.6
    600  return   -582.1
    700  return   -858.4
    800  return   -679.8
    900  return   -557.0
   1000  return   -719.8
   1100  return   -983.4
   1200  return   -567.7
   1300  return   -939.5
   1400  return   -618.3

Two columns, two different verdicts. The shaped-score column is what the training loop saw; the landing column is what you wanted. Read them against each other.

def ar(text: str) -> str:
    """Join Arabic letters into their connected forms, then reorder for LTR drawing."""
    return get_display(arabic_reshaper.reshape(text))


def evaluate(net, reward_fn, seed, n):
    """Two numbers per agent, deliberately.

    `shaped` is what the training loop optimised. `landed` is the job. Keeping
    them apart is the entire point: an agent can move one without the other,
    and reporting a single 'score' would hide exactly the effect being taught.
    """
    environment = gym.make("LunarLander-v3")
    shaped_total, landed = 0.0, 0
    for i in range(n):
        obs, _ = environment.reset(seed=10_000 + seed + i)
        done = False
        while not done:
            with torch.no_grad():
                logits, _ = net(torch.as_tensor(obs, dtype=torch.float32))
            action = int(torch.argmax(logits))
            next_obs, raw_reward, terminated, truncated, _ = environment.step(action)
            shaped_total += reward_fn(raw_reward, obs, next_obs, terminated or truncated)
            obs = next_obs
            done = terminated or truncated
        # Both legs down and the lander at rest: gymnasium pays +100 on a
        # successful landing, so the final raw reward is the honest verdict.
        if terminated and raw_reward >= 100:
            landed += 1
    environment.close()
    return shaped_total / n, landed / n


n_eval = env.cfg["evalEpisodes"]
results = {}
for name, net in agents.items():
    # Every agent is scored on the HACKED reward as well as its own, so the
    # columns are comparable — otherwise each agent is graded on its own exam.
    own_shaped, landed = evaluate(net, REWARDS[name], env.cfg["seed"], n_eval)
    hacked_shaped, _ = evaluate(net, reward_hacked, env.cfg["seed"], n_eval)
    results[name] = {"own": own_shaped, "hacked_scale": hacked_shaped, "landed": landed}

header = f"{'agent':10}{'own reward':>14}{'hacked reward':>16}{'landed':>10}"
print("\n" + header)
print("-" * len(header))
for name, r in results.items():
    print(f"{name:10}{r['own']:>14.1f}{r['hacked_scale']:>16.1f}{r['landed']:>9.0%}")

# The card's thumbnail, and the finding in one picture: the bar that wins the
# written objective is the bar that never lands. A table says it; a chart makes
# it impossible to miss at index size.
import matplotlib.pyplot as plt

fig, ax = plt.subplots(figsize=(6, 3))
names = list(results)
ax.bar(names, [results[n]["landed"] * 100 for n in names], color=["#2a9d8f", "#457b9d", "#e76f51"])
for i, n in enumerate(names):
    ax.text(
        i, results[n]["landed"] * 100 + 2, f"{results[n]['landed']:.0%}", ha="center", fontsize=9
    )
ax.set_ylabel("landed (%)" if env.lang == "en" else ar("نسبة الهبوط (%)"))
ax.set_ylim(0, 105)
ax.spines[["top", "right"]].set_visible(False)
fig.tight_layout()
plt.show()

honest_landing = results["honest"]["landed"]
hacked_landing = results["hacked"]["landed"]
# A DIFFERENCE, not a ratio. Returns here are negative — the shaping term is
# -distance summed over the episode — and a ratio over signed quantities is
# meaningless: -742.9 / 1372.6 came out as -0.54 and read as "the exploit lost"
# when it had in fact won by 630 points. Ratios need a positive denominator
# and a meaningful zero; neither holds for a return.
shaped_advantage = results["hacked"]["hacked_scale"] - results["honest"]["hacked_scale"]
# Landing rates ARE ratios of counts: bounded, non-negative, zero means zero.
landing_ratio = hacked_landing / max(honest_landing, 1e-9)
Workshop code
agent         own reward   hacked reward    landed
--------------------------------------------------
honest             -34.2         -1654.1      24%
shaped             -82.5         -5592.5       0%
hacked            -808.3          -808.3       0%

One trajectory each, compared. LENGTH is the tell — read it before you read anything else.

# One trajectory from the hacked agent, summarised. The scoreboard says it
# scores well and lands rarely; this says what it is doing instead.
environment = gym.make("LunarLander-v3")
obs, _ = environment.reset(seed=99)
altitudes, steps, done = [], 0, False
while not done and steps < 1000:
    with torch.no_grad():
        logits, _ = agents["hacked"](torch.as_tensor(obs, dtype=torch.float32))
    obs, _, terminated, truncated, _ = environment.step(int(torch.argmax(logits)))
    altitudes.append(float(obs[1]))
    steps += 1
    done = terminated or truncated
environment.close()

# Compare against an honest episode, and let the NUMBERS say what happened.
# The first version of this cell asserted hovering — "income every step" —
# and was exactly wrong: `potential` is -distance, so it is always negative,
# the bonus is a per-step TAX, and the fastest way to stop paying it is to end
# the episode. The agent did not learn to loiter. It learned to die.
environment = gym.make("LunarLander-v3")
obs, _ = environment.reset(seed=99)
honest_steps, done = 0, False
while not done and honest_steps < 1000:
    with torch.no_grad():
        logits, _ = agents["honest"](torch.as_tensor(obs, dtype=torch.float32))
    obs, _, terminated, truncated, _ = environment.step(int(torch.argmax(logits)))
    honest_steps += 1
    done = terminated or truncated
environment.close()

if env.lang == "ar":
    print(f"المخترِق: {steps} خطوة · ارتفاع وسيط {np.median(altitudes):.2f}")
    print(f"الأمين:  {honest_steps} خطوة")
else:
    print(f"hacked: {steps} steps · median altitude {np.median(altitudes):.2f}")
    print(f"honest: {honest_steps} steps")

if steps < honest_steps * 0.7:
    verdict_en = (
        "The exploit is a SHORT episode. `potential` is negative everywhere, so the "
        "bonus is a tax charged every step, and the cheapest policy is to stop "
        "paying it — end the episode. The agent is not confused; it found the fastest "
        "exit from a reward you wrote."
    )
    verdict_ar = (
        "الثغرة حلقة قصيرة. دالة الجهد سالبة في كل مكان، فالمكافأة ضريبة تُجبى كل "
        "خطوة، وأرخص سياسة أن تكفّ عن دفعها — أي أن تنهي الحلقة. الوكيل ليس مرتبكاً؛ "
        "بل وجد أسرع مخرج من مكافأة كتبتَها أنت."
    )
elif steps > honest_steps * 1.3:
    verdict_en = (
        "The exploit is a LONG episode: it loiters where the bonus is largest and "
        "never risks the landing. Same mechanism, opposite sign."
    )
    verdict_ar = (
        "الثغرة حلقة طويلة: يتسكّع حيث المكافأة أكبر ولا يخاطر بالهبوط قط. الآلية "
        "ذاتها بإشارة معاكسة."
    )
else:
    verdict_en = "Episode lengths are similar — look at where it spends its altitude instead."
    verdict_ar = "أطوال الحلقات متقاربة — انظر إلى أين ينفق ارتفاعه بدلاً من ذلك."

print("\n" + (verdict_ar if env.lang == "ar" else verdict_en))
Workshop code
hacked: 103 steps · median altitude 0.41
honest: 148 steps

The exploit is a SHORT episode. `potential` is negative everywhere, so the bonus is a tax charged every step, and the cheapest policy is to stop paying it — end the episode. The agent is not confused; it found the fastest exit from a reward you wrote.

Exercise

Repair the mis-specified reward without deleting it. is the known-safe form: reward γ·Φ(s′) − Φ(s) rather than Φ(s) itself. Convert the proximity bonus into that form and retrain.

Then answer the question the fix raises. If potential-based shaping provably cannot change the , what did you actually buy by adding it? Compare not just the final landing rate but how many episodes each agent needed to get there.

# YOUR TURN.
#
# Convert the proximity bonus to potential-based form and retrain. The safe
# shaping is already written above as `reward_shaped` — the exercise is to
# understand WHY it is safe, then answer what it actually bought you.
#
# Compare against the numbers above, and look at episodes-to-competence, not
# just the final landing rate.
YOUR_REWARD = reward_shaped  # try your own

fixed = train(YOUR_REWARD, env.cfg["seed"])
fixed_shaped, fixed_landed = evaluate(fixed, YOUR_REWARD, env.cfg["seed"], n_eval)
print(
    f"yours: landed {fixed_landed:.0%}  ·  honest {honest_landing:.0%}  ·  hacked {hacked_landing:.0%}"
)
Workshop code

A hint is available in the notebook — env.hint(1)

Both checks. The second is a MAXIMUM: the exploit must land less, or nothing was demonstrated.

# The control first. A comparison against an agent that never learned the task
# is not a comparison, and 0/0 quietly satisfies "the exploit lands less".
control_ok = env.check("honest-agent-works", honest_landing)
if not control_ok:
    if env.lang == "ar":
        print("  الشاهد لم يتعلّم الهبوط — ارفع episodes، ولا تخفض العتبة.")
    else:
        print("  the control never learned to land — raise `episodes`, do not lower the bar.")

hacked_ok = env.check("hacked-scores-higher", shaped_advantage)
landing_ok = env.check("hacked-lands-less", landing_ratio)
Workshop code
✗ Landing rate of the control agent, before any comparison is meaningful: 0.24 (needs ≥ 0.3)
  the control never learned to land — raise `episodes`, do not lower the bar.
✓ How far the mis-specified agent out-scores the honest one, on the written objective: 845.8 (needs ≥ 50)
✓ Mis-specified agent's landing rate, relative to the honest agent's: 0 (needs ≤ 0.5)
receipt = env.receipt()
Workshop code
Not complete yet — failed: honest-agent-works
Last verified: 2026-08-28 · unknown · PyTorch unknown · Python 3.13.5 · e884576

Terms in this workshop