Agents & ControlintermediateGPU optional~45 minColab
Reward shaping, and how it goes wrong
تشكيل المكافأة — وكيف ينحرف
The agent optimises what you wrote, not what you meant
Sparse rewards are hard to learn from: an that only scores on success spends most of training seeing nothing at all. So you help it. You add a term for progress — distance closed, fuel saved, altitude held — and learning speeds up immediately. That term is a shaping reward, and it is the most common intervention in applied RL.
It is also where the failures live. The agent does not read your intent; it reads the sum. If some behaviour you never considered scores higher than the behaviour you wanted, it will find that behaviour, and it will look like success on your own metric while doing the opposite of the job.
The goal
Train the same agent on three reward functions — the environment's own, a correct potential-based shaping, and one plausible mis-specification — then show that the mis-specified agent scores best on the shaped reward and worst on the thing you actually wanted.
Colab opens a read-only copy. Save a copy to Drive to keep your edits.
The notebook needs a keyboard — best opened on a desktop.
The papers behind this
The exploit you are about to see is not the one most people predict. Ask yourself, before running anything, what an agent does when the bonus you added is negative at every point in the state space.
LunarLander gives you a clean case to break. Its own reward already balances several things — distance to the pad, fuel spent, a large bonus for landing and a large penalty for crashing — and an agent trained on it learns to land.
We will keep that agent as the control, then add two shaping terms to it. One is provably safe. The other is the kind of thing a reasonable engineer writes on a Tuesday afternoon.
Preflight. No GPU, no downloads — the is generated, not fetched.
import arabic_reshaper
import azimuth_nb as azimuth
from bidi.algorithm import get_display
env = azimuth.setup(SLUG, lang=LANG, profile=PROFILE)Read the three functions before you run anything. Decide now which one you think will land most often — the point is partly to be wrong.
# THE THREE REWARDS. Read them before running anything.
#
# LunarLander's observation is
# [x, y, vx, vy, angle, angular_velocity, leg_left_contact, leg_right_contact]
# so `-(x**2 + y**2) ** 0.5` is distance to the pad, which sits at the origin.
GAMMA = env.cfg["gamma"]
def potential(obs):
"""Φ(s): closer to the pad is higher. Used by BOTH shaped rewards, so the
difference between them is purely HOW it is applied, not what it measures."""
return -((obs[0] ** 2 + obs[1] ** 2) ** 0.5)
def reward_honest(reward, obs, next_obs, done):
"""The environment's own reward, untouched. The control."""
return reward
def reward_shaped(reward, obs, next_obs, done):
"""POTENTIAL-BASED shaping (Ng, Harada & Russell, 1999).
Rewards the CHANGE in potential, γ·Φ(s′) − Φ(s). Because the added terms
telescope over an episode, they cannot change which policy is optimal —
they only make the gradient less sparse. This is the safe form, and it is
the only shaping with a proof attached.
"""
return reward + 10.0 * (GAMMA * potential(next_obs) - potential(obs))
def reward_hacked(reward, obs, next_obs, done):
"""The mis-specification. Rewards the STATE, not the change.
This is not a strawman. "Give it points for being close to the target" is
the single most natural thing to write, it reads as obviously helpful, and
it is wrong for a reason that is invisible until you watch the agent: a
bonus paid every step for BEING somewhere is a bonus for STAYING there.
Landing ends the episode, and ending the episode ends the income.
"""
return reward + 10.0 * potential(next_obs)
REWARDS = {
"honest": reward_honest,
"shaped": reward_shaped,
"hacked": reward_hacked,
}
if env.lang == "ar":
print("ثلاث دوال مكافأة:", " · ".join(REWARDS))
print("أيّها تظنّه سيهبط أكثر؟ قرّر قبل التشغيل.")
else:
print("three reward functions:", " · ".join(REWARDS))
print("which do you think lands most often? decide before you run.")three reward functions: honest · shaped · hacked
which do you think lands most often? decide before you run.Three agents, one architecture, one seed. Only the reward differs.
import gymnasium as gym
import numpy as np
import torch
import torch.nn as nn
class Policy(nn.Module):
"""A small categorical policy with a value baseline."""
def __init__(self, n_obs, n_actions, hidden):
super().__init__()
self.body = nn.Sequential(
nn.Linear(n_obs, hidden), nn.Tanh(), nn.Linear(hidden, hidden), nn.Tanh()
)
self.actor = nn.Linear(hidden, n_actions)
self.critic = nn.Linear(hidden, 1)
def forward(self, obs):
h = self.body(obs)
return self.actor(h), self.critic(h).squeeze(-1)
def train(reward_fn, seed):
"""One agent, one reward function. Same seed for all three, so the only
thing that differs between the runs is the function itself."""
torch.manual_seed(seed)
environment = gym.make("LunarLander-v3")
net = Policy(
environment.observation_space.shape[0], environment.action_space.n, env.cfg["hidden"]
)
optimizer = torch.optim.Adam(net.parameters(), lr=env.cfg["learningRate"])
for episode in range(env.cfg["episodes"]):
obs, _ = environment.reset(seed=seed + episode)
log_probs, values, rewards = [], [], []
done = False
while not done:
obs_t = torch.as_tensor(obs, dtype=torch.float32)
logits, value = net(obs_t)
dist = torch.distributions.Categorical(logits=logits)
action = dist.sample()
next_obs, raw_reward, terminated, truncated, _ = environment.step(action.item())
done = terminated or truncated
log_probs.append(dist.log_prob(action))
values.append(value)
# THE ONLY LINE THAT DIFFERS between the three agents.
rewards.append(reward_fn(raw_reward, obs, next_obs, done))
obs = next_obs
returns, running = [], 0.0
for r in reversed(rewards):
running = r + GAMMA * running
returns.append(running)
returns = torch.tensor(list(reversed(returns)), dtype=torch.float32)
returns = (returns - returns.mean()) / (returns.std() + 1e-8)
values_t = torch.stack(values)
advantage = returns - values_t.detach()
loss = -(torch.stack(log_probs) * advantage).sum() + nn.functional.mse_loss(
values_t, returns, reduction="sum"
)
optimizer.zero_grad()
loss.backward()
optimizer.step()
if episode % 100 == 0:
print(f" {episode:5d} return {sum(rewards):8.1f}")
environment.close()
return net
agents = {}
for name, fn in REWARDS.items():
print(f"\n{name}:")
agents[name] = train(fn, env.cfg["seed"])
episodes_run = env.cfg["episodes"] * len(REWARDS)honest:
0 return -264.6
100 return -132.8
200 return -102.0
300 return -65.5
400 return -7.1
500 return -22.2
600 return -99.2
700 return -83.4
800 return 8.3
900 return 93.6
1000 return -2.7
1100 return -40.8
1200 return 52.2
1300 return -100.0
1400 return 277.0
shaped:
0 return -248.9
100 return -159.0
200 return -98.8
300 return -63.4
400 return -0.6
500 return 25.3
600 return -37.4
700 return -334.1
800 return -9.0
900 return -67.7
1000 return -116.3
1100 return 50.5
1200 return -1.7
1300 return 237.0
1400 return 166.9
hacked:
0 return -1887.8
100 return -1665.9
200 return -1426.3
300 return -1058.9
400 return -1399.0
500 return -683.6
600 return -582.1
700 return -858.4
800 return -679.8
900 return -557.0
1000 return -719.8
1100 return -983.4
1200 return -567.7
1300 return -939.5
1400 return -618.3Two columns, two different verdicts. The shaped-score column is what the training loop saw; the landing column is what you wanted. Read them against each other.
def ar(text: str) -> str:
"""Join Arabic letters into their connected forms, then reorder for LTR drawing."""
return get_display(arabic_reshaper.reshape(text))
def evaluate(net, reward_fn, seed, n):
"""Two numbers per agent, deliberately.
`shaped` is what the training loop optimised. `landed` is the job. Keeping
them apart is the entire point: an agent can move one without the other,
and reporting a single 'score' would hide exactly the effect being taught.
"""
environment = gym.make("LunarLander-v3")
shaped_total, landed = 0.0, 0
for i in range(n):
obs, _ = environment.reset(seed=10_000 + seed + i)
done = False
while not done:
with torch.no_grad():
logits, _ = net(torch.as_tensor(obs, dtype=torch.float32))
action = int(torch.argmax(logits))
next_obs, raw_reward, terminated, truncated, _ = environment.step(action)
shaped_total += reward_fn(raw_reward, obs, next_obs, terminated or truncated)
obs = next_obs
done = terminated or truncated
# Both legs down and the lander at rest: gymnasium pays +100 on a
# successful landing, so the final raw reward is the honest verdict.
if terminated and raw_reward >= 100:
landed += 1
environment.close()
return shaped_total / n, landed / n
n_eval = env.cfg["evalEpisodes"]
results = {}
for name, net in agents.items():
# Every agent is scored on the HACKED reward as well as its own, so the
# columns are comparable — otherwise each agent is graded on its own exam.
own_shaped, landed = evaluate(net, REWARDS[name], env.cfg["seed"], n_eval)
hacked_shaped, _ = evaluate(net, reward_hacked, env.cfg["seed"], n_eval)
results[name] = {"own": own_shaped, "hacked_scale": hacked_shaped, "landed": landed}
header = f"{'agent':10}{'own reward':>14}{'hacked reward':>16}{'landed':>10}"
print("\n" + header)
print("-" * len(header))
for name, r in results.items():
print(f"{name:10}{r['own']:>14.1f}{r['hacked_scale']:>16.1f}{r['landed']:>9.0%}")
# The card's thumbnail, and the finding in one picture: the bar that wins the
# written objective is the bar that never lands. A table says it; a chart makes
# it impossible to miss at index size.
import matplotlib.pyplot as plt
fig, ax = plt.subplots(figsize=(6, 3))
names = list(results)
ax.bar(names, [results[n]["landed"] * 100 for n in names], color=["#2a9d8f", "#457b9d", "#e76f51"])
for i, n in enumerate(names):
ax.text(
i, results[n]["landed"] * 100 + 2, f"{results[n]['landed']:.0%}", ha="center", fontsize=9
)
ax.set_ylabel("landed (%)" if env.lang == "en" else ar("نسبة الهبوط (%)"))
ax.set_ylim(0, 105)
ax.spines[["top", "right"]].set_visible(False)
fig.tight_layout()
plt.show()
honest_landing = results["honest"]["landed"]
hacked_landing = results["hacked"]["landed"]
# A DIFFERENCE, not a ratio. Returns here are negative — the shaping term is
# -distance summed over the episode — and a ratio over signed quantities is
# meaningless: -742.9 / 1372.6 came out as -0.54 and read as "the exploit lost"
# when it had in fact won by 630 points. Ratios need a positive denominator
# and a meaningful zero; neither holds for a return.
shaped_advantage = results["hacked"]["hacked_scale"] - results["honest"]["hacked_scale"]
# Landing rates ARE ratios of counts: bounded, non-negative, zero means zero.
landing_ratio = hacked_landing / max(honest_landing, 1e-9)agent own reward hacked reward landed
--------------------------------------------------
honest -34.2 -1654.1 24%
shaped -82.5 -5592.5 0%
hacked -808.3 -808.3 0%
One trajectory each, compared. LENGTH is the tell — read it before you read anything else.
# One trajectory from the hacked agent, summarised. The scoreboard says it
# scores well and lands rarely; this says what it is doing instead.
environment = gym.make("LunarLander-v3")
obs, _ = environment.reset(seed=99)
altitudes, steps, done = [], 0, False
while not done and steps < 1000:
with torch.no_grad():
logits, _ = agents["hacked"](torch.as_tensor(obs, dtype=torch.float32))
obs, _, terminated, truncated, _ = environment.step(int(torch.argmax(logits)))
altitudes.append(float(obs[1]))
steps += 1
done = terminated or truncated
environment.close()
# Compare against an honest episode, and let the NUMBERS say what happened.
# The first version of this cell asserted hovering — "income every step" —
# and was exactly wrong: `potential` is -distance, so it is always negative,
# the bonus is a per-step TAX, and the fastest way to stop paying it is to end
# the episode. The agent did not learn to loiter. It learned to die.
environment = gym.make("LunarLander-v3")
obs, _ = environment.reset(seed=99)
honest_steps, done = 0, False
while not done and honest_steps < 1000:
with torch.no_grad():
logits, _ = agents["honest"](torch.as_tensor(obs, dtype=torch.float32))
obs, _, terminated, truncated, _ = environment.step(int(torch.argmax(logits)))
honest_steps += 1
done = terminated or truncated
environment.close()
if env.lang == "ar":
print(f"المخترِق: {steps} خطوة · ارتفاع وسيط {np.median(altitudes):.2f}")
print(f"الأمين: {honest_steps} خطوة")
else:
print(f"hacked: {steps} steps · median altitude {np.median(altitudes):.2f}")
print(f"honest: {honest_steps} steps")
if steps < honest_steps * 0.7:
verdict_en = (
"The exploit is a SHORT episode. `potential` is negative everywhere, so the "
"bonus is a tax charged every step, and the cheapest policy is to stop "
"paying it — end the episode. The agent is not confused; it found the fastest "
"exit from a reward you wrote."
)
verdict_ar = (
"الثغرة حلقة قصيرة. دالة الجهد سالبة في كل مكان، فالمكافأة ضريبة تُجبى كل "
"خطوة، وأرخص سياسة أن تكفّ عن دفعها — أي أن تنهي الحلقة. الوكيل ليس مرتبكاً؛ "
"بل وجد أسرع مخرج من مكافأة كتبتَها أنت."
)
elif steps > honest_steps * 1.3:
verdict_en = (
"The exploit is a LONG episode: it loiters where the bonus is largest and "
"never risks the landing. Same mechanism, opposite sign."
)
verdict_ar = (
"الثغرة حلقة طويلة: يتسكّع حيث المكافأة أكبر ولا يخاطر بالهبوط قط. الآلية "
"ذاتها بإشارة معاكسة."
)
else:
verdict_en = "Episode lengths are similar — look at where it spends its altitude instead."
verdict_ar = "أطوال الحلقات متقاربة — انظر إلى أين ينفق ارتفاعه بدلاً من ذلك."
print("\n" + (verdict_ar if env.lang == "ar" else verdict_en))hacked: 103 steps · median altitude 0.41
honest: 148 steps
The exploit is a SHORT episode. `potential` is negative everywhere, so the bonus is a tax charged every step, and the cheapest policy is to stop paying it — end the episode. The agent is not confused; it found the fastest exit from a reward you wrote.Exercise
Repair the mis-specified reward without deleting it. is the known-safe form: reward γ·Φ(s′) − Φ(s) rather than Φ(s) itself. Convert the proximity bonus into that form and retrain.
Then answer the question the fix raises. If potential-based shaping provably cannot change the , what did you actually buy by adding it? Compare not just the final landing rate but how many episodes each agent needed to get there.
# YOUR TURN.
#
# Convert the proximity bonus to potential-based form and retrain. The safe
# shaping is already written above as `reward_shaped` — the exercise is to
# understand WHY it is safe, then answer what it actually bought you.
#
# Compare against the numbers above, and look at episodes-to-competence, not
# just the final landing rate.
YOUR_REWARD = reward_shaped # try your own
fixed = train(YOUR_REWARD, env.cfg["seed"])
fixed_shaped, fixed_landed = evaluate(fixed, YOUR_REWARD, env.cfg["seed"], n_eval)
print(
f"yours: landed {fixed_landed:.0%} · honest {honest_landing:.0%} · hacked {hacked_landing:.0%}"
)A hint is available in the notebook — env.hint(1)
Both checks. The second is a MAXIMUM: the exploit must land less, or nothing was demonstrated.
# The control first. A comparison against an agent that never learned the task
# is not a comparison, and 0/0 quietly satisfies "the exploit lands less".
control_ok = env.check("honest-agent-works", honest_landing)
if not control_ok:
if env.lang == "ar":
print(" الشاهد لم يتعلّم الهبوط — ارفع episodes، ولا تخفض العتبة.")
else:
print(" the control never learned to land — raise `episodes`, do not lower the bar.")
hacked_ok = env.check("hacked-scores-higher", shaped_advantage)
landing_ok = env.check("hacked-lands-less", landing_ratio)✗ Landing rate of the control agent, before any comparison is meaningful: 0.24 (needs ≥ 0.3)
the control never learned to land — raise `episodes`, do not lower the bar.
✓ How far the mis-specified agent out-scores the honest one, on the written objective: 845.8 (needs ≥ 50)
✓ Mis-specified agent's landing rate, relative to the honest agent's: 0 (needs ≤ 0.5)receipt = env.receipt()Not complete yet — failed: honest-agent-worksTerms in this workshop
- Reward Shapingصياغة وهندسة دالة المكافأة
- Reward Hackingاختراق المكافأة
- Policyالسياسة