الوكلاء والتحكمintermediateمعالج رسوميات اختياري~45 دقيقةColab
تشكيل المكافأة — وكيف ينحرف
Reward shaping, and how it goes wrong
الوكيل يُحسِّن ما كتبتَه، لا ما قصدتَه
التعلّم من صعب. الذي لا يُكافأ إلا عند النجاح يقضي معظم التدريب دون أن تصله أي إشارة. فتقرّر أن تساعده: تضيف حدًّا يكافئ التقدّم، كالمسافة التي قطعها نحو الهدف، أو الوقود الذي وفّره، أو الارتفاع الذي حافظ عليه. ويتسارع التعلّم فورًا. بهذا الحدّ الإضافي تكون قد مارستَ ، وهو أكثر التدخّلات شيوعًا في تطبيقات .
وهنا بالضبط تبدأ الإخفاقات. الوكيل لا يقرأ نيّتك، بل يقرأ مجموع المكافآت كما كتبتَه. فإن كان هناك سلوك لم يخطر لك يحصد درجة أعلى من السلوك الذي أردتَه، فسيعثر عليه. وسيبدو ناجحًا بمقياسك أنت، بينما يفعل عكس المطلوب تمامًا.
الهدف
ندرّب الوكيل نفسه ثلاث مرّات، ولا يتغيّر بينها إلا <Term en="Reward Function">دالة المكافأة</Term>: مرّة على مكافأة البيئة كما هي، ومرّة مع <Term en="Potential-Based Shaping">تشكيل صحيح قائم على الجهد</Term>، ومرّة مع صياغة ثالثة تبدو معقولة لكنها خاطئة. ثم نُثبت أن الوكيل الثالث يحصد أعلى درجة وفق المكافأة المُشكَّلة، وأسوأ نتيجة فيما أردناه فعلًا.
يفتح Colab نسخة للقراءة فقط. احفظ نسخة في Drive للاحتفاظ بتعديلاتك.
يحتاج الدفتر إلى لوحة مفاتيح — يُفضَّل فتحه على حاسوب مكتبي.
الأوراق وراء هذه الورشة
الثغرة التي ستراها هنا ليست ما يتوقّعه معظم الناس. قبل أن تشغّل أي شيء، اسأل نفسك: ماذا يفعل الوكيل حين تكون المكافأة التي أضفتَها سالبة في كل نقطة من فضاء الحالات؟
LunarLander مثال واضح يسهل أن نكسره عمدًا. مكافأتها الأصلية توازن بين عدّة عوامل: المسافة إلى منصّة الهبوط، والوقود المستهلَك، ومكافأة كبيرة عند الهبوط، وعقوبة كبيرة عند التحطّم. والوكيل الذي يتدرّب عليها يتعلّم الهبوط فعلًا.
سنُبقي هذا الوكيل شاهدًا نقارن به، ثم نضيف إلى مكافأته حدَّين من التشكيل. أحدهما آمن ببرهان رياضي. والآخر من النوع الذي يكتبه أي مهندس كفء في يوم عمل عادي.
تهيئة أوّلية. لا حاجة لمعالج رسوميّات ولا لتنزيل شيء — البيئة تُولَّد محليًّا.
import arabic_reshaper
import azimuth_nb as azimuth
from bidi.algorithm import get_display
env = azimuth.setup(SLUG, lang=LANG, profile=PROFILE)تشكيل المكافأة وكيف ينحرف
بدون معالج رسوميات · ذاكرة 12.7 غ.ب · PyTorch 2.11.0+cpu
الملف: free
جاهز · episodes=1500, evalEpisodes=50, hidden=128, learningRate=0.0007, gamma=0.99, seed=17
الشيفرة · 359973e6e1900567اقرأ الدوال الثلاث قبل أن تُشغّل أي خليّة. وقرّر الآن: أيّها سيُنتج وكيلًا يهبط أكثر من غيره؟ جزء من الفكرة أن يأتي توقّعك خاطئًا.
# THE THREE REWARDS. Read them before running anything.
#
# LunarLander's observation is
# [x, y, vx, vy, angle, angular_velocity, leg_left_contact, leg_right_contact]
# so `-(x**2 + y**2) ** 0.5` is distance to the pad, which sits at the origin.
GAMMA = env.cfg["gamma"]
def potential(obs):
"""Φ(s): closer to the pad is higher. Used by BOTH shaped rewards, so the
difference between them is purely HOW it is applied, not what it measures."""
return -((obs[0] ** 2 + obs[1] ** 2) ** 0.5)
def reward_honest(reward, obs, next_obs, done):
"""The environment's own reward, untouched. The control."""
return reward
def reward_shaped(reward, obs, next_obs, done):
"""POTENTIAL-BASED shaping (Ng, Harada & Russell, 1999).
Rewards the CHANGE in potential, γ·Φ(s′) − Φ(s). Because the added terms
telescope over an episode, they cannot change which policy is optimal —
they only make the gradient less sparse. This is the safe form, and it is
the only shaping with a proof attached.
"""
return reward + 10.0 * (GAMMA * potential(next_obs) - potential(obs))
def reward_hacked(reward, obs, next_obs, done):
"""The mis-specification. Rewards the STATE, not the change.
This is not a strawman. "Give it points for being close to the target" is
the single most natural thing to write, it reads as obviously helpful, and
it is wrong for a reason that is invisible until you watch the agent: a
bonus paid every step for BEING somewhere is a bonus for STAYING there.
Landing ends the episode, and ending the episode ends the income.
"""
return reward + 10.0 * potential(next_obs)
REWARDS = {
"honest": reward_honest,
"shaped": reward_shaped,
"hacked": reward_hacked,
}
if env.lang == "ar":
print("ثلاث دوال مكافأة:", " · ".join(REWARDS))
print("أيّها تظنّه سيهبط أكثر؟ قرّر قبل التشغيل.")
else:
print("three reward functions:", " · ".join(REWARDS))
print("which do you think lands most often? decide before you run.")ثلاث دوال مكافأة: honest · shaped · hacked
أيّها تظنّه سيهبط أكثر؟ قرّر قبل التشغيل.ثلاثة وكلاء، معمارية واحدة، بذرة عشوائية واحدة. ما يختلف هو دالة المكافأة فقط.
import gymnasium as gym
import numpy as np
import torch
import torch.nn as nn
class Policy(nn.Module):
"""A small categorical policy with a value baseline."""
def __init__(self, n_obs, n_actions, hidden):
super().__init__()
self.body = nn.Sequential(
nn.Linear(n_obs, hidden), nn.Tanh(), nn.Linear(hidden, hidden), nn.Tanh()
)
self.actor = nn.Linear(hidden, n_actions)
self.critic = nn.Linear(hidden, 1)
def forward(self, obs):
h = self.body(obs)
return self.actor(h), self.critic(h).squeeze(-1)
def train(reward_fn, seed):
"""One agent, one reward function. Same seed for all three, so the only
thing that differs between the runs is the function itself."""
torch.manual_seed(seed)
environment = gym.make("LunarLander-v3")
net = Policy(
environment.observation_space.shape[0], environment.action_space.n, env.cfg["hidden"]
)
optimizer = torch.optim.Adam(net.parameters(), lr=env.cfg["learningRate"])
for episode in range(env.cfg["episodes"]):
obs, _ = environment.reset(seed=seed + episode)
log_probs, values, rewards = [], [], []
done = False
while not done:
obs_t = torch.as_tensor(obs, dtype=torch.float32)
logits, value = net(obs_t)
dist = torch.distributions.Categorical(logits=logits)
action = dist.sample()
next_obs, raw_reward, terminated, truncated, _ = environment.step(action.item())
done = terminated or truncated
log_probs.append(dist.log_prob(action))
values.append(value)
# THE ONLY LINE THAT DIFFERS between the three agents.
rewards.append(reward_fn(raw_reward, obs, next_obs, done))
obs = next_obs
returns, running = [], 0.0
for r in reversed(rewards):
running = r + GAMMA * running
returns.append(running)
returns = torch.tensor(list(reversed(returns)), dtype=torch.float32)
returns = (returns - returns.mean()) / (returns.std() + 1e-8)
values_t = torch.stack(values)
advantage = returns - values_t.detach()
loss = -(torch.stack(log_probs) * advantage).sum() + nn.functional.mse_loss(
values_t, returns, reduction="sum"
)
optimizer.zero_grad()
loss.backward()
optimizer.step()
if episode % 100 == 0:
print(f" {episode:5d} return {sum(rewards):8.1f}")
environment.close()
return net
agents = {}
for name, fn in REWARDS.items():
print(f"\n{name}:")
agents[name] = train(fn, env.cfg["seed"])
episodes_run = env.cfg["episodes"] * len(REWARDS)honest:
0 return -264.6
100 return -132.8
200 return -102.0
300 return -65.5
400 return -7.1
500 return -22.2
600 return -99.2
700 return -83.4
800 return -13.2
900 return 50.1
1000 return 7.8
1100 return -15.0
1200 return 154.8
1300 return 177.8
1400 return 24.7
shaped:
0 return -248.9
100 return -159.0
200 return -98.8
300 return -63.4
400 return -0.6
500 return 25.3
600 return -37.4
700 return -334.1
800 return 3.4
900 return 132.5
1000 return 177.5
1100 return -10.7
1200 return 174.7
1300 return 134.4
1400 return 152.7
hacked:
0 return -1887.8
100 return -1665.9
200 return -1426.3
300 return -1058.9
400 return -1399.0
500 return -683.6
600 return -582.1
700 return -858.4
800 return -679.8
900 return -591.8
1000 return -840.4
1100 return -1082.3
1200 return -612.8
1300 return -904.3
1400 return -635.5عمودان يحكيان قصّتين مختلفتين. عمود الدرجة المُشكَّلة هو كل ما رآه الوكيل أثناء التدريب، وعمود الهبوط هو ما أردتَه أنت. اقرأ كلًّا منهما في ضوء الآخر.
def ar(text: str) -> str:
"""Join Arabic letters into their connected forms, then reorder for LTR drawing."""
return get_display(arabic_reshaper.reshape(text))
def evaluate(net, reward_fn, seed, n):
"""Two numbers per agent, deliberately.
`shaped` is what the training loop optimised. `landed` is the job. Keeping
them apart is the entire point: an agent can move one without the other,
and reporting a single 'score' would hide exactly the effect being taught.
"""
environment = gym.make("LunarLander-v3")
shaped_total, landed = 0.0, 0
for i in range(n):
obs, _ = environment.reset(seed=10_000 + seed + i)
done = False
while not done:
with torch.no_grad():
logits, _ = net(torch.as_tensor(obs, dtype=torch.float32))
action = int(torch.argmax(logits))
next_obs, raw_reward, terminated, truncated, _ = environment.step(action)
shaped_total += reward_fn(raw_reward, obs, next_obs, terminated or truncated)
obs = next_obs
done = terminated or truncated
# Both legs down and the lander at rest: gymnasium pays +100 on a
# successful landing, so the final raw reward is the honest verdict.
if terminated and raw_reward >= 100:
landed += 1
environment.close()
return shaped_total / n, landed / n
n_eval = env.cfg["evalEpisodes"]
results = {}
for name, net in agents.items():
# Every agent is scored on the HACKED reward as well as its own, so the
# columns are comparable — otherwise each agent is graded on its own exam.
own_shaped, landed = evaluate(net, REWARDS[name], env.cfg["seed"], n_eval)
hacked_shaped, _ = evaluate(net, reward_hacked, env.cfg["seed"], n_eval)
results[name] = {"own": own_shaped, "hacked_scale": hacked_shaped, "landed": landed}
header = f"{'agent':10}{'own reward':>14}{'hacked reward':>16}{'landed':>10}"
print("\n" + header)
print("-" * len(header))
for name, r in results.items():
print(f"{name:10}{r['own']:>14.1f}{r['hacked_scale']:>16.1f}{r['landed']:>9.0%}")
# The card's thumbnail, and the finding in one picture: the bar that wins the
# written objective is the bar that never lands. A table says it; a chart makes
# it impossible to miss at index size.
import matplotlib.pyplot as plt
fig, ax = plt.subplots(figsize=(6, 3))
names = list(results)
ax.bar(names, [results[n]["landed"] * 100 for n in names], color=["#2a9d8f", "#457b9d", "#e76f51"])
for i, n in enumerate(names):
ax.text(
i, results[n]["landed"] * 100 + 2, f"{results[n]['landed']:.0%}", ha="center", fontsize=9
)
ax.set_ylabel("landed (%)" if env.lang == "en" else ar("نسبة الهبوط (%)"))
ax.set_ylim(0, 105)
ax.spines[["top", "right"]].set_visible(False)
fig.tight_layout()
plt.show()
honest_landing = results["honest"]["landed"]
hacked_landing = results["hacked"]["landed"]
# A DIFFERENCE, not a ratio. Returns here are negative — the shaping term is
# -distance summed over the episode — and a ratio over signed quantities is
# meaningless: -742.9 / 1372.6 came out as -0.54 and read as "the exploit lost"
# when it had in fact won by 630 points. Ratios need a positive denominator
# and a meaningful zero; neither holds for a return.
shaped_advantage = results["hacked"]["hacked_scale"] - results["honest"]["hacked_scale"]
# Landing rates ARE ratios of counts: bounded, non-negative, zero means zero.
landing_ratio = hacked_landing / max(honest_landing, 1e-9)agent own reward hacked reward landed
--------------------------------------------------
honest 128.0 -2027.5 58%
shaped 158.3 -3428.4 70%
hacked -753.6 -753.6 0%
مسار واحد لكلٍّ من الوكيلين المُخترِق والأمين، جنبًا إلى جنب. طول الحلقة هو المفتاح، فانظر إليه قبل أي رقم آخر.
# One trajectory from the hacked agent, summarised. The scoreboard says it
# scores well and lands rarely; this says what it is doing instead.
environment = gym.make("LunarLander-v3")
obs, _ = environment.reset(seed=99)
altitudes, steps, done = [], 0, False
while not done and steps < 1000:
with torch.no_grad():
logits, _ = agents["hacked"](torch.as_tensor(obs, dtype=torch.float32))
obs, _, terminated, truncated, _ = environment.step(int(torch.argmax(logits)))
altitudes.append(float(obs[1]))
steps += 1
done = terminated or truncated
environment.close()
# Compare against an honest episode, and let the NUMBERS say what happened.
# The first version of this cell asserted hovering — "income every step" —
# and was exactly wrong: `potential` is -distance, so it is always negative,
# the bonus is a per-step TAX, and the fastest way to stop paying it is to end
# the episode. The agent did not learn to loiter. It learned to die.
environment = gym.make("LunarLander-v3")
obs, _ = environment.reset(seed=99)
honest_steps, done = 0, False
while not done and honest_steps < 1000:
with torch.no_grad():
logits, _ = agents["honest"](torch.as_tensor(obs, dtype=torch.float32))
obs, _, terminated, truncated, _ = environment.step(int(torch.argmax(logits)))
honest_steps += 1
done = terminated or truncated
environment.close()
if env.lang == "ar":
print(f"المخترِق: {steps} خطوة · ارتفاع وسيط {np.median(altitudes):.2f}")
print(f"الأمين: {honest_steps} خطوة")
else:
print(f"hacked: {steps} steps · median altitude {np.median(altitudes):.2f}")
print(f"honest: {honest_steps} steps")
if steps < honest_steps * 0.7:
verdict_en = (
"The exploit is a SHORT episode. `potential` is negative everywhere, so the "
"bonus is a tax charged every step, and the cheapest policy is to stop "
"paying it — end the episode. The agent is not confused; it found the fastest "
"exit from a reward you wrote."
)
verdict_ar = (
"الثغرة حلقة قصيرة. دالة الجهد سالبة في كل مكان، فالمكافأة ضريبة تُجبى كل "
"خطوة، وأرخص سياسة أن تكفّ عن دفعها — أي أن تنهي الحلقة. الوكيل ليس مرتبكاً؛ "
"بل وجد أسرع مخرج من مكافأة كتبتَها أنت."
)
elif steps > honest_steps * 1.3:
verdict_en = (
"The exploit is a LONG episode: it loiters where the bonus is largest and "
"never risks the landing. Same mechanism, opposite sign."
)
verdict_ar = (
"الثغرة حلقة طويلة: يتسكّع حيث المكافأة أكبر ولا يخاطر بالهبوط قط. الآلية "
"ذاتها بإشارة معاكسة."
)
else:
verdict_en = "Episode lengths are similar — look at where it spends its altitude instead."
verdict_ar = "أطوال الحلقات متقاربة — انظر إلى أين ينفق ارتفاعه بدلاً من ذلك."
print("\n" + (verdict_ar if env.lang == "ar" else verdict_en))المخترِق: 64 خطوة · ارتفاع وسيط 0.98
الأمين: 1000 خطوة
الثغرة حلقة قصيرة. دالة الجهد سالبة في كل مكان، فالمكافأة ضريبة تُجبى كل خطوة، وأرخص سياسة أن تكفّ عن دفعها — أي أن تنهي الحلقة. الوكيل ليس مرتبكاً؛ بل وجد أسرع مخرج من مكافأة كتبتَها أنت.تمرين
أصلِح المكافأة الخاطئة من غير أن تحذفها. الصيغة المعروفة بأمانها هي التشكيل القائم على الجهد. الفكرة أن تكافئ الوكيل على تغيُّر الجهد من خطوة إلى التي تليها، لا على قيمة الجهد حيث يقف. بالرموز: كافئ γ·Φ(s′) − Φ(s) بدلًا من Φ(s) نفسها، حيث γ هو . حوِّل مكافأة القرب إلى هذه الصيغة وأعِد التدريب.
ثم فكّر في السؤال الذي يطرحه هذا الإصلاح. إن كان التشكيل القائم على الجهد لا يستطيع، بالبرهان، أن يغيّر ، فماذا كسبتَ فعلًا بإضافته؟ لا تقارن نسبة الهبوط النهائية فحسب، بل قارن أيضًا عدد الحلقات التي احتاجها كل وكيل ليصل إليها.
# YOUR TURN.
#
# Convert the proximity bonus to potential-based form and retrain. The safe
# shaping is already written above as `reward_shaped` — the exercise is to
# understand WHY it is safe, then answer what it actually bought you.
#
# Compare against the numbers above, and look at episodes-to-competence, not
# just the final landing rate.
YOUR_REWARD = reward_shaped # try your own
fixed = train(YOUR_REWARD, env.cfg["seed"])
fixed_shaped, fixed_landed = evaluate(fixed, YOUR_REWARD, env.cfg["seed"], n_eval)
print(
f"yours: landed {fixed_landed:.0%} · honest {honest_landing:.0%} · hacked {hacked_landing:.0%}"
)يتوفّر تلميح في الدفتر — env.hint(1)
الفحصان معًا. الثاني حدّ أقصى: يجب أن يهبط الوكيل المُخترِق أقلّ من الأمين، وإلّا فلم نُثبت شيئًا.
# The control first. A comparison against an agent that never learned the task
# is not a comparison, and 0/0 quietly satisfies "the exploit lands less".
control_ok = env.check("honest-agent-works", honest_landing)
if not control_ok:
if env.lang == "ar":
print(" الشاهد لم يتعلّم الهبوط — ارفع episodes، ولا تخفض العتبة.")
else:
print(" the control never learned to land — raise `episodes`, do not lower the bar.")
hacked_ok = env.check("hacked-scores-higher", shaped_advantage)
landing_ok = env.check("hacked-lands-less", landing_ratio)✓ معدّل هبوط الوكيل الشاهد، قبل أن تكون أي مقارنة ذات معنى: 0.58 (المطلوب ≥ 0.3)
✓ بكم يتفوّق الوكيل سيئ التحديد على الأمين في الهدف كما كُتب: 1274 (المطلوب ≥ 50)
✓ معدّل هبوط الوكيل سيئ التحديد منسوباً إلى الوكيل الأمين: 0 (المطلوب ≤ 0.5)receipt = env.receipt()اكتملت الورشة.
رمز الإتمام: AZ-██████████
الصقه في صفحة الورشة على أزيموث لتسجيل إتمامها.مصطلحات هذه الورشة
- صياغة وهندسة دالة المكافأةReward Shaping
- اختراق المكافأةReward Hacking
- السياسةPolicy