الوكلاء والتحكمintermediateمعالج رسوميات اختياري~45 دقيقةColab

تشكيل المكافأة — وكيف ينحرف

Reward shaping, and how it goes wrong

الوكيل يُحسِّن ما كتبتَه، لا ما قصدتَه

التعلّم من صعب. الذي لا يُكافأ إلا عند النجاح يقضي معظم التدريب دون أن تصله أي إشارة. فتقرّر أن تساعده: تضيف حدًّا يكافئ التقدّم، كالمسافة التي قطعها نحو الهدف، أو الوقود الذي وفّره، أو الارتفاع الذي حافظ عليه. ويتسارع التعلّم فورًا. بهذا الحدّ الإضافي تكون قد مارستَ ، وهو أكثر التدخّلات شيوعًا في تطبيقات .

وهنا بالضبط تبدأ الإخفاقات. الوكيل لا يقرأ نيّتك، بل يقرأ مجموع المكافآت كما كتبتَه. فإن كان هناك سلوك لم يخطر لك يحصد درجة أعلى من السلوك الذي أردتَه، فسيعثر عليه. وسيبدو ناجحًا بمقياسك أنت، بينما يفعل عكس المطلوب تمامًا.

الهدف

ندرّب الوكيل نفسه ثلاث مرّات، ولا يتغيّر بينها إلا <Term en="Reward Function">دالة المكافأة</Term>: مرّة على مكافأة البيئة كما هي، ومرّة مع <Term en="Potential-Based Shaping">تشكيل صحيح قائم على الجهد</Term>، ومرّة مع صياغة ثالثة تبدو معقولة لكنها خاطئة. ثم نُثبت أن الوكيل الثالث يحصد أعلى درجة وفق المكافأة المُشكَّلة، وأسوأ نتيجة فيما أردناه فعلًا.

يفتح Colab نسخة للقراءة فقط. احفظ نسخة في Drive للاحتفاظ بتعديلاتك.

يحتاج الدفتر إلى لوحة مفاتيح — يُفضَّل فتحه على حاسوب مكتبي.

الأوراق وراء هذه الورشة

الثغرة التي ستراها هنا ليست ما يتوقّعه معظم الناس. قبل أن تشغّل أي شيء، اسأل نفسك: ماذا يفعل الوكيل حين تكون المكافأة التي أضفتَها سالبة في كل نقطة من فضاء الحالات؟

LunarLander مثال واضح يسهل أن نكسره عمدًا. مكافأتها الأصلية توازن بين عدّة عوامل: المسافة إلى منصّة الهبوط، والوقود المستهلَك، ومكافأة كبيرة عند الهبوط، وعقوبة كبيرة عند التحطّم. والوكيل الذي يتدرّب عليها يتعلّم الهبوط فعلًا.

سنُبقي هذا الوكيل شاهدًا نقارن به، ثم نضيف إلى مكافأته حدَّين من التشكيل. أحدهما آمن ببرهان رياضي. والآخر من النوع الذي يكتبه أي مهندس كفء في يوم عمل عادي.

تهيئة أوّلية. لا حاجة لمعالج رسوميّات ولا لتنزيل شيء — البيئة تُولَّد محليًّا.

import arabic_reshaper
import azimuth_nb as azimuth
from bidi.algorithm import get_display

env = azimuth.setup(SLUG, lang=LANG, profile=PROFILE)
شيفرة الورشة
تشكيل المكافأة وكيف ينحرف
بدون معالج رسوميات · ذاكرة 12.7 غ.ب · PyTorch 2.11.0+cpu
الملف: ⁦free⁩
جاهز · ⁦episodes=1500, evalEpisodes=50, hidden=128, learningRate=0.0007, gamma=0.99, seed=17⁩
الشيفرة · ⁦359973e6e1900567⁩

اقرأ الدوال الثلاث قبل أن تُشغّل أي خليّة. وقرّر الآن: أيّها سيُنتج وكيلًا يهبط أكثر من غيره؟ جزء من الفكرة أن يأتي توقّعك خاطئًا.

# THE THREE REWARDS. Read them before running anything.
#
# LunarLander's observation is
#   [x, y, vx, vy, angle, angular_velocity, leg_left_contact, leg_right_contact]
# so `-(x**2 + y**2) ** 0.5` is distance to the pad, which sits at the origin.

GAMMA = env.cfg["gamma"]


def potential(obs):
    """Φ(s): closer to the pad is higher. Used by BOTH shaped rewards, so the
    difference between them is purely HOW it is applied, not what it measures."""
    return -((obs[0] ** 2 + obs[1] ** 2) ** 0.5)


def reward_honest(reward, obs, next_obs, done):
    """The environment's own reward, untouched. The control."""
    return reward


def reward_shaped(reward, obs, next_obs, done):
    """POTENTIAL-BASED shaping (Ng, Harada & Russell, 1999).

    Rewards the CHANGE in potential, γ·Φ(s′) − Φ(s). Because the added terms
    telescope over an episode, they cannot change which policy is optimal —
    they only make the gradient less sparse. This is the safe form, and it is
    the only shaping with a proof attached.
    """
    return reward + 10.0 * (GAMMA * potential(next_obs) - potential(obs))


def reward_hacked(reward, obs, next_obs, done):
    """The mis-specification. Rewards the STATE, not the change.

    This is not a strawman. "Give it points for being close to the target" is
    the single most natural thing to write, it reads as obviously helpful, and
    it is wrong for a reason that is invisible until you watch the agent: a
    bonus paid every step for BEING somewhere is a bonus for STAYING there.
    Landing ends the episode, and ending the episode ends the income.
    """
    return reward + 10.0 * potential(next_obs)


REWARDS = {
    "honest": reward_honest,
    "shaped": reward_shaped,
    "hacked": reward_hacked,
}

if env.lang == "ar":
    print("ثلاث دوال مكافأة:", " · ".join(REWARDS))
    print("أيّها تظنّه سيهبط أكثر؟ قرّر قبل التشغيل.")
else:
    print("three reward functions:", " · ".join(REWARDS))
    print("which do you think lands most often? decide before you run.")
شيفرة الورشة
ثلاث دوال مكافأة: honest · shaped · hacked
أيّها تظنّه سيهبط أكثر؟ قرّر قبل التشغيل.

ثلاثة وكلاء، معمارية واحدة، بذرة عشوائية واحدة. ما يختلف هو دالة المكافأة فقط.

import gymnasium as gym
import numpy as np
import torch
import torch.nn as nn


class Policy(nn.Module):
    """A small categorical policy with a value baseline."""

    def __init__(self, n_obs, n_actions, hidden):
        super().__init__()
        self.body = nn.Sequential(
            nn.Linear(n_obs, hidden), nn.Tanh(), nn.Linear(hidden, hidden), nn.Tanh()
        )
        self.actor = nn.Linear(hidden, n_actions)
        self.critic = nn.Linear(hidden, 1)

    def forward(self, obs):
        h = self.body(obs)
        return self.actor(h), self.critic(h).squeeze(-1)


def train(reward_fn, seed):
    """One agent, one reward function. Same seed for all three, so the only
    thing that differs between the runs is the function itself."""
    torch.manual_seed(seed)
    environment = gym.make("LunarLander-v3")
    net = Policy(
        environment.observation_space.shape[0], environment.action_space.n, env.cfg["hidden"]
    )
    optimizer = torch.optim.Adam(net.parameters(), lr=env.cfg["learningRate"])

    for episode in range(env.cfg["episodes"]):
        obs, _ = environment.reset(seed=seed + episode)
        log_probs, values, rewards = [], [], []
        done = False
        while not done:
            obs_t = torch.as_tensor(obs, dtype=torch.float32)
            logits, value = net(obs_t)
            dist = torch.distributions.Categorical(logits=logits)
            action = dist.sample()
            next_obs, raw_reward, terminated, truncated, _ = environment.step(action.item())
            done = terminated or truncated

            log_probs.append(dist.log_prob(action))
            values.append(value)
            # THE ONLY LINE THAT DIFFERS between the three agents.
            rewards.append(reward_fn(raw_reward, obs, next_obs, done))
            obs = next_obs

        returns, running = [], 0.0
        for r in reversed(rewards):
            running = r + GAMMA * running
            returns.append(running)
        returns = torch.tensor(list(reversed(returns)), dtype=torch.float32)
        returns = (returns - returns.mean()) / (returns.std() + 1e-8)

        values_t = torch.stack(values)
        advantage = returns - values_t.detach()
        loss = -(torch.stack(log_probs) * advantage).sum() + nn.functional.mse_loss(
            values_t, returns, reduction="sum"
        )
        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

        if episode % 100 == 0:
            print(f"  {episode:5d}  return {sum(rewards):8.1f}")

    environment.close()
    return net


agents = {}
for name, fn in REWARDS.items():
    print(f"\n{name}:")
    agents[name] = train(fn, env.cfg["seed"])

episodes_run = env.cfg["episodes"] * len(REWARDS)
شيفرة الورشة
honest:
      0  return   -264.6
    100  return   -132.8
    200  return   -102.0
    300  return    -65.5
    400  return     -7.1
    500  return    -22.2
    600  return    -99.2
    700  return    -83.4
    800  return    -13.2
    900  return     50.1
   1000  return      7.8
   1100  return    -15.0
   1200  return    154.8
   1300  return    177.8
   1400  return     24.7

shaped:
      0  return   -248.9
    100  return   -159.0
    200  return    -98.8
    300  return    -63.4
    400  return     -0.6
    500  return     25.3
    600  return    -37.4
    700  return   -334.1
    800  return      3.4
    900  return    132.5
   1000  return    177.5
   1100  return    -10.7
   1200  return    174.7
   1300  return    134.4
   1400  return    152.7

hacked:
      0  return  -1887.8
    100  return  -1665.9
    200  return  -1426.3
    300  return  -1058.9
    400  return  -1399.0
    500  return   -683.6
    600  return   -582.1
    700  return   -858.4
    800  return   -679.8
    900  return   -591.8
   1000  return   -840.4
   1100  return  -1082.3
   1200  return   -612.8
   1300  return   -904.3
   1400  return   -635.5

عمودان يحكيان قصّتين مختلفتين. عمود الدرجة المُشكَّلة هو كل ما رآه الوكيل أثناء التدريب، وعمود الهبوط هو ما أردتَه أنت. اقرأ كلًّا منهما في ضوء الآخر.

def ar(text: str) -> str:
    """Join Arabic letters into their connected forms, then reorder for LTR drawing."""
    return get_display(arabic_reshaper.reshape(text))


def evaluate(net, reward_fn, seed, n):
    """Two numbers per agent, deliberately.

    `shaped` is what the training loop optimised. `landed` is the job. Keeping
    them apart is the entire point: an agent can move one without the other,
    and reporting a single 'score' would hide exactly the effect being taught.
    """
    environment = gym.make("LunarLander-v3")
    shaped_total, landed = 0.0, 0
    for i in range(n):
        obs, _ = environment.reset(seed=10_000 + seed + i)
        done = False
        while not done:
            with torch.no_grad():
                logits, _ = net(torch.as_tensor(obs, dtype=torch.float32))
            action = int(torch.argmax(logits))
            next_obs, raw_reward, terminated, truncated, _ = environment.step(action)
            shaped_total += reward_fn(raw_reward, obs, next_obs, terminated or truncated)
            obs = next_obs
            done = terminated or truncated
        # Both legs down and the lander at rest: gymnasium pays +100 on a
        # successful landing, so the final raw reward is the honest verdict.
        if terminated and raw_reward >= 100:
            landed += 1
    environment.close()
    return shaped_total / n, landed / n


n_eval = env.cfg["evalEpisodes"]
results = {}
for name, net in agents.items():
    # Every agent is scored on the HACKED reward as well as its own, so the
    # columns are comparable — otherwise each agent is graded on its own exam.
    own_shaped, landed = evaluate(net, REWARDS[name], env.cfg["seed"], n_eval)
    hacked_shaped, _ = evaluate(net, reward_hacked, env.cfg["seed"], n_eval)
    results[name] = {"own": own_shaped, "hacked_scale": hacked_shaped, "landed": landed}

header = f"{'agent':10}{'own reward':>14}{'hacked reward':>16}{'landed':>10}"
print("\n" + header)
print("-" * len(header))
for name, r in results.items():
    print(f"{name:10}{r['own']:>14.1f}{r['hacked_scale']:>16.1f}{r['landed']:>9.0%}")

# The card's thumbnail, and the finding in one picture: the bar that wins the
# written objective is the bar that never lands. A table says it; a chart makes
# it impossible to miss at index size.
import matplotlib.pyplot as plt

fig, ax = plt.subplots(figsize=(6, 3))
names = list(results)
ax.bar(names, [results[n]["landed"] * 100 for n in names], color=["#2a9d8f", "#457b9d", "#e76f51"])
for i, n in enumerate(names):
    ax.text(
        i, results[n]["landed"] * 100 + 2, f"{results[n]['landed']:.0%}", ha="center", fontsize=9
    )
ax.set_ylabel("landed (%)" if env.lang == "en" else ar("نسبة الهبوط (%)"))
ax.set_ylim(0, 105)
ax.spines[["top", "right"]].set_visible(False)
fig.tight_layout()
plt.show()

honest_landing = results["honest"]["landed"]
hacked_landing = results["hacked"]["landed"]
# A DIFFERENCE, not a ratio. Returns here are negative — the shaping term is
# -distance summed over the episode — and a ratio over signed quantities is
# meaningless: -742.9 / 1372.6 came out as -0.54 and read as "the exploit lost"
# when it had in fact won by 630 points. Ratios need a positive denominator
# and a meaningful zero; neither holds for a return.
shaped_advantage = results["hacked"]["hacked_scale"] - results["honest"]["hacked_scale"]
# Landing rates ARE ratios of counts: bounded, non-negative, zero means zero.
landing_ratio = hacked_landing / max(honest_landing, 1e-9)
شيفرة الورشة
agent         own reward   hacked reward    landed
--------------------------------------------------
honest             128.0         -2027.5      58%
shaped             158.3         -3428.4      70%
hacked            -753.6          -753.6       0%

مسار واحد لكلٍّ من الوكيلين المُخترِق والأمين، جنبًا إلى جنب. طول الحلقة هو المفتاح، فانظر إليه قبل أي رقم آخر.

# One trajectory from the hacked agent, summarised. The scoreboard says it
# scores well and lands rarely; this says what it is doing instead.
environment = gym.make("LunarLander-v3")
obs, _ = environment.reset(seed=99)
altitudes, steps, done = [], 0, False
while not done and steps < 1000:
    with torch.no_grad():
        logits, _ = agents["hacked"](torch.as_tensor(obs, dtype=torch.float32))
    obs, _, terminated, truncated, _ = environment.step(int(torch.argmax(logits)))
    altitudes.append(float(obs[1]))
    steps += 1
    done = terminated or truncated
environment.close()

# Compare against an honest episode, and let the NUMBERS say what happened.
# The first version of this cell asserted hovering — "income every step" —
# and was exactly wrong: `potential` is -distance, so it is always negative,
# the bonus is a per-step TAX, and the fastest way to stop paying it is to end
# the episode. The agent did not learn to loiter. It learned to die.
environment = gym.make("LunarLander-v3")
obs, _ = environment.reset(seed=99)
honest_steps, done = 0, False
while not done and honest_steps < 1000:
    with torch.no_grad():
        logits, _ = agents["honest"](torch.as_tensor(obs, dtype=torch.float32))
    obs, _, terminated, truncated, _ = environment.step(int(torch.argmax(logits)))
    honest_steps += 1
    done = terminated or truncated
environment.close()

if env.lang == "ar":
    print(f"المخترِق: {steps} خطوة · ارتفاع وسيط {np.median(altitudes):.2f}")
    print(f"الأمين:  {honest_steps} خطوة")
else:
    print(f"hacked: {steps} steps · median altitude {np.median(altitudes):.2f}")
    print(f"honest: {honest_steps} steps")

if steps < honest_steps * 0.7:
    verdict_en = (
        "The exploit is a SHORT episode. `potential` is negative everywhere, so the "
        "bonus is a tax charged every step, and the cheapest policy is to stop "
        "paying it — end the episode. The agent is not confused; it found the fastest "
        "exit from a reward you wrote."
    )
    verdict_ar = (
        "الثغرة حلقة قصيرة. دالة الجهد سالبة في كل مكان، فالمكافأة ضريبة تُجبى كل "
        "خطوة، وأرخص سياسة أن تكفّ عن دفعها — أي أن تنهي الحلقة. الوكيل ليس مرتبكاً؛ "
        "بل وجد أسرع مخرج من مكافأة كتبتَها أنت."
    )
elif steps > honest_steps * 1.3:
    verdict_en = (
        "The exploit is a LONG episode: it loiters where the bonus is largest and "
        "never risks the landing. Same mechanism, opposite sign."
    )
    verdict_ar = (
        "الثغرة حلقة طويلة: يتسكّع حيث المكافأة أكبر ولا يخاطر بالهبوط قط. الآلية "
        "ذاتها بإشارة معاكسة."
    )
else:
    verdict_en = "Episode lengths are similar — look at where it spends its altitude instead."
    verdict_ar = "أطوال الحلقات متقاربة — انظر إلى أين ينفق ارتفاعه بدلاً من ذلك."

print("\n" + (verdict_ar if env.lang == "ar" else verdict_en))
شيفرة الورشة
المخترِق: 64 خطوة · ارتفاع وسيط 0.98
الأمين:  1000 خطوة

الثغرة حلقة قصيرة. دالة الجهد سالبة في كل مكان، فالمكافأة ضريبة تُجبى كل خطوة، وأرخص سياسة أن تكفّ عن دفعها — أي أن تنهي الحلقة. الوكيل ليس مرتبكاً؛ بل وجد أسرع مخرج من مكافأة كتبتَها أنت.

تمرين

أصلِح المكافأة الخاطئة من غير أن تحذفها. الصيغة المعروفة بأمانها هي التشكيل القائم على الجهد. الفكرة أن تكافئ الوكيل على تغيُّر الجهد من خطوة إلى التي تليها، لا على قيمة الجهد حيث يقف. بالرموز: كافئ γ·Φ(s′) − Φ(s) بدلًا من Φ(s) نفسها، حيث γ هو . حوِّل مكافأة القرب إلى هذه الصيغة وأعِد التدريب.

ثم فكّر في السؤال الذي يطرحه هذا الإصلاح. إن كان التشكيل القائم على الجهد لا يستطيع، بالبرهان، أن يغيّر ، فماذا كسبتَ فعلًا بإضافته؟ لا تقارن نسبة الهبوط النهائية فحسب، بل قارن أيضًا عدد الحلقات التي احتاجها كل وكيل ليصل إليها.

# YOUR TURN.
#
# Convert the proximity bonus to potential-based form and retrain. The safe
# shaping is already written above as `reward_shaped` — the exercise is to
# understand WHY it is safe, then answer what it actually bought you.
#
# Compare against the numbers above, and look at episodes-to-competence, not
# just the final landing rate.
YOUR_REWARD = reward_shaped  # try your own

fixed = train(YOUR_REWARD, env.cfg["seed"])
fixed_shaped, fixed_landed = evaluate(fixed, YOUR_REWARD, env.cfg["seed"], n_eval)
print(
    f"yours: landed {fixed_landed:.0%}  ·  honest {honest_landing:.0%}  ·  hacked {hacked_landing:.0%}"
)
شيفرة الورشة

يتوفّر تلميح في الدفتر — env.hint(1)

الفحصان معًا. الثاني حدّ أقصى: يجب أن يهبط الوكيل المُخترِق أقلّ من الأمين، وإلّا فلم نُثبت شيئًا.

# The control first. A comparison against an agent that never learned the task
# is not a comparison, and 0/0 quietly satisfies "the exploit lands less".
control_ok = env.check("honest-agent-works", honest_landing)
if not control_ok:
    if env.lang == "ar":
        print("  الشاهد لم يتعلّم الهبوط — ارفع episodes، ولا تخفض العتبة.")
    else:
        print("  the control never learned to land — raise `episodes`, do not lower the bar.")

hacked_ok = env.check("hacked-scores-higher", shaped_advantage)
landing_ok = env.check("hacked-lands-less", landing_ratio)
شيفرة الورشة
✓ معدّل هبوط الوكيل الشاهد، قبل أن تكون أي مقارنة ذات معنى: 0.58 (المطلوب ≥ 0.3)
✓ بكم يتفوّق الوكيل سيئ التحديد على الأمين في الهدف كما كُتب: 1274 (المطلوب ≥ 50)
✓ معدّل هبوط الوكيل سيئ التحديد منسوباً إلى الوكيل الأمين: 0 (المطلوب ≤ 0.5)
receipt = env.receipt()
شيفرة الورشة
اكتملت الورشة.

رمز الإتمام: ⁦AZ-██████████⁩
الصقه في صفحة الورشة على أزيموث لتسجيل إتمامها.
آخر تحقّق: 2026-08-28 · unknown · PyTorch unknown · Python 3.13.5 · e884576

مصطلحات هذه الورشة