Dual Reward Model
نموذج المكافأة المزدوج
نهج يستخدم نموذجَي مكافأة منفصلَين — واحد للمفيدية وآخر للسلامة — بدلاً من نموذج واحد يوازن بين الهدفين. يتجنّب التوتر حيث قد يُضحّي نموذج وحيد بأحد الهدفين لصالح الآخر.
An approach using two separate reward models — one for helpfulness and one for safety — instead of a single model balancing both objectives. Avoids the tension where a single model might sacrifice one objective for the other.
Also translated asنموذجا مكافأة منفصلان
First appears in this corpus in: Llama 2: Open Foundation and Fine-Tuned Chat Models (2023)
Appears in these papers