{"repo":"reward-scope-ai/reward-scope","free":true,"listed":false,"github":"https://github.com/reward-scope-ai/reward-scope","clone":"git clone https://github.com/reward-scope-ai/reward-scope.git","description":"Real-time reward debugging and hacking detection for reinforcement learning","language":"Python","stars":20,"topics":["debugging","ml-tools","reward-hacking","robotics","ai-safety","gymnasium","monitoring","observability","rlhf","stable-baselines3"],"license":"MIT","category":"analytics","readme_excerpt":"RewardScope 🔬 Your agent's reward is going up, but its behavior is broken. RewardScope shows you why. RewardScope detects reward hacking during RL training, saving you from wasting hours on a broken policy. It tracks reward components, flags exploitation patterns, and provides a live dashboard to show exactly how your agent is learning. Try It Now Demo Watch RewardScope detect reward hacking in real-time during Overcooked multi-agent training https://github.com/user-attachments/assets/6ca3ba70-ad0c-418e-8146-5c9616669215 Why This Matters Reward hacking, where agents exploit gaps between proxy rewards and true objectives, is a well-documented problem. Research shows that designing unhackable proxy rewards is nearly impossible in general settings, making detection during training essential. Left unchecked, reward hacking leads to policies that score well but behave poorly. Recent research from Anthropic found that models learning to exploit reward functions also develop alignment faking and deceptive behaviors, patterns that generalize beyond the original hacking. RewardScope takes a detection-first approach. Monitor for exploitation patterns in real-time rather than trying to craft perfect rewards. Features - 🎯 Reward Decomposition - Track individual reward components separately - 🚨 Hacking Detection - 5 detectors for common exploitation patterns - 🧠 Adaptive Baselines - Learns \"normal\" patterns per training run to reduce false positives - 📊 Live Dashboard - Real-time vis","default_branch":null,"files":null,"tree":[],"storefront":"/r/reward-scope-ai","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/reward-scope-ai/reward-scope/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}