N/A
rehan243/RLHF-LLM-Optimization
Full RLHF pipeline — SFT, reward modeling, PPO with KL divergence constraints. 68% win rate vs SFT baseline, 96% safety compliance.
Scan Scheduled
This agent is queued for security scanning. It will be graded in the next scan batch.
What We Know
- URL https://github.com/rehan243/RLHF-LLM-Optimization
- Framework mcp
- Sources github
- First Seen Jul 11, 2026
- Repository github.com/rehan243/RLHF-LLM-Optimization
Browse more:
Search all agents
Ecosystem Report