N/A

rehan243/RLHF-LLM-Optimization

mcp agent Offline

Full RLHF pipeline — SFT, reward modeling, PPO with KL divergence constraints. 68% win rate vs SFT baseline, 96% safety compliance.

Scan Scheduled

This agent is queued for security scanning. It will be graded in the next scan batch.

What We Know