N/A

SaiArja/LLM Evaluation Harness MCP Server

mcp agent Offline

Evaluates RAG outputs on faithfulness, answer relevancy, and context precision using an LLM-as-a-Judge backend. Exposes tools for running evaluations, scoring individual samples, and checking thresholds, enabling CI gating and on-demand assessment via MCP.

Scan Scheduled

This agent is queued for security scanning. It will be graded in the next scan batch.

What We Know