N/A

Basaltlabs-app/Gauntlet

mcp agent Offline

Deterministic behavioral benchmark for LLMs. 109 probes, 16 categories, no LLM-as-judge. Tests sycophancy resistance, instruction decay, temporal coherence, confidence calibration, and more.

Scan Scheduled

This agent is queued for security scanning. It will be graded in the next scan batch.

What We Know