← Back to match
DeepSafe
All-in-one safety evaluation framework for LLMs and multimodal models, with 25+ safety datasets.
Pluginfreeglobal
DeepSafe is an all-in-one AI safety evaluation framework integrating 25+ safety datasets and a specialized ProGuard evaluation model for full-modal LLM and VLM assessment, with published leaderboards covering GPT, Claude, Gemini, DeepSeek, Qwen, Llama, and Mistral. It's for AI safety researchers and labs who want a standardized, comprehensive evaluation suite instead of assembling ad hoc safety benchmarks themselves.
Categories
AI safety toolsbenchmarkssecurity tools

