Post
1413
π¨ New Agent Benchmark π¨
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
ai-safety-institute/AgentHarm
Collaboration between UK AI Safety Institute and Gray Swan AI to create a dataset for measuring harmfulness of LLM agents.
The benchmark contains both harmful and benign sets of 11 categories with varied difficulty levels and detailed evaluation, not only testing success rate but also tool level accuracy.
We provide refusal and accuracy metrics across a wide range of models in both no attack and prompt attack scenarios.
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents (2410.09024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
ai-safety-institute/AgentHarm
Collaboration between UK AI Safety Institute and Gray Swan AI to create a dataset for measuring harmfulness of LLM agents.
The benchmark contains both harmful and benign sets of 11 categories with varied difficulty levels and detailed evaluation, not only testing success rate but also tool level accuracy.
We provide refusal and accuracy metrics across a wide range of models in both no attack and prompt attack scenarios.
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents (2410.09024)