arxiv:2402.13494

GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis

Published on Feb 21, 2024

Authors:

Abstract

Large Language Models (LLMs) face threats from jailbreak prompts. Existing methods for detecting jailbreak prompts are primarily online moderation APIs or finetuned LLMs. These strategies, however, often require extensive and resource-intensive data collection and training processes. In this study, we propose GradSafe, which effectively detects jailbreak prompts by scrutinizing the gradients of safety-critical parameters in LLMs. Our method is grounded in a pivotal observation: the gradients of an LLM's loss for jailbreak prompts paired with compliance response exhibit similar patterns on certain safety-critical parameters. In contrast, safe prompts lead to different gradient patterns. Building on this observation, GradSafe analyzes the gradients from prompts (paired with <PRE_TAG>compliance responses</POST_TAG>) to accurately detect jailbreak prompts. We show that GradSafe, applied to Llama-2 without further training, outperforms Llama Guard, despite its extensive finetuning with a large dataset, in detecting jailbreak prompts. This superior performance is consistent across both zero-shot and adaptation scenarios, as evidenced by our evaluations on ToxicChat and XSTest. The source code is available at https://github.com/xyq7/GradSafe.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment

No model linking this paper

Cite arxiv.org/abs/2402.13494 in a model README.md to link it from this page.

No dataset linking this paper

Cite arxiv.org/abs/2402.13494 in a dataset README.md to link it from this page.

No Space linking this paper

Cite arxiv.org/abs/2402.13494 in a Space README.md to link it from this page.

No Collection including this paper

Add this paper to a collection to link it from this page.