Back to results

GCP Vertex AI Jailbreak Prompt Heuristics

Detects Vertex AI GenerateContent prompts that match common jailbreak / instruction-override heuristics (for example ignore previous instructions, DAN or developer mode, unrestricted-AI language, or explicit…

Description

Detects Vertex AI GenerateContent prompts that match common jailbreak / instruction-override heuristics (for example ignore previous instructions, DAN or developer mode, unrestricted-AI language, or explicit safety-filter bypass). Typical prompt_response_logs exports do not populate HARM_CATEGORY_JAILBREAK (unlike hate/harassment/dangerous ratings); Model Armor findings remain the preferred first-party jailbreak signal when that API is enabled.

Detection logic

Its licence does not clear it for publishing here

Sunturai publishes a detection's own text where the licence it arrived under has been reviewed and permits it, and Elastic License 2.0 has not. The query as its source wrote it, its canonical form and the hash that pins this revision are in the workspace record.

Detection requirements

Platform
ContainersESXiIaaSIdentity ProviderLinuxmacOSNetwork DevicesOffice SuiteWindows

The rule states no platform. This is derived from the ATT&CK technique it maps to.

Known benign triggers

  • Approved red-team or evaluation prompts that deliberately include jailbreak strings. Exclude known test principals or lower severity for those models.

Detections can measure how the public catalogue is used — which detections people look for, and which pages bring them here. It sets a cookie that recognises this browser for 180 days. It is never linked to an account and never follows you to other sites. Privacy notice