GCP Vertex AI Jailbreak Prompt Heuristics
Detects Vertex AI GenerateContent prompts that match common jailbreak / instruction-override heuristics (for example ignore previous instructions, DAN or developer mode, unrestricted-AI language, or explicit…
Description
Detects Vertex AI GenerateContent prompts that match common jailbreak / instruction-override heuristics (for example ignore previous instructions, DAN or developer mode, unrestricted-AI language, or explicit safety-filter bypass). Typical prompt_response_logs exports do not populate HARM_CATEGORY_JAILBREAK (unlike hate/harassment/dangerous ratings); Model Armor findings remain the preferred first-party jailbreak signal when that API is enabled.
Detection logic
Its licence does not clear it for publishing here
Sunturai publishes a detection's own text where the licence it arrived under has been reviewed and permits it, and Elastic License 2.0 has not. The query as its source wrote it, its canonical form and the hash that pins this revision are in the workspace record.
Detection requirements
- Platform
- ContainersESXiIaaSIdentity ProviderLinuxmacOSNetwork DevicesOffice SuiteWindows
The rule states no platform. This is derived from the ATT&CK technique it maps to.
Known benign triggers
- Approved red-team or evaluation prompts that deliberately include jailbreak strings. Exclude known test principals or lower severity for those models.
MITRE ATT&CK mappings
0 exclusive techniques.This is coverage no other published rule has; it is not this rule's total technique count.
References
| Reference | Cited by |
|---|---|
| cloud.google.com/blog/products/identity-security/introducing-ai-protection-security-for-the-ai-era | Only this detection cites it |
| github.com/elastic/integrations/issues/20740 | 2 |
| www.elastic.co/docs/reference/integrations/gcp_vertexai | 9 |
From the source
- At source
- Open at source
- Upstream identifier
- 42cc9ad9-4ad1-4702-8abc-7d55096f6ee0
- Tagged by the source as
- Mitre Atlas: AML.T0051Mitre Atlas: AML.T0054Data Source: GCPData Source: GCP Vertex AIData Source: Google Cloud PlatformDomain: CloudDomain: GenAIPlatform: GCPResources: Investigation GuideRule Type: ES|QLService: GCP Vertex AITactic: Defense EvasionThreat: Unauthorized AI UsageUse Case: Threat Detection
Licence
- Published under
- Elastic License 2.0Read the licence
- Attribution
- Required
Authorship
- Written by
- Published
- Oct 8, 2026
- Version
- 1