GCP Vertex AI High Request and Token Volume
Detects a high number of Vertex AI prompt-response log rows and high total token usage for one model in the lookback window. Totals use provider usage_metadata.total_token_count (includes thinking tokens when…
Description
Detects a high number of Vertex AI prompt-response log rows and high total token usage for one model in the lookback window. Totals use provider usage_metadata.total_token_count (includes thinking tokens when billed). High request and token volume can indicate model extraction, quota abuse, or a compromised caller. Tune thresholds for your baseline.
Detection logic
Its licence does not clear it for publishing here
Sunturai publishes a detection's own text where the licence it arrived under has been reviewed and permits it, and Elastic License 2.0 has not. The query as its source wrote it, its canonical form and the hash that pins this revision are in the workspace record.
MITRE ATT&CK mappings
0 exclusive techniques.This is coverage no other published rule has; it is not this rule's total technique count.
References
| Reference | Cited by |
|---|---|
| www.elastic.co/docs/reference/integrations/gcp_vertexai | 9 |
From the source
- At source
- Open at source
- Upstream identifier
- 9bd7c550-5d7f-43ed-969a-76eb50f9e553
- Tagged by the source as
- Mitre Atlas: AML.T0029Mitre Atlas: AML.T0034Data Source: GCPData Source: GCP Vertex AIData Source: Google Cloud PlatformDomain: CloudDomain: GenAIPlatform: GCPResources: Investigation GuideRule Type: ES|QLService: GCP Vertex AITactic: ImpactThreat: LLMjackingThreat: Unauthorized AI UsageUse Case: Potential OverloadUse Case: Threat Detection
Licence
- Published under
- Elastic License 2.0Read the licence
- Attribution
- Required
Authorship
- Written by
- Published
- Oct 8, 2026
- Version
- 1