Back to results

GCP Vertex AI High Request and Token Volume

Detects a high number of Vertex AI prompt-response log rows and high total token usage for one model in the lookback window. Totals use provider usage_metadata.total_token_count (includes thinking tokens when…

Description

Detects a high number of Vertex AI prompt-response log rows and high total token usage for one model in the lookback window. Totals use provider usage_metadata.total_token_count (includes thinking tokens when billed). High request and token volume can indicate model extraction, quota abuse, or a compromised caller. Tune thresholds for your baseline.

Detection logic

Its licence does not clear it for publishing here

Sunturai publishes a detection's own text where the licence it arrived under has been reviewed and permits it, and Elastic License 2.0 has not. The query as its source wrote it, its canonical form and the hash that pins this revision are in the workspace record.

Detection requirements

Platform
ContainersIaaSLinuxmacOSSaaSWindows

The rule states no platform. This is derived from the ATT&CK technique it maps to.

Known benign triggers

  • Approved batch evaluation or high-traffic applications. Raise thresholds above the normal peak per model.

Detections can measure how the public catalogue is used — which detections people look for, and which pages bring them here. It sets a cookie that recognises this browser for 180 days. It is never linked to an account and never follows you to other sites. Privacy notice