Run penetration tests from your CI/CD pipeline
Public preview
CI/CD pipeline integration is in preview and is subject to change.
You can run an AWS Security Agent penetration test as a post-deployment gate in your continuous integration and continuous delivery (CI/CD) pipeline. You add the integration as a pipeline step that runs after a deployment. The step tests the change that was just deployed and can block promotion to the next stage (for example, to production) when a finding is at or above a severity threshold that you set.
Each deploy gets security coverage automatically. You don’t have to define the test scope or wait for a scheduled assessment, because AWS Security Agent scopes each run to the change that the pipeline deployed.
This is a post-deployment gate, not a pull request check
This integration tests code after it is deployed to a running environment, and it scopes the test to the deployed change — a commit range that can span more than one pull request. It is different from pull request code review, which scans the changes in an individual pull request before they merge. If you want to scan pull requests as they are opened, see Review code security findings in pull requests and Run a differential code scan with S3 instead. Use this CI/CD pipeline integration when you want to dynamically test a deployed application and gate promotion on the result.
AWS Security Agent provides an integration for common CI/CD platforms:
-
GitHub Actions
-
GitLab CI/CD
-
Bitbucket Pipelines
-
Azure DevOps Pipelines
All testing runs server-side in your AWS account. The pipeline step only starts the job, waits for the result, and reports the outcome to your pipeline; no penetration testing logic runs on your CI/CD runner.
Before you start
The pipeline integration starts a job for a penetration test that you have already configured. It does not create the target or the test for you. Complete this one-time setup first:
-
Verify the target domain you want to test. See Enable an application domain for penetration testing.
-
Create a penetration test in your Agent Space and configure its target, credentials, and network scope. See Create a penetration test. The pipeline tests whatever this penetration test is configured to test.
-
Enable CI/CD mode on the penetration test, and note the Agent Space ID (
as-…) and penetration test ID (pt-…). -
Connect your repository as a source control integration so the pipeline can resolve your repository to its AWS Security Agent integration.
-
Set up AWS authentication with OIDC and add the pipeline step, as described in the following sections. For what differs between the four providers, see How the integration differs by provider.
Note
AWS Security Agent is available in a subset of AWS Regions. Configure your pipeline to use a Region where your Agent Space and penetration test are located.
How the pipeline integration scopes the test
When the pipeline step runs, AWS Security Agent determines the commit range to test and scopes the penetration test to that range:
-
Head – the commit being gated. The integration derives it from the event that triggered the pipeline.
-
Base – the previous commit the gate already cleared. The integration tracks this automatically and falls back to a per-trigger source (such as the previous deployment) when there is no usable record. It never falls back to a full-scope test.
You don’t normally supply commit identifiers. AWS Security Agent computes the difference server-side and validates the range, so the base that the pipeline tracks cannot be used to widen or evade the configured scope. Each integration does accept explicit base and head commit inputs for the case where you need to override the automatic range; pinning a head commit also leaves the recorded baseline unchanged, as What advances the baseline, and why that matters describes.
When the test completes, AWS Security Agent returns a scope decision that tells the pipeline how to proceed:
-
In scope – the change modifies the tested attack surface. Findings are evaluated against your severity threshold, and any finding at or above it blocks promotion.
-
Scoped out – the change is within the configured scope but does not modify the attack surface. There is nothing to test, and the pipeline passes.
-
Scope conflict – the change falls outside the penetration test’s configured scope, so it was not tested. You choose whether this blocks the pipeline. For more information, see Gate on scope conflicts.
Every run records its decision. The Scoping tab of a run in the AWS Security Agent console shows the decision, the reasoning behind it, and the commit range that was compared:
How the pipeline integration authenticates
The pipeline integration does not use long-lived AWS access keys. Instead, it uses OpenID Connect (OIDC) federation: your CI/CD provider issues a short-lived OIDC token that identifies the pipeline, and AWS Security Token Service (AWS STS) exchanges that token for temporary credentials by having the pipeline assume an IAM role that you create in your account. The temporary credentials are used only for the duration of the pipeline step.
The security of the integration depends on two things you control:
-
The IAM role’s trust policy, which decides which pipelines can assume the role.
-
The IAM role’s permissions policy, which decides what the pipeline can do after it assumes the role.
The following sections describe how to scope both to the principle of least privilege.
Scope the trust policy to your pipeline
First, create an IAM OIDC identity provider for your CI/CD provider. The provider URL and audience differ by provider, and on Bitbucket and Azure DevOps the URL contains an identifier specific to your workspace or organization. The following section gives the value for each provider.
Then create the IAM role with a trust policy that federates that provider and restricts which repository can assume the role by matching on the token’s sub (subject) claim. Never wildcard the part of the claim that identifies the repository, project, or service connection: a sub condition such as repo:* lets any repository on that provider — including forks and repositories in other organizations — assume your role and start penetration tests billed to your account.
A trailing wildcard that covers only the ref or step portion of the claim, such as repo:my-org/my-app:*, is a different matter. The repository is still fixed, so the policy admits only that one repository. Tighten it to a specific branch or environment where your provider’s claim allows it, and note that not every provider does. On Bitbucket you cannot scope to a branch at all, and repository scope, optionally plus a deployment environment, is as tight as the claim allows.
The following example trust policy allows only the main branch of the my-org/my-app repository on GitHub Actions to assume the role. Replace the account ID, organization, repository, and branch with your own values.
{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": { "Federated": "arn:aws:iam::111122223333:oidc-provider/token.actions.githubusercontent.com" }, "Action": "sts:AssumeRoleWithWebIdentity", "Condition": { "StringEquals": { "token.actions.githubusercontent.com:aud": "sts.amazonaws.com" }, "StringLike": { "token.actions.githubusercontent.com:sub": "repo:my-org/my-app:ref:refs/heads/main" } } } ] }
The claim name and the format of the sub value differ by provider. The following table gives both for each of the four.
| CI/CD provider | OIDC provider URL and audience | OIDC subject (sub) scoping |
|---|---|---|
|
GitHub Actions |
|
Scope by repository, for example |
|
GitLab CI/CD |
|
Scope by project path, for example |
|
Bitbucket Pipelines |
|
Scope by repository UUID, with braces: |
|
Azure DevOps Pipelines |
|
The subject is the service connection, not a repository: |
Immutable subject claims on GitHub
Repositories that GitHub created or renamed on or after July 15, 2026 (and older repositories that opt in) emit an immutable
sub claim that appends numeric organization and repository IDs after each name, for example repo:my-org@123456/my-repo@789012:ref:refs/heads/main. If your trust policy matches only the name-only form, AssumeRoleWithWebIdentity fails. Confirm the exact sub value that your workflow emits before you write the trust policy.
Grant least-privilege permissions
Attach a permissions policy that allows only the actions the pipeline step performs: starting the penetration test job, polling its status, reading findings, stopping an abandoned job during cleanup, and resolving the repository integration.
{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "securityagent:StartPentestJob", "securityagent:BatchGetPentestJobs", "securityagent:StopPentestJob", "securityagent:ListFindings", "securityagent:BatchGetFindings", "securityagent:ListIntegratedResources" ], "Resource": "*" } ] }
Each action supports a specific part of the pipeline step:
| Action | Why the pipeline step needs it |
|---|---|
|
|
Start the penetration test job for the deployed commit range. |
|
|
Poll the job until it reaches a terminal state. |
|
|
Stop an abandoned job if the pipeline is cancelled or times out before the job finishes. |
|
|
Read findings to evaluate them against your severity threshold. |
|
|
Resolve your repository to its AWS Security Agent integration. |
Note
Some of these actions are list operations that do not support resource-level permissions, so the example uses "Resource": "*". Where an action supports resource-level permissions, you can scope it to your specific Agent Space or penetration test ARN in a separate statement. For which actions support resource-level permissions, see the AWS Security Agent API Reference. Do not store AWS access keys or session tokens as CI/CD variables for this integration; OIDC federation removes the need for stored credentials.
These six actions are the complete set: the pipeline step makes no other AWS Security Agent call, and removing any one of them breaks a step that the gate performs. Do not add an sts: action to this policy. sts:GetCallerIdentity, which the credential step uses to confirm which identity it assumed, requires no permission at all, and sts:AssumeRoleWithWebIdentity is granted by the role’s trust policy rather than by what the role can do once assumed.
How the integration differs by provider
The remaining setup is specific to your CI/CD provider: the OIDC provider URL and audience, how the pipeline obtains credentials, and how the gate records the commit it last cleared. The settings you pass the gate are the same everywhere, but each provider spells their names differently.
| CI/CD provider | How the integration runs, and where it keeps the baseline |
|---|---|
|
GitHub Actions |
Runs as an action. A prior |
|
GitLab CI/CD |
Runs as a component. The component performs the credential exchange itself, and the baseline is a tag in your project — one tag per cleared commit, rather than a single moving record. Four separate timeouts have to agree; see Cancellation and cleanup. |
|
Bitbucket Pipelines |
Runs as a pipe on Bitbucket Cloud. The step must export the role ARN and write the OIDC token to a file before the pipe runs, and the baseline is a repository variable rather than anything in git history. |
|
Azure DevOps Pipelines |
Runs as a task from the AWS Toolkit for Azure DevOps, which an organization administrator installs. Credentials come from an AWS service connection, and the baseline is a repository ref for an Azure Repos source or a build tag for a repository hosted elsewhere. |
Whichever provider you use, run the gate in a step that runs after your pre-production deployment, and gate promotion to production on its result.
Note
Pin the integration to an immutable version — a release tag or commit SHA for a GitHub Action, a pinned component version for GitLab, or a pinned tag or image digest for a container-based integration — so that a change to the upstream integration cannot alter what runs in your pipeline without your review.
Defaults the integration ships with
The settings are named slightly differently on each provider, but they map to the same concepts and share the same defaults. The defaults are identical across all four integrations.
| Setting | Default | What the default means |
|---|---|---|
|
Severity threshold |
|
Findings at |
|
Timeout |
60 minutes |
The step waits up to an hour for the penetration test job to reach a terminal state. |
|
On scope conflict |
|
A change that falls outside the penetration test’s configured scope fails the step, because it was not tested. |
|
Fail on infrastructure error |
|
The shipped default is fail-open. If the job cannot start or complete — for a reason that is not a finding and not a misconfiguration the integration recognizes — the step passes with a warning. Set it to |
|
Dry run |
|
The step starts a real, billable penetration test job. |
Important
Because the default is fail-open on infrastructure errors, a gate you have not validated can pass without testing anything. Run a dry run first, as described in Validate your setup before you trust the gate, and set fail-on-error to true if you rely on the gate to hold a release.
Validate your setup before you trust the gate
Because the gate can be configured to pass the pipeline on infrastructure errors (see Choose how the pipeline responds), a misconfiguration — a bad OIDC trust policy, the wrong Region, or a repository that is not connected as an integration — can otherwise pass silently. Run the integration once in dry-run mode to confirm the setup end to end before you rely on it to gate promotion.
In dry-run mode, the integration checks AWS authentication, the Region, and that your repository resolves to its integration, then exits without starting a billable penetration test job. If validation fails, the step fails regardless of your error-handling setting, so the misconfiguration is surfaced.
Choose how the pipeline responds
Decide how the pipeline step should behave for each outcome. The following table summarizes the outcomes and the settings that control them:
| Outcome | Pipeline behavior |
|---|---|
|
Findings at or above your severity threshold |
The step fails and blocks promotion. Set the threshold to |
|
No findings at or above the threshold, or a scoped-out change |
The step passes. |
|
Any finding, with the threshold set to |
The step passes. The penetration test still runs and every finding is still collected, published, and counted, but no finding fails the step. See Report-only runs. |
|
Scope conflict (change outside the configured scope) |
You choose: block the pipeline, or pass with a warning. The default is to block. See Gate on scope conflicts. |
|
Infrastructure error (the job could not start or complete) |
You choose fail-closed (fail the step) or fail-open (pass the step). The default is fail-open. For changes that require security sign-off, choose fail-closed. A fail-open run also advances the baseline on most providers; see What advances the baseline, and why that matters. |
|
AWS Security Agent rejects the request as unauthorized (HTTP 401 or 403) |
The step always fails, regardless of your error-handling setting, so a permissions misconfiguration cannot silently allow a promotion. All four integrations behave this way. |
|
A misconfiguration the integration recognizes, such as a repository that is not connected |
The step always fails, regardless of your error-handling setting. |
Note
The always-fail rule covers rejections from the AWS Security Agent control plane. A credential problem that fails earlier — when your pipeline exchanges its OIDC token for AWS credentials, before the gate calls AWS Security Agent — is handled by your CI/CD provider’s own credential step, and what it does to the pipeline depends on the provider.
Three settings fail the step independently
The three settings that decide whether the step fails are independent of each other, and each covers a different kind of outcome:
-
Severity threshold decides what happens to findings.
-
On scope conflict decides what happens when the change was not tested because it falls outside the configured scope.
-
Fail on infrastructure error decides what happens when the job could not run.
Setting the severity threshold to NONE relaxes only the first. A scope conflict still fails the step while on-scope-conflict is block, and an infrastructure error still fails it while fail-on-error is true. To make the step report without ever failing the pipeline, relax all three: set the threshold to NONE, on-scope-conflict to warn, and fail-on-error to false.
What advances the baseline, and why that matters
The gate records the commit it last cleared, and the next run tests only what changed after that commit. Any run that passes normally advances that record — including, on most providers, a run that passed without assessing anything.
That is intended. Relaxing a setting is you accepting the outcome it allows: a report-only run passes because you chose not to gate on findings, and a fail-open run passes because you chose to accept an infrastructure error. In both cases the commits in that run’s range are recorded as cleared.
The consequence is the same in both cases, and it is easy to miss: those commits are not tested again. If you later tighten the threshold, or set fail-on-error to true, the gate assesses only commits after the last recorded one. Raising the threshold does not revisit what a report-only run already accepted.
| Run outcome | Advances the baseline |
|---|---|
|
Passes with no finding at or above the threshold, or the change was scoped out |
Yes. |
|
Passes under threshold |
Yes, on all four providers. |
|
Passes with a scope conflict while on-scope-conflict is |
Yes on GitHub Actions, GitLab CI/CD, and Azure DevOps Pipelines. Not on Bitbucket Pipelines, for the same reason as the next row: the change was never tested. |
|
Passes fail-open after an infrastructure error |
Yes on GitHub Actions, GitLab CI/CD, and Azure DevOps Pipelines. Not on Bitbucket Pipelines, which advances the baseline only when AWS Security Agent actually reached a verdict on the range, and warns when it does not. |
|
Fails for any reason — a finding at or above the threshold, a scope conflict while on-scope-conflict is |
No. |
|
A dry run, or a run whose head commit you pinned explicitly |
No. |
To make a blocking threshold reassess a range that an earlier run already cleared, delete the baseline record before you switch. Where it is stored differs by provider; see How the integration differs by provider.
On GitLab, the committed settings are not the gate’s contract
On GitLab CI/CD, a pipeline variable overrides the value a job sets in its own variables: block. The gate’s settings are passed that way, so if you can run a pipeline, you can supply a pipeline variable that lowers the severity threshold, turns on dry-run mode, or relaxes error handling — and the gate then does what the variable says, not what your .gitlab-ci.yml says.
The settings you commit are therefore the gate’s defaults rather than its contract, unless you also restrict who may override variables. In your project’s CI/CD settings, set the minimum role required to override variables on a pipeline to a role you trust with that decision. Restrict it before you rely on the gate to hold a release.
Report-only runs
Setting the severity threshold to NONE gives you a report-only gate: the penetration test still runs and every finding is still collected, published, and counted, but no finding fails the step. This is a useful way to see what the integration finds on real deployments before you let it block a release.
NONE relaxes findings only. A scope conflict and an infrastructure error still fail the step on their own settings, as described in Three settings fail the step independently.
Warning
A report-only run that passes advances the baseline, so the commits it reported on are not assessed again if you later raise the threshold. See What advances the baseline, and why that matters.
Gate on scope conflicts
AWS Security Agent scopes each pipeline-triggered penetration test to the change that the pipeline deployed. When the change falls outside the penetration test’s configured scope, AWS Security Agent reports a scope conflict and does not test the change, rather than silently testing the wrong surface.
You choose how strict to be. With the pipeline integration, you can treat a scope conflict as a pipeline-blocking condition — fail the step and require someone to review the scope — or as a non-blocking signal that records the result and lets the pipeline continue. For a change that requires security sign-off, treat a scope conflict as blocking.
Cancellation and cleanup
A penetration test job runs server-side and is billable, and it keeps running until it finishes even if your pipeline stops waiting for it. When the pipeline step reaches its own timeout, the integration asks AWS Security Agent to stop the abandoned job. This cleanup uses the securityagent:StopPentestJob and securityagent:BatchGetPentestJobs permissions in the policy above, and it never fails the pipeline step.
Cancelling the pipeline does not reliably stop the job
How much the integration can do when you cancel a pipeline depends on what your CI/CD provider lets a step do once cancellation begins, and it is not the same on all four:
| CI/CD provider | What happens when you cancel |
|---|---|
|
GitHub Actions |
The action runs a post-job step during GitHub’s cancellation grace period, which stops the job. |
|
GitLab CI/CD |
The component’s |
|
Bitbucket Pipelines |
There is no post-step hook for a pipe. The pipe asks to stop the job from inside the step when it is signalled, but a cancelled pipeline has been observed to leave the job running to completion. Treat a cancellation as leaving a billable job running. |
|
Azure DevOps Pipelines |
There is no post-step hook. The task stops the job from inside the step, on both cancellation and its own timeout. |
Where cancellation cannot stop the job, stop it yourself from the penetration test’s page in the AWS Security Agent console. Until it reaches a terminal state it continues to run and to accrue cost.
The credentials the pipeline assumes do not refresh while the step runs, so their lifetime has to cover the whole penetration test rather than only the call that starts it. Set the IAM role’s maximum session duration to at least (timeout + 10) * 60 seconds — 4,200 seconds for the default 60-minute timeout — and, on the providers that let you request a lifetime, request that much. AWS STS rejects a request for more than the role allows rather than reducing it, so a role whose maximum is too low fails the credential exchange outright. A credential that expires partway through instead makes an otherwise healthy run fail as though the service were at fault, and it leaves the cleanup unable to stop the abandoned job.
Where you request that lifetime differs by provider: GitLab’s component derives it from the gate timeout, GitHub takes it on the credential step, and Azure DevOps takes it as a pipeline variable. On GitLab the role’s maximum session duration is one of four durations that have to agree — the gate timeout, the gate job’s own timeout, the project’s maximum job timeout, and the role’s session ceiling — and the project maximum caps the other two.
Protect credentials on the runner
The temporary AWS credentials the pipeline obtains, and the OIDC token used to obtain them, are sensitive. Anyone who can read them while the pipeline step runs can act as your pipeline. Follow these practices on your CI/CD runner:
-
Keep debug and verbose tracing off for steps that handle credentials. Debug or shell-trace modes (for example, GitLab CI’s
CI_DEBUG_TRACE) echo command arguments and environment values to the job log in plaintext, which can expose a token or temporary credentials. -
Do not persist source control credentials in the workspace. When your pipeline checks out source, disable credential persistence (for example, set
persist-credentials: falseon the checkout step in GitHub Actions) so that a token is not written to the runner’s.gitconfiguration where a later step could read it. -
Do not expose the OIDC token or credentials to untrusted steps. Grant permission to request the OIDC token only to the jobs that need it, and do not run untrusted code (such as a build from a fork) in a job that has access to the token or the assumed role.
Worked examples by provider
The following topics walk through a complete, reproducible setup for each supported CI/CD platform, using an intentionally vulnerable sample application so that you can see the gate block a real finding:
Related security topics
-
Create a penetration test – configure the penetration test that the pipeline starts, including target URLs, credentials, and network scope.
-
Security best practices for AWS Security Agent – general best practices, including testing against non-production environments and validating AI-generated findings.
-
Identity and access management for AWS Security Agent – how AWS Security Agent works with IAM.
-
Cross-service confused deputy prevention – how AWS Security Agent helps prevent the confused deputy problem, where a less-privileged caller tricks a more-privileged service into acting on its behalf.