TABLE OF CONTENTS

Introduction

When the Aqua scanner runs in GitHub Actions and Azure DevOps against a repository that contains Maven projects, the scan can fail with an error like this:

remote Maven repository returned 429 Too Many Requests for
https://repo.maven.apache.org/maven2/com/fasterxml/jackson/jackson-bom/2.13.5/jackson-bom-2.13.5.pom
Retry-After: 1800


The scanner uses Trivy, which reads every pom.xml in the repository. To resolve parent POMs and imported BOMs, Trivy downloads them from Maven Central [1]. GitHub-hosted runners share IP addresses, so Maven Central rate-limits these requests and returns HTTP 429.

This article explains how to avoid the error by building a local Maven repository in the workflow, caching it between runs, and making the scanner read from it instead of Maven Central.


Why this happens

The scanner uses Trivy, which reads every pom.xml in the repository. To resolve parent POMs and imported BOMs, Trivy downloads them from Maven Central [1].

Maven Central enforces consumption limits on high-volume traffic and returns HTTP 429 when they are exceeded [6]. The limit is applied to aggregate traffic, not to a single request. GitHub-hosted runners share outgoing IP addresses with many other builds and scanners, so a scan can be blocked even when the repository itself makes few requests. Once the block is in place, Maven Central rejects all further requests from that IP until the block clears.

For more information, refer to the Maven Central documentation on 429 errors and consumption limits [6].

How to prevent it

There are two ways to handle the error:

  • Wait for the limit to reset. The Retry-After value in the error is the wait time in seconds (1800 in the example above). This clears the block, but it will come back while the scan keeps downloading from Maven Central. Sonatype also states that repeated blocks can last longer [6].
  • Pre-populate the local Maven cache before the scan. This is the fix the error message itself recommends in its last line: run mvn dependency:resolve and cache ~/.m2 in CI. Trivy then uses the dependency data stored locally and does not download it from Maven Central on every run.

This article explains how to implement the second option. The workflow builds a local Maven repository, caches it between runs, and makes the scanner read from it.

Applicability

Aqua SaaS Edition > Supply Chain Module, when code repositories are scanned with the aquasec/aqua-scanner image [5] in GitHub Actions.


Prerequisites

  • A GitHub repository with an Aqua scanner workflow already working (the AQUA_KEY and AQUA_SECRET secrets are configured).
  • Permission to edit the workflow file in .github/workflows/.
  • GitHub-hosted runners, or self-hosted runners on version 2.327.1 or newer with Python 3 and Maven installed.
  • If the projects use a private Maven repository (Nexus, Artifactory), the credentials to access it.


Steps/Procedure

1. Replace the workflow file. Open .github/workflows/aqua.yml in the repository and replace its content with the workflow below. Keep your own AQUA_URL and CSPM_URL values if your tenant is in a different region.

GitHub Actions:

name: Aqua
on:
  pull_request:
  workflow_dispatch:

env:
  # runner.temp/_github_home is mounted as /github/home inside container
  # actions (docker://...), so Trivy sees this folder as ~/.m2/repository
  M2_SUBDIR: _github_home/.m2/repository

jobs:
  aqua:
    name: Aqua scanner
    runs-on: ubuntu-latest
    steps:
      - name: Checkout code
        uses: actions/checkout@v7

      # 1. Find every pom.xml and work out the minimal set of Maven entry points
      - name: Find Maven projects (pom.xml)
        id: poms
        shell: bash
        run: |
          python3 - <<'PY'
          import os, xml.etree.ElementTree as ET

          SKIP_DIRS = {".git", "target", "node_modules", ".m2", ".idea", ".mvn", "build", "out"}

          # 1) every pom.xml in the repo (ignoring build output and tooling folders)
          all_poms = []
          for root, dirs, files in os.walk("."):
              dirs[:] = [d for d in dirs if d not in SKIP_DIRS]
              if "pom.xml" in files:
                  all_poms.append(os.path.normpath(os.path.join(root, "pom.xml")))

          def modules_of(pom):
              """Paths of the modules declared by a POM (<modules> and <profiles>)."""
              try:
                  tree = ET.parse(pom)
              except ET.ParseError as e:
                  print(f"::warning file={pom}::Could not parse {pom}: {e}")
                  return []
              base = os.path.dirname(pom)
              result = []
              for el in tree.iter():
                  if el.tag.split("}")[-1] == "module" and el.text and el.text.strip():
                      path = os.path.normpath(os.path.join(base, el.text.strip()))
                      if os.path.isdir(path):
                          path = os.path.join(path, "pom.xml")
                      result.append(os.path.normpath(path))
              return result

          def reactor(pom, seen):
              """The POM itself plus all of its modules, recursively."""
              if pom in seen:
                  return
              seen.add(pom)
              for m in modules_of(pom):
                  if os.path.isfile(m):
                      reactor(m, seen)

          # 2) shallowest POMs first; skip any POM already covered by an earlier reactor
          covered, entry_points = set(), []
          for pom in sorted(all_poms, key=lambda p: (p.count(os.sep), p)):
              if pom in covered:
                  continue
              entry_points.append(pom)
              reactor(pom, covered)

          print(f"pom.xml files found: {len(all_poms)}")
          for p in sorted(all_poms):
              print(f"  {p}")
          print(f"Maven entry points (one mvn run each): {len(entry_points)}")
          for p in entry_points:
              print(f"  {p}")

          with open("maven-entry-points.txt", "w") as f:
              f.write("\n".join(entry_points) + ("\n" if entry_points else ""))
          with open(os.environ["GITHUB_OUTPUT"], "a") as out:
              out.write(f"found={'true' if entry_points else 'false'}\n")
              out.write(f"count={len(entry_points)}\n")
          PY

      # 2. Restore / save the local Maven repository between runs
      - name: Set up Java
        if: steps.poms.outputs.found == 'true'
        uses: actions/setup-java@v5
        with:
          distribution: temurin
          java-version: '17'

      - name: Cache Maven repository
        if: steps.poms.outputs.found == 'true'
        uses: actions/cache@v6
        with:
          path: ${{ runner.temp }}/_github_home/.m2/repository
          key: m2-${{ runner.os }}-${{ hashFiles('**/pom.xml') }}
          restore-keys: |
            m2-${{ runner.os }}-

      # 3. Download dependencies, parent POMs and BOMs for every project
      - name: Resolve Maven dependencies (fills local cache)
        if: steps.poms.outputs.found == 'true'
        shell: bash
        run: |
          REPO="${RUNNER_TEMP}/${M2_SUBDIR}"
          mkdir -p "$REPO"
          failed=0
          while IFS= read -r pom; do
            [ -z "$pom" ] && continue
            echo "::group::mvn dependency:resolve -f $pom"
            if ! mvn -B -ntp -fae -f "$pom" dependency:resolve -Dmaven.repo.local="$REPO"; then
              echo "::warning file=$pom::Could not resolve all dependencies for $pom (the scan will continue with what is cached)"
              failed=$((failed+1))
            fi
            echo "::endgroup::"
          done < maven-entry-points.txt
          rm -f maven-entry-points.txt
          echo "Projects with resolution problems: $failed"
          echo "Local Maven repository size: $(du -sh "$REPO" | cut -f1)"

      # 4. Scan: --offline-scan makes Trivy read POMs from ~/.m2/repository
      - name: Run Aqua scanner
        uses: docker://aquasec/aqua-scanner
        with:
          # Add --debug to see where Trivy reads each POM from
          args: trivy fs --scanners misconfig,vuln,secret --offline-scan .
        env:
          AQUA_KEY: '${{ secrets.AQUA_KEY }}'
          AQUA_SECRET: ${{ secrets.AQUA_SECRET }}
          GITHUB_TOKEN: ${{ github.token }}
          AQUA_URL: https://api.eu-1.supply-chain.cloud.aquasec.com
          CSPM_URL: https://eu-1.api.cloudsploit.com
          TRIVY_RUN_AS_PLUGIN: 'aqua'


Azure Devops:

# Aqua scanner for Azure Pipelines, with a cached local Maven repository.
#
# Avoids "429 Too Many Requests" from Maven Central: Maven dependencies, parent
# POMs and BOMs are downloaded once, cached between runs, and the scanner reads
# them from the local repository (--offline-scan).
#
# Required secret pipeline variables: AQUA_KEY, AQUA_SECRET, AZURE_TOKEN

trigger: none

pool:
  vmImage: ubuntu-latest

variables:
  # Outside the sources folder, so the scanner does not scan the cached JARs
  MAVEN_CACHE_FOLDER: $(Pipeline.Workspace)/.m2/repository

# The scanner image is declared as a container resource and used only by the
# scan step (target: aqua). The other steps run on the agent itself, where
# Python, Java and Maven are available.
resources:
  containers:
    - container: aqua
      image: aquasec/aqua-scanner
      env:
        AQUA_KEY: $(AQUA_KEY)
        AQUA_SECRET: $(AQUA_SECRET)
        AZURE_TOKEN: $(AZURE_TOKEN)
        AQUA_URL: https://api.eu-1.supply-chain.cloud.aquasec.com
        CSPM_URL: https://eu-1.api.cloudsploit.com
        TRIVY_RUN_AS_PLUGIN: aqua
        # For http/https proxy configuration add env vars: HTTP_PROXY/HTTPS_PROXY, CA-CRET (path to CA certificate)

steps:
  - checkout: self
    fetchDepth: 0

  # ------------------------------------------------------------------
  # 1. Find every pom.xml in the repository and work out the minimal set
  #    of "entry points" to run Maven on:
  #    - a multi-module project is resolved once, from its top POM
  #      (its <modules>, including modules declared in <profiles>, are
  #      covered by that single run);
  #    - any pom.xml that is not a module of another POM (standalone
  #      projects, POMs in unrelated folders) gets its own run.
  # ------------------------------------------------------------------
  - bash: |
      python3 - <<'PY'
      import os, xml.etree.ElementTree as ET

      SKIP_DIRS = {".git", "target", "node_modules", ".m2", ".idea", ".mvn", "build", "out"}

      # 1) every pom.xml in the repo (ignoring build output and tooling folders)
      all_poms = []
      for root, dirs, files in os.walk("."):
          dirs[:] = [d for d in dirs if d not in SKIP_DIRS]
          if "pom.xml" in files:
              all_poms.append(os.path.normpath(os.path.join(root, "pom.xml")))

      def modules_of(pom):
          """Paths of the modules declared by a POM (<modules> and <profiles>)."""
          try:
              tree = ET.parse(pom)
          except ET.ParseError as e:
              print(f"##vso[task.logissue type=warning;sourcepath={pom}]Could not parse {pom}: {e}")
              return []
          base = os.path.dirname(pom)
          result = []
          for el in tree.iter():
              if el.tag.split("}")[-1] == "module" and el.text and el.text.strip():
                  path = os.path.normpath(os.path.join(base, el.text.strip()))
                  if os.path.isdir(path):
                      path = os.path.join(path, "pom.xml")
                  result.append(os.path.normpath(path))
          return result

      def reactor(pom, seen):
          """The POM itself plus all of its modules, recursively."""
          if pom in seen:
              return
          seen.add(pom)
          for m in modules_of(pom):
              if os.path.isfile(m):
                  reactor(m, seen)

      # 2) shallowest POMs first; skip any POM already covered by an earlier reactor
      covered, entry_points = set(), []
      for pom in sorted(all_poms, key=lambda p: (p.count(os.sep), p)):
          if pom in covered:
              continue
          entry_points.append(pom)
          reactor(pom, covered)

      print(f"pom.xml files found: {len(all_poms)}")
      for p in sorted(all_poms):
          print(f"  {p}")
      print(f"Maven entry points (one mvn run each): {len(entry_points)}")
      for p in entry_points:
          print(f"  {p}")

      # Kept in the agent temp folder so it is not part of the scanned sources
      with open(os.path.join(os.environ["AGENT_TEMPDIRECTORY"], "maven-entry-points.txt"), "w") as f:
          f.write("\n".join(entry_points) + ("\n" if entry_points else ""))
      print(f"##vso[task.setvariable variable=MAVEN_PROJECTS_FOUND]{'true' if entry_points else 'false'}")
      PY
    displayName: Find Maven projects (pom.xml)
    workingDirectory: $(Build.SourcesDirectory)

  # ------------------------------------------------------------------
  # 2. Restore / save the local Maven repository between runs
  # ------------------------------------------------------------------
  - task: JavaToolInstaller@0
    displayName: Set up Java
    condition: and(succeeded(), eq(variables['MAVEN_PROJECTS_FOUND'], 'true'))
    inputs:
      versionSpec: '17'
      jdkArchitectureOption: x64
      jdkSourceOption: PreInstalled

  - task: Cache@2
    displayName: Cache Maven repository
    condition: and(succeeded(), eq(variables['MAVEN_PROJECTS_FOUND'], 'true'))
    inputs:
      key: 'maven | "$(Agent.OS)" | **/pom.xml'
      restoreKeys: |
        maven | "$(Agent.OS)"
      path: $(MAVEN_CACHE_FOLDER)

  # ------------------------------------------------------------------
  # 3. Download dependencies, parent POMs and BOMs for every project
  #    (anything already in the cache is not downloaded again)
  # ------------------------------------------------------------------
  - bash: |
      REPO="$(MAVEN_CACHE_FOLDER)"
      LIST="$(Agent.TempDirectory)/maven-entry-points.txt"
      mkdir -p "$REPO"
      failed=0
      while IFS= read -r pom; do
        [ -z "$pom" ] && continue
        echo "##[group]mvn dependency:resolve -f $pom"
        if ! mvn -B -ntp -fae -f "$pom" dependency:resolve -Dmaven.repo.local="$REPO"; then
          echo "##vso[task.logissue type=warning;sourcepath=$pom]Could not resolve all dependencies for $pom (the scan will continue with what is cached)"
          failed=$((failed+1))
        fi
        echo "##[endgroup]"
      done < "$LIST"
      echo "Projects with resolution problems: $failed"
      echo "Local Maven repository size: $(du -sh "$REPO" | cut -f1)"
    displayName: Resolve Maven dependencies (fills local cache)
    condition: and(succeeded(), eq(variables['MAVEN_PROJECTS_FOUND'], 'true'))
    workingDirectory: $(Build.SourcesDirectory)

  # ------------------------------------------------------------------
  # 4. Scan, inside the scanner container.
  #    Trivy looks for POMs in $HOME/.m2/repository, so the cached
  #    repository is linked there first. --offline-scan then makes Trivy
  #    read from it instead of Maven Central (no more 429).
  # ------------------------------------------------------------------
  - script: |
      REPO=""
      for candidate in "$(MAVEN_CACHE_FOLDER)" "$PIPELINE_WORKSPACE/.m2/repository"; do
        if [ -d "$candidate" ]; then REPO="$candidate"; break; fi
      done
      if [ -n "$REPO" ]; then
        mkdir -p "$HOME/.m2"
        rm -rf "$HOME/.m2/repository"
        ln -s "$REPO" "$HOME/.m2/repository"
        echo "Local Maven repository linked: $HOME/.m2/repository -> $REPO"
      else
        echo "No local Maven repository found (no Maven projects in this repository)"
      fi
      # Add --debug to see where Trivy reads each POM from
      trivy fs --scanners misconfig,vuln,secret --offline-scan .
      # To customize which severities to scan for, add the following flag: --severity UNKNOWN,LOW,MEDIUM,HIGH,CRITICAL
      # To enable SAST scanning, add: --sast
      # To enable reachability scanning, add: --reachability
      # To enable npm/dotnet/gradle non-lock file scanning, add: --package-json / --dotnet-proj / --gradle
    displayName: Aqua scanner
    target: aqua
    workingDirectory: $(Build.SourcesDirectory)




2. Understand what the workflow does. It adds four steps before the scan:

  1. Find Maven projects. Locates every pom.xml in the repository, wherever it is, and reads each POM's <modules> (including modules declared in <profiles>). A multi-module project is resolved once, from its top-level POM. A pom.xml that is not a module of another POM gets its own Maven run. The folders target/, build/, out/, .git/, node_modules/, .idea/ and .mvn/ are ignored.
  2. Set up Java. Installs a JDK with actions/setup-java [3] so Maven can run.
  3. Cache Maven repository. Restores and saves the local Maven repository with actions/cache [2]. The cache key is a hash of all pom.xml files, so the cache is refreshed when dependencies change.
  4. Resolve Maven dependencies. Runs mvn dependency:resolve [4] for each project found. This downloads only what the projects use: declared dependencies, transitive dependencies, parent POMs and BOMs.

The scan step then runs with --offline-scan, so Trivy reads the POMs from the local repository instead of Maven Central [1].

3. Commit the change and push it to a branch.

4. Run the workflow twice. Open a pull request, or go to Actions → Aqua → Run workflow.

  • First run: the "Cache Maven repository" step reports a cache miss, and Maven downloads the dependencies.
  • Second run: the step reports that the cache was restored, and Maven downloads little or nothing.

5. Confirm the result. Check that the "Run Aqua scanner" step finishes without 429 Too Many Requests errors. The log of the "Find Maven projects" step lists every pom.xml found and the projects Maven was run on.



Tips and Tricks

  • Why the cache path matters. docker://aquasec/aqua-scanner runs in a container where HOME is /github/home, not the runner's home directory. A regular ~/.m2 cache is invisible to Trivy. The workflow stores the repository in ${{ runner.temp }}/_github_home/.m2/repository, which GitHub mounts into the container as /github/home/.m2/repository.

  • Keep the cache outside the workspace. If the local repository were inside the checkout folder, trivy fs . would scan every cached JAR, which slows the scan and produces false positives.

  • Repositories without Maven projects. If no pom.xml is found, the Java, cache and Maven steps are skipped and the scan runs as usual. The same workflow can be used across all repositories.

  • Private Maven repositories. Configure the credentials in settings.xml, for example with the server-id, server-username and server-password inputs of actions/setup-java [3]. A <mirror> pointing to your proxy also removes the dependency on repo.maven.apache.org.

  • Java version. dependency:resolve does not compile code, so Java 17 works for most projects. Change java-version if a project uses Maven extensions that need another version.


Troubleshooting

SymptomCauseSolution
The scan still reports 429 errorsTrivy is not finding the POMs in the local repositoryAdd --debug to args to see where each POM is read from. Check that the cache path and -Dmaven.repo.local were not changed.
Warning "Could not resolve all dependencies for …"Maven failed for that project (invalid POM, or a private repository without credentials)Open the collapsed mvn dependency:resolve group in the log to see the Maven error. The scan continues with what is cached, but dependencies that were not downloaded are not analysed.
Warning "Could not parse …/pom.xml"The file is not valid XMLFix the POM, or ignore the warning if the file is a test fixture.
Some dependencies are missing from the scan resultsWith --offline-scan, Trivy skips anything that is not in the local repositoryCheck the "Resolve Maven dependencies" step for warnings and fix the failing project.
"Set up Java" or "Cache" fails on a self-hosted runneractions/setup-java@v5 and actions/cache@v6 run on Node 24Upgrade the runner to version 2.327.1 or newer.
mvn: command not found on a self-hosted runnerMaven is not installed on the runnerInstall Maven on the runner, or use the project's Maven Wrapper.
The cache is never restoredThe cache was created on another branch that the current branch cannot readRun the workflow once on the default branch so pull requests can reuse its cache [2].


Conclusion

The 429 error comes from Maven Central rate-limiting the shared IP addresses of GitHub-hosted runners while Trivy resolves parent POMs and BOMs. Pre-populating a local Maven repository, caching it between runs and scanning with --offline-scan removes those requests. The workflow finds the Maven projects by itself, so it can be reused across repositories without changes.


Additional Resources

[1] https://trivy.dev/latest/docs/coverage/language/java/
[2] https://github.com/actions/cache
[3] https://github.com/actions/setup-java
[4] https://maven.apache.org/plugins/maven-dependency-plugin/resolve-mojo.html
[5] https://hub.docker.com/r/aquasec/aqua-scanner
[6] 429 Too Many Requests and Maven Central consumption limits: https://central.sonatype.org/faq/429-error/
[7] What Are the Maven Central Publishing Limits?: https://central.sonatype.org/publish/maven-central-publishing-limits/#what-are-the-maven-central-publishing-limits

image