How to avoid the "429 Too Many Requests" error in Maven Central when creating an M2 cache as a Maven repository.
TABLE OF CONTENTS
- Introduction
- Applicability
- Prerequisites
- Steps/Procedure
- Tips and Tricks
- Troubleshooting
- Conclusion
- Additional Resources
Introduction
When the Aqua scanner runs in GitHub Actions and Azure DevOps against a repository that contains Maven projects, the scan can fail with an error like this:
remote Maven repository returned 429 Too Many Requests for https://repo.maven.apache.org/maven2/com/fasterxml/jackson/jackson-bom/2.13.5/jackson-bom-2.13.5.pom Retry-After: 1800
The scanner uses Trivy, which reads every pom.xml in the repository. To resolve parent POMs and imported BOMs, Trivy downloads them from Maven Central [1]. GitHub-hosted runners share IP addresses, so Maven Central rate-limits these requests and returns HTTP 429.
This article explains how to avoid the error by building a local Maven repository in the workflow, caching it between runs, and making the scanner read from it instead of Maven Central.
Why this happens
The scanner uses Trivy, which reads every pom.xml in the repository. To resolve parent POMs and imported BOMs, Trivy downloads them from Maven Central [1].
Maven Central enforces consumption limits on high-volume traffic and returns HTTP 429 when they are exceeded [6]. The limit is applied to aggregate traffic, not to a single request. GitHub-hosted runners share outgoing IP addresses with many other builds and scanners, so a scan can be blocked even when the repository itself makes few requests. Once the block is in place, Maven Central rejects all further requests from that IP until the block clears.
For more information, refer to the Maven Central documentation on 429 errors and consumption limits [6].
How to prevent it
There are two ways to handle the error:
- Wait for the limit to reset. The
Retry-Aftervalue in the error is the wait time in seconds (1800 in the example above). This clears the block, but it will come back while the scan keeps downloading from Maven Central. Sonatype also states that repeated blocks can last longer [6]. - Pre-populate the local Maven cache before the scan. This is the fix the error message itself recommends in its last line: run
mvn dependency:resolveand cache~/.m2in CI. Trivy then uses the dependency data stored locally and does not download it from Maven Central on every run.
This article explains how to implement the second option. The workflow builds a local Maven repository, caches it between runs, and makes the scanner read from it.
Applicability
Aqua SaaS Edition > Supply Chain Module, when code repositories are scanned with the aquasec/aqua-scanner image [5] in GitHub Actions.
Prerequisites
- A GitHub repository with an Aqua scanner workflow already working (the
AQUA_KEYandAQUA_SECRETsecrets are configured). - Permission to edit the workflow file in
.github/workflows/. - GitHub-hosted runners, or self-hosted runners on version 2.327.1 or newer with Python 3 and Maven installed.
- If the projects use a private Maven repository (Nexus, Artifactory), the credentials to access it.
Steps/Procedure
1. Replace the workflow file. Open .github/workflows/aqua.yml in the repository and replace its content with the workflow below. Keep your own AQUA_URL and CSPM_URL values if your tenant is in a different region.
GitHub Actions:
name: Aqua
on:
pull_request:
workflow_dispatch:
env:
# runner.temp/_github_home is mounted as /github/home inside container
# actions (docker://...), so Trivy sees this folder as ~/.m2/repository
M2_SUBDIR: _github_home/.m2/repository
jobs:
aqua:
name: Aqua scanner
runs-on: ubuntu-latest
steps:
- name: Checkout code
uses: actions/checkout@v7
# 1. Find every pom.xml and work out the minimal set of Maven entry points
- name: Find Maven projects (pom.xml)
id: poms
shell: bash
run: |
python3 - <<'PY'
import os, xml.etree.ElementTree as ET
SKIP_DIRS = {".git", "target", "node_modules", ".m2", ".idea", ".mvn", "build", "out"}
# 1) every pom.xml in the repo (ignoring build output and tooling folders)
all_poms = []
for root, dirs, files in os.walk("."):
dirs[:] = [d for d in dirs if d not in SKIP_DIRS]
if "pom.xml" in files:
all_poms.append(os.path.normpath(os.path.join(root, "pom.xml")))
def modules_of(pom):
"""Paths of the modules declared by a POM (<modules> and <profiles>)."""
try:
tree = ET.parse(pom)
except ET.ParseError as e:
print(f"::warning file={pom}::Could not parse {pom}: {e}")
return []
base = os.path.dirname(pom)
result = []
for el in tree.iter():
if el.tag.split("}")[-1] == "module" and el.text and el.text.strip():
path = os.path.normpath(os.path.join(base, el.text.strip()))
if os.path.isdir(path):
path = os.path.join(path, "pom.xml")
result.append(os.path.normpath(path))
return result
def reactor(pom, seen):
"""The POM itself plus all of its modules, recursively."""
if pom in seen:
return
seen.add(pom)
for m in modules_of(pom):
if os.path.isfile(m):
reactor(m, seen)
# 2) shallowest POMs first; skip any POM already covered by an earlier reactor
covered, entry_points = set(), []
for pom in sorted(all_poms, key=lambda p: (p.count(os.sep), p)):
if pom in covered:
continue
entry_points.append(pom)
reactor(pom, covered)
print(f"pom.xml files found: {len(all_poms)}")
for p in sorted(all_poms):
print(f" {p}")
print(f"Maven entry points (one mvn run each): {len(entry_points)}")
for p in entry_points:
print(f" {p}")
with open("maven-entry-points.txt", "w") as f:
f.write("\n".join(entry_points) + ("\n" if entry_points else ""))
with open(os.environ["GITHUB_OUTPUT"], "a") as out:
out.write(f"found={'true' if entry_points else 'false'}\n")
out.write(f"count={len(entry_points)}\n")
PY
# 2. Restore / save the local Maven repository between runs
- name: Set up Java
if: steps.poms.outputs.found == 'true'
uses: actions/setup-java@v5
with:
distribution: temurin
java-version: '17'
- name: Cache Maven repository
if: steps.poms.outputs.found == 'true'
uses: actions/cache@v6
with:
path: ${{ runner.temp }}/_github_home/.m2/repository
key: m2-${{ runner.os }}-${{ hashFiles('**/pom.xml') }}
restore-keys: |
m2-${{ runner.os }}-
# 3. Download dependencies, parent POMs and BOMs for every project
- name: Resolve Maven dependencies (fills local cache)
if: steps.poms.outputs.found == 'true'
shell: bash
run: |
REPO="${RUNNER_TEMP}/${M2_SUBDIR}"
mkdir -p "$REPO"
failed=0
while IFS= read -r pom; do
[ -z "$pom" ] && continue
echo "::group::mvn dependency:resolve -f $pom"
if ! mvn -B -ntp -fae -f "$pom" dependency:resolve -Dmaven.repo.local="$REPO"; then
echo "::warning file=$pom::Could not resolve all dependencies for $pom (the scan will continue with what is cached)"
failed=$((failed+1))
fi
echo "::endgroup::"
done < maven-entry-points.txt
rm -f maven-entry-points.txt
echo "Projects with resolution problems: $failed"
echo "Local Maven repository size: $(du -sh "$REPO" | cut -f1)"
# 4. Scan: --offline-scan makes Trivy read POMs from ~/.m2/repository
- name: Run Aqua scanner
uses: docker://aquasec/aqua-scanner
with:
# Add --debug to see where Trivy reads each POM from
args: trivy fs --scanners misconfig,vuln,secret --offline-scan .
env:
AQUA_KEY: '${{ secrets.AQUA_KEY }}'
AQUA_SECRET: ${{ secrets.AQUA_SECRET }}
GITHUB_TOKEN: ${{ github.token }}
AQUA_URL: https://api.eu-1.supply-chain.cloud.aquasec.com
CSPM_URL: https://eu-1.api.cloudsploit.com
TRIVY_RUN_AS_PLUGIN: 'aqua'Azure Devops:
# Aqua scanner for Azure Pipelines, with a cached local Maven repository.
#
# Avoids "429 Too Many Requests" from Maven Central: Maven dependencies, parent
# POMs and BOMs are downloaded once, cached between runs, and the scanner reads
# them from the local repository (--offline-scan).
#
# Required secret pipeline variables: AQUA_KEY, AQUA_SECRET, AZURE_TOKEN
trigger: none
pool:
vmImage: ubuntu-latest
variables:
# Outside the sources folder, so the scanner does not scan the cached JARs
MAVEN_CACHE_FOLDER: $(Pipeline.Workspace)/.m2/repository
# The scanner image is declared as a container resource and used only by the
# scan step (target: aqua). The other steps run on the agent itself, where
# Python, Java and Maven are available.
resources:
containers:
- container: aqua
image: aquasec/aqua-scanner
env:
AQUA_KEY: $(AQUA_KEY)
AQUA_SECRET: $(AQUA_SECRET)
AZURE_TOKEN: $(AZURE_TOKEN)
AQUA_URL: https://api.eu-1.supply-chain.cloud.aquasec.com
CSPM_URL: https://eu-1.api.cloudsploit.com
TRIVY_RUN_AS_PLUGIN: aqua
# For http/https proxy configuration add env vars: HTTP_PROXY/HTTPS_PROXY, CA-CRET (path to CA certificate)
steps:
- checkout: self
fetchDepth: 0
# ------------------------------------------------------------------
# 1. Find every pom.xml in the repository and work out the minimal set
# of "entry points" to run Maven on:
# - a multi-module project is resolved once, from its top POM
# (its <modules>, including modules declared in <profiles>, are
# covered by that single run);
# - any pom.xml that is not a module of another POM (standalone
# projects, POMs in unrelated folders) gets its own run.
# ------------------------------------------------------------------
- bash: |
python3 - <<'PY'
import os, xml.etree.ElementTree as ET
SKIP_DIRS = {".git", "target", "node_modules", ".m2", ".idea", ".mvn", "build", "out"}
# 1) every pom.xml in the repo (ignoring build output and tooling folders)
all_poms = []
for root, dirs, files in os.walk("."):
dirs[:] = [d for d in dirs if d not in SKIP_DIRS]
if "pom.xml" in files:
all_poms.append(os.path.normpath(os.path.join(root, "pom.xml")))
def modules_of(pom):
"""Paths of the modules declared by a POM (<modules> and <profiles>)."""
try:
tree = ET.parse(pom)
except ET.ParseError as e:
print(f"##vso[task.logissue type=warning;sourcepath={pom}]Could not parse {pom}: {e}")
return []
base = os.path.dirname(pom)
result = []
for el in tree.iter():
if el.tag.split("}")[-1] == "module" and el.text and el.text.strip():
path = os.path.normpath(os.path.join(base, el.text.strip()))
if os.path.isdir(path):
path = os.path.join(path, "pom.xml")
result.append(os.path.normpath(path))
return result
def reactor(pom, seen):
"""The POM itself plus all of its modules, recursively."""
if pom in seen:
return
seen.add(pom)
for m in modules_of(pom):
if os.path.isfile(m):
reactor(m, seen)
# 2) shallowest POMs first; skip any POM already covered by an earlier reactor
covered, entry_points = set(), []
for pom in sorted(all_poms, key=lambda p: (p.count(os.sep), p)):
if pom in covered:
continue
entry_points.append(pom)
reactor(pom, covered)
print(f"pom.xml files found: {len(all_poms)}")
for p in sorted(all_poms):
print(f" {p}")
print(f"Maven entry points (one mvn run each): {len(entry_points)}")
for p in entry_points:
print(f" {p}")
# Kept in the agent temp folder so it is not part of the scanned sources
with open(os.path.join(os.environ["AGENT_TEMPDIRECTORY"], "maven-entry-points.txt"), "w") as f:
f.write("\n".join(entry_points) + ("\n" if entry_points else ""))
print(f"##vso[task.setvariable variable=MAVEN_PROJECTS_FOUND]{'true' if entry_points else 'false'}")
PY
displayName: Find Maven projects (pom.xml)
workingDirectory: $(Build.SourcesDirectory)
# ------------------------------------------------------------------
# 2. Restore / save the local Maven repository between runs
# ------------------------------------------------------------------
- task: JavaToolInstaller@0
displayName: Set up Java
condition: and(succeeded(), eq(variables['MAVEN_PROJECTS_FOUND'], 'true'))
inputs:
versionSpec: '17'
jdkArchitectureOption: x64
jdkSourceOption: PreInstalled
- task: Cache@2
displayName: Cache Maven repository
condition: and(succeeded(), eq(variables['MAVEN_PROJECTS_FOUND'], 'true'))
inputs:
key: 'maven | "$(Agent.OS)" | **/pom.xml'
restoreKeys: |
maven | "$(Agent.OS)"
path: $(MAVEN_CACHE_FOLDER)
# ------------------------------------------------------------------
# 3. Download dependencies, parent POMs and BOMs for every project
# (anything already in the cache is not downloaded again)
# ------------------------------------------------------------------
- bash: |
REPO="$(MAVEN_CACHE_FOLDER)"
LIST="$(Agent.TempDirectory)/maven-entry-points.txt"
mkdir -p "$REPO"
failed=0
while IFS= read -r pom; do
[ -z "$pom" ] && continue
echo "##[group]mvn dependency:resolve -f $pom"
if ! mvn -B -ntp -fae -f "$pom" dependency:resolve -Dmaven.repo.local="$REPO"; then
echo "##vso[task.logissue type=warning;sourcepath=$pom]Could not resolve all dependencies for $pom (the scan will continue with what is cached)"
failed=$((failed+1))
fi
echo "##[endgroup]"
done < "$LIST"
echo "Projects with resolution problems: $failed"
echo "Local Maven repository size: $(du -sh "$REPO" | cut -f1)"
displayName: Resolve Maven dependencies (fills local cache)
condition: and(succeeded(), eq(variables['MAVEN_PROJECTS_FOUND'], 'true'))
workingDirectory: $(Build.SourcesDirectory)
# ------------------------------------------------------------------
# 4. Scan, inside the scanner container.
# Trivy looks for POMs in $HOME/.m2/repository, so the cached
# repository is linked there first. --offline-scan then makes Trivy
# read from it instead of Maven Central (no more 429).
# ------------------------------------------------------------------
- script: |
REPO=""
for candidate in "$(MAVEN_CACHE_FOLDER)" "$PIPELINE_WORKSPACE/.m2/repository"; do
if [ -d "$candidate" ]; then REPO="$candidate"; break; fi
done
if [ -n "$REPO" ]; then
mkdir -p "$HOME/.m2"
rm -rf "$HOME/.m2/repository"
ln -s "$REPO" "$HOME/.m2/repository"
echo "Local Maven repository linked: $HOME/.m2/repository -> $REPO"
else
echo "No local Maven repository found (no Maven projects in this repository)"
fi
# Add --debug to see where Trivy reads each POM from
trivy fs --scanners misconfig,vuln,secret --offline-scan .
# To customize which severities to scan for, add the following flag: --severity UNKNOWN,LOW,MEDIUM,HIGH,CRITICAL
# To enable SAST scanning, add: --sast
# To enable reachability scanning, add: --reachability
# To enable npm/dotnet/gradle non-lock file scanning, add: --package-json / --dotnet-proj / --gradle
displayName: Aqua scanner
target: aqua
workingDirectory: $(Build.SourcesDirectory)
2. Understand what the workflow does. It adds four steps before the scan:
- Find Maven projects. Locates every
pom.xmlin the repository, wherever it is, and reads each POM's<modules>(including modules declared in<profiles>). A multi-module project is resolved once, from its top-level POM. Apom.xmlthat is not a module of another POM gets its own Maven run. The folderstarget/,build/,out/,.git/,node_modules/,.idea/and.mvn/are ignored. - Set up Java. Installs a JDK with
actions/setup-java[3] so Maven can run. - Cache Maven repository. Restores and saves the local Maven repository with
actions/cache[2]. The cache key is a hash of allpom.xmlfiles, so the cache is refreshed when dependencies change. - Resolve Maven dependencies. Runs
mvn dependency:resolve[4] for each project found. This downloads only what the projects use: declared dependencies, transitive dependencies, parent POMs and BOMs.
The scan step then runs with --offline-scan, so Trivy reads the POMs from the local repository instead of Maven Central [1].
3. Commit the change and push it to a branch.
4. Run the workflow twice. Open a pull request, or go to Actions → Aqua → Run workflow.
- First run: the "Cache Maven repository" step reports a cache miss, and Maven downloads the dependencies.
- Second run: the step reports that the cache was restored, and Maven downloads little or nothing.
5. Confirm the result. Check that the "Run Aqua scanner" step finishes without 429 Too Many Requests errors. The log of the "Find Maven projects" step lists every pom.xml found and the projects Maven was run on.
Tips and Tricks
- Why the cache path matters.
docker://aquasec/aqua-scannerruns in a container whereHOMEis/github/home, not the runner's home directory. A regular~/.m2cache is invisible to Trivy. The workflow stores the repository in${{ runner.temp }}/_github_home/.m2/repository, which GitHub mounts into the container as/github/home/.m2/repository. - Keep the cache outside the workspace. If the local repository were inside the checkout folder,
trivy fs .would scan every cached JAR, which slows the scan and produces false positives. - Repositories without Maven projects. If no
pom.xmlis found, the Java, cache and Maven steps are skipped and the scan runs as usual. The same workflow can be used across all repositories. - Private Maven repositories. Configure the credentials in
settings.xml, for example with theserver-id,server-usernameandserver-passwordinputs ofactions/setup-java[3]. A<mirror>pointing to your proxy also removes the dependency onrepo.maven.apache.org. - Java version.
dependency:resolvedoes not compile code, so Java 17 works for most projects. Changejava-versionif a project uses Maven extensions that need another version.
Troubleshooting
| Symptom | Cause | Solution |
|---|---|---|
| The scan still reports 429 errors | Trivy is not finding the POMs in the local repository | Add --debug to args to see where each POM is read from. Check that the cache path and -Dmaven.repo.local were not changed. |
| Warning "Could not resolve all dependencies for …" | Maven failed for that project (invalid POM, or a private repository without credentials) | Open the collapsed mvn dependency:resolve group in the log to see the Maven error. The scan continues with what is cached, but dependencies that were not downloaded are not analysed. |
| Warning "Could not parse …/pom.xml" | The file is not valid XML | Fix the POM, or ignore the warning if the file is a test fixture. |
| Some dependencies are missing from the scan results | With --offline-scan, Trivy skips anything that is not in the local repository | Check the "Resolve Maven dependencies" step for warnings and fix the failing project. |
| "Set up Java" or "Cache" fails on a self-hosted runner | actions/setup-java@v5 and actions/cache@v6 run on Node 24 | Upgrade the runner to version 2.327.1 or newer. |
mvn: command not found on a self-hosted runner | Maven is not installed on the runner | Install Maven on the runner, or use the project's Maven Wrapper. |
| The cache is never restored | The cache was created on another branch that the current branch cannot read | Run the workflow once on the default branch so pull requests can reuse its cache [2]. |
Conclusion
The 429 error comes from Maven Central rate-limiting the shared IP addresses of GitHub-hosted runners while Trivy resolves parent POMs and BOMs. Pre-populating a local Maven repository, caching it between runs and scanning with --offline-scan removes those requests. The workflow finds the Maven projects by itself, so it can be reused across repositories without changes.
Additional Resources
[1] https://trivy.dev/latest/docs/coverage/language/java/
[2] https://github.com/actions/cache
[3] https://github.com/actions/setup-java
[4] https://maven.apache.org/plugins/maven-dependency-plugin/resolve-mojo.html
[5] https://hub.docker.com/r/aquasec/aqua-scanner
[6] 429 Too Many Requests and Maven Central consumption limits: https://central.sonatype.org/faq/429-error/
[7] What Are the Maven Central Publishing Limits?: https://central.sonatype.org/publish/maven-central-publishing-limits/#what-are-the-maven-central-publishing-limits

Did you find it helpful? Yes No
Send feedback