Compare commits

...
Author SHA1 Message Date
mintlify[bot] c5e8f7e176 Merge remote-tracking branch 'origin/master' into mintlify/55c1f1a5 2026-09-28 07:07:36 +00:00
Alejandro Bailo c2b8092461 feat(ui): hint the resource re-check beside Last seen (#12883) 2026-09-25 14:55:21 +02:00
Pedro Martín 26d9d24e5d docs(introduction): list UI and API support for Okta (#12881) 2026-09-25 10:16:53 +02:00
Alejandro Bailo ee59e35bc2 fix(ui): stop offering reports and compliance for partial scans (#12880) 2026-09-25 10:09:33 +02:00
César Arroba e0fa23b9ee fix(ui): show API error message when mute rule creation fails (#12853) 2026-09-24 15:35:19 +02:00
Alejandro Bailo 4dbc3c7e74 feat(ui): re-check a resource with a partial scan from the findings actions (#12879) 2026-09-24 13:54:15 +02:00
César Arroba 576433d85d fix(api): reap orphaned attack paths temp Neo4j databases (#12832) 2026-09-24 13:23:36 +02:00
Rubén De la Torre Vico bf179212a5 fix(api): return the new scan id when a scan is created (#12878) 2026-09-24 13:12:44 +02:00
César Arroba 60f936a10b fix(ui): wait for the full provider connection check before reporting a result (#12869) 2026-09-24 11:23:56 +02:00
César Arroba 15630f54d2 fix(aws): reuse the STS region that answered and add PROWLER_AWS_BOTO3_RETRIES_MAX_ATTEMPTS (#12870) 2026-09-24 10:24:44 +02:00
César Arroba 706603fe4d feat(api): add an explicit endpoint for S3-compatible scan output storage (#12871) 2026-09-24 10:24:19 +02:00
Pedro Martínandalejandrobailo dc67fe4f37 feat(ui): connect and test AWS accounts in one step (#12876)
Co-authored-by: alejandrobailo <alejandrobailo94@gmail.com>
2026-09-24 09:57:59 +02:00
Prowler Botandprowler-bot 2c233c2f6c chore(release): Bump versions to v5.44.0 (#12860)
Co-authored-by: prowler-bot <179230569+prowler-bot@users.noreply.github.com>
2026-09-23 13:11:45 +02:00
Alejandro Bailo 69e1d19abe feat(ui): open the paid plan upgrade modal from report downloads (#12875) 2026-09-23 13:08:49 +02:00
Alejandro Bailo 859421b0ec fix(ui): read the persisted sidebar mode without a hydration mismatch (#12873) 2026-09-23 13:03:28 +02:00
Pedro Martínandalejandrobailo ea36f12a01 feat(ui): connect AWS accounts in a single wizard step (#12852)
Co-authored-by: alejandrobailo <alejandrobailo94@gmail.com>
2026-09-22 18:21:55 +02:00
Pujitha Paladugu 50a9138bea fix(api): sign report download URLs with SigV4 when the bucket region is set (#12746)
When DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION is set, get_s3_presign_client() signs download URLs with SigV4, path-style, against the regional host, so SSE-KMS buckets no longer reject them with InvalidArgument. Without a region or a public endpoint, URLs are signed as before, and get_s3_client() is unchanged.

Fixes #12734
2026-09-22 16:12:44 +02:00
tejas_0007 09821e6328 fix(api): log Celery worker failures and restart compose services (#12465)
Declare the Celery worker, kombu, billiard and amqp loggers so fatal worker errors are no longer silenced by disable_existing_loggers, and log task failures from celery.app.trace at WARNING. Add restart: unless-stopped to the long-running services in docker-compose.yml.

Refs #12461
2026-09-22 12:23:51 +02:00
César Arroba 94899e20fd ci(ui): pass E2E AWS credentials through env vars (#12864) 2026-09-22 10:42:40 +02:00
Pedro Martín 79da676c86 chore(changelog): v5.43.0 highlights (#12839) 2026-09-22 09:22:26 +02:00
Rubén De la Torre Vico 8b35b69731 fix(api): make every endpoint resolve the same latest scan for a provider (#12858) 2026-09-21 17:38:11 +02:00
Prowler Botandprowler-bot 81d90fc31e chore(changelog): v5.43.0 (#12856)
Co-authored-by: prowler-bot <179230569+prowler-bot@users.noreply.github.com>
2026-09-21 15:30:24 +02:00
César Arroba 8b5f1250a9 chore(changelog): reclassify FedRAMP 20x pilot removal as changed (#12855) 2026-09-21 15:26:00 +02:00
98d2db4e13 feat(aws): Update regions for AWS services (#12846)
Co-authored-by: prowler-bot <179230569+prowler-bot@users.noreply.github.com>
Co-authored-by: César Arroba <19954079+cesararroba@users.noreply.github.com>
2026-09-21 13:00:52 +02:00
César Arroba c9068515b2 fix(deps): bump anyio to 4.14.2 and accept unfixable CPython CVE-2026-82049 (#12848) 2026-09-21 12:51:41 +02:00
Alejandro Bailo 3823186914 fix(ui): hide Registry when it is unavailable to the deployment (#12847) 2026-09-21 12:27:22 +02:00
mintlify[bot] 2dbdcf2b50 docs: add missing blank line before Registry heading in environment variables page 2026-09-21 07:09:03 +00:00
mintlify[bot] 4f145a3c0d Merge remote-tracking branch 'origin/master' into mintlify/55c1f1a5 2026-09-21 07:07:46 +00:00
Alejandro BailoandClaude Fable 5.1 03cb59c20d feat(ui): install Registry checks artifacts on the API's verdict (#12843)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 16:04:58 +02:00
Pedro Martín d819639f0e fix(cloudflare): request every permission the checks need (#12842) 2026-09-18 13:13:30 +02:00
Pedro Martín 6c8d6994bb fix(api): sign report URLs against public storage host (#12552) 2026-09-18 10:16:13 +02:00
Pedro Martín 07d48ab15d fix(huaweicloud): remove mock data from SMN service (#12836) 2026-09-18 09:20:35 +02:00
Alejandro Bailo 5fe1a6713b revert: remove the onboarding profile step (#12818) (#12838) 2026-09-17 18:01:41 +02:00
Alejandro Bailo 61d13f078c fix(ui): hide the Registry tab in Add Provider when Registry is unavailable (#12837) 2026-09-17 17:22:21 +02:00
bbb297aee7 feat(providers/huaweicloud): add smn_topic_subscriptions check (#12186)
Co-authored-by: tomitobio <tomitobio@users.noreply.github.com>
Co-authored-by: Hugo P.Brito <hugopbrit@gmail.com>
2026-09-17 13:31:08 +02:00
Haitao Zheng d0d29108e0 fix(sdk): report all-users 2sv override failures (#12700) 2026-09-17 13:19:11 +02:00
Pedro Martín f382400037 fix(invitations): expire lapsed invitations on re-invite (#12831) 2026-09-17 12:12:14 +02:00
Alejandro BailoandClaude Opus 5 396ccf56bb feat(ui): offer an invite-your-team step before the onboarding checkpoint (#12819)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 11:58:53 +02:00
Alejandro BailoandClaude Opus 5 3069486549 feat: record the tenant profile declared at onboarding (#12818)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 10:03:59 +02:00
Pedro Martín 9f616a5d43 feat(ui): allow disabling self-registration in Cloud (#12815) 2026-09-16 14:00:54 +02:00
Alan Buscagliaandalejandrobailo 2198ba2d84 feat(ui): complete Registry provider onboarding for Private Cloud (#12494)
Co-authored-by: alejandrobailo <alejandrobailo94@gmail.com>
2026-09-16 12:21:25 +02:00
César Arroba 974f4251dd chore(trivy): suppress fast-uri CVE-2026-75931 from Teams SPDX manifest (#12823) 2026-09-16 10:29:50 +02:00
César Arroba 75c22df63b fix(azure): use the selected cloud endpoints in Defender and Key Vault (#12813) 2026-09-16 09:25:08 +02:00
César Arrobaandpedrooot 757cd44ecb fix(aws): try the rest of the partition when the bootstrap region is unreachable (#12799)
Co-authored-by: pedrooot <pedromarting3@gmail.com>
2026-09-16 09:24:41 +02:00
Pedro Martínandalejandrobailo c0fdd5bdf3 fix(ui): add FedRAMP 20x cross-provider mappers (#12810)
Co-authored-by: alejandrobailo <alejandrobailo94@gmail.com>
2026-09-15 11:07:31 +02:00
Pedro Martín 61ef44a03b fix(m365): skip defender preset policies without rules (#12809) 2026-09-15 09:56:50 +02:00
Pedro Martín 682353e054 fix(container): bump PowerShell to 7.5.11 in SDK and API (#12811) 2026-09-15 09:44:32 +02:00
Pedro Martín 3860cd3dce feat(compliance): FedRAMP 20x Class C FRR + AWS checks (#12808) 2026-09-14 16:35:05 +02:00
1b228d590b feat(compliance): replace FedRAMP 20x pilot KSI with Consolidated Rules 2026.06.24.01 (#11701)
Co-authored-by: Ethan Troy <ethanolivertroy@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: pedrooot <pedromarting3@gmail.com>
2026-09-14 11:53:19 +02:00
Alejandro Bailo 8b265a8314 fix(ui): defer billing onboarding and add AWS button styling (#12803) 2026-09-14 11:18:31 +02:00
Pedro Martín ba258a5346 fix(container): patch high Debian CVEs in SDK and API (#12804) 2026-09-14 11:01:15 +02:00
mintlify[bot] 538d5273c7 docs: fix formatting and punctuation in AWS partition and boto3 pages 2026-09-14 07:08:51 +00:00
mintlify[bot] e0379250ae Merge remote-tracking branch 'origin/master' into mintlify/55c1f1a5 2026-09-14 07:07:02 +00:00
Prowler Botandprowler-bot 282fe5b46b chore(release): Bump versions to v5.43.0 (#12797)
Co-authored-by: prowler-bot <179230569+prowler-bot@users.noreply.github.com>
2026-09-11 13:23:08 +02:00
Pedro MartínandPepe Fagoaga c08cb65d84 chore(changelog): v5.42.0 highlights (#12795)
Co-authored-by: Pepe Fagoaga <pepe@prowler.com>
2026-09-11 12:11:44 +02:00
Prowler Botandprowler-bot 4727da7ca7 chore(changelog): v5.42.0 (#12794)
Co-authored-by: prowler-bot <179230569+prowler-bot@users.noreply.github.com>
2026-09-11 10:21:17 +02:00
Pedro Martín b378f15798 fix(aws): guard checks reading iam roles when unlisted (#12785) 2026-09-11 08:37:20 +02:00
StylusFrostandpedrooot f9c02da90a feat(aws): support the ISO partitions for region resolution and scanning (#12759)
Co-authored-by: pedrooot <pedromarting3@gmail.com>
2026-09-10 17:15:13 +02:00
Pedro Martín 865eebe7fb fix(aws): configurable boto3 timeouts, 10s connect default (#12774) 2026-09-10 08:21:27 +02:00
Alejandro Bailo 8270979ec8 fix(ui): align scan filters and actions (#12781) 2026-09-09 23:05:09 +02:00
César Arrobaandpedrooot 369f852837 fix(aws): lead the partition bootstrap regions with the configured region (#12764)
Co-authored-by: pedrooot <pedromarting3@gmail.com>
2026-09-09 18:07:22 +02:00
César Arroba 1e8454a3cb fix(mcp): patch the six high libuuid CVEs in the container image (#12780) 2026-09-09 17:21:27 +02:00
César Arrobaandalejandrobailo 6f6ae88a66 fix(ui): patch the Next.js and sharp image-handling vulnerabilities (#12778)
Co-authored-by: alejandrobailo <alejandrobailo94@gmail.com>
2026-09-09 17:11:43 +02:00
César Arrobaandpedrooot c71f226e5c fix(image): honour TRIVY_CACHE_DIR when it is set (#12773)
Co-authored-by: pedrooot <pedromarting3@gmail.com>
2026-09-09 17:00:50 +02:00
Pedro Martín bbf5e1fa9f fix(tests): isolate secretsmanager policy test (#12782) 2026-09-09 16:58:35 +02:00
César Arroba 9cab9b8653 fix(ci): suppress grpc xDS DoS CVE from the Trivy binary (#12777) 2026-09-09 14:30:22 +02:00
César Arroba 806be2d061 chore(ci): bump agilepathway/label-checker to v1.6.66 (#12760) 2026-09-08 11:58:06 +02:00
Alejandro Bailo 2769cb9876 fix(ui): patch dependency vulnerabilities flagged by dependabot and pnpm audit (#12758) 2026-09-08 11:42:09 +02:00
César Arroba 623dc3125a chore(codeowners): consolidate retired teams under engineering (#12755) 2026-09-07 19:11:56 +02:00
mintlify[bot] a7cf5690e8 docs: fix typos, grammar, and formatting across docs 2026-09-07 07:12:39 +00:00
Pedro Martínandalejandrobailo 1edcf6e5de fix(jira): fix connection check timeout (#12742)
Co-authored-by: alejandrobailo <alejandrobailo94@gmail.com>
2026-09-04 15:51:09 +02:00
Pedro Martín 8bdb597921 perf(api): speed up compliance overview ingestion (#12738) 2026-09-04 11:13:25 +02:00
Pedro Martín 1746e1052b fix(ci): suppress unfixed x/crypto CVEs from Trivy binary (#12740) 2026-09-04 10:09:17 +02:00
Pedro Martín 6827eef347 fix(ci): don't fail setup-python-uv on empty grep match (#12737) 2026-09-04 09:12:15 +02:00
90fc815d3c chore(release): Bump versions to v5.42.0 (#12713)
Co-authored-by: prowler-bot <179230569+prowler-bot@users.noreply.github.com>
Co-authored-by: Pepe Fagoaga <pepe@prowler.com>
Co-authored-by: pedrooot <pedromarting3@gmail.com>
2026-09-03 14:23:50 +02:00
Pedro Martín 8f35fff255 fix(compliance): remove duplicate ids and stale check refs (#12717) 2026-09-03 13:39:12 +02:00
Pedro Martín ab51d09543 fix(tests): isolate mock class attrs leaking across tests (#12728) 2026-09-03 12:20:33 +02:00
Pedro Martín 12faeb9aa2 chore(trivy): suppress fast-uri CVEs from Teams SPDX manifest (#12727) 2026-09-03 11:09:13 +02:00
Alejandro Bailo 9ed07de610 feat(ui): separate PostHog hosts and enable Toolbar in development (#12582) 2026-09-03 10:36:47 +02:00
Alejandro Bailo 8007501574 fix(ui): avoid missing selector in scan tour (#12705) 2026-09-03 10:23:24 +02:00
Pedro Martín 36514534cb chore(trivy): suppress CVE-2026-84304 in embedded grpc (#12720) 2026-09-03 10:17:14 +02:00
Pedro Martín 9621bdfb9c fix(tests): isolate provider mock in agentcore passrole (#12724) 2026-09-03 10:15:49 +02:00
f013a5e1ad chore(changelog): v5.41.0 highlights (#12709)
Co-authored-by: Daniel Barranquero <danielbo2001@gmail.com>
Co-authored-by: Pepe Fagoaga <pepe@prowler.com>
2026-09-02 13:51:45 +02:00
Prowler Botandprowler-bot 18453e592e chore(changelog): v5.41.0 (#12707)
Co-authored-by: prowler-bot <179230569+prowler-bot@users.noreply.github.com>
2026-09-02 11:23:18 +02:00
Jonathan Nguyen 86e4408f29 feat(ecr): assess enhanced scanning on registries holding repositories (#12660) 2026-09-02 08:53:52 +02:00
Jonathan Nguyen ae43d21efb fix(sagemaker): read DirectInternetAccess instead of RootAccess on notebook instances (#12659) 2026-09-01 18:22:41 +02:00
Pedro Martín 7c84822fa3 fix(image): skip non-image OCI artifacts in registry scan (#12695) 2026-09-01 17:18:47 +02:00
Daniel Barranquero 51c5fa7168 fix(checks): report MANUAL instead of FAIL on permission and data-availability errors (#12645) 2026-09-01 17:16:56 +02:00
Jonathan Nguyen e9121f5f1a feat(cloudwatch): add agentcore log group data protection policy check (#12662) 2026-09-01 17:04:16 +02:00
Jonathan Nguyen 821fe43efd feat(iam): scope AgentCore PassRole and workload token grants, and flag unbound service trust (#12664) 2026-09-01 17:02:05 +02:00
Jonathan Nguyen 9ffbb4b758 fix(ecr): read each registry scanning rule's frequency instead of assuming scan on push (#12560) 2026-09-01 16:53:58 +02:00
Jonathan Nguyen 9c5285adc3 fix(cloudwatch): skip metric filters whose log group was not retrieved (#12561) 2026-09-01 14:00:41 +02:00
Pedro Martín f295d290dd feat(image): private network allowlist for SSRF guard (#12678) 2026-09-01 14:00:04 +02:00
Jonathan Nguyen e5df95c259 feat(eks): assess Kubernetes network policy enforcement in the Amazon VPC CNI add-on (#12661) 2026-09-01 13:40:45 +02:00
Jonathan Nguyen b6a8af3c54 feat(guardduty): assess unified Runtime Monitoring and AI Protection (#12564) 2026-09-01 13:18:57 +02:00
Pedro MartínandLydia Vilchez fb7064401b feat(compliance): add CIS 1.4 google workspace compliance (#12513)
Co-authored-by: Lydia Vilchez <lydiavilchezlopez@gmail.com>
2026-09-01 09:35:39 +02:00
Josema Camacho 8d60f9703a fix(api): limit mute rules to current and future scans (#12681) 2026-09-01 09:33:15 +02:00
Pablo Fernandez Guerra (PFE) ceb601028e fix(ui): apply Slack integration design feedback (#12677) 2026-08-31 18:27:11 +02:00
870 changed files with 66116 additions and 7378 deletions
+17 -1
View File
@@ -110,11 +110,27 @@ DJANGO_OUTPUT_S3_AWS_SECRET_ACCESS_KEY=""
DJANGO_OUTPUT_S3_AWS_SESSION_TOKEN=""
# The AWS region where your S3 bucket is located (e.g., "us-east-1")
# Required if the bucket uses SSE-KMS: download URLs are then signed with SigV4, which
# is scoped to this region, so it must match the bucket's
DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION=""
# The name of the S3 bucket where scan output should be stored
DJANGO_OUTPUT_S3_AWS_OUTPUT_BUCKET=""
# The storage endpoint the API and Celery workers use to upload and list scan output
# (e.g. "http://minio:9000"). Leave empty on AWS S3. Set it when scan output is stored on
# S3-compatible object storage such as MinIO instead of real S3.
# If set without DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL below, report download URLs are
# signed against this internal host, and a browser outside the container network cannot open them.
DJANGO_OUTPUT_S3_AWS_ENDPOINT_URL=""
# The storage address the browser can reach, used only to sign report download URLs
# (e.g. "https://storage.example.com"). Leave empty on AWS S3. Set it when storage is
# only reachable inside the container network, such as MinIO on "http://minio:9000".
# The reverse proxy in front of it must forward the Host header unchanged: SigV4 signs
# Host, so rewriting it to the internal name invalidates the signature.
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL=""
# Django settings
DJANGO_ALLOWED_HOSTS=localhost,127.0.0.1,prowler-api
DJANGO_BIND_ADDRESS=0.0.0.0
@@ -158,7 +174,7 @@ SENTRY_RELEASE=local
# REO_DEV_CLIENT_ID=
#### Prowler release version ####
NEXT_PUBLIC_PROWLER_RELEASE_VERSION=v5.41.0
NEXT_PUBLIC_PROWLER_RELEASE_VERSION=v5.44.0
# Social login credentials
SOCIAL_GOOGLE_OAUTH_CALLBACK_URL="${AUTH_URL}/api/auth/callback/google"
+13 -13
View File
@@ -1,23 +1,23 @@
# SDK
/* @prowler-cloud/detection-remediation
/prowler/ @prowler-cloud/detection-remediation
/tests/ @prowler-cloud/detection-remediation
/dashboard/ @prowler-cloud/detection-remediation
/docs/ @prowler-cloud/detection-remediation
/examples/ @prowler-cloud/detection-remediation
/util/ @prowler-cloud/detection-remediation
/contrib/ @prowler-cloud/detection-remediation
/permissions/ @prowler-cloud/detection-remediation
/codecov.yml @prowler-cloud/detection-remediation @prowler-cloud/api
/* @prowler-cloud/engineering
/prowler/ @prowler-cloud/engineering
/tests/ @prowler-cloud/engineering
/dashboard/ @prowler-cloud/engineering
/docs/ @prowler-cloud/engineering
/examples/ @prowler-cloud/engineering
/util/ @prowler-cloud/engineering
/contrib/ @prowler-cloud/engineering
/permissions/ @prowler-cloud/engineering
/codecov.yml @prowler-cloud/engineering
# API
/api/ @prowler-cloud/api
/api/ @prowler-cloud/engineering
# UI
/ui/ @prowler-cloud/ui
/ui/ @prowler-cloud/engineering
# AI
/mcp_server/ @prowler-cloud/detection-remediation
/mcp_server/ @prowler-cloud/engineering
# Platform
/.github/ @prowler-cloud/platform
@@ -46,6 +46,17 @@ runs:
env:
GITHUB_TOKEN: ${{ github.token }}
run: |
if grep -q "prowler-cloud/prowler" uv.lock; then
:
else
status=$?
if [ "$status" -ne 1 ]; then
echo "::error::grep failed reading uv.lock (exit code $status)."
exit "$status"
fi
echo "No prowler-cloud/prowler entry in uv.lock, nothing to update."
exit 0
fi
LATEST_COMMIT=$(curl -sf --retry 3 --retry-all-errors --retry-delay 2 --retry-max-time 60 \
-H "Authorization: Bearer ${GITHUB_TOKEN}" \
-H "Accept: application/vnd.github+json" \
@@ -66,6 +77,17 @@ runs:
env:
GITHUB_TOKEN: ${{ github.token }}
run: |
if grep -q "prowler-cloud/prowler" uv.lock; then
:
else
status=$?
if [ "$status" -ne 1 ]; then
echo "::error::grep failed reading uv.lock (exit code $status)."
exit "$status"
fi
echo "No prowler-cloud/prowler entry in uv.lock, nothing to update."
exit 0
fi
LATEST_COMMIT=$(curl -sf --retry 3 --retry-all-errors --retry-delay 2 --retry-max-time 60 \
-H "Authorization: Bearer ${GITHUB_TOKEN}" \
-H "Accept: application/vnd.github+json" \
+11
View File
@@ -451,6 +451,17 @@ modules:
e2e:
- ui/tests/home/**
- name: ui-registry
match:
- ui/actions/registry/**
- ui/app/**/registry/**
- ui/components/registry/**
- ui/lib/registry/**
- ui/tests/registry/**
tests: []
e2e:
- ui/tests/registry/**
- name: ui-shadcn
match:
- ui/components/shadcn/**
+1 -1
View File
@@ -39,7 +39,7 @@ jobs:
- name: Check labels
id: label_check
uses: agilepathway/label-checker@c3d16ad512e7cea5961df85ff2486bb774caf3c5 # v1.6.65
uses: agilepathway/label-checker@c324842522fbd012e4f590afe3b4e591301322ed # v1.6.66
with:
allow_failure: true
prefix_mode: true
@@ -44,7 +44,10 @@ jobs:
cache: 'pip'
- name: Install dependencies
run: pip install boto3
# Pinned to the versions in pyproject.toml: the ISO partitions region
# data comes from the endpoints.json bundled with botocore, so the
# botocore version is itself a data source and must be deterministic
run: pip install boto3==1.40.61 botocore==1.40.61
- name: Configure AWS credentials
uses: aws-actions/configure-aws-credentials@d979d5b3a71173a29b74b5b88418bfda9437d885 # v6.1.1
+45 -45
View File
@@ -10,12 +10,13 @@ on:
- master
- "v5.*"
paths:
- '.github/workflows/ui-e2e-tests-v2.yml'
- '.github/test-impact.yml'
- 'ui/**'
- 'api/**' # API changes can affect UI E2E
- '!ui/CHANGELOG.md'
- '!api/CHANGELOG.md'
- ".github/workflows/ui-e2e-tests-v2.yml"
- ".github/workflows/test-impact-analysis.yml"
- ".github/test-impact.yml"
- "ui/**"
- "api/**" # API changes can affect UI E2E
- "!ui/CHANGELOG.md"
- "!api/CHANGELOG.md"
concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }}
@@ -40,11 +41,11 @@ jobs:
(needs.impact-analysis.outputs.has-ui-e2e == 'true' || needs.impact-analysis.outputs.run-all == 'true')
runs-on: ubuntu-latest
env:
AUTH_SECRET: 'fallback-ci-secret-for-testing'
AUTH_SECRET: "fallback-ci-secret-for-testing"
AUTH_TRUST_HOST: true
NEXTAUTH_URL: 'http://localhost:3000'
AUTH_URL: 'http://localhost:3000'
UI_API_BASE_URL: 'http://localhost:8080/api/v1'
NEXTAUTH_URL: "http://localhost:3000"
AUTH_URL: "http://localhost:3000"
UI_API_BASE_URL: "http://localhost:8080/api/v1"
E2E_ADMIN_USER: ${{ secrets.E2E_ADMIN_USER }}
E2E_ADMIN_PASSWORD: ${{ secrets.E2E_ADMIN_PASSWORD }}
E2E_AWS_PROVIDER_ACCOUNT_ID: ${{ secrets.E2E_AWS_PROVIDER_ACCOUNT_ID }}
@@ -60,7 +61,7 @@ jobs:
E2E_M365_SECRET_ID: ${{ secrets.E2E_M365_SECRET_ID }}
E2E_M365_TENANT_ID: ${{ secrets.E2E_M365_TENANT_ID }}
E2E_M365_CERTIFICATE_CONTENT: ${{ secrets.E2E_M365_CERTIFICATE_CONTENT }}
E2E_KUBERNETES_CONTEXT: 'kind-kind'
E2E_KUBERNETES_CONTEXT: "kind-kind"
E2E_KUBERNETES_KUBECONFIG_PATH: /home/runner/.kube/config
E2E_GCP_BASE64_SERVICE_ACCOUNT_KEY: ${{ secrets.E2E_GCP_BASE64_SERVICE_ACCOUNT_KEY }}
E2E_GCP_PROJECT_ID: ${{ secrets.E2E_GCP_PROJECT_ID }}
@@ -237,8 +238,8 @@ jobs:
- name: Add AWS credentials for testing
run: |
echo "AWS_ACCESS_KEY_ID=${{ secrets.E2E_AWS_PROVIDER_ACCESS_KEY }}" >> .env
echo "AWS_SECRET_ACCESS_KEY=${{ secrets.E2E_AWS_PROVIDER_SECRET_KEY }}" >> .env
echo "AWS_ACCESS_KEY_ID=${E2E_AWS_PROVIDER_ACCESS_KEY}" >> .env
echo "AWS_SECRET_ACCESS_KEY=${E2E_AWS_PROVIDER_SECRET_KEY}" >> .env
- name: Build API image from current code
# docker-compose.yml references prowlercloud/prowler-api:latest from the registry,
@@ -292,7 +293,7 @@ jobs:
- name: Setup Node.js
uses: actions/setup-node@53b83947a5a98c8d113130e565377fae1a50d02f # v6.3.0
with:
node-version-file: 'ui/.nvmrc'
node-version-file: "ui/.nvmrc"
- name: Setup pnpm
uses: pnpm/action-setup@fc06bc1257f339d1d5d8b3a19a8cae5388b55320 # v5.0.0
@@ -337,60 +338,59 @@ jobs:
if: steps.playwright-cache.outputs.cache-hit != 'true'
run: pnpm run test:e2e:install
- name: Run E2E tests
- name: Run standard E2E tests
id: standard-e2e
working-directory: ./ui
run: |
if [[ "${RUN_ALL_TESTS}" == "true" ]]; then
echo "Running ALL E2E tests..."
echo "Running all standard E2E tests..."
pnpm run test:e2e
else
echo "Running targeted E2E tests: ${E2E_TEST_PATHS}"
# Convert glob patterns to playwright test paths
# e.g., "ui/tests/providers/**" -> "tests/providers"
echo "Running targeted standard E2E tests: ${E2E_TEST_PATHS}"
TEST_PATHS="${E2E_TEST_PATHS}"
# Remove ui/ prefix and convert ** to empty (playwright handles recursion)
TEST_PATHS=$(echo "$TEST_PATHS" | sed 's|ui/||g' | sed 's|\*\*||g' | tr ' ' '\n' | sort -u)
# Drop auth setup helpers (not runnable test suites)
TEST_PATHS=$(echo "$TEST_PATHS" | grep -v '^tests/setups/')
# Safety net: if bare "tests/" appears (from broad patterns like ui/tests/**),
# expand to specific subdirs to avoid Playwright discovering setup files
TEST_PATHS=$(echo "$TEST_PATHS" | grep -vE '^tests/(setups|registry)/' || true)
if echo "$TEST_PATHS" | grep -qx 'tests/'; then
echo "Expanding bare 'tests/' to specific subdirs (excluding setups)..."
SPECIFIC_DIRS=""
for dir in tests/*/; do
[[ "$dir" == "tests/setups/" ]] && continue
[[ "$dir" == "tests/setups/" || "$dir" == "tests/registry/" ]] && continue
SPECIFIC_DIRS="${SPECIFIC_DIRS}${dir}"$'\n'
done
# Replace "tests/" with specific dirs, keep other paths
TEST_PATHS=$(echo "$TEST_PATHS" | grep -vx 'tests/')
TEST_PATHS=$(echo "$TEST_PATHS" | grep -vx 'tests/' || true)
TEST_PATHS="${TEST_PATHS}"$'\n'"${SPECIFIC_DIRS}"
TEST_PATHS=$(echo "$TEST_PATHS" | grep -v '^$' | sort -u)
fi
if [[ -z "$TEST_PATHS" ]]; then
echo "No runnable E2E test paths after filtering setups"
exit 0
fi
# Filter out directories that don't contain any test files
VALID_PATHS=""
while IFS= read -r p; do
[[ -z "$p" ]] && continue
if find "$p" -name '*.spec.ts' -o -name '*.test.ts' 2>/dev/null | head -1 | grep -q .; then
VALID_PATHS="${VALID_PATHS}${p}"$'\n'
while IFS= read -r path; do
[[ -z "$path" ]] && continue
if find "$path" -name '*.spec.ts' -o -name '*.test.ts' 2>/dev/null | head -1 | grep -q .; then
VALID_PATHS="${VALID_PATHS}${path}"$'\n'
else
echo "Skipping empty test directory: $p"
echo "Skipping empty test directory: $path"
fi
done <<< "$TEST_PATHS"
VALID_PATHS=$(echo "$VALID_PATHS" | grep -v '^$' || true)
if [[ -z "$VALID_PATHS" ]]; then
echo "No test files found in any resolved paths — skipping E2E"
exit 0
if [[ -n "$VALID_PATHS" ]]; then
TEST_PATHS=$(echo "$VALID_PATHS" | tr '\n' ' ')
echo "Resolved standard test paths: $TEST_PATHS"
read -ra test_paths <<< "$TEST_PATHS"
pnpm exec playwright test "${test_paths[@]}"
else
echo "No standard E2E test paths selected."
fi
TEST_PATHS=$(echo "$VALID_PATHS" | tr '\n' ' ')
echo "Resolved test paths: $TEST_PATHS"
read -ra test_paths <<< "$TEST_PATHS"
pnpm exec playwright test "${test_paths[@]}"
fi
- name: Run Registry fixture E2E tests
if: |
!cancelled() &&
(steps.standard-e2e.outcome == 'success' || steps.standard-e2e.outcome == 'failure') &&
(env.RUN_ALL_TESTS == 'true' || contains(format(' {0} ', env.E2E_TEST_PATHS), ' ui/tests/registry/'))
working-directory: ./ui
run: pnpm run test:e2e:registry
- name: Upload test reports
uses: actions/upload-artifact@bbbca2ddaa5d8feaa63e36b76fdaad77386f024f # v7.0.0
if: failure()
+45
View File
@@ -17,6 +17,40 @@ ignore:
- vulnerability: CVE-2026-71556
package:
name: github.com/go-git/go-git/v5
# CVE-2026-84304 is the same temporary exception documented in .trivyignore.yaml:
# Trivy 0.74.0 still embeds grpc 1.82.1, while the 1.83.1 fix is not in any release.
# Prowler only runs `trivy image` / `trivy fs`, never client/server mode, so no gRPC
# endpoint exists in the image. Pinned to the embedded version so the rule stops
# matching on its own once Trivy bumps grpc. Remove with the Trivy exception by 2026-10-15.
# https://github.com/aquasecurity/trivy/pull/11176
- vulnerability: CVE-2026-84304
package:
name: google.golang.org/grpc
version: v1.82.1
# CVE-2026-84445 is the same temporary exception documented in .trivyignore.yaml:
# Trivy 0.74.0 still embeds grpc 1.82.1, while the 1.82.2 / 1.83.2 fix is not in any
# release. The panic needs a gRPC server built with `xds.NewGRPCServer()`; Prowler only
# runs `trivy image` / `trivy fs`, so the image serves no gRPC at all. Pinned to the
# embedded version so the rule stops matching on its own once Trivy bumps grpc. Remove
# with the Trivy exception by 2026-10-15.
# https://github.com/advisories/GHSA-2v4p-qf9q-27wj
- vulnerability: CVE-2026-84445
package:
name: google.golang.org/grpc
version: v1.82.1
# CVE-2026-56855 / CVE-2026-78662 are the same temporary exception documented in
# .trivyignore.yaml: Trivy 0.74.0 still embeds golang.org/x/crypto v0.55.0, while the
# 0.56.0 fix (published 2026-09-02) hasn't reached any Trivy release, or even Trivy
# main, yet. Pinned to the embedded version so the rule stops matching on its own once
# Trivy bumps it. Remove with the Trivy exception by 2026-10-15.
- vulnerability: CVE-2026-56855
package:
name: golang.org/x/crypto
version: v0.55.0
- vulnerability: CVE-2026-78662
package:
name: golang.org/x/crypto
version: v0.55.0
- vulnerability: CVE-2026-56852
package:
name: golang.org/x/text
@@ -81,3 +115,14 @@ ignore:
- vulnerability: CVE-2026-9669
package:
name: python
# CVE-2026-82049 (tarfile data/tar filter bypass via a hard link to a symlink) has no
# fixed CPython release on any branch: the fix is merged on main and 3.13 only, and the
# 3.12 backport is still open. Grype records 3.14.0b1 as the fix, so only-fixed does not
# drop it, yet python:3.12.14-slim-trixie reports it too. Prowler never extracts tar
# archives to disk: the ECR image inspection reads members in memory with extractfile().
# Remove once the base image ships a 3.12 release that includes the backport.
# https://github.com/python/cpython/issues/157190
# https://github.com/python/cpython/pull/157454
- vulnerability: CVE-2026-82049
package:
name: python
+68 -26
View File
@@ -113,40 +113,82 @@ vulnerabilities:
purls:
- "pkg:npm/fast-uri"
expired_at: 2027-01-31
- id: CVE-2026-75899
purls:
- "pkg:npm/fast-uri"
expired_at: 2027-01-31
- id: CVE-2026-75975
purls:
- "pkg:npm/fast-uri"
expired_at: 2027-01-31
- id: CVE-2026-76172
purls:
- "pkg:npm/fast-uri"
expired_at: 2027-01-31
- id: CVE-2026-75931
purls:
- "pkg:npm/fast-uri"
expired_at: 2027-01-31
- id: CVE-2026-69192
purls:
- "pkg:npm/ip-address"
expired_at: 2027-01-31
# CVE-2026-62901 is a DoS in System.Net.WebSockets (unchecked input for loop condition,
# CWE-606), fixed in .NET 9.0.19 / 10.0.11 (published 2026-08-11). The vulnerable runtime
# ships inside the PowerShell tarball the Dockerfile pins: 7.5.9 is the latest 7.5.x and
# bundles .NET 9.0.18; 7.6.4 bundles .NET 10.0.x < 10.0.11, so no published PowerShell
# release contains the fix yet. Prowler only invokes pwsh locally to run M365 module
# cmdlets; the image does not accept inbound WebSocket connections, so the DoS path is
# not reachable from the network. Remove this temporary suppression as soon as a
# PowerShell release shipping .NET 9.0.19+ is available.
- id: CVE-2026-62901
# CVE-2026-84304 is a DoS in grpc-go <= 1.83.0: a peer fragments a gRPC stream into
# millions of tiny HTTP/2 DATA frames until the receiver runs out of heap. Fixed in
# 1.83.1 (published 2026-09-01). Trivy 0.74.0, the latest published release and the
# version the images ship, pins 1.82.1 as an indirect dependency:
# https://github.com/aquasecurity/trivy/blob/v0.74.0/go.mod
# Upstream bump still open: https://github.com/aquasecurity/trivy/pull/11176
# Trivy only speaks gRPC in client/server mode (`trivy server`, `--server`). Prowler
# invokes it exclusively as `trivy image` and `trivy fs` on a local path, so no gRPC
# listener or connection ever exists in the image and the affected path is not
# reachable. Remove this temporary suppression as soon as a Trivy release pins
# grpc >= 1.83.1.
- id: CVE-2026-84304
purls:
- "pkg:nuget/Microsoft.NETCore.App.Runtime.linux-x64"
- "pkg:nuget/Microsoft.NETCore.App.Runtime.linux-arm64"
expired_at: 2026-09-15
- "pkg:golang/google.golang.org/grpc"
expired_at: 2026-10-15
# Modules compiled into the Trivy binary the images ship. The binary is pinned by version
# and verified by checksum in the Dockerfile; only a rebuild by its vendor moves these.
# CVE-2026-71556 affects go-git worktree operations that can follow symlinks outside a
# cloned repository. Trivy 0.73.0, the latest published release and the version the
# images ship, still pins that vulnerable version:
# https://github.com/aquasecurity/trivy/blob/v0.73.0/go.mod#L46
# Trivy main already contains the 5.19.2 fix, but no published release includes it yet:
# https://github.com/aquasecurity/trivy/commit/a2edba9a03987ba0d2ebc8212c1a9a1e6979497b
# Prowler invokes Trivy only with `fs` on an existing local path or with `image`; it does
# not ask Trivy to clone or mutate a Git worktree, so the affected path is not reachable.
# Remove this temporary suppression as soon as a fixed Trivy release is available.
- id: CVE-2026-71556
# CVE-2026-84445 is a DoS in grpc-go servers built with `xds.NewGRPCServer()`: a request
# carrying neither `:authority` nor `Host` reaches the xDS routing interceptor, which
# indexes an empty slice of authorities and panics. The per-RPC goroutine does not
# recover, so the whole server process dies. Fixed in 1.82.2 and 1.83.2 (published
# 2026-09-08). Trivy 0.74.0, the latest published release and the version the images
# ship, pins 1.82.1 as an indirect dependency:
# https://github.com/aquasecurity/trivy/blob/v0.74.0/go.mod
# Trivy main already carries 1.83.2, but no published release includes it yet.
# The reachability argument is the one made for CVE-2026-84304 above, only narrower:
# this panic needs an xDS-managed gRPC server. Prowler invokes Trivy exclusively as
# `trivy image` and `trivy fs` on a local path, never `trivy server`, so the image runs
# no gRPC server at all, xDS or otherwise. Remove this temporary suppression as soon as
# a Trivy release pins grpc >= 1.83.2.
# https://github.com/advisories/GHSA-2v4p-qf9q-27wj
- id: CVE-2026-84445
purls:
- "pkg:golang/github.com/go-git/go-git/v5"
expired_at: 2026-09-15
- "pkg:golang/google.golang.org/grpc@v1.82.1"
expired_at: 2026-10-15
# CVE-2026-56855 and CVE-2026-78662 are DoS deadlocks in x/crypto/ssh: a malicious peer
# can flood or misuse channel messages (RFC 4254) to block the whole connection.
# Fixed in golang.org/x/crypto v0.56.0 (published 2026-09-02). Trivy 0.74.0, the latest
# published release and the version the images ship, still pins v0.55.0, and Trivy main
# has not bumped it either:
# https://github.com/aquasecurity/trivy/blob/v0.74.0/go.mod
# x/crypto/ssh is pulled in transitively through go-git's ssh transport, the same
# dependency chain as the CVE-2026-71556 entry above. Prowler invokes Trivy only with
# `fs` on an existing local path or with `image`; it never asks Trivy to clone over SSH
# or to run `trivy server`, so no SSH connection -- as client or server -- ever exists in
# the image and the affected code path is not reachable. Remove this temporary
# suppression as soon as a fixed Trivy release is available.
- id: CVE-2026-56855
purls:
- "pkg:golang/golang.org/x/crypto@v0.55.0"
expired_at: 2026-10-15
- id: CVE-2026-78662
purls:
- "pkg:golang/golang.org/x/crypto@v0.55.0"
expired_at: 2026-10-15
- id: CVE-2026-56852
purls:
+13 -8
View File
@@ -3,7 +3,7 @@ FROM python:3.12.13-slim-trixie@sha256:57cd7c3a7a273101a6485ba99423ee56815788280
LABEL maintainer="https://github.com/prowler-cloud/prowler"
LABEL org.opencontainers.image.source="https://github.com/prowler-cloud/prowler"
ARG POWERSHELL_VERSION=7.5.9
ARG POWERSHELL_VERSION=7.5.11
ENV POWERSHELL_VERSION=${POWERSHELL_VERSION}
# Opt out of PowerShell telemetry (Application Insights -> dc.services.visualstudio.com)
ENV POWERSHELL_TELEMETRY_OPTOUT=1
@@ -17,25 +17,30 @@ ENV ZIZMOR_VERSION=${ZIZMOR_VERSION}
# Pinned here, not fetched with the artefact: a compromised release ships its own checksum.
ARG TRIVY_SHA256_AMD64=2ae6fe3ee734b7fdf11335663e18c75ea12dccc76062f09f164a3b0f8be4371a
ARG TRIVY_SHA256_ARM64=b94ce1976bbf3c15b514b605ee88be7c6d94a29be2302847ff01cb794d47aad5
ARG POWERSHELL_SHA256_AMD64=492ff26bb958336bf61e597ce19e07648b4003bd2a08659e02f0e3e0446ebfe0
ARG POWERSHELL_SHA256_ARM64=2503b71da3e83635592b092df59a0aca4c3606b4d9b068217bb00be989cb0d56
ARG POWERSHELL_SHA256_AMD64=82a8b13d92b0f3ae48e56cf2f3f7961679371736ca90145ca71617c2913ba9d8
ARG POWERSHELL_SHA256_ARM64=830ebda118c731ece3fa7e6b7e8573a21346387cbbca5b2f5e3b9bfe24f96672
ARG ZIZMOR_SHA256_AMD64=a8000f3c683319a523d3b20df0e75457ba591f049cfcbfa98966631b56733c03
ARG ZIZMOR_SHA256_ARM64=d66e37ef8a375fb07939c630ebf9709a6e0f20242bdc3faf672a7ed97e0b768d
# High CVEs fixed in Debian trixie-security but not yet in the pinned base image:
# High CVEs fixed in Debian trixie but not yet in the pinned base image:
# openssl/libssl3t64/openssl-provider-legacy 3.5.7-1~deb13u2 CVE-2026-14456,
# -14457, -18798, -54874, -63072, -63073, -63074, -63075, -63076, -75803
# (image ships 3.5.6-1~deb13u2)
# libsqlite3-0 3.46.1-7+deb13u2 CVE-2026-11822, -11824
# gzip 1.13-1+deb13u1 CVE-2026-41992
# perl-base 5.40.1-6+deb13u1 CVE-2026-42497, -48962, -57432
# libssh2-1t64 1.11.1-1+deb13u2 CVE-2026-58050
# libpcre2-8-0 10.46-1~deb13u2 CVE-2026-86145, -89161
# Taken as a targeted --only-upgrade rather than by moving the digest: the newest
# published python:3.12-slim-trixie carries the same vulnerable version. The three
# packages are all built from openssl and are flagged separately, so all are named.
# Drop them once the base image ships 3.5.7-1~deb13u2 or later.
# published python:3.12-slim-trixie carries the same vulnerable versions. The three
# openssl packages are flagged separately, so all are named.
# Drop each one once the base image ships its fixed version.
# hadolint ignore=DL3008
RUN apt-get update && apt-get install -y --no-install-recommends \
wget libicu76 libunwind8 libssl3 libcurl4 ca-certificates apt-transport-https gnupg \
build-essential pkg-config libzstd-dev zlib1g-dev \
&& apt-get install -y --no-install-recommends --only-upgrade \
util-linux libssl3t64 openssl openssl-provider-legacy \
libsqlite3-0 gzip perl-base libssh2-1t64 libpcre2-8-0 \
&& rm -rf /var/lib/apt/lists/*
# Install PowerShell
+6 -6
View File
@@ -126,12 +126,12 @@ Every AWS provider scan will enqueue an Attack Paths ingestion job automatically
| Provider | Checks | Services | [Compliance Frameworks](https://docs.prowler.com/user-guide/compliance/tutorials/compliance) | [Categories](https://docs.prowler.com/user-guide/cli/tutorials/misc#categories) | Support | Interface |
|---|---|---|---|---|---|---|
| AWS | 639 | 86 | 47 | 19 | Official | UI, API, CLI |
| Azure | 191 | 22 | 21 | 16 | Official | UI, API, CLI |
| GCP | 109 | 20 | 19 | 12 | Official | UI, API, CLI |
| Kubernetes | 92 | 7 | 8 | 11 | Official | UI, API, CLI |
| AWS | 662 | 86 | 50 | 19 | Official | UI, API, CLI |
| Azure | 191 | 22 | 25 | 16 | Official | UI, API, CLI |
| GCP | 110 | 20 | 22 | 12 | Official | UI, API, CLI |
| Kubernetes | 92 | 7 | 11 | 11 | Official | UI, API, CLI |
| GitHub | 24 | 3 | 2 | 5 | Official | UI, API, CLI |
| M365 | 143 | 10 | 6 | 10 | Official | UI, API, CLI |
| M365 | 144 | 10 | 9 | 10 | Official | UI, API, CLI |
| OCI | 52 | 14 | 5 | 10 | Official | UI, API, CLI |
| Alibaba Cloud | 63 | 9 | 6 | 9 | Official | UI, API, CLI |
| Cloudflare | 29 | 3 | 2 | 5 | Official | UI, API, CLI |
@@ -139,7 +139,7 @@ Every AWS provider scan will enqueue an Attack Paths ingestion job automatically
| MongoDB Atlas | 10 | 3 | 1 | 8 | Official | UI, API, CLI |
| LLM | [See `promptfoo` docs.](https://www.promptfoo.dev/docs/red-team/plugins/) | N/A | N/A | N/A | Official | CLI |
| Image | N/A | N/A | N/A | N/A | Official | UI, API, CLI |
| Google Workspace | 65 | 11 | 3 | 6 | Official | UI, API, CLI |
| Google Workspace | 65 | 11 | 4 | 6 | Official | UI, API, CLI |
| OpenStack | 34 | 5 | 1 | 9 | Official | UI, API, CLI |
| Vercel | 26 | 6 | 1 | 8 | Official | UI, API, CLI |
| Okta | 29 | 8 | 2 | 2 | Official | UI, API, CLI |
+36
View File
@@ -4,6 +4,42 @@ All notable changes to the **Prowler API** are documented in this file.
<!-- changelog: release notes start -->
## [1.44.0] (Prowler v5.43.0)
### 🐞 Fixed
- Report download URLs can be signed against a browser-reachable storage host via `DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL`, so downloads complete on deployments where storage is only reachable inside the container network [(#12552)](https://github.com/prowler-cloud/prowler/pull/12552)
- A scan report download no longer fails with a server error when `DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION` is unset, which is common on storage with no meaningful region [(#12552)](https://github.com/prowler-cloud/prowler/pull/12552)
- Lapsed pending invitations are reported as expired and no longer block a new invitation for the same email [(#12831)](https://github.com/prowler-cloud/prowler/pull/12831)
### 🔐 Security
- `libsqlite3-0`, `gzip`, `perl-base` and `libpcre2-8-0` upgraded in the API container image, patching high Debian CVEs [(#12804)](https://github.com/prowler-cloud/prowler/pull/12804)
- PowerShell from 7.5.9 to 7.5.11 in the API container image, bundling .NET runtime 9.0.20 and patching CVE-2026-62901 [(#12811)](https://github.com/prowler-cloud/prowler/pull/12811)
- Bumped `anyio` to 4.14.2 to resolve CVE-2026-63374 [(#12848)](https://github.com/prowler-cloud/prowler/pull/12848)
---
## [1.43.0] (Prowler v5.42.0)
### 🔄 Changed
- Speed up compliance overview ingestion by reading ThreatScore mappings from the compliance template instead of each finding, generating time-ordered `uuid7` row ids and grouping inserted rows by framework and requirement [(#12738)](https://github.com/prowler-cloud/prowler/pull/12738)
---
## [1.42.0] (Prowler v5.41.0)
### 🚀 Added
- Jira issues created from Prowler Cloud now carry the `prowler`, `prowler-<provider>`, `prowler-<severity>`, `prowler-<check-id>`, and `prowler-finding-<finding-uid>` labels, a link back to the finding when `DJANGO_UI_BASE_URL` is configured, and the tenant name [(#12540)](https://github.com/prowler-cloud/prowler/pull/12540)
### 🐞 Fixed
- `POST /api/v1/mute-rules` now updates only each affected provider's latest completed scan and future scans, preventing historical reaggregation from flooding Celery queues [(#12681)](https://github.com/prowler-cloud/prowler/pull/12681)
---
## [1.41.0] (Prowler v5.40.0)
### 🐞 Fixed
+13 -8
View File
@@ -2,7 +2,7 @@ FROM python:3.12.13-slim-trixie@sha256:57cd7c3a7a273101a6485ba99423ee56815788280
LABEL maintainer="https://github.com/prowler-cloud/api"
ARG POWERSHELL_VERSION=7.5.9
ARG POWERSHELL_VERSION=7.5.11
ENV POWERSHELL_VERSION=${POWERSHELL_VERSION}
# Opt out of PowerShell telemetry (Application Insights -> dc.services.visualstudio.com)
ENV POWERSHELL_TELEMETRY_OPTOUT=1
@@ -16,19 +16,23 @@ ENV ZIZMOR_VERSION=${ZIZMOR_VERSION}
# Pinned here, not fetched with the artefact: a compromised release ships its own checksum.
ARG TRIVY_SHA256_AMD64=2ae6fe3ee734b7fdf11335663e18c75ea12dccc76062f09f164a3b0f8be4371a
ARG TRIVY_SHA256_ARM64=b94ce1976bbf3c15b514b605ee88be7c6d94a29be2302847ff01cb794d47aad5
ARG POWERSHELL_SHA256_AMD64=492ff26bb958336bf61e597ce19e07648b4003bd2a08659e02f0e3e0446ebfe0
ARG POWERSHELL_SHA256_ARM64=2503b71da3e83635592b092df59a0aca4c3606b4d9b068217bb00be989cb0d56
ARG POWERSHELL_SHA256_AMD64=82a8b13d92b0f3ae48e56cf2f3f7961679371736ca90145ca71617c2913ba9d8
ARG POWERSHELL_SHA256_ARM64=830ebda118c731ece3fa7e6b7e8573a21346387cbbca5b2f5e3b9bfe24f96672
ARG ZIZMOR_SHA256_AMD64=a8000f3c683319a523d3b20df0e75457ba591f049cfcbfa98966631b56733c03
ARG ZIZMOR_SHA256_ARM64=d66e37ef8a375fb07939c630ebf9709a6e0f20242bdc3faf672a7ed97e0b768d
# High CVEs fixed in Debian trixie-security but not yet in the pinned base image:
# High CVEs fixed in Debian trixie but not yet in the pinned base image:
# openssl/libssl3t64/openssl-provider-legacy 3.5.7-1~deb13u2 CVE-2026-14456,
# -14457, -18798, -54874, -63072, -63073, -63074, -63075, -63076, -75803
# (image ships 3.5.6-1~deb13u2)
# libsqlite3-0 3.46.1-7+deb13u2 CVE-2026-11822, -11824
# gzip 1.13-1+deb13u1 CVE-2026-41992
# perl-base 5.40.1-6+deb13u1 CVE-2026-42497, -48962, -57432
# libssh2-1t64 1.11.1-1+deb13u2 CVE-2026-58050
# libpcre2-8-0 10.46-1~deb13u2 CVE-2026-86145, -89161
# Taken as a targeted --only-upgrade rather than by moving the digest: the newest
# published python:3.12-slim-trixie carries the same vulnerable version. The three
# packages are all built from openssl and are flagged separately, so all are named.
# Drop them once the base image ships 3.5.7-1~deb13u2 or later.
# published python:3.12-slim-trixie carries the same vulnerable versions. The three
# openssl packages are flagged separately, so all are named.
# Drop each one once the base image ships its fixed version.
# hadolint ignore=DL3008
RUN apt-get update && apt-get install -y --no-install-recommends \
wget \
@@ -46,6 +50,7 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
python3-dev \
&& apt-get install -y --no-install-recommends --only-upgrade \
util-linux libssl3t64 openssl openssl-provider-legacy \
libsqlite3-0 gzip perl-base libssh2-1t64 libpcre2-8-0 \
&& rm -rf /var/lib/apt/lists/*
# Install PowerShell
@@ -0,0 +1 @@
Adds a periodic sweep that drops orphaned Attack Paths temp Neo4j scan databases left behind when a worker or Neo4j crashes mid-scan, before they accumulate unbounded
@@ -0,0 +1 @@
Resources no longer keep a stale failed findings count forever when a scoped or imported scan for the same provider completes after a full scan, which used to make the full scan skip its own cleanup
@@ -1 +0,0 @@
Jira issues created from Prowler Cloud now carry the `prowler`, `prowler-<provider>`, `prowler-<severity>`, `prowler-<check-id>`, and `prowler-finding-<finding-uid>` labels, a link back to the finding when `DJANGO_UI_BASE_URL` is configured, and the tenant name
@@ -0,0 +1 @@
Providers whose most recent completed scan has no `completed_at` timestamp are no longer missing from every endpoint that reports a provider's latest scan, which now falls back to scan creation order instead of skipping the provider
@@ -0,0 +1 @@
Unify how every endpoint resolves a provider latest completed scan, so overlapping scans no longer make findings, compliance and mute rules read from different scans
@@ -0,0 +1 @@
Scan output uploads and downloads can now target S3-compatible object storage such as MinIO directly via `DJANGO_OUTPUT_S3_AWS_ENDPOINT_URL`, instead of relying on process-wide AWS environment variables that also hijacked unrelated AWS API calls
@@ -0,0 +1 @@
Scan report downloads from an S3 bucket with default SSE-KMS encryption no longer fail with an `InvalidArgument` error: when `DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION` is set, presigned download URLs are signed with AWS Signature Version 4 for that region
@@ -0,0 +1 @@
`POST /api/v1/scans` again returns the new scan id in the response `task_args`, which had been empty since the scan broker publish moved to transaction commit
@@ -0,0 +1 @@
Celery loggers are now declared explicitly in `custom_logging.py` so fatal worker errors are no longer silenced by `disable_existing_loggers=True`. All long-running services in `docker-compose.yml` now have `restart: unless-stopped` so containers recover automatically after unexpected crashes.
+2 -2
View File
@@ -71,7 +71,7 @@ name = "prowler-api"
package-mode = false
# Needed for the SDK compatibility
requires-python = ">=3.11,<3.13"
version = "1.42.0"
version = "1.45.0"
# Shared ruff baseline (kept in sync with mcp_server/pyproject.toml).
# target-version tracks this project's lowest supported Python.
@@ -137,7 +137,7 @@ constraint-dependencies = [
"aliyun-log-fastpb==0.2.0",
"amqp==5.3.1",
"annotated-types==0.7.0",
"anyio==4.12.1",
"anyio==4.14.2",
"applicationinsights==0.11.10",
"apscheduler==3.11.2",
"argcomplete==3.5.3",
@@ -207,6 +207,11 @@ def drop_database(database: str) -> None:
sink_module.get_backend().drop_database(database)
def list_databases() -> list[str]:
"""List database names on the ingest cluster. Temp scan DBs always live here."""
return ingest.list_databases()
def drop_subgraph(database: str, provider_id: str) -> int:
return sink_module.get_backend().drop_subgraph(database, provider_id)
@@ -13,6 +13,7 @@ from api.attack_paths.ingest.driver import (
get_session,
get_uri,
init_driver,
list_databases,
run_cypher,
)
@@ -25,5 +26,6 @@ __all__ = [
"get_session",
"get_uri",
"init_driver",
"list_databases",
"run_cypher",
]
@@ -165,6 +165,14 @@ def drop_database(database: str) -> None:
session.run(f"DROP DATABASE `{database}` IF EXISTS DESTROY DATA")
def list_databases() -> list[str]:
"""List every database name on the Neo4j temp-database cluster."""
# A cluster returns one row per hosting server, so dedupe on name
with get_session() as session:
result = session.run("SHOW DATABASES YIELD name RETURN DISTINCT name")
return [record["name"] for record in result]
def clear_cache(database: str) -> None:
"""Best-effort cache clear for a Neo4j database."""
from api.attack_paths.database import GraphDatabaseQueryException
+1 -1
View File
@@ -409,7 +409,7 @@ def batch_delete(tenant_id, queryset, batch_size=settings.DJANGO_DELETION_BATCH_
Args:
tenant_id (str): Tenant ID the queryset belongs to.
queryset (QuerySet): The queryset of objects to delete.
queryset: The queryset of objects to delete.
batch_size (int): The number of objects to delete in each batch.
Returns:
+14 -2
View File
@@ -1439,8 +1439,20 @@ class InvitationFilter(FilterSet):
inserted_at = DateFilter(field_name="inserted_at", lookup_expr="date")
updated_at = DateFilter(field_name="updated_at", lookup_expr="date")
expires_at = DateFilter(field_name="expires_at", lookup_expr="date")
state = ChoiceFilter(choices=Invitation.State.choices)
state__in = ChoiceInFilter(choices=Invitation.State.choices, lookup_expr="in")
state = ChoiceFilter(choices=Invitation.State.choices, method="filter_state")
state__in = ChoiceInFilter(
choices=Invitation.State.choices, lookup_expr="in", method="filter_state_in"
)
def filter_state(self, queryset, name, value):
return self.filter_state_in(queryset, name, [value])
def filter_state_in(self, queryset, name, value):
lapsed = Invitation.lapsed_q()
query = Q(state__in=value) & ~lapsed
if Invitation.State.EXPIRED in value:
query |= lapsed
return queryset.filter(query)
class Meta:
model = Invitation
@@ -0,0 +1,116 @@
import uuid
import api.rls
import django.db.models.deletion
from django.conf import settings
from django.db import migrations, models
class Migration(migrations.Migration):
dependencies = [
("api", "0097_attack_paths_scan_db_defaults"),
migrations.swappable_dependency(settings.AUTH_USER_MODEL),
]
operations = [
migrations.CreateModel(
name="TenantOnboardingProfile",
fields=[
(
"id",
models.UUIDField(
default=uuid.uuid4,
editable=False,
primary_key=True,
serialize=False,
),
),
("inserted_at", models.DateTimeField(auto_now_add=True)),
(
"declared_cloud_accounts",
models.CharField(
blank=True,
choices=[
("1", "1"),
("2-10", "2-10"),
("11-50", "11-50"),
("51-200", "51-200"),
("200+", "200+"),
],
max_length=16,
null=True,
),
),
(
"declared_role",
models.CharField(
blank=True,
choices=[
("security", "Security"),
("devops_platform", "DevOps / Platform"),
("developer", "Developer"),
("compliance_grc", "Compliance / GRC"),
("other", "Other"),
],
max_length=32,
null=True,
),
),
(
"declared_seniority",
models.CharField(
blank=True,
choices=[
("practitioner", "Practitioner / IC"),
("lead", "Team lead / Manager"),
("director", "Director / Head of"),
("executive", "VP / C-level"),
("founder", "Founder / Owner"),
],
max_length=32,
null=True,
),
),
("skipped", models.BooleanField(default=False)),
(
"submitted_by",
models.ForeignKey(
blank=True,
null=True,
on_delete=django.db.models.deletion.SET_NULL,
related_name="tenant_onboarding_profiles",
related_query_name="tenant_onboarding_profile",
to=settings.AUTH_USER_MODEL,
),
),
(
"tenant",
models.ForeignKey(
on_delete=django.db.models.deletion.CASCADE, to="api.tenant"
),
),
],
options={
"db_table": "tenant_onboarding_profiles",
"abstract": False,
},
),
migrations.AddConstraint(
model_name="tenantonboardingprofile",
constraint=models.UniqueConstraint(
fields=("tenant_id",), name="unique_tenant_onboarding_profile"
),
),
migrations.AddConstraint(
model_name="tenantonboardingprofile",
# `statements` written out explicitly: RowLevelSecurityConstraint
# .deconstruct() does not serialize it, so an autogenerated
# migration falls back to ["SELECT"] and leaves the table without
# INSERT/UPDATE/DELETE policies.
constraint=api.rls.RowLevelSecurityConstraint(
"tenant_id",
name="rls_on_tenantonboardingprofile",
statements=["SELECT", "INSERT", "UPDATE", "DELETE"],
),
),
]
@@ -0,0 +1,15 @@
from django.db import migrations
class Migration(migrations.Migration):
# The onboarding profile step was reverted after 0098 had been merged, so
# the table goes away through a new migration rather than by deleting 0098.
dependencies = [
("api", "0098_tenant_onboarding_profile"),
]
operations = [
migrations.DeleteModel(
name="TenantOnboardingProfile",
),
]
@@ -0,0 +1,48 @@
from django.db import migrations
TASK_NAME = "attack-paths-reap-orphaned-tmp-databases"
INTERVAL_HOURS = 6
def create_periodic_task(apps, schema_editor):
IntervalSchedule = apps.get_model("django_celery_beat", "IntervalSchedule")
PeriodicTask = apps.get_model("django_celery_beat", "PeriodicTask")
schedule, _ = IntervalSchedule.objects.get_or_create(
every=INTERVAL_HOURS,
period="hours",
)
PeriodicTask.objects.update_or_create(
name=TASK_NAME,
defaults={
"task": TASK_NAME,
"interval": schedule,
"enabled": True,
},
)
def delete_periodic_task(apps, schema_editor):
IntervalSchedule = apps.get_model("django_celery_beat", "IntervalSchedule")
PeriodicTask = apps.get_model("django_celery_beat", "PeriodicTask")
PeriodicTask.objects.filter(name=TASK_NAME).delete()
# Clean up the schedule if no other task references it
IntervalSchedule.objects.filter(
every=INTERVAL_HOURS,
period="hours",
periodictask__isnull=True,
).delete()
class Migration(migrations.Migration):
dependencies = [
("api", "0099_delete_tenant_onboarding_profile"),
("django_celery_beat", "0019_alter_periodictasks_options"),
]
operations = [
migrations.RunPython(create_periodic_task, delete_periodic_task),
]
+74 -2
View File
@@ -617,9 +617,66 @@ class Task(RowLevelSecurityProtectedModel):
resource_name = "tasks"
class ScanQuerySet(models.QuerySet):
"""Shared selectors for "the latest scan of a provider".
The queryset must already be scoped by the caller: manager, tenant, RBAC,
providers and database alias.
"""
# How "which completed scan is the provider's current one" is ordered.
LATEST_ORDER_BY = (
models.F("completed_at").desc(nulls_last=True),
models.F("inserted_at").desc(),
models.F("id").desc(),
)
def _eligible_for_latest(self) -> "ScanQuerySet":
"""Restrict to the scans that may be a provider's latest.
Returns:
ScanQuerySet: The completed scans.
"""
return self.filter(state=StateChoices.COMPLETED)
def latest_per_provider(self) -> "ScanQuerySet":
"""Pick each provider's latest scan with `DISTINCT ON (provider_id)`.
Returns:
ScanQuerySet: One scan per provider, the latest one.
"""
return (
self._eligible_for_latest()
.order_by("provider_id", *self.LATEST_ORDER_BY)
.distinct("provider_id")
)
def latest_ids_per_provider(self) -> list[UUID]:
"""Evaluate `latest_per_provider` and return the scan ids.
The ids are materialised so callers can pass them as a literal `IN`
list; as a subquery Postgres misestimates the row count and picks a
slow nested loop.
Returns:
list[UUID]: The id of each provider's latest scan.
"""
return list(self.latest_per_provider().values_list("id", flat=True))
def latest_first(self) -> "ScanQuerySet":
"""Order eligible scans newest first, without deduplicating per provider.
Expects the queryset to be already filtered to a single provider.
Returns:
ScanQuerySet: The eligible scans, latest first.
"""
return self._eligible_for_latest().order_by(*self.LATEST_ORDER_BY)
class Scan(RowLevelSecurityProtectedModel):
objects = ActiveProviderManager()
all_objects = models.Manager()
objects = ActiveProviderManager.from_queryset(ScanQuerySet)()
all_objects = ScanQuerySet.as_manager()
_SCOPING_SCANNER_ARG_KEYS_CACHE: tuple[str, ...] | None = None
@@ -726,6 +783,12 @@ class Scan(RowLevelSecurityProtectedModel):
name="scans_prov_state_ins_desc_idx",
),
# TODO This might replace `scans_prov_state_ins_desc_idx` completely. Review usage
# Since `ScanQuerySet`, no code path reads a provider's
# completed scans by `-inserted_at`. The only query left that
# matches this index (and `scans_prov_state_ins_desc_idx` above)
# is `GET /scans?filter[provider]=…&filter[state]=completed` with
# the default sort. Both are candidates to drop in a follow-up
# once production `pg_stat_user_indexes.idx_scan` confirms it.
models.Index(
fields=["tenant_id", "provider_id", "-inserted_at"],
condition=Q(state=StateChoices.COMPLETED),
@@ -1380,6 +1443,15 @@ class Invitation(RowLevelSecurityProtectedModel):
self.email = self.email.strip().lower()
super().save(*args, **kwargs)
@classmethod
def lapsed_q(cls):
"""Pending invitations whose expiry date has already passed."""
return Q(state=cls.State.PENDING, expires_at__lte=datetime.now(UTC))
@property
def is_lapsed(self):
return self.state == self.State.PENDING and self.expires_at <= datetime.now(UTC)
class Meta(RowLevelSecurityProtectedModel.Meta):
db_table = "invitations"
+1 -1
View File
@@ -1,7 +1,7 @@
openapi: 3.0.3
info:
title: Prowler API
version: 1.42.0
version: 1.45.0
description: |-
Prowler API specification.
@@ -187,6 +187,27 @@ class TestRoutingByDatabasePrefix:
sink_backend_stub.drop_database.assert_called_once_with("db-tenant-abc")
mock_ingest.drop_database.assert_not_called()
def test_list_databases_always_routes_to_ingest(self, sink_backend_stub):
with patch("api.attack_paths.database.ingest") as mock_ingest:
mock_ingest.list_databases.return_value = ["db-tmp-scan-uuid-1"]
assert db_module.list_databases() == ["db-tmp-scan-uuid-1"]
mock_ingest.list_databases.assert_called_once_with()
def test_ingest_list_databases_dedupes_cluster_rows(self):
from api.attack_paths.ingest import driver as ingest_driver
with patch.object(ingest_driver, "get_session") as mock_get_session:
session = mock_get_session.return_value.__enter__.return_value
session.run.return_value = [{"name": "db-tmp-scan-uuid-1"}]
assert ingest_driver.list_databases() == ["db-tmp-scan-uuid-1"]
session.run.assert_called_once_with(
"SHOW DATABASES YIELD name RETURN DISTINCT name"
)
def test_clear_cache_routes_temp_to_ingest(self, sink_backend_stub):
with patch("api.attack_paths.database.ingest") as mock_ingest:
db_module.clear_cache("db-tmp-scan-uuid-1")
+226 -1
View File
@@ -1,14 +1,16 @@
from datetime import UTC, datetime
from datetime import UTC, datetime, timedelta
import pytest
from allauth.socialaccount.models import SocialApp
from api.db_router import MainRouter
from api.models import (
Provider,
ProviderComplianceScore,
Resource,
ResourceTag,
SAMLConfiguration,
SAMLDomainIndex,
Scan,
StateChoices,
StatusChoices,
TenantComplianceSummary,
@@ -524,3 +526,226 @@ class TestTenantComplianceSummaryModel:
assert summary1.id != summary2.id
assert summary1.requirements_passed != summary2.requirements_passed
def _latest_scan_fixture(tenant, provider, *, completed_at, inserted_at=None, **kwargs):
scan = Scan.objects.create(
tenant_id=tenant.id,
provider=provider,
trigger=kwargs.pop("trigger", Scan.TriggerChoices.MANUAL),
state=kwargs.pop("state", StateChoices.COMPLETED),
completed_at=completed_at,
**kwargs,
)
if inserted_at is not None:
# `inserted_at` is auto_now_add, so it has to be forced after the fact.
Scan.all_objects.filter(pk=scan.pk).update(inserted_at=inserted_at)
scan.refresh_from_db()
return scan
@pytest.mark.django_db
class TestScanQuerySetOrdering:
def test_completed_later_wins_over_inserted_later(
self, tenants_fixture, aws_provider
):
"""The scan that FINISHED last is current, not the one that started last."""
tenant, *_ = tenants_fixture
now = datetime.now(UTC)
finished_last = _latest_scan_fixture(
tenant,
aws_provider,
inserted_at=now - timedelta(hours=3),
completed_at=now,
)
_latest_scan_fixture(
tenant,
aws_provider,
inserted_at=now - timedelta(hours=1),
completed_at=now - timedelta(hours=1),
)
assert Scan.all_objects.filter(
tenant_id=tenant.id
).latest_ids_per_provider() == [finished_last.id]
def test_null_completed_at_provider_is_still_returned(
self, tenants_fixture, aws_provider
):
"""NULLS LAST, not `completed_at__isnull=False`.
Excluding NULL `completed_at` would drop the provider from every
"latest" endpoint instead of falling back to `inserted_at`.
"""
tenant, *_ = tenants_fixture
only_scan = _latest_scan_fixture(tenant, aws_provider, completed_at=None)
assert Scan.all_objects.filter(
tenant_id=tenant.id
).latest_ids_per_provider() == [only_scan.id]
def test_null_completed_at_never_outranks_a_finished_scan(
self, tenants_fixture, aws_provider
):
"""Postgres sorts NULLs first under DESC; NULLS LAST is what fixes it."""
tenant, *_ = tenants_fixture
now = datetime.now(UTC)
finished = _latest_scan_fixture(
tenant,
aws_provider,
inserted_at=now - timedelta(hours=2),
completed_at=now - timedelta(hours=2),
)
_latest_scan_fixture(
tenant,
aws_provider,
inserted_at=now,
completed_at=None,
)
assert Scan.all_objects.filter(
tenant_id=tenant.id
).latest_ids_per_provider() == [finished.id]
def test_id_breaks_an_exact_timestamp_tie_deterministically(
self, tenants_fixture, aws_provider
):
tenant, *_ = tenants_fixture
now = datetime.now(UTC)
scans = [
_latest_scan_fixture(
tenant, aws_provider, inserted_at=now, completed_at=now
)
for _ in range(3)
]
expected = max(scan.id for scan in scans)
picks = {
Scan.all_objects.filter(tenant_id=tenant.id).latest_ids_per_provider()[0]
for _ in range(5)
}
assert picks == {expected}
@pytest.mark.django_db
class TestScanQuerySetEligibility:
def test_unfinished_scans_are_excluded(self, tenants_fixture, aws_provider):
tenant, *_ = tenants_fixture
_latest_scan_fixture(
tenant,
aws_provider,
completed_at=None,
state=StateChoices.EXECUTING,
)
assert (
Scan.all_objects.filter(tenant_id=tenant.id).latest_ids_per_provider() == []
)
@pytest.mark.django_db
class TestScanQuerySetManagerChoice:
def test_active_manager_hides_soft_deleted_providers(
self, tenants_fixture, aws_provider
):
"""`Scan.objects` drops soft-deleted providers, `all_objects` keeps them.
The queryset must not decide this for the caller.
"""
tenant, *_ = tenants_fixture
scan = _latest_scan_fixture(
tenant, aws_provider, completed_at=datetime.now(UTC)
)
Provider.all_objects.filter(pk=aws_provider.pk).update(is_deleted=True)
assert Scan.all_objects.filter(
tenant_id=tenant.id
).latest_ids_per_provider() == [scan.id]
assert Scan.objects.filter(tenant_id=tenant.id).latest_ids_per_provider() == []
@pytest.mark.django_db
class TestScanQuerySetPerProviderScoping:
def test_one_scan_per_provider(self, tenants_fixture, aws_provider_pair):
tenant, *_ = tenants_fixture
provider_one, provider_two = aws_provider_pair
now = datetime.now(UTC)
newest_one = _latest_scan_fixture(tenant, provider_one, completed_at=now)
_latest_scan_fixture(tenant, provider_one, completed_at=now - timedelta(days=1))
newest_two = _latest_scan_fixture(tenant, provider_two, completed_at=now)
assert set(
Scan.all_objects.filter(tenant_id=tenant.id).latest_ids_per_provider()
) == {newest_one.id, newest_two.id}
def test_caller_filters_are_preserved(self, tenants_fixture, aws_provider_pair):
tenant, *_ = tenants_fixture
provider_one, provider_two = aws_provider_pair
now = datetime.now(UTC)
scan_one = _latest_scan_fixture(tenant, provider_one, completed_at=now)
_latest_scan_fixture(tenant, provider_two, completed_at=now)
assert Scan.all_objects.filter(
tenant_id=tenant.id, provider__in=[provider_one]
).latest_ids_per_provider() == [scan_one.id]
def test_latest_first_is_ordered_not_deduplicated(
self, tenants_fixture, aws_provider
):
tenant, *_ = tenants_fixture
now = datetime.now(UTC)
newest = _latest_scan_fixture(tenant, aws_provider, completed_at=now)
older = _latest_scan_fixture(
tenant, aws_provider, completed_at=now - timedelta(days=1)
)
ordered = list(
Scan.all_objects.filter(
tenant_id=tenant.id, provider_id=aws_provider.id
).latest_first()
)
assert [scan.id for scan in ordered] == [newest.id, older.id]
def test_tenant_isolation(self, tenants_fixture, aws_provider):
tenant, other_tenant, *_ = tenants_fixture
_latest_scan_fixture(tenant, aws_provider, completed_at=datetime.now(UTC))
assert (
Scan.all_objects.filter(tenant_id=other_tenant.id).latest_ids_per_provider()
== []
)
def test_empty_queryset_returns_empty_list(self, tenants_fixture):
tenant, *_ = tenants_fixture
assert (
Scan.all_objects.filter(tenant_id=tenant.id).latest_ids_per_provider() == []
)
@pytest.mark.django_db
class TestScanQuerySetPerProviderQuerysetShape:
def test_returns_a_queryset_not_a_list(self, tenants_fixture, aws_provider):
tenant, *_ = tenants_fixture
_latest_scan_fixture(tenant, aws_provider, completed_at=datetime.now(UTC))
qs = Scan.all_objects.filter(tenant_id=tenant.id).latest_per_provider()
# Callers chain .values(...) / .values_list(...) onto this.
assert qs.values_list("provider_id", flat=True).count() == 1
@pytest.mark.django_db
class TestScanQuerySetRelatedManager:
def test_reverse_relation_exposes_the_methods(self, tenants_fixture, aws_provider):
tenant, *_ = tenants_fixture
now = datetime.now(UTC)
newest = _latest_scan_fixture(tenant, aws_provider, completed_at=now)
_latest_scan_fixture(tenant, aws_provider, completed_at=now - timedelta(days=1))
assert aws_provider.scans.latest_first().first().id == newest.id
+338 -28
View File
@@ -82,7 +82,7 @@ from django.db import close_old_connections, connection, connections
from django.db.models import Count
from django.db.models.signals import pre_delete
from django.http import JsonResponse
from django.test import RequestFactory
from django.test import RequestFactory, override_settings
from django.test.utils import CaptureQueriesContext
from django.urls import reverse
from django_celery_results.models import TaskResult
@@ -3951,6 +3951,43 @@ class TestScanViewSet:
mock_enqueue_scan_execution.assert_called_once()
# assert scan.scanner_args == expected_scanner_args
@patch("api.v1.views.enqueue_scan_execution_on_commit")
def test_scans_create_returns_the_scan_id_in_task_args(
self,
mock_enqueue_scan_execution,
authenticated_client,
okta_provider,
):
"""The 202 is a task, so `task_args` is the only place the scan id is.
It is serialized before the on_commit publish that would otherwise fill
the kwargs, so the record has to carry them from the start.
"""
payload = {
"data": {
"type": "scans",
"attributes": {"name": "New Scan"},
"relationships": {
"provider": {
"data": {"type": "providers", "id": str(okta_provider.id)}
}
},
}
}
response = authenticated_client.post(
reverse("scan-list"),
data=payload,
content_type=API_JSON_CONTENT_TYPE,
)
assert response.status_code == status.HTTP_202_ACCEPTED
scan = Scan.objects.get()
assert response.json()["data"]["attributes"]["task_args"] == {
"scan_id": str(scan.id),
"provider_id": str(okta_provider.id),
}
@patch("tasks.tasks.perform_scan_task.apply_async")
def test_scans_create_queues_scan_when_provider_has_active_scan(
self,
@@ -4540,6 +4577,52 @@ class TestScanViewSet:
assert response.status_code == status.HTTP_302_FOUND
assert response["Location"] == presigned_url
@override_settings(
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="https://storage.example.com",
DJANGO_OUTPUT_S3_AWS_ACCESS_KEY_ID="access-key",
DJANGO_OUTPUT_S3_AWS_SECRET_ACCESS_KEY="secret-key",
DJANGO_OUTPUT_S3_AWS_SESSION_TOKEN="",
DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION="eu-west-1",
)
def test_report_s3_redirects_to_the_public_storage_host(
self, authenticated_client, scans_fixture, monkeypatch
):
"""The object is looked up internally but the redirect the browser follows is public."""
scan = scans_fixture[0]
bucket = "test-bucket"
key = "report.zip"
scan.output_location = f"s3://{bucket}/{key}"
scan.state = StateChoices.COMPLETED
scan.save()
monkeypatch.setattr(
"api.v1.views.env",
type("env", (), {"str": lambda self, *_args, **_kwargs: bucket})(),
)
head_calls = []
class InternalS3Client:
def head_object(self, Bucket, Key):
head_calls.append((Bucket, Key))
return {}
def generate_presigned_url(self, *_args, **_kwargs):
raise AssertionError("the internal client must not sign the redirect")
monkeypatch.setattr("api.v1.views.get_s3_client", lambda: InternalS3Client())
url = reverse("scan-report", kwargs={"pk": scan.id})
response = authenticated_client.get(url)
assert response.status_code == status.HTTP_302_FOUND
assert head_calls == [(bucket, key)]
location = urlparse(response["Location"])
assert location.netloc == "storage.example.com"
assert location.path == f"/{bucket}/{key}"
assert "X-Amz-Signature" in parse_qs(location.query)
def test_report_s3_success_no_local_files(
self, authenticated_client, scans_fixture, monkeypatch
):
@@ -8784,6 +8867,190 @@ class TestInvitationViewSet:
user.id
)
@staticmethod
def _invitation_create_payload(email, role):
return json.dumps(
{
"data": {
"type": "invitations",
"attributes": {"email": email},
"relationships": {
"roles": {"data": [{"type": "roles", "id": str(role.id)}]}
},
}
}
)
@staticmethod
def _create_lapsed_invitation(email, tenant, inviter):
return Invitation.objects.create(
email=email,
state=Invitation.State.PENDING,
expires_at=datetime.now(UTC) - timedelta(days=1),
inviter=inviter,
tenant=tenant,
)
def test_invitations_create_with_lapsed_pending_invitation_for_same_email(
self,
authenticated_client,
create_test_user,
tenants_fixture,
invitations_fixture,
roles_fixture,
):
lapsed_invitation, expired_invitation = invitations_fixture
lapsed_invitation.expires_at = datetime.now(UTC) - timedelta(days=1)
lapsed_invitation.save()
other_email_lapsed_invitation = self._create_lapsed_invitation(
"other@prowler.com", tenants_fixture[0], create_test_user
)
response = authenticated_client.post(
reverse("invitation-list"),
data=self._invitation_create_payload(
lapsed_invitation.email, roles_fixture[0]
),
content_type="application/vnd.api+json",
)
assert response.status_code == status.HTTP_201_CREATED
new_invitation = Invitation.objects.get(id=response.json()["data"]["id"])
assert new_invitation.email == lapsed_invitation.email
assert new_invitation.state == Invitation.State.PENDING
lapsed_invitation.refresh_from_db()
assert lapsed_invitation.state == Invitation.State.EXPIRED
expired_invitation.refresh_from_db()
assert expired_invitation.state == Invitation.State.EXPIRED
other_email_lapsed_invitation.refresh_from_db()
assert other_email_lapsed_invitation.state == Invitation.State.PENDING
def test_invitations_create_with_active_pending_invitation_for_same_email(
self,
authenticated_client,
create_test_user,
tenants_fixture,
invitations_fixture,
roles_fixture,
):
active_invitation, _ = invitations_fixture
self._create_lapsed_invitation(
active_invitation.email, tenants_fixture[0], create_test_user
)
invitation_count = Invitation.objects.count()
response = authenticated_client.post(
reverse("invitation-list"),
data=self._invitation_create_payload(
active_invitation.email, roles_fixture[0]
),
content_type="application/vnd.api+json",
)
assert response.status_code == status.HTTP_400_BAD_REQUEST
assert (
response.json()["errors"][0]["source"]["pointer"]
== "/data/attributes/email"
)
assert Invitation.objects.count() == invitation_count
active_invitation.refresh_from_db()
assert active_invitation.state == Invitation.State.PENDING
def test_invitations_create_ignores_pending_invitations_from_other_tenants(
self, authenticated_client, create_test_user, tenants_fixture, roles_fixture
):
email = "cross_tenant@prowler.com"
other_tenant = tenants_fixture[1]
other_tenant_lapsed_invitation = self._create_lapsed_invitation(
email, other_tenant, create_test_user
)
Invitation.objects.create(
email=email, inviter=create_test_user, tenant=other_tenant
)
response = authenticated_client.post(
reverse("invitation-list"),
data=self._invitation_create_payload(email, roles_fixture[0]),
content_type="application/vnd.api+json",
)
assert response.status_code == status.HTTP_201_CREATED
other_tenant_lapsed_invitation.refresh_from_db()
assert other_tenant_lapsed_invitation.state == Invitation.State.PENDING
def test_invitations_report_lapsed_pending_invitation_as_expired(
self,
authenticated_client,
create_test_user,
tenants_fixture,
invitations_fixture,
):
active_invitation, expired_invitation = invitations_fixture
lapsed_invitation = self._create_lapsed_invitation(
"lapsed@prowler.com", tenants_fixture[0], create_test_user
)
list_response = authenticated_client.get(reverse("invitation-list"))
retrieve_response = authenticated_client.get(
reverse("invitation-detail", kwargs={"pk": lapsed_invitation.id})
)
assert list_response.status_code == status.HTTP_200_OK
assert retrieve_response.status_code == status.HTTP_200_OK
assert {
invitation["id"]: invitation["attributes"]["state"]
for invitation in list_response.json()["data"]
} == {
str(active_invitation.id): Invitation.State.PENDING.value,
str(expired_invitation.id): Invitation.State.EXPIRED.value,
str(lapsed_invitation.id): Invitation.State.EXPIRED.value,
}
assert (
retrieve_response.json()["data"]["attributes"]["state"]
== Invitation.State.EXPIRED.value
)
@pytest.mark.parametrize(
"filter_name, filter_value, expected_invitations",
[
("state", "pending", {"active"}),
("state", "expired", {"expired", "lapsed"}),
("state", "accepted", set()),
("state__in", "pending", {"active"}),
("state__in", "expired", {"expired", "lapsed"}),
("state__in", "pending,expired", {"active", "expired", "lapsed"}),
("state__in", "accepted,revoked", set()),
],
)
def test_invitations_filter_state_treats_lapsed_pending_as_expired(
self,
authenticated_client,
create_test_user,
tenants_fixture,
invitations_fixture,
filter_name,
filter_value,
expected_invitations,
):
active_invitation, expired_invitation = invitations_fixture
lapsed_invitation = self._create_lapsed_invitation(
"lapsed@prowler.com", tenants_fixture[0], create_test_user
)
invitation_ids = {
"active": str(active_invitation.id),
"expired": str(expired_invitation.id),
"lapsed": str(lapsed_invitation.id),
}
response = authenticated_client.get(
reverse("invitation-list"), {f"filter[{filter_name}]": filter_value}
)
assert response.status_code == status.HTTP_200_OK
assert {invitation["id"] for invitation in response.json()["data"]} == {
invitation_ids[name] for name in expected_invitations
}
@pytest.mark.parametrize(
"email",
[
@@ -8791,8 +9058,10 @@ class TestInvitationViewSet:
"invalid_email@",
# There is a pending invitation with this email
"testing@prowler.com",
"TESTING@prowler.com",
# User is already a member of the tenant
TEST_USER,
TEST_USER.upper(),
],
)
def test_invitations_create_invalid_email(
@@ -9047,6 +9316,56 @@ class TestInvitationViewSet:
== "This invitation cannot be revoked."
)
def test_invitations_delete_lapsed_invitation(
self, authenticated_client, invitations_fixture
):
invitation, *_ = invitations_fixture
invitation.expires_at = datetime.now(UTC) - timedelta(days=1)
invitation.save()
response = authenticated_client.delete(
reverse("invitation-detail", kwargs={"pk": str(invitation.id)})
)
assert response.status_code == status.HTTP_400_BAD_REQUEST
assert (
response.json()["errors"][0]["detail"]
== "This invitation cannot be revoked."
)
invitation.refresh_from_db()
assert invitation.state == Invitation.State.PENDING
def test_invitations_partial_update_lapsed_invitation(
self, authenticated_client, invitations_fixture
):
invitation, *_ = invitations_fixture
invitation.expires_at = datetime.now(UTC) - timedelta(days=1)
invitation.save()
data = {
"data": {
"id": str(invitation.id),
"type": "invitations",
"attributes": {
"email": invitation.email,
"expires_at": self.TOMORROW_ISO,
},
}
}
response = authenticated_client.patch(
reverse("invitation-detail", kwargs={"pk": str(invitation.id)}),
data=json.dumps(data),
content_type="application/vnd.api+json",
)
assert response.status_code == status.HTTP_400_BAD_REQUEST
assert (
response.json()["errors"][0]["detail"]
== "This invitation cannot be updated."
)
invitation.refresh_from_db()
assert invitation.is_lapsed
def test_invitations_accept_invitation_new_user(self, client, invitations_fixture):
invitation, *_ = invitations_fixture
@@ -18333,19 +18652,14 @@ class TestMuteRuleViewSet:
assert len(data) == 2
assert data[0]["id"] == str(mute_rules_fixture[first_index].id)
@patch("api.v1.views.chain")
@patch("api.v1.views.reaggregate_all_finding_group_summaries_task.si")
@patch("api.v1.views.mute_historical_findings_task.si")
@patch("api.v1.views.mute_findings_in_latest_scans_task.apply_async")
@patch("api.v1.views.transaction.on_commit", side_effect=lambda fn: fn())
def test_mute_rules_create_valid(
self,
_mock_on_commit,
mock_mute_signature,
mock_reaggregate_signature,
mock_chain,
mock_mute_task,
authenticated_client,
findings_fixture,
create_test_user,
):
"""Test creating a valid mute rule."""
finding_ids = [str(findings_fixture[0].id)]
@@ -18372,24 +18686,20 @@ class TestMuteRuleViewSet:
assert response_data["attributes"]["name"] == "New Mute Rule"
assert response_data["attributes"]["reason"] == "Security exception approved"
# Verify the finding was immediately muted
from api.models import Finding
finding = Finding.objects.get(id=findings_fixture[0].id)
assert finding.muted is True
assert finding.muted_at is not None
assert finding.muted_reason == "Security exception approved"
assert finding.muted is False
assert finding.muted_at is None
assert finding.muted_reason is None
# Verify background task chain was called: mute → reaggregate all
mock_mute_signature.assert_called_once()
mock_reaggregate_signature.assert_called_once()
mock_chain.assert_called_once_with(
mock_mute_signature.return_value,
mock_reaggregate_signature.return_value,
mock_mute_task.assert_called_once_with(
kwargs={
"tenant_id": str(finding.tenant_id),
"mute_rule_id": response_data["id"],
"provider_ids": [str(finding.scan.provider_id)],
}
)
mock_chain.return_value.apply_async.assert_called_once()
@patch("tasks.tasks.mute_historical_findings_task.apply_async")
@patch("api.v1.views.mute_findings_in_latest_scans_task.apply_async")
def test_mute_rules_create_converts_finding_ids_to_uids(
self,
mock_task,
@@ -18425,7 +18735,7 @@ class TestMuteRuleViewSet:
]
assert set(mute_rule.finding_uids) == set(expected_uids)
@patch("tasks.tasks.mute_historical_findings_task.apply_async")
@patch("api.v1.views.mute_findings_in_latest_scans_task.apply_async")
def test_mute_rules_deduplicates_uids(
self,
mock_task,
@@ -18492,10 +18802,10 @@ class TestMuteRuleViewSet:
finding1.refresh_from_db()
finding2.refresh_from_db()
assert finding1.muted is True
assert finding2.muted is True
assert finding1.muted is False
assert finding2.muted is False
@patch("tasks.tasks.mute_historical_findings_task.apply_async")
@patch("api.v1.views.mute_findings_in_latest_scans_task.apply_async")
def test_mute_rules_create_overlap_detection_active(
self,
mock_task,
@@ -18528,7 +18838,7 @@ class TestMuteRuleViewSet:
"already muted" in error_detail.lower() or "overlap" in error_detail.lower()
)
@patch("tasks.tasks.mute_historical_findings_task.apply_async")
@patch("api.v1.views.mute_findings_in_latest_scans_task.apply_async")
def test_mute_rules_create_no_overlap_with_inactive(
self,
mock_task,
@@ -18584,7 +18894,7 @@ class TestMuteRuleViewSet:
== "/data/attributes/finding_ids"
)
@patch("tasks.tasks.mute_historical_findings_task.apply_async")
@patch("api.v1.views.mute_findings_in_latest_scans_task.apply_async")
def test_mute_rules_create_invalid_finding_ids(
self, mock_task, authenticated_client
):
+14 -3
View File
@@ -2149,6 +2149,12 @@ class InvitationSerializer(RLSSerializer):
if tenant_id is not None:
self.fields["roles"].queryset = Role.objects.filter(tenant_id=tenant_id)
def to_representation(self, instance):
data = super().to_representation(instance)
if instance.is_lapsed:
data["state"] = Invitation.State.EXPIRED.value
return data
class Meta:
model = Invitation
fields = [
@@ -2175,6 +2181,7 @@ class InvitationBaseWriteSerializer(BaseWriteSerializer):
self.fields["roles"].queryset = Role.objects.filter(tenant_id=tenant_id)
def validate_email(self, value):
value = value.strip().lower()
user = User.objects.filter(email=value).first()
tenant_id = self.context["tenant_id"]
if user and Membership.objects.filter(user=user, tenant=tenant_id).exists():
@@ -2182,9 +2189,13 @@ class InvitationBaseWriteSerializer(BaseWriteSerializer):
"The user may already be a member of the tenant or there was an issue with the "
"email provided."
)
if Invitation.objects.filter(
email=value, state=Invitation.State.PENDING
).exists():
pending_invitations = Invitation.objects.filter(
tenant_id=tenant_id, email=value, state=Invitation.State.PENDING
)
pending_invitations.filter(Invitation.lapsed_q()).update(
state=Invitation.State.EXPIRED
)
if pending_invitations.filter(expires_at__gt=datetime.now(UTC)).exists():
raise ValidationError(
"Unable to process your request. Please check the information provided and "
"try again."
+67 -104
View File
@@ -244,7 +244,6 @@ from api.v1.serializers import (
UserUpdateSerializer,
)
from botocore.exceptions import ClientError, NoCredentialsError, ParamValidationError
from celery import chain
from celery.result import AsyncResult
from config.custom_logging import BackendLogger
from config.env import env
@@ -327,7 +326,7 @@ from rest_framework_simplejwt.token_blacklist.models import (
)
from tasks.beat import schedule_provider_scan
from tasks.jobs.attack_paths import db_utils as attack_paths_db_utils
from tasks.jobs.export import get_s3_client
from tasks.jobs.export import get_s3_client, get_s3_presign_client
from tasks.tasks import (
QUEUED_SCAN_TASK_STATE,
backfill_compliance_summaries_task,
@@ -342,8 +341,7 @@ from tasks.tasks import (
enqueue_scan_execution_on_commit,
get_active_provider_scan,
jira_integration_task,
mute_historical_findings_task,
reaggregate_all_finding_group_summaries_task,
mute_findings_in_latest_scans_task,
refresh_lighthouse_provider_models_task,
)
@@ -2409,7 +2407,8 @@ class ScanViewSet(ProviderVisibilityMixin, BaseRLSViewSet):
}
if content_type:
params["ResponseContentType"] = content_type
url = client.generate_presigned_url(
# The browser follows this URL, so it is signed against the public host.
url = (get_s3_presign_client() or client).generate_presigned_url(
"get_object",
Params=params,
ExpiresIn=300,
@@ -2823,6 +2822,15 @@ class ScanViewSet(ProviderVisibilityMixin, BaseRLSViewSet):
tenant_id=self.request.tenant_id,
task_id=pre_task_id,
task_status=(QUEUED_SCAN_TASK_STATE if active_scan else None),
# This response is serialized before the on_commit publish,
# so without these the caller gets a task id and no scan id.
# Kept in step with what `enqueue_scan_execution_on_commit`
# publishes below.
task_kwargs={
"tenant_id": str(self.request.tenant_id),
"scan_id": str(scan.id),
"provider_id": str(scan.provider_id),
},
)
if not active_scan:
@@ -3479,12 +3487,8 @@ class ResourceViewSet(PaginateByPkMixin, BaseRLSViewSet):
filtered_queryset = self.filter_queryset(self.get_queryset())
latest_scans = (
Scan.all_objects.filter(
tenant_id=tenant_id,
state=StateChoices.COMPLETED,
)
.order_by("provider_id", "-inserted_at")
.distinct("provider_id")
Scan.all_objects.filter(tenant_id=tenant_id)
.latest_per_provider()
.values("provider_id")
)
@@ -3616,11 +3620,9 @@ class ResourceViewSet(PaginateByPkMixin, BaseRLSViewSet):
tenant_id = request.tenant_id
query_params = request.query_params
latest_scans_queryset = (
Scan.all_objects.filter(tenant_id=tenant_id, state=StateChoices.COMPLETED)
.order_by("provider_id", "-inserted_at")
.distinct("provider_id")
)
latest_scans_queryset = Scan.all_objects.filter(
tenant_id=tenant_id
).latest_per_provider()
queryset = ResourceScanSummary.objects.filter(
tenant_id=tenant_id,
@@ -4207,12 +4209,9 @@ class FindingViewSet(PaginateByPkMixin, BaseRLSViewSet):
tenant_id = request.tenant_id
filtered_queryset = self.filter_queryset(self.get_queryset())
latest_scan_ids = list(
Scan.all_objects.filter(tenant_id=tenant_id, state=StateChoices.COMPLETED)
.order_by("provider_id", "-inserted_at")
.distinct("provider_id")
.values_list("id", flat=True)
)
latest_scan_ids = Scan.all_objects.filter(
tenant_id=tenant_id
).latest_ids_per_provider()
filtered_queryset = filtered_queryset.filter(
tenant_id=tenant_id, scan_id__in=latest_scan_ids
)
@@ -4235,11 +4234,9 @@ class FindingViewSet(PaginateByPkMixin, BaseRLSViewSet):
tenant_id = request.tenant_id
query_params = request.query_params
latest_scans_queryset = (
Scan.all_objects.filter(tenant_id=tenant_id, state=StateChoices.COMPLETED)
.order_by("provider_id", "-inserted_at")
.distinct("provider_id")
)
latest_scans_queryset = Scan.all_objects.filter(
tenant_id=tenant_id
).latest_per_provider()
raw_latest_scans_ids = list(
latest_scans_queryset.values_list("id", "unique_resource_count")
)
@@ -4473,7 +4470,7 @@ class InvitationViewSet(BaseRLSViewSet):
def partial_update(self, request, *args, **kwargs):
instance = self.get_object()
if instance.state != Invitation.State.PENDING:
if instance.state != Invitation.State.PENDING or instance.is_lapsed:
raise ValidationError(detail="This invitation cannot be updated.")
serializer = self.get_serializer(
instance,
@@ -4487,7 +4484,7 @@ class InvitationViewSet(BaseRLSViewSet):
def destroy(self, request, *args, **kwargs):
instance = self.get_object()
if instance.state != Invitation.State.PENDING:
if instance.state != Invitation.State.PENDING or instance.is_lapsed:
raise ValidationError(detail="This invitation cannot be revoked.")
instance.state = Invitation.State.REVOKED
instance.save()
@@ -4980,11 +4977,7 @@ class ComplianceOverviewViewSet(
if provider_filters:
scans = scans.filter(**provider_filters)
return list(
scans.order_by("provider_id", "-inserted_at")
.distinct("provider_id")
.values_list("id", flat=True)
)
return scans.latest_ids_per_provider()
def _filtered_queryset_for_latest_provider_scans(self, latest_scan_ids=None):
if latest_scan_ids is None:
@@ -5726,14 +5719,9 @@ class OverviewViewSet(ProviderFilterParamsMixin, BaseRLSViewSet):
else {}
)
latest_scan_ids = (
Scan.all_objects.filter(
tenant_id=tenant_id, state=StateChoices.COMPLETED, **provider_filter
)
.order_by("provider_id", "-inserted_at")
.distinct("provider_id")
.values_list("id", flat=True)
)
latest_scan_ids = Scan.all_objects.filter(
tenant_id=tenant_id, **provider_filter
).latest_ids_per_provider()
return filtered_queryset.filter(
tenant_id=tenant_id, scan_id__in=latest_scan_ids
@@ -5760,16 +5748,10 @@ class OverviewViewSet(ProviderFilterParamsMixin, BaseRLSViewSet):
def _latest_scan_ids_for_allowed_providers(self, tenant_id, provider_filters=None):
provider_filter = self._get_provider_filter()
queryset = Scan.all_objects.filter(
tenant_id=tenant_id, state=StateChoices.COMPLETED, **provider_filter
)
queryset = Scan.all_objects.filter(tenant_id=tenant_id, **provider_filter)
if provider_filters:
queryset = queryset.filter(**provider_filters)
return (
queryset.order_by("provider_id", "-inserted_at")
.distinct("provider_id")
.values_list("id", flat=True)
)
return queryset.latest_ids_per_provider()
@action(detail=False, methods=["get"], url_name="providers")
def providers(self, request):
@@ -5781,14 +5763,9 @@ class OverviewViewSet(ProviderFilterParamsMixin, BaseRLSViewSet):
else {}
)
latest_scan_ids = (
Scan.all_objects.filter(
tenant_id=tenant_id, state=StateChoices.COMPLETED, **provider_filter
)
.order_by("provider_id", "-inserted_at")
.distinct("provider_id")
.values_list("id", flat=True)
)
latest_scan_ids = Scan.all_objects.filter(
tenant_id=tenant_id, **provider_filter
).latest_ids_per_provider()
findings_aggregated = (
queryset.filter(scan_id__in=latest_scan_ids)
@@ -7551,35 +7528,28 @@ class MuteRuleViewSet(BaseRLSViewSet):
serializer = self.get_serializer(data=request.data)
serializer.is_valid(raise_exception=True)
# Create the mute rule
tenant_id = str(request.tenant_id)
finding_ids = serializer.validated_data["finding_ids"]
provider_ids = list(
dict.fromkeys(
Finding.all_objects.filter(
id__in=finding_ids, tenant_id=tenant_id
).values_list("scan__provider_id", flat=True)
)
)
mute_rule = serializer.save()
tenant_id = str(request.tenant_id)
finding_ids = request.data.get("finding_ids", [])
# Immediately mute the selected findings
Finding.all_objects.filter(
id__in=finding_ids, tenant_id=tenant_id, muted=False
).update(
muted=True,
muted_at=mute_rule.inserted_at,
muted_reason=mute_rule.reason,
)
# Launch background task for historical muting + reaggregation
transaction.on_commit(
lambda: chain(
mute_historical_findings_task.si(
tenant_id=tenant_id,
mute_rule_id=str(mute_rule.id),
),
reaggregate_all_finding_group_summaries_task.si(
tenant_id=tenant_id,
),
).apply_async()
lambda: mute_findings_in_latest_scans_task.apply_async(
kwargs={
"tenant_id": tenant_id,
"mute_rule_id": str(mute_rule.id),
"provider_ids": [str(provider_id) for provider_id in provider_ids],
}
)
)
# Return the created mute rule
serializer = self.get_serializer(mute_rule)
return Response(
data=serializer.data,
@@ -7942,18 +7912,13 @@ class FindingGroupViewSet(JsonApiFilterMixin, BaseRLSViewSet):
def _get_latest_findings_per_provider(self, filtered_queryset):
"""Keep only findings from each provider's most recent completed scan."""
# Materialize to a literal IN list. Left as a subquery, Postgres can't
# estimate the match count and picks a serial nested loop on
# resource_finding_mappings when one scan dominates findings
latest_scan_ids = list(
Scan.objects.filter(
tenant_id=self.request.tenant_id,
state=StateChoices.COMPLETED,
)
.order_by("provider_id", "-completed_at", "-inserted_at")
.distinct("provider_id")
.values_list("id", flat=True)
)
# `latest_ids_per_provider` materializes to a literal IN list.
# Left as a subquery, Postgres can't estimate the match count and picks
# a serial nested loop on resource_finding_mappings when one scan
# dominates findings
latest_scan_ids = Scan.objects.filter(
tenant_id=self.request.tenant_id
).latest_ids_per_provider()
return filtered_queryset.filter(scan_id__in=latest_scan_ids)
def _post_process_aggregation(self, aggregated_data):
@@ -8899,16 +8864,14 @@ class FindingGroupViewSet(JsonApiFilterMixin, BaseRLSViewSet):
tenant_id = request.tenant_id
queryset = self._get_finding_queryset()
# Order by -completed_at (matching the /latest summary path and the
# daily summary upsert keyed on midnight(completed_at)) so that
# overlapping scans do not make /resources and /latest read from
# different scans and report diverging counts.
latest_scan_ids = (
Scan.objects.filter(tenant_id=tenant_id, state=StateChoices.COMPLETED)
.order_by("provider_id", "-completed_at", "-inserted_at")
.distinct("provider_id")
.values_list("id", flat=True)
)
# The shared selector orders by -completed_at (matching the /latest
# summary path and the daily summary upsert keyed on
# midnight(completed_at)) so that overlapping scans do not make
# /resources and /latest read from different scans and report
# diverging counts.
latest_scan_ids = Scan.objects.filter(
tenant_id=tenant_id
).latest_ids_per_provider()
normalized_params = self._normalize_jsonapi_params(request.query_params)
# Remove date filters since we're using latest
+56
View File
@@ -231,6 +231,62 @@ LOGGING = {
"level": LEVEL,
"propagate": False,
},
# Celery loggers must be declared explicitly because
# disable_existing_loggers=True silences any logger that exists at
# dictConfig time but is not named here. Without these, fatal worker
# errors (e.g. celery.worker CRITICAL) produce no output.
# "celery" must keep propagating: get_task_logger() parents task
# loggers under celery.task, so blocking here hides them from root.
"celery": {
"level": LEVEL,
"propagate": True,
},
"celery.worker": {
"handlers": ["tasks_console"],
"level": LEVEL,
"propagate": False,
},
"celery.worker.consumer": {
"handlers": ["tasks_console"],
"level": LEVEL,
"propagate": False,
},
"celery.worker.consumer.consumer": {
"handlers": ["tasks_console"],
"level": LEVEL,
"propagate": False,
},
"kombu": {
"handlers": ["tasks_console"],
"level": LEVEL,
"propagate": False,
},
"kombu.transport.redis": {
"handlers": ["tasks_console"],
"level": LEVEL,
"propagate": False,
},
"billiard": {
"handlers": ["tasks_console"],
"level": LEVEL,
"propagate": False,
},
"amqp": {
"handlers": ["tasks_console"],
"level": LEVEL,
"propagate": False,
},
# WARNING keeps task failures but skips one "succeeded" line per task.
"celery.app.trace": {
"handlers": ["tasks_console"],
"level": "WARNING",
"propagate": False,
},
"celery.beat": {
"handlers": ["tasks_console"],
"level": LEVEL,
"propagate": False,
},
},
# Gunicorn required configuration
"root": {
+20
View File
@@ -7,6 +7,7 @@ from config.settings.eventstream import * # noqa
from config.settings.partitions import * # noqa
from config.settings.sentry import * # noqa
from config.settings.social_login import * # noqa
from django.core.exceptions import ImproperlyConfigured
SECRET_KEY = env("SECRET_KEY", default="secret")
DEBUG = env.bool("DJANGO_DEBUG", default=False)
@@ -295,6 +296,14 @@ DJANGO_OUTPUT_S3_AWS_SECRET_ACCESS_KEY = env.str(
)
DJANGO_OUTPUT_S3_AWS_SESSION_TOKEN = env.str("DJANGO_OUTPUT_S3_AWS_SESSION_TOKEN", "")
DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION = env.str("DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION", "")
# Storage endpoint the API and Celery workers use to talk to S3-compatible object storage
# such as MinIO. Empty means the real AWS S3 endpoint, which is unaffected.
DJANGO_OUTPUT_S3_AWS_ENDPOINT_URL = env.str("DJANGO_OUTPUT_S3_AWS_ENDPOINT_URL", "")
# Browser-reachable storage host used to sign download URLs. Empty means sign against the
# same endpoint the API talks to, which is what Prowler Cloud on S3 does.
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL = env.str(
"DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL", ""
)
# HTTP Security Headers
SECURE_CONTENT_TYPE_NOSNIFF = True
@@ -320,6 +329,17 @@ ATTACK_PATHS_SCAN_STALE_THRESHOLD_MINUTES = env.int(
"ATTACK_PATHS_SCAN_STALE_THRESHOLD_MINUTES", 960
) # 16h
# Minimum age (of the scan row, or of the scan id itself when the row is gone) before
# the periodic reaper will drop an orphaned temp Neo4j database. Keeps a scan that is
# still legitimately in flight from ever losing its staging database mid-run.
ATTACK_PATHS_TMP_DB_REAP_SAFETY_MARGIN_HOURS = env.int(
"ATTACK_PATHS_TMP_DB_REAP_SAFETY_MARGIN_HOURS", 6
)
if ATTACK_PATHS_TMP_DB_REAP_SAFETY_MARGIN_HOURS <= 0:
raise ImproperlyConfigured(
"ATTACK_PATHS_TMP_DB_REAP_SAFETY_MARGIN_HOURS must be a positive number of hours"
)
# Selects where the persistent attack-paths graph is stored. The scan
# temporary database is always Neo4j; only the sink is configurable.
# Valid values: "neo4j" (default, OSS and local dev), "neptune" (hosted).
@@ -0,0 +1,107 @@
"""Periodic reaper for orphaned temp Neo4j scan databases.
`scan.py` creates a throw-away `db-tmp-scan-<attack_paths_scan_id>` database per
scan and drops it once the scan finishes, success or failure. When the worker
or Neo4j itself dies mid-scan, that drop never runs and nothing else ever
revisits the database - it sits there forever. This sweep lists every temp
database on the ingest cluster and drops the ones whose scan is gone or has
been finished for longer than the configured safety margin.
"""
from datetime import UTC, datetime, timedelta
from api.attack_paths import database as graph_database
from api.db_router import MainRouter
from api.models import AttackPathsScan, StateChoices
from api.uuid_utils import datetime_from_uuid7
from celery.utils.log import get_task_logger
from config.django.base import ATTACK_PATHS_TMP_DB_REAP_SAFETY_MARGIN_HOURS
from uuid6 import UUID as UUID7
logger = get_task_logger(__name__)
TERMINAL_STATES = (
StateChoices.COMPLETED,
StateChoices.FAILED,
StateChoices.CANCELLED,
)
def reap_orphaned_tmp_databases() -> dict:
"""Drop temp Neo4j scan databases whose scan is gone or long finished.
A failure listing databases aborts the whole sweep (nothing to iterate).
A failure reaping one database is logged and skipped so the rest of the
sweep still runs.
"""
now = datetime.now(tz=UTC)
safety_margin = timedelta(hours=ATTACK_PATHS_TMP_DB_REAP_SAFETY_MARGIN_HOURS)
try:
databases = graph_database.list_databases()
except Exception:
logger.exception("Failed to list ingest Neo4j databases for temp-db reap")
return {"dropped_count": 0, "databases": []}
tmp_databases = [
name for name in databases if name.startswith(graph_database.TEMP_DB_PREFIX)
]
dropped: list[str] = []
for database in tmp_databases:
try:
if _is_orphaned(database, now, safety_margin):
graph_database.drop_database(database)
dropped.append(database)
logger.info(f"Dropped orphaned temp Neo4j database `{database}`")
except Exception:
logger.exception(f"Failed to reap temp Neo4j database `{database}`")
logger.info(f"Temp Neo4j database reap: {len(dropped)} dropped")
return {"dropped_count": len(dropped), "databases": dropped}
def _is_orphaned(database: str, now: datetime, safety_margin: timedelta) -> bool:
"""Decide whether a temp database is safe to drop.
No scan row: the row was hard-deleted (tenant/provider cleanup) or was
never created. Falls back to the scan id's own UUIDv7 timestamp so a
database created moments ago is never touched even without a row to check.
Scan row present: only reapable once it reached a terminal state and has
been finished for longer than the safety margin, so a scan still
legitimately executing is never touched.
"""
scan_id = database[len(graph_database.TEMP_DB_PREFIX) :]
try:
scan_uuid = UUID7(scan_id)
except ValueError:
logger.warning(
f"Temp database `{database}` has an unparseable scan id, skipping"
)
return False
# Global sweep with no tenant context: admin_db bypasses RLS on purpose, the same
# way cleanup_stale_attack_paths_scans finds stale scans across every tenant.
scan = (
AttackPathsScan.all_objects.using(MainRouter.admin_db)
.filter(id=scan_uuid)
.first()
)
if scan is None:
if scan_uuid.version != 7:
logger.warning(
f"Temp database `{database}` has no scan row and a non-UUIDv7 id, "
"skipping"
)
return False
return now - datetime_from_uuid7(scan_uuid) >= safety_margin
if scan.state not in TERMINAL_STATES:
return False
# `mark_scan_finished` does not touch `updated_at`, so prefer `completed_at`
finished_at = scan.completed_at or scan.updated_at
return now - finished_at >= safety_margin
+6 -3
View File
@@ -499,10 +499,13 @@ def backfill_provider_compliance_scores(tenant_id: str) -> dict:
provider_id__in=existing_providers
)
# `completed_scans` keeps its own `completed_at__isnull=False`: this
# task writes a *dated* ProviderComplianceScore row, so unlike the read
# paths it genuinely cannot use a scan without a `completed_at`.
scan_info = list(
completed_scans.order_by("provider_id", "-completed_at")
.distinct("provider_id")
.values("id", "provider_id", "completed_at")
completed_scans.latest_per_provider().values(
"id", "provider_id", "completed_at"
)
)
if not scan_info:
+63 -4
View File
@@ -6,6 +6,7 @@ import boto3
import config.django.base as base
from api.db_utils import rls_transaction
from api.models import Scan
from botocore.config import Config
from botocore.exceptions import ClientError, NoCredentialsError, ParamValidationError
from celery.utils.log import get_task_logger
from django.conf import settings
@@ -206,15 +207,21 @@ def get_s3_client():
This function attempts to initialize an S3 client by reading the AWS access key, secret key,
session token, and region from environment variables. It then validates the client by listing
available S3 buckets. If an error occurs during this process (for example, due to missing or
invalid credentials), it falls back to creating an S3 client without explicitly provided credentials,
which may rely on other configuration sources (e.g., IAM roles).
invalid credentials), it falls back to creating an S3 client without explicitly provided
credentials, which may rely on other configuration sources (e.g., IAM roles).
That fallback is only safe when no explicit endpoint is configured: with an endpoint set, the
explicit client already targets the intended S3-compatible storage, and the fallback client
would go to the AWS default provider chain instead, an unrelated real-AWS account reachable
from the host. So when an endpoint is configured, the original error propagates instead.
Returns:
boto3.client: A configured S3 client instance.
Raises:
ClientError, NoCredentialsError, or ParamValidationError if both attempts to create a client fail.
ClientError, NoCredentialsError, or ParamValidationError if the client cannot be created.
"""
endpoint = settings.DJANGO_OUTPUT_S3_AWS_ENDPOINT_URL
s3_client = None
try:
s3_client = boto3.client(
@@ -222,16 +229,68 @@ def get_s3_client():
aws_access_key_id=settings.DJANGO_OUTPUT_S3_AWS_ACCESS_KEY_ID,
aws_secret_access_key=settings.DJANGO_OUTPUT_S3_AWS_SECRET_ACCESS_KEY,
aws_session_token=settings.DJANGO_OUTPUT_S3_AWS_SESSION_TOKEN,
region_name=settings.DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION,
# Storage that has no meaningful region, MinIO among it, is usually configured
# without one, and botocore rejects an empty region before any request is made.
region_name=settings.DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION or "us-east-1",
endpoint_url=endpoint or None,
)
s3_client.list_buckets()
except (ClientError, NoCredentialsError, ParamValidationError, ValueError):
if endpoint:
raise
s3_client = boto3.client("s3")
s3_client.list_buckets()
return s3_client
def get_s3_presign_client():
"""Return a client that signs download URLs with SigV4.
It is used when a public or internal storage host is configured, or when the bucket's
region is: boto3 otherwise presigns S3 URLs with SigV2, which S3 rejects for SSE-KMS
objects. None means none of those is set and the caller should presign with its own
client, which leaves those deployments with the URL they get today.
The public endpoint wins when both are set: the internal endpoint may only be reachable
from inside the cluster, and a URL signed against it would not open in a browser.
"""
endpoint = (
settings.DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL
or settings.DJANGO_OUTPUT_S3_AWS_ENDPOINT_URL
)
if not endpoint and not settings.DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION:
return None
# Blank keys are signed as-is (empty credential scope) instead of deferring to the
# provider chain, so static credentials are only passed when they are set.
credentials = {}
if (
settings.DJANGO_OUTPUT_S3_AWS_ACCESS_KEY_ID
and settings.DJANGO_OUTPUT_S3_AWS_SECRET_ACCESS_KEY
):
credentials = {
"aws_access_key_id": settings.DJANGO_OUTPUT_S3_AWS_ACCESS_KEY_ID,
"aws_secret_access_key": settings.DJANGO_OUTPUT_S3_AWS_SECRET_ACCESS_KEY,
# An empty string is a token as far as botocore is concerned: it appends an
# empty X-Amz-Security-Token that storage counts when it recomputes the signature.
"aws_session_token": settings.DJANGO_OUTPUT_S3_AWS_SESSION_TOKEN or None,
}
return boto3.client(
"s3",
**credentials,
# SigV4 puts the region in the credential scope, and MinIO answers to us-east-1
# unless it was told otherwise, so an empty region would sign an unusable URL.
region_name=settings.DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION or "us-east-1",
endpoint_url=endpoint or None,
# The signature covers the host, so the addressing style has to be pinned rather
# than guessed from the endpoint: MinIO serves path-style, and on AWS it keeps the
# regional host instead of the global one, which redirects for new buckets.
config=Config(signature_version="s3v4", s3={"addressing_style": "path"}),
)
def _upload_to_s3(
tenant_id: str, scan_id: str, local_path: str, relative_key: str
) -> str | None:
+79 -46
View File
@@ -1,63 +1,96 @@
from collections.abc import Iterable
from api.db_utils import rls_transaction
from api.models import Finding, MuteRule
from api.models import Finding, MuteRule, Scan
from celery.utils.log import get_task_logger
from config.django.base import DJANGO_FINDINGS_BATCH_SIZE
from tasks.utils import batched
logger = get_task_logger(__name__)
def mute_historical_findings(tenant_id: str, mute_rule_id: str):
"""
Mute historical findings that match the given mute rule.
def _mute_findings_for_rule(
*,
tenant_id: str,
scan_id: str,
finding_uids: Iterable[str],
muted_at,
muted_reason: str,
) -> int:
finding_uids = list(finding_uids)
if not finding_uids:
return 0
This function processes findings in batches, updating their muted status
and adding the mute reason.
return Finding.all_objects.filter(
tenant_id=tenant_id,
scan_id=scan_id,
uid__in=finding_uids,
muted=False,
).update(
muted=True,
muted_at=muted_at,
muted_reason=muted_reason,
)
Args:
tenant_id (str): The tenant ID for RLS context
mute_rule_id (str): The ID of the mute rule to apply
Returns:
dict: Summary of the muting operation with findings_muted count
"""
findings_muted_count = 0
def mute_findings_in_latest_scans(
tenant_id: str, mute_rule_id: str, provider_ids: list[str]
) -> dict:
"""Apply a mute rule to the latest completed scan of each provider."""
provider_ids = list(dict.fromkeys(provider_ids))
# Get the list of UIDs to mute and the reason
with rls_transaction(tenant_id):
mute_rule = MuteRule.objects.get(id=mute_rule_id, tenant_id=tenant_id)
finding_uids = mute_rule.finding_uids
mute_reason = mute_rule.reason
muted_at = mute_rule.inserted_at
latest_scans = Scan.objects.filter(
tenant_id=tenant_id, provider_id__in=provider_ids
).latest_ids_per_provider()
# Query findings that match the UIDs and are not already muted
with rls_transaction(tenant_id):
findings_to_mute = Finding.objects.filter(
tenant_id=tenant_id, uid__in=finding_uids, muted=False
)
total_findings = findings_to_mute.count()
logger.info(
f"Processing {total_findings} findings for mute rule {mute_rule_id}"
)
if total_findings > 0:
for batch, is_last in batched(
findings_to_mute.iterator(), DJANGO_FINDINGS_BATCH_SIZE
):
batch_ids = [f.id for f in batch]
updated_count = Finding.all_objects.filter(
id__in=batch_ids, tenant_id=tenant_id
).update(
muted=True,
muted_at=muted_at,
muted_reason=mute_reason,
)
findings_muted_count += updated_count
logger.info(f"Muted {findings_muted_count} findings for rule {mute_rule_id}")
changed_scan_ids = []
findings_muted = 0
for scan_id in latest_scans:
updated = _mute_findings_for_rule(
tenant_id=tenant_id,
scan_id=str(scan_id),
finding_uids=mute_rule.finding_uids,
muted_at=mute_rule.inserted_at,
muted_reason=mute_rule.reason,
)
if updated:
findings_muted += updated
changed_scan_ids.append(str(scan_id))
logger.info(
"Muted %d findings in %d latest scans for rule %s",
findings_muted,
len(changed_scan_ids),
mute_rule_id,
)
return {
"findings_muted": findings_muted_count,
"findings_muted": findings_muted,
"rule_id": mute_rule_id,
"scan_ids": changed_scan_ids,
}
def reconcile_scan_mute_rules(tenant_id: str, scan_id: str) -> dict:
"""Apply the current enabled mute rules to one completed scan."""
findings_muted = 0
with rls_transaction(tenant_id):
mute_rules = MuteRule.objects.filter(tenant_id=tenant_id, enabled=True).values(
"finding_uids", "reason", "inserted_at"
)
for mute_rule in mute_rules:
findings_muted += _mute_findings_for_rule(
tenant_id=tenant_id,
scan_id=scan_id,
finding_uids=mute_rule["finding_uids"],
muted_at=mute_rule["inserted_at"],
muted_reason=mute_rule["reason"],
)
logger.info(
"Reconciled mute rules for scan %s; muted %d findings",
scan_id,
findings_muted,
)
return {"findings_muted": findings_muted, "scan_id": str(scan_id)}
@@ -84,6 +84,7 @@ _SKIP_RECOVERY = {
"scan-perform-scheduled",
"attack-paths-scan-perform",
"attack-paths-cleanup-stale-scans",
"attack-paths-reap-orphaned-tmp-databases",
"reconcile-orphan-tasks",
}
+87 -49
View File
@@ -5,7 +5,6 @@ import json
import random
import re
import time
import uuid
from collections import defaultdict
from collections.abc import Callable, Iterable
from datetime import UTC, datetime
@@ -73,6 +72,7 @@ from tasks.jobs.queries import (
COMPLIANCE_UPSERT_TENANT_SUMMARY_SQL,
)
from tasks.utils import CustomEncoder, batched
from uuid6 import uuid7
logger = get_task_logger(__name__)
@@ -1756,32 +1756,27 @@ def aggregate_findings(tenant_id: str, scan_id: str):
def _aggregate_findings_by_region(
tenant_id: str, scan_id: str, modeled_threatscore_compliance_id: str
tenant_id: str,
scan_id: str,
normalized_threatscore_id: str,
threatscore_requirements_by_check: dict[str, list[str]],
) -> tuple[dict, dict]:
"""
Aggregate findings by region using streaming, column-scoped ORM reads.
Reads only the consumed columns as tuples via ``values_list`` and streams
them with ``.iterator()``, using the denormalized ``resource_regions`` array
instead of ``prefetch_related("resources")``. ``resource_regions`` mirrors the
regions of a finding's related resources, so it yields the same per-region
tally without joining the resource table.
Args:
tenant_id: Tenant UUID
scan_id: Scan UUID
modeled_threatscore_compliance_id: ID for ThreatScore compliance framework
instead of ``prefetch_related("resources")``. ThreatScore requirement ids
are resolved per ``check_id`` from ``threatscore_requirements_by_check``.
Returns:
tuple: (check_status_by_region, findings_count_by_compliance)
- check_status_by_region: {region: {check_id: status}}
- findings_count_by_compliance: {region: {normalized_id: {requirement_id: {total, pass}}}}
- findings_count_by_compliance: {region: {normalized_threatscore_id: {requirement_id: {total, pass}}}}
"""
check_status_by_region: dict = {}
findings_count_by_compliance: dict = {}
normalized_id = re.sub(r"[^a-z0-9]", "", modeled_threatscore_compliance_id.lower())
with rls_transaction(tenant_id, using=READ_REPLICA_ALIAS):
findings = (
Finding.all_objects.filter(
@@ -1790,14 +1785,12 @@ def _aggregate_findings_by_region(
muted=False,
status__in=["PASS", "FAIL"],
)
.values_list("check_id", "status", "resource_regions", "compliance")
.values_list("check_id", "status", "resource_regions")
.iterator(chunk_size=DJANGO_FINDINGS_BATCH_SIZE)
)
for check_id, status, resource_regions, compliance in findings:
threatscore_requirements = (compliance or {}).get(
modeled_threatscore_compliance_id
)
for check_id, status, resource_regions in findings:
threatscore_requirements = threatscore_requirements_by_check.get(check_id)
for region in resource_regions or ():
# Priority: FAIL > any other status
@@ -1809,7 +1802,7 @@ def _aggregate_findings_by_region(
if threatscore_requirements:
compliance_key = findings_count_by_compliance.setdefault(
region, {}
).setdefault(normalized_id, {})
).setdefault(normalized_threatscore_id, {})
for requirement_id in threatscore_requirements:
requirement_stats = compliance_key.setdefault(
@@ -1848,15 +1841,28 @@ def create_compliance_requirements(tenant_id: str, scan_id: str):
compliance_template = PROWLER_COMPLIANCE_OVERVIEW_TEMPLATE[
provider_instance.provider
]
modeled_threatscore_compliance_id = "ProwlerThreatScore-1.0"
normalized_threatscore_id = _normalized_compliance_key(
"ProwlerThreatScore", "1.0"
)
requirement_lookup: dict[str, list[tuple[str, str]]] = {}
threatscore_requirements_by_check: dict[str, list[str]] = {}
for compliance_id, compliance in compliance_template.items():
is_threatscore = (
_normalized_compliance_key(
compliance["framework"], compliance["version"]
)
== normalized_threatscore_id
)
for requirement_id, requirement in compliance["requirements"].items():
for check_id in requirement["checks"].keys():
requirement_lookup.setdefault(check_id, []).append(
(compliance_id, requirement_id)
)
if is_threatscore:
threatscore_requirements_by_check.setdefault(
check_id, []
).append(requirement_id)
regions = []
requirements_created = 0
@@ -1869,7 +1875,10 @@ def create_compliance_requirements(tenant_id: str, scan_id: str):
# Aggregate findings by region using SQL for optimal performance
check_status_by_region, findings_count_by_compliance = (
_aggregate_findings_by_region(
tenant_id, scan_id, modeled_threatscore_compliance_id
tenant_id,
scan_id,
normalized_threatscore_id,
threatscore_requirements_by_check,
)
)
@@ -1934,23 +1943,35 @@ def create_compliance_requirements(tenant_id: str, scan_id: str):
# Yield rows lazily (consumed batch-by-batch by COPY) so peak memory
# stays bounded; tally requirement_statuses in the same pass. The
# ORM fallback re-iterates from scratch, so the tally resets first.
# Region is the innermost loop so consecutive rows share the leading
# columns of the table's secondary indexes.
def _iter_compliance_requirement_rows():
requirement_statuses.clear()
for region in regions:
region_stats = region_requirement_stats.get(region, {})
region_findings = findings_count_by_compliance.get(region, {})
for (
compliance_id,
framework,
version,
modeled_compliance_id,
requirements,
) in compliance_plan:
compliance_stats = region_stats.get(compliance_id, {})
compliance_findings = region_findings.get(
modeled_compliance_id, {}
for (
compliance_id,
framework,
version,
modeled_compliance_id,
requirements,
) in compliance_plan:
stats_by_region = [
(
region,
region_requirement_stats.get(region, {}).get(
compliance_id, {}
),
findings_count_by_compliance.get(region, {}).get(
modeled_compliance_id, {}
),
)
for requirement_id, description, total_checks in requirements:
for region in regions
]
for requirement_id, description, total_checks in requirements:
for (
region,
compliance_stats,
compliance_findings,
) in stats_by_region:
stats = compliance_stats.get(requirement_id)
if stats:
passed_checks = stats["passed_checks"]
@@ -1981,7 +2002,7 @@ def create_compliance_requirements(tenant_id: str, scan_id: str):
requirement_statuses[key]["pass_count"] += 1
yield {
"id": uuid.uuid4(),
"id": uuid7(),
"tenant_id": tenant_id_str,
"inserted_at": utc_datetime_now,
"compliance_id": compliance_id,
@@ -2679,20 +2700,37 @@ def reset_ephemeral_resource_findings_count(tenant_id: str, scan_id: str) -> dic
# refreshed). Wiping based on the older scan would zero counts the newer
# scan just set. Skip and let the newer scan's reset task do the work; if
# this task was delayed in the queue, that's the correct outcome.
# `completed_at__isnull=False` is required: Postgres orders NULL first in
# DESC, so a sibling COMPLETED scan with a missing completed_at would sort
# as "newest" and incorrectly cause us to skip.
#
# The comparison must be against the newest *full-scope* scan, which is
# what this variable has always been named after but did not use to be:
# the query filtered nothing about scope, so any newer scan that is not
# full-scope (an imported one, for instance) made the full-scope scan
# skip its own cleanup and leave ephemeral resources with a stale
# failed_findings_count permanently.
#
# `is_full_scope()` reads `trigger` plus the scoping keys inside
# `scanner_args`, which is not expressible as a WHERE clause, so the
# candidates are walked newest-first in Python until the first full-scope
# one. The walk needs no cap: `scan` is itself a full-scope candidate, so
# it stops at `scan` at the latest, after reading only the scans newer than
# it. A fixed window would return None once more newer scoped scans had
# landed than it inspected, and skip the cleanup exactly like the bug above.
#
# NULL `completed_at` no longer needs an explicit filter here: the shared
# ordering in `ScanQuerySet.LATEST_ORDER_BY` sorts NULLs last
# rather than excluding them, which also fixes the case where a provider
# whose completed scans all have a NULL `completed_at` resolved to None and
# therefore never ran the reset at all.
with rls_transaction(tenant_id):
latest_full_scope_scan_id = (
Scan.objects.filter(
tenant_id=tenant_id,
provider_id=scan.provider_id,
state=StateChoices.COMPLETED,
completed_at__isnull=False,
)
.order_by("-completed_at", "-inserted_at")
.values_list("id", flat=True)
.first()
candidates = (
Scan.objects.filter(tenant_id=tenant_id, provider_id=scan.provider_id)
.latest_first()
.only("id", "trigger", "scanner_args")
.iterator(chunk_size=100)
)
latest_full_scope_scan_id = next(
(candidate.id for candidate in candidates if candidate.is_full_scope()),
None,
)
if latest_full_scope_scan_id != scan.id:
logger.info(
+70 -100
View File
@@ -1,3 +1,4 @@
import json
import os
from datetime import UTC, datetime, timedelta
from pathlib import Path
@@ -42,6 +43,7 @@ from tasks.jobs.attack_paths import (
)
from tasks.jobs.attack_paths import db_utils as attack_paths_db_utils
from tasks.jobs.attack_paths.cleanup import cleanup_stale_attack_paths_scans
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
from tasks.jobs.backfill import (
aggregate_scan_category_summaries,
aggregate_scan_resource_group_summaries,
@@ -73,7 +75,10 @@ from tasks.jobs.lighthouse_providers import (
check_lighthouse_provider_connection,
refresh_lighthouse_provider_models,
)
from tasks.jobs.muting import mute_historical_findings
from tasks.jobs.muting import (
mute_findings_in_latest_scans,
reconcile_scan_mute_rules,
)
from tasks.jobs.orphan_recovery import reconcile_orphans
from tasks.jobs.report import (
STALE_TMP_OUTPUT_MAX_AGE_HOURS,
@@ -160,13 +165,28 @@ def create_scan_task_record(
task_id: str,
task_name: str = "scan-perform",
task_status: str | None = states.PENDING,
task_kwargs: dict | None = None,
) -> Task:
"""Pre-create the TaskResult + Task rows for a pre-generated task id.
Pass ``task_kwargs`` when the response built from this record is serialized
before the broker publish. ``task_kwargs`` is otherwise only written by the
``before_task_publish`` signal (``api/signals.py``), and the scan publish is
deferred to ``on_commit``, so the 202 would carry an empty ``task_args`` and
the caller would have no way to learn the scan id it was just handed a task
for. The publish later overwrites the field with the same kwargs as a Python
repr; both forms decode to the same dict (``decode_celery_field``).
"""
if task_status is None:
task_status = states.PENDING
defaults = {"status": task_status, "task_name": task_name}
if task_kwargs is not None:
defaults["task_kwargs"] = json.dumps(task_kwargs)
task_result, _ = TaskResult.objects.update_or_create(
task_id=str(task_id),
defaults={"status": task_status, "task_name": task_name},
defaults=defaults,
)
prowler_task, _ = Task.objects.update_or_create(
id=str(task_id),
@@ -526,6 +546,7 @@ def perform_scan_task(
provider_id=provider_id,
checks_to_execute=checks_to_execute,
)
reconcile_scan_mute_rules(tenant_id, scan_id)
_perform_scan_complete_tasks(tenant_id, scan_id, provider_id)
return result
finally:
@@ -635,6 +656,7 @@ def perform_scheduled_scan_task(self, tenant_id: str, provider_id: str):
scan_id=str(scan_instance.id),
provider_id=provider_id,
)
reconcile_scan_mute_rules(tenant_id, str(scan_instance.id))
_perform_scan_complete_tasks(tenant_id, str(scan_instance.id), provider_id)
return result
finally:
@@ -706,6 +728,13 @@ def cleanup_stale_attack_paths_scans_task():
return cleanup_stale_attack_paths_scans()
@shared_task(
name="attack-paths-reap-orphaned-tmp-databases", queue="attack-paths-scans"
)
def reap_orphaned_attack_paths_tmp_databases_task():
return reap_orphaned_tmp_databases()
@shared_task(name="reconcile-orphan-tasks", queue="celery")
def reconcile_orphan_tasks_task():
"""Periodic watchdog: recover tasks whose worker is gone (deploys, crashes)."""
@@ -1188,85 +1217,48 @@ def aggregate_finding_group_summaries_task(tenant_id: str, scan_id: str):
return aggregate_finding_group_summaries(tenant_id=tenant_id, scan_id=scan_id)
@shared_task(
base=RLSTask, name="reaggregate-all-finding-group-summaries", queue="overview"
)
@set_tenant(keep_tenant=True)
def reaggregate_all_finding_group_summaries_task(tenant_id: str):
"""Reaggregate every pre-aggregated summary table for this tenant.
def _dispatch_scan_summary_reaggregation(tenant_id: str, scan_ids: list[str]) -> None:
if not scan_ids:
return
Mirrors the unbounded scope of `mute_historical_findings_task`: that task
rewrites every Finding row whose UID matches a mute rule, with no time
limit. To keep the pre-aggregated tables consistent with that update,
this task re-runs the same per-scan aggregation pipeline that scan
completion runs on the latest completed scan of every (provider, day)
pair, rebuilding the tables that power the read endpoints:
- `ScanSummary` and `DailySeveritySummary` -> `/overviews/findings`,
`/overviews/findings-severity`, `/overviews/services`.
- `FindingGroupDailySummary` -> `/finding-groups` and
`/finding-groups/latest`.
- `ScanGroupSummary` -> `/overviews/resource-groups` (resource
inventory).
- `ScanCategorySummary` -> `/overviews/categories`.
- `AttackSurfaceOverview` -> `/overviews/attack-surfaces`.
Per-scan pipelines are dispatched in parallel via a Celery group so
wallclock scales with the worker pool.
"""
completed_scans = list(
Scan.objects.filter(
tenant_id=tenant_id,
state=StateChoices.COMPLETED,
completed_at__isnull=False,
)
.order_by("-completed_at")
.values("id", "completed_at", "provider_id")
logger.info(
"Reaggregating overview/finding summaries for %d latest scans",
len(scan_ids),
)
# Keep the latest scan per (provider, day) pair so the daily summary row
# the aggregator writes is the most recent snapshot of that day for that
# provider. Iterating from most recent to oldest means the first scan we
# see for a given key wins.
latest_scans: dict[tuple, str] = {}
for scan in completed_scans:
key = (scan["provider_id"], scan["completed_at"].date())
if key not in latest_scans:
latest_scans[key] = str(scan["id"])
scan_ids = list(latest_scans.values())
if scan_ids:
logger.info(
"Reaggregating overview/finding summaries for %d scans (provider x day)",
len(scan_ids),
)
# DailySeveritySummary reads from ScanSummary, so ScanSummary must be
# recomputed first; the other aggregators read Finding directly and
# can run in parallel with the severity step.
group(
chain(
perform_scan_summary_task.si(tenant_id=tenant_id, scan_id=scan_id),
group(
aggregate_daily_severity_task.si(
tenant_id=tenant_id, scan_id=scan_id
),
aggregate_finding_group_summaries_task.si(
tenant_id=tenant_id, scan_id=scan_id
),
aggregate_scan_resource_group_summaries_task.si(
tenant_id=tenant_id, scan_id=scan_id
),
aggregate_scan_category_summaries_task.si(
tenant_id=tenant_id, scan_id=scan_id
),
aggregate_attack_surface_task.si(
tenant_id=tenant_id, scan_id=scan_id
),
group(
chain(
perform_scan_summary_task.si(tenant_id=tenant_id, scan_id=scan_id),
group(
aggregate_daily_severity_task.si(tenant_id=tenant_id, scan_id=scan_id),
aggregate_finding_group_summaries_task.si(
tenant_id=tenant_id, scan_id=scan_id
),
)
for scan_id in scan_ids
).apply_async()
return {"scans_reaggregated": len(scan_ids)}
aggregate_scan_resource_group_summaries_task.si(
tenant_id=tenant_id, scan_id=scan_id
),
aggregate_scan_category_summaries_task.si(
tenant_id=tenant_id, scan_id=scan_id
),
aggregate_attack_surface_task.si(tenant_id=tenant_id, scan_id=scan_id),
),
)
for scan_id in scan_ids
).apply_async()
@shared_task(base=RLSTask, name="findings-mute-latest-scans", queue="overview")
@set_tenant(keep_tenant=True)
def mute_findings_in_latest_scans_task(
tenant_id: str, mute_rule_id: str, provider_ids: list[str]
):
"""Apply a mute rule to current scans and rebuild only changed summaries."""
result = mute_findings_in_latest_scans(
tenant_id=tenant_id,
mute_rule_id=mute_rule_id,
provider_ids=provider_ids,
)
_dispatch_scan_summary_reaggregation(tenant_id, result["scan_ids"])
return result
@shared_task(base=RLSTask, name="lighthouse-connection-check")
@@ -1467,25 +1459,3 @@ def generate_compliance_reports_task(tenant_id: str, scan_id: str, provider_id:
generate_csa=True,
generate_cis=True,
)
@shared_task(name="findings-mute-historical")
def mute_historical_findings_task(tenant_id: str, mute_rule_id: str):
"""
Background task to mute all historical findings matching a mute rule.
This task processes findings in batches to avoid memory issues with large datasets.
It updates the Finding.muted, Finding.muted_at, and Finding.muted_reason fields
for all findings whose UID is in the mute rule's finding_uids list.
Args:
tenant_id (str): The tenant ID for RLS context.
mute_rule_id (str): The primary key of the MuteRule to apply.
Returns:
dict: A dictionary containing:
- 'findings_muted' (int): Total number of findings muted.
- 'rule_id' (str): The mute rule ID.
- 'status' (str): Final status ('completed').
"""
return mute_historical_findings(tenant_id, mute_rule_id)
@@ -0,0 +1,286 @@
from datetime import UTC, datetime, timedelta
from unittest.mock import patch
from uuid import uuid4
import pytest
from api.attack_paths.database import TEMP_DB_PREFIX
from api.models import AttackPathsScan, StateChoices
from api.uuid_utils import datetime_to_uuid7
from config.django.base import ATTACK_PATHS_TMP_DB_REAP_SAFETY_MARGIN_HOURS
MARGIN = timedelta(hours=ATTACK_PATHS_TMP_DB_REAP_SAFETY_MARGIN_HOURS)
def _tmp_db_name(scan_uuid) -> str:
return f"{TEMP_DB_PREFIX}{scan_uuid}"
@pytest.mark.django_db
class TestReapOrphanedTmpDatabases:
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.drop_database")
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_ignores_databases_without_the_temp_prefix(self, mock_list, mock_drop):
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
mock_list.return_value = ["db-tenant-abc123", "system", "neo4j"]
result = reap_orphaned_tmp_databases()
assert result == {"dropped_count": 0, "databases": []}
mock_drop.assert_not_called()
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.drop_database")
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_drops_temp_db_with_no_scan_row_past_safety_margin(
self, mock_list, mock_drop
):
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
old_scan_id = datetime_to_uuid7(
datetime.now(tz=UTC) - MARGIN - timedelta(hours=1)
)
database = _tmp_db_name(old_scan_id)
mock_list.return_value = [database]
result = reap_orphaned_tmp_databases()
assert result == {"dropped_count": 1, "databases": [database]}
mock_drop.assert_called_once_with(database)
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.drop_database")
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_preserves_temp_db_with_no_scan_row_inside_safety_margin(
self, mock_list, mock_drop
):
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
recent_scan_id = datetime_to_uuid7(datetime.now(tz=UTC) - timedelta(minutes=5))
database = _tmp_db_name(recent_scan_id)
mock_list.return_value = [database]
result = reap_orphaned_tmp_databases()
assert result == {"dropped_count": 0, "databases": []}
mock_drop.assert_not_called()
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.drop_database")
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_preserves_temp_db_with_unparseable_scan_id(self, mock_list, mock_drop):
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
database = f"{TEMP_DB_PREFIX}not-a-uuid"
mock_list.return_value = [database]
result = reap_orphaned_tmp_databases()
assert result == {"dropped_count": 0, "databases": []}
mock_drop.assert_not_called()
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.drop_database")
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_drops_terminal_scan_past_safety_margin(
self, mock_list, mock_drop, tenants_fixture, aws_provider
):
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
tenant = tenants_fixture[0]
old_updated_at = datetime.now(tz=UTC) - MARGIN - timedelta(hours=1)
scan = AttackPathsScan.objects.create(
tenant_id=tenant.id,
provider=aws_provider,
state=StateChoices.COMPLETED,
)
AttackPathsScan.objects.filter(id=scan.id).update(updated_at=old_updated_at)
database = _tmp_db_name(scan.id)
mock_list.return_value = [database]
result = reap_orphaned_tmp_databases()
assert result == {"dropped_count": 1, "databases": [database]}
mock_drop.assert_called_once_with(database)
@pytest.mark.parametrize(
"state",
[StateChoices.FAILED, StateChoices.CANCELLED],
)
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.drop_database")
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_drops_other_terminal_states_past_safety_margin(
self, mock_list, mock_drop, tenants_fixture, aws_provider, state
):
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
tenant = tenants_fixture[0]
old_updated_at = datetime.now(tz=UTC) - MARGIN - timedelta(hours=1)
scan = AttackPathsScan.objects.create(
tenant_id=tenant.id,
provider=aws_provider,
state=state,
)
AttackPathsScan.objects.filter(id=scan.id).update(updated_at=old_updated_at)
database = _tmp_db_name(scan.id)
mock_list.return_value = [database]
result = reap_orphaned_tmp_databases()
assert result["dropped_count"] == 1
mock_drop.assert_called_once_with(database)
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.drop_database")
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_preserves_terminal_scan_inside_safety_margin(
self, mock_list, mock_drop, tenants_fixture, aws_provider
):
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
tenant = tenants_fixture[0]
scan = AttackPathsScan.objects.create(
tenant_id=tenant.id,
provider=aws_provider,
state=StateChoices.COMPLETED,
)
database = _tmp_db_name(scan.id)
mock_list.return_value = [database]
result = reap_orphaned_tmp_databases()
assert result == {"dropped_count": 0, "databases": []}
mock_drop.assert_not_called()
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.drop_database")
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_margin_counts_from_completed_at_not_updated_at(
self, mock_list, mock_drop, tenants_fixture, aws_provider
):
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
tenant = tenants_fixture[0]
old = datetime.now(tz=UTC) - MARGIN - timedelta(hours=1)
scan = AttackPathsScan.objects.create(
tenant_id=tenant.id,
provider=aws_provider,
state=StateChoices.COMPLETED,
)
AttackPathsScan.objects.filter(id=scan.id).update(
updated_at=old, completed_at=datetime.now(tz=UTC)
)
database = _tmp_db_name(scan.id)
mock_list.return_value = [database]
result = reap_orphaned_tmp_databases()
assert result == {"dropped_count": 0, "databases": []}
mock_drop.assert_not_called()
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.drop_database")
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_drops_scan_completed_past_safety_margin(
self, mock_list, mock_drop, tenants_fixture, aws_provider
):
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
tenant = tenants_fixture[0]
old = datetime.now(tz=UTC) - MARGIN - timedelta(hours=1)
scan = AttackPathsScan.objects.create(
tenant_id=tenant.id,
provider=aws_provider,
state=StateChoices.FAILED,
)
AttackPathsScan.objects.filter(id=scan.id).update(
updated_at=old, completed_at=old
)
database = _tmp_db_name(scan.id)
mock_list.return_value = [database]
result = reap_orphaned_tmp_databases()
assert result == {"dropped_count": 1, "databases": [database]}
mock_drop.assert_called_once_with(database)
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.drop_database")
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_never_drops_an_executing_scan_regardless_of_age(
self, mock_list, mock_drop, tenants_fixture, aws_provider
):
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
tenant = tenants_fixture[0]
very_old = datetime.now(tz=UTC) - timedelta(days=30)
scan = AttackPathsScan.objects.create(
tenant_id=tenant.id,
provider=aws_provider,
state=StateChoices.EXECUTING,
)
AttackPathsScan.objects.filter(id=scan.id).update(updated_at=very_old)
database = _tmp_db_name(scan.id)
mock_list.return_value = [database]
result = reap_orphaned_tmp_databases()
assert result == {"dropped_count": 0, "databases": []}
mock_drop.assert_not_called()
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.drop_database")
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_one_failed_drop_does_not_stop_the_rest_of_the_sweep(
self, mock_list, mock_drop
):
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
old_time = datetime.now(tz=UTC) - MARGIN - timedelta(hours=1)
failing_scan_id = datetime_to_uuid7(old_time)
succeeding_scan_id = datetime_to_uuid7(old_time)
failing_db = _tmp_db_name(failing_scan_id)
succeeding_db = _tmp_db_name(succeeding_scan_id)
mock_list.return_value = [failing_db, succeeding_db]
mock_drop.side_effect = [Exception("boom"), None]
result = reap_orphaned_tmp_databases()
assert result == {"dropped_count": 1, "databases": [succeeding_db]}
assert mock_drop.call_count == 2
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_returns_empty_result_when_listing_databases_fails(self, mock_list):
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
mock_list.side_effect = Exception("neo4j unreachable")
result = reap_orphaned_tmp_databases()
assert result == {"dropped_count": 0, "databases": []}
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.drop_database")
@patch("tasks.jobs.attack_paths.tmp_db_reaper.graph_database.list_databases")
def test_preserves_temp_db_with_random_uuid_and_no_row(self, mock_list, mock_drop):
"""A non-UUIDv7 id with no matching row has no reliable timestamp, so it
must be left alone rather than guessed at."""
from tasks.jobs.attack_paths.tmp_db_reaper import reap_orphaned_tmp_databases
database = _tmp_db_name(uuid4())
mock_list.return_value = [database]
result = reap_orphaned_tmp_databases()
assert result == {"dropped_count": 0, "databases": []}
mock_drop.assert_not_called()
class TestReapOrphanedTmpDatabasesTask:
@patch(
"tasks.tasks.reap_orphaned_tmp_databases",
return_value={"dropped_count": 2, "databases": ["db-tmp-scan-a"]},
)
def test_task_invokes_the_reaper(self, mock_reap):
from tasks.tasks import reap_orphaned_attack_paths_tmp_databases_task
result = reap_orphaned_attack_paths_tmp_databases_task.run()
assert result == {"dropped_count": 2, "databases": ["db-tmp-scan-a"]}
mock_reap.assert_called_once_with()
+213 -1
View File
@@ -3,16 +3,20 @@ import uuid
import zipfile
from datetime import datetime
from pathlib import Path
from unittest.mock import MagicMock, patch
from unittest.mock import MagicMock, call, patch
from urllib.parse import parse_qs, urlparse
import boto3
import pytest
from botocore.exceptions import ClientError
from django.test import override_settings
from tasks.jobs.export import (
_compress_output_files,
_generate_compliance_output_directory,
_generate_output_directory,
_upload_to_s3,
get_s3_client,
get_s3_presign_client,
)
@@ -47,15 +51,59 @@ class TestOutputs:
assert client is not None
client_mock.list_buckets.assert_called()
@patch("tasks.jobs.export.boto3.client")
@override_settings(
DJANGO_OUTPUT_S3_AWS_ACCESS_KEY_ID="access-key",
DJANGO_OUTPUT_S3_AWS_SECRET_ACCESS_KEY="secret-key",
DJANGO_OUTPUT_S3_AWS_SESSION_TOKEN="",
DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION="",
)
def test_get_s3_client_without_a_region_uses_a_default(self, mock_boto_client):
"""botocore rejects an empty region up front, and the download views do not catch it."""
get_s3_client()
assert mock_boto_client.call_args.kwargs["region_name"] == "us-east-1"
@patch("tasks.jobs.export.boto3.client")
@override_settings(DJANGO_OUTPUT_S3_AWS_ENDPOINT_URL="http://minio:9000")
def test_get_s3_client_passes_the_endpoint_when_set(self, mock_boto_client):
get_s3_client()
assert mock_boto_client.call_args.kwargs["endpoint_url"] == "http://minio:9000"
@patch("tasks.jobs.export.boto3.client")
@override_settings(DJANGO_OUTPUT_S3_AWS_ENDPOINT_URL="")
def test_get_s3_client_endpoint_empty_by_default(self, mock_boto_client):
"""Empty keeps today's behavior: no endpoint override, real S3 is used."""
get_s3_client()
assert mock_boto_client.call_args.kwargs["endpoint_url"] is None
@patch("tasks.jobs.export.boto3.client")
@override_settings(DJANGO_OUTPUT_S3_AWS_ENDPOINT_URL="http://minio:9000")
def test_get_s3_client_does_not_fall_back_when_endpoint_set(self, mock_boto_client):
"""A configured endpoint means the explicit client failed talking to it. The fallback
goes to the default provider chain (e.g. an EC2 instance role) against real AWS, so it
must not be used: the original error propagates instead."""
error = ClientError({"Error": {"Code": "403"}}, "ListBuckets")
mock_boto_client.side_effect = error
with pytest.raises(ClientError):
get_s3_client()
mock_boto_client.assert_called_once()
@patch("tasks.jobs.export.boto3.client")
@patch("tasks.jobs.export.settings")
def test_get_s3_client_fallback(self, mock_settings, mock_boto_client):
mock_settings.DJANGO_OUTPUT_S3_AWS_ENDPOINT_URL = ""
mock_boto_client.side_effect = [
ClientError({"Error": {"Code": "403"}}, "ListBuckets"),
MagicMock(),
]
client = get_s3_client()
assert client is not None
assert mock_boto_client.call_args_list[1] == call("s3")
@patch("tasks.jobs.export.get_s3_client")
@patch("tasks.jobs.export.base")
@@ -243,3 +291,167 @@ class TestOutputs:
assert os.path.isdir(os.path.dirname(ens))
assert threatscore.endswith(f"aws-test-check-{expected_timestamp}")
assert ens.endswith(f"aws-test-check-{expected_timestamp}")
PRESIGN_SETTINGS = {
"DJANGO_OUTPUT_S3_AWS_ACCESS_KEY_ID": "access-key",
"DJANGO_OUTPUT_S3_AWS_SECRET_ACCESS_KEY": "secret-key",
"DJANGO_OUTPUT_S3_AWS_SESSION_TOKEN": "",
"DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION": "eu-west-1",
}
def _presign(client):
return client.generate_presigned_url(
"get_object",
Params={"Bucket": "output-bucket", "Key": "tenant/scan/report.zip"},
ExpiresIn=300,
)
class TestS3PresignClient:
@override_settings(
**{**PRESIGN_SETTINGS, "DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION": ""},
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="",
)
def test_no_public_endpoint_and_no_region_returns_none(self):
# Without a region, SigV4 would have to guess one and break other regions.
assert get_s3_presign_client() is None
@override_settings(**PRESIGN_SETTINGS, DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="")
def test_region_without_public_endpoint_signs_sigv4_on_the_regional_host(self):
# SSE-KMS objects reject the SigV2 URLs boto3 presigns by default, and the
# global host redirects for new buckets, which breaks a SigV4 signature.
url = urlparse(_presign(get_s3_presign_client()))
query = parse_qs(url.query)
assert url.netloc == "s3.eu-west-1.amazonaws.com"
assert url.path == "/output-bucket/tenant/scan/report.zip"
assert query["X-Amz-Algorithm"] == ["AWS4-HMAC-SHA256"]
assert "/eu-west-1/s3/aws4_request" in query["X-Amz-Credential"][0]
@override_settings(
**{
**PRESIGN_SETTINGS,
"DJANGO_OUTPUT_S3_AWS_ACCESS_KEY_ID": "",
"DJANGO_OUTPUT_S3_AWS_SECRET_ACCESS_KEY": "",
},
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="",
)
def test_region_without_static_keys_signs_with_the_default_chain(self, monkeypatch):
# An ECS task role reaches boto3 through the default chain, like the env here.
# A fresh default session keeps these keys from being cached for later tests.
monkeypatch.setattr(boto3, "DEFAULT_SESSION", None)
monkeypatch.setenv("AWS_ACCESS_KEY_ID", "role-access-key")
monkeypatch.setenv("AWS_SECRET_ACCESS_KEY", "role-secret-key")
monkeypatch.setenv("AWS_DEFAULT_REGION", "us-east-1")
query = parse_qs(urlparse(_presign(get_s3_presign_client())).query)
assert query["X-Amz-Credential"][0].startswith("role-access-key/")
assert "/eu-west-1/s3/aws4_request" in query["X-Amz-Credential"][0]
@override_settings(
**{**PRESIGN_SETTINGS, "DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION": ""},
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="",
DJANGO_OUTPUT_S3_AWS_ENDPOINT_URL="http://minio:9000",
)
def test_internal_endpoint_without_public_endpoint_signs_against_it(self):
# No browser-reachable host was configured, so the internal one is the best
# available target instead of falling through to the real AWS host.
url = urlparse(_presign(get_s3_presign_client()))
assert url.netloc == "minio:9000"
assert url.path == "/output-bucket/tenant/scan/report.zip"
@override_settings(
**PRESIGN_SETTINGS,
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="https://storage.example.com",
DJANGO_OUTPUT_S3_AWS_ENDPOINT_URL="http://minio:9000",
)
def test_public_endpoint_wins_over_the_internal_endpoint(self):
url = urlparse(_presign(get_s3_presign_client()))
assert url.netloc == "storage.example.com"
@override_settings(
**PRESIGN_SETTINGS,
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="https://storage.example.com",
)
def test_url_targets_the_public_host_in_path_style(self):
url = urlparse(_presign(get_s3_presign_client()))
assert url.netloc == "storage.example.com"
assert url.path == "/output-bucket/tenant/scan/report.zip"
@override_settings(
**PRESIGN_SETTINGS,
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="https://storage.example.com",
)
def test_signature_covers_the_public_host(self):
query = parse_qs(urlparse(_presign(get_s3_presign_client())).query)
assert query["X-Amz-SignedHeaders"] == ["host"]
assert "/eu-west-1/s3/aws4_request" in query["X-Amz-Credential"][0]
@override_settings(
**PRESIGN_SETTINGS,
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="https://storage.example.com",
)
def test_signature_is_bound_to_the_host_it_was_signed_against(self):
"""Rewriting the host afterwards cannot work, which is why the endpoint is a setting."""
public = parse_qs(urlparse(_presign(get_s3_presign_client())).query)
with override_settings(
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="http://minio:9000"
):
internal = parse_qs(urlparse(_presign(get_s3_presign_client())).query)
assert public["X-Amz-Signature"] != internal["X-Amz-Signature"]
@override_settings(
**{**PRESIGN_SETTINGS, "DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION": ""},
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="https://storage.example.com",
)
def test_region_falls_back_to_the_minio_default(self):
query = parse_qs(urlparse(_presign(get_s3_presign_client())).query)
assert "/us-east-1/s3/aws4_request" in query["X-Amz-Credential"][0]
@override_settings(
**PRESIGN_SETTINGS,
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="https://storage.example.com",
)
def test_unset_session_token_is_left_out_of_the_url(self):
"""An empty token still reaches the URL as a blank param that storage signs over."""
url = _presign(get_s3_presign_client())
query = parse_qs(urlparse(url).query, keep_blank_values=True)
assert "X-Amz-Security-Token" not in query
@override_settings(
**{**PRESIGN_SETTINGS, "DJANGO_OUTPUT_S3_AWS_SESSION_TOKEN": "session-token"},
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="https://storage.example.com",
)
def test_session_token_is_forwarded_when_set(self):
query = parse_qs(urlparse(_presign(get_s3_presign_client())).query)
assert query["X-Amz-Security-Token"] == ["session-token"]
@override_settings(
**{
**PRESIGN_SETTINGS,
"DJANGO_OUTPUT_S3_AWS_ACCESS_KEY_ID": "",
"DJANGO_OUTPUT_S3_AWS_SECRET_ACCESS_KEY": "",
},
DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL="https://storage.example.com",
)
def test_blank_static_credentials_defer_to_the_provider_chain(self, monkeypatch):
"""Empty keys would otherwise be signed as-is, yielding a blank credential scope."""
monkeypatch.setenv("AWS_ACCESS_KEY_ID", "chain-key")
monkeypatch.setenv("AWS_SECRET_ACCESS_KEY", "chain-secret")
monkeypatch.delenv("AWS_SESSION_TOKEN", raising=False)
query = parse_qs(urlparse(_presign(get_s3_presign_client())).query)
assert query["X-Amz-Credential"][0].startswith("chain-key/")
+176 -502
View File
@@ -1,531 +1,205 @@
from datetime import UTC, datetime
from datetime import UTC, datetime, timedelta
from uuid import uuid4
import pytest
from api.models import Finding, MuteRule
from django.core.exceptions import ObjectDoesNotExist
from api.models import Finding, MuteRule, Scan, StateChoices
from prowler.lib.check.models import Severity
from prowler.lib.outputs.finding import Status
from tasks.jobs.muting import mute_historical_findings
from tasks.jobs.muting import (
mute_findings_in_latest_scans,
reconcile_scan_mute_rules,
)
def _create_finding(scan: Scan, uid: str) -> Finding:
return Finding.objects.create(
tenant_id=scan.tenant_id,
uid=uid,
scan=scan,
status=Status.FAIL,
status_extended="Test finding",
impact=Severity.high,
severity=Severity.high,
raw_result={},
check_id="test_check",
check_metadata={"CheckId": "test_check"},
muted=False,
)
def _create_mute_rule(tenant_id, user, finding_uids, *, enabled=True) -> MuteRule:
return MuteRule.objects.create(
tenant_id=tenant_id,
name=f"Mute rule {uuid4()}",
reason="Approved exception",
enabled=enabled,
created_by=user,
finding_uids=finding_uids,
)
@pytest.mark.django_db
class TestMuteHistoricalFindings:
"""
Test suite for the mute_historical_findings function.
class TestMuteFindingsInLatestScans:
def test_mutes_latest_scan_and_leaves_older_scan_unchanged(
self, scans_fixture, create_test_user
):
latest_scan = scans_fixture[0]
older_scan = Scan.objects.create(
tenant_id=latest_scan.tenant_id,
provider=latest_scan.provider,
name="Older scan",
trigger=Scan.TriggerChoices.MANUAL,
state=StateChoices.COMPLETED,
started_at=datetime.now(UTC) - timedelta(days=1),
completed_at=datetime.now(UTC) - timedelta(days=1),
)
uid = "latest-scan-only"
older_finding = _create_finding(older_scan, uid)
latest_finding = _create_finding(latest_scan, uid)
mute_rule = _create_mute_rule(latest_scan.tenant_id, create_test_user, [uid])
This class tests the batch processing of findings to update their muted status
based on MuteRule criteria.
"""
result = mute_findings_in_latest_scans(
str(latest_scan.tenant_id),
str(mute_rule.id),
[str(latest_scan.provider_id)],
)
@pytest.fixture(scope="function")
def test_user(self, create_test_user):
"""Create a test user for mute rule creation."""
return create_test_user
older_finding.refresh_from_db()
latest_finding.refresh_from_db()
assert older_finding.muted is False
assert latest_finding.muted is True
assert latest_finding.muted_at == mute_rule.inserted_at
assert latest_finding.muted_reason == mute_rule.reason
assert result == {
"findings_muted": 1,
"rule_id": str(mute_rule.id),
"scan_ids": [str(latest_scan.id)],
}
@pytest.fixture(scope="function")
def mute_rule_with_findings(self, tenants_fixture, findings_fixture, test_user):
"""
Create a mute rule that targets the first finding in the fixture.
"""
def test_mutes_one_latest_scan_per_provider(self, scans_fixture, create_test_user):
first_scan, second_scan, _ = scans_fixture
uid = "shared-selected-uid"
first_finding = _create_finding(first_scan, uid)
second_finding = _create_finding(second_scan, uid)
mute_rule = _create_mute_rule(first_scan.tenant_id, create_test_user, [uid])
result = mute_findings_in_latest_scans(
str(first_scan.tenant_id),
str(mute_rule.id),
[str(first_scan.provider_id), str(second_scan.provider_id)],
)
first_finding.refresh_from_db()
second_finding.refresh_from_db()
assert first_finding.muted is True
assert second_finding.muted is True
assert result["findings_muted"] == 2
assert set(result["scan_ids"]) == {str(first_scan.id), str(second_scan.id)}
def test_provider_without_completed_scan_does_nothing(
self, tenants_fixture, provider_factory, create_test_user
):
tenant = tenants_fixture[0]
finding = findings_fixture[0]
mute_rule = MuteRule.objects.create(
tenant_id=tenant.id,
name="Test Mute Rule",
reason="Testing mute functionality",
enabled=True,
created_by=test_user,
finding_uids=[finding.uid],
provider = provider_factory()
mute_rule = _create_mute_rule(
tenant.id, create_test_user, ["future-scan-finding"]
)
return mute_rule
result = mute_findings_in_latest_scans(
str(tenant.id), str(mute_rule.id), [str(provider.id)]
)
@pytest.fixture(scope="function")
def mute_rule_multiple_findings(self, scans_fixture, test_user):
"""
Create multiple unmuted findings and a mute rule targeting all of them.
"""
assert result == {
"findings_muted": 0,
"rule_id": str(mute_rule.id),
"scan_ids": [],
}
def test_retry_does_not_report_changed_scans_twice(
self, scans_fixture, create_test_user
):
scan = scans_fixture[0]
tenant_id = scan.tenant_id
# Create 5 unmuted findings
finding_uids = []
for i in range(5):
finding = Finding.objects.create(
tenant_id=tenant_id,
uid=f"test_finding_uid_mute_{i}",
scan=scan,
status=Status.FAIL,
status_extended=f"Test status {i}",
impact=Severity.high,
severity=Severity.high,
raw_result={
"status": Status.FAIL,
"impact": Severity.high,
"severity": Severity.high,
},
check_id=f"test_check_id_{i}",
check_metadata={
"CheckId": f"test_check_id_{i}",
"Description": f"Test description {i}",
},
muted=False,
)
finding_uids.append(finding.uid)
# Create mute rule targeting all findings
mute_rule = MuteRule.objects.create(
tenant_id=tenant_id,
name="Test Multiple Findings Mute Rule",
reason="Testing batch muting",
enabled=True,
created_by=test_user,
finding_uids=finding_uids,
finding = _create_finding(scan, "idempotent-mute")
mute_rule = _create_mute_rule(scan.tenant_id, create_test_user, [finding.uid])
args = (
str(scan.tenant_id),
str(mute_rule.id),
[str(scan.provider_id)],
)
return mute_rule, finding_uids
first_result = mute_findings_in_latest_scans(*args)
second_result = mute_findings_in_latest_scans(*args)
@pytest.fixture(scope="function")
def mute_rule_already_muted(self, findings_fixture, test_user):
"""
Create a mute rule that targets an already-muted finding.
"""
tenant_id = findings_fixture[1].tenant_id
already_muted_finding = findings_fixture[1]
assert first_result["scan_ids"] == [str(scan.id)]
assert second_result["findings_muted"] == 0
assert second_result["scan_ids"] == []
mute_rule = MuteRule.objects.create(
tenant_id=tenant_id,
name="Test Already Muted Rule",
reason="Testing already muted findings",
enabled=True,
created_by=test_user,
finding_uids=[already_muted_finding.uid],
def test_does_not_cross_tenant_boundary(
self, tenants_fixture, provider_factory, create_test_user
):
tenant = tenants_fixture[0]
other_tenant = tenants_fixture[2]
other_provider = provider_factory(tenant=other_tenant)
other_scan = Scan.objects.create(
tenant_id=other_tenant.id,
provider=other_provider,
name="Other tenant scan",
trigger=Scan.TriggerChoices.MANUAL,
state=StateChoices.COMPLETED,
started_at=datetime.now(UTC),
completed_at=datetime.now(UTC),
)
other_finding = _create_finding(other_scan, "tenant-isolated-uid")
mute_rule = _create_mute_rule(tenant.id, create_test_user, [other_finding.uid])
result = mute_findings_in_latest_scans(
str(tenant.id), str(mute_rule.id), [str(other_provider.id)]
)
return mute_rule
other_finding.refresh_from_db()
assert other_finding.muted is False
assert result["scan_ids"] == []
@pytest.fixture(scope="function")
def mute_rule_mixed_findings(self, scans_fixture, test_user):
"""
Create a mute rule with a mix of muted and unmuted findings.
"""
def test_nonexistent_rule_raises(self, tenants_fixture):
with pytest.raises(MuteRule.DoesNotExist):
mute_findings_in_latest_scans(str(tenants_fixture[0].id), str(uuid4()), [])
@pytest.mark.django_db
class TestReconcileScanMuteRules:
def test_applies_only_enabled_rules_to_requested_scan(
self, scans_fixture, create_test_user
):
scan = scans_fixture[0]
tenant_id = scan.tenant_id
# Create 3 unmuted findings
unmuted_uids = []
for i in range(3):
finding = Finding.objects.create(
tenant_id=tenant_id,
uid=f"unmuted_finding_{i}",
scan=scan,
status=Status.FAIL,
status_extended=f"Unmuted status {i}",
impact=Severity.medium,
severity=Severity.medium,
raw_result={
"status": Status.FAIL,
"impact": Severity.medium,
"severity": Severity.medium,
},
check_id=f"unmuted_check_{i}",
check_metadata={
"CheckId": f"unmuted_check_{i}",
"Description": f"Unmuted description {i}",
},
muted=False,
)
unmuted_uids.append(finding.uid)
# Create 2 already muted findings
muted_uids = []
for i in range(2):
finding = Finding.objects.create(
tenant_id=tenant_id,
uid=f"muted_finding_{i}",
scan=scan,
status=Status.FAIL,
status_extended=f"Muted status {i}",
impact=Severity.low,
severity=Severity.low,
raw_result={
"status": Status.FAIL,
"impact": Severity.low,
"severity": Severity.low,
},
check_id=f"muted_check_{i}",
check_metadata={
"CheckId": f"muted_check_{i}",
"Description": f"Muted description {i}",
},
muted=True,
muted_at=datetime.now(UTC),
muted_reason="Already muted",
)
muted_uids.append(finding.uid)
# Create mute rule targeting all findings
all_uids = unmuted_uids + muted_uids
mute_rule = MuteRule.objects.create(
tenant_id=tenant_id,
name="Test Mixed Findings Rule",
reason="Testing mixed muted/unmuted findings",
enabled=True,
created_by=test_user,
finding_uids=all_uids,
active_finding = _create_finding(scan, "active-rule-uid")
disabled_finding = _create_finding(scan, "disabled-rule-uid")
active_rule = _create_mute_rule(
scan.tenant_id, create_test_user, [active_finding.uid]
)
return mute_rule, unmuted_uids, muted_uids
@pytest.fixture(scope="function")
def mute_rule_batch_test(self, scans_fixture, test_user):
"""
Create enough findings to test batch processing (>1000 for default batch size).
"""
scan = scans_fixture[0]
tenant_id = scan.tenant_id
# Create 1500 findings to exceed default batch size of 1000
finding_uids = []
for i in range(1500):
finding = Finding.objects.create(
tenant_id=tenant_id,
uid=f"batch_test_finding_{i}",
scan=scan,
status=Status.FAIL,
status_extended=f"Batch test status {i}",
impact=Severity.critical,
severity=Severity.critical,
raw_result={
"status": Status.FAIL,
"impact": Severity.critical,
"severity": Severity.critical,
},
check_id=f"batch_test_check_{i}",
check_metadata={
"CheckId": f"batch_test_check_{i}",
"Description": f"Batch test description {i}",
},
muted=False,
)
finding_uids.append(finding.uid)
# Create mute rule targeting all findings
mute_rule = MuteRule.objects.create(
tenant_id=tenant_id,
name="Test Batch Processing Rule",
reason="Testing batch processing functionality",
enabled=True,
created_by=test_user,
finding_uids=finding_uids,
_create_mute_rule(
scan.tenant_id,
create_test_user,
[disabled_finding.uid],
enabled=False,
)
return mute_rule, finding_uids
def test_mute_historical_findings_single_finding(
self, mute_rule_with_findings, findings_fixture
):
"""
Test muting a single historical finding.
"""
mute_rule = mute_rule_with_findings
tenant_id = str(mute_rule.tenant_id)
finding = findings_fixture[0]
# Ensure the finding is not muted before execution
finding.refresh_from_db()
assert finding.muted is False
assert finding.muted_at is None
assert finding.muted_reason is None
# Execute the muting function
result = mute_historical_findings(tenant_id, str(mute_rule.id))
# Verify return value
assert result["findings_muted"] == 1
assert result["rule_id"] == str(mute_rule.id)
# Verify the finding was muted
finding.refresh_from_db()
assert finding.muted is True
assert finding.muted_at == mute_rule.inserted_at
assert finding.muted_reason == mute_rule.reason
def test_mute_historical_findings_multiple_findings(
self, mute_rule_multiple_findings
):
"""
Test muting multiple historical findings.
"""
mute_rule, finding_uids = mute_rule_multiple_findings
tenant_id = str(mute_rule.tenant_id)
# Verify all findings are unmuted
findings = Finding.objects.filter(tenant_id=tenant_id, uid__in=finding_uids)
assert findings.count() == 5
for finding in findings:
assert finding.muted is False
# Execute the muting function
result = mute_historical_findings(tenant_id, str(mute_rule.id))
# Verify return value
assert result["findings_muted"] == 5
assert result["rule_id"] == str(mute_rule.id)
# Verify all findings were muted
findings = Finding.objects.filter(tenant_id=tenant_id, uid__in=finding_uids)
for finding in findings:
assert finding.muted is True
assert finding.muted_at == mute_rule.inserted_at
assert finding.muted_reason == mute_rule.reason
def test_mute_historical_findings_already_muted(
self, mute_rule_already_muted, findings_fixture
):
"""
Test that already-muted findings are not counted or updated.
"""
mute_rule = mute_rule_already_muted
tenant_id = str(mute_rule.tenant_id)
finding = findings_fixture[1]
# Verify the finding is already muted
finding.refresh_from_db()
assert finding.muted is True
original_muted_at = finding.muted_at
original_muted_reason = finding.muted_reason
# Execute the muting function
result = mute_historical_findings(tenant_id, str(mute_rule.id))
# Verify no findings were muted
assert result["findings_muted"] == 0
assert result["rule_id"] == str(mute_rule.id)
# Verify the finding's mute status did not change
finding.refresh_from_db()
assert finding.muted is True
assert finding.muted_at == original_muted_at
assert finding.muted_reason == original_muted_reason
def test_mute_historical_findings_mixed_status(self, mute_rule_mixed_findings):
"""
Test muting when some findings are already muted and others are not.
"""
mute_rule, unmuted_uids, muted_uids = mute_rule_mixed_findings
tenant_id = str(mute_rule.tenant_id)
# Execute the muting function
result = mute_historical_findings(tenant_id, str(mute_rule.id))
# Verify only unmuted findings were counted
assert result["findings_muted"] == 3
assert result["rule_id"] == str(mute_rule.id)
# Verify unmuted findings are now muted
unmuted_findings = Finding.objects.filter(
tenant_id=tenant_id, uid__in=unmuted_uids
older_scan = Scan.objects.create(
tenant_id=scan.tenant_id,
provider=scan.provider,
name="Older matching scan",
trigger=Scan.TriggerChoices.MANUAL,
state=StateChoices.COMPLETED,
started_at=datetime.now(UTC) - timedelta(days=1),
completed_at=datetime.now(UTC) - timedelta(days=1),
)
for finding in unmuted_findings:
assert finding.muted is True
assert finding.muted_at == mute_rule.inserted_at
assert finding.muted_reason == mute_rule.reason
older_finding = _create_finding(older_scan, active_finding.uid)
# Verify already-muted findings remained unchanged
already_muted_findings = Finding.objects.filter(
tenant_id=tenant_id, uid__in=muted_uids
)
for finding in already_muted_findings:
assert finding.muted is True
assert finding.muted_reason == "Already muted"
result = reconcile_scan_mute_rules(str(scan.tenant_id), str(scan.id))
def test_mute_historical_findings_nonexistent_rule(self, tenants_fixture):
"""
Test that a nonexistent mute rule raises ObjectDoesNotExist.
"""
tenant_id = str(tenants_fixture[0].id)
nonexistent_rule_id = str(uuid4())
with pytest.raises(ObjectDoesNotExist):
mute_historical_findings(tenant_id, nonexistent_rule_id)
def test_mute_historical_findings_no_matching_findings(
self, tenants_fixture, test_user
):
"""
Test muting when no findings match the rule's UIDs.
"""
tenant_id = str(tenants_fixture[0].id)
# Create a mute rule with non-existent finding UIDs
mute_rule = MuteRule.objects.create(
tenant_id=tenant_id,
name="Test No Match Rule",
reason="Testing no matching findings",
enabled=True,
created_by=test_user,
finding_uids=[
"nonexistent_uid_1",
"nonexistent_uid_2",
"nonexistent_uid_3",
],
)
# Execute the muting function
result = mute_historical_findings(tenant_id, str(mute_rule.id))
# Verify no findings were muted
assert result["findings_muted"] == 0
assert result["rule_id"] == str(mute_rule.id)
def test_mute_historical_findings_batch_processing(self, mute_rule_batch_test):
"""
Test that large numbers of findings are processed in batches correctly.
"""
mute_rule, finding_uids = mute_rule_batch_test
tenant_id = str(mute_rule.tenant_id)
# Verify all findings exist and are unmuted
findings = Finding.objects.filter(tenant_id=tenant_id, uid__in=finding_uids)
assert findings.count() == 1500
for finding in findings:
assert finding.muted is False
# Execute the muting function
result = mute_historical_findings(tenant_id, str(mute_rule.id))
# Verify return value
assert result["findings_muted"] == 1500
assert result["rule_id"] == str(mute_rule.id)
# Verify all findings were muted
findings = Finding.objects.filter(tenant_id=tenant_id, uid__in=finding_uids)
for finding in findings:
assert finding.muted is True
assert finding.muted_at == mute_rule.inserted_at
assert finding.muted_reason == mute_rule.reason
def test_mute_historical_findings_preserves_muted_at_timestamp(
self, mute_rule_with_findings, findings_fixture
):
"""
Test that muted_at is set to the rule's inserted_at, not the current time.
"""
mute_rule = mute_rule_with_findings
tenant_id = str(mute_rule.tenant_id)
finding = findings_fixture[0]
# Execute the muting function
result = mute_historical_findings(tenant_id, str(mute_rule.id))
# Verify the finding was muted
assert result["findings_muted"] == 1
# Verify muted_at matches the rule's inserted_at timestamp
finding.refresh_from_db()
assert finding.muted_at == mute_rule.inserted_at
assert finding.muted_at is not None
def test_mute_historical_findings_partial_match(self, scans_fixture, test_user):
"""
Test muting when only some of the rule's UIDs exist as findings.
"""
scan = scans_fixture[0]
tenant_id = str(scan.tenant_id)
# Create 3 findings
existing_uids = []
for i in range(3):
finding = Finding.objects.create(
tenant_id=tenant_id,
uid=f"partial_match_finding_{i}",
scan=scan,
status=Status.FAIL,
status_extended=f"Partial match status {i}",
impact=Severity.high,
severity=Severity.high,
raw_result={
"status": Status.FAIL,
"impact": Severity.high,
"severity": Severity.high,
},
check_id=f"partial_match_check_{i}",
check_metadata={
"CheckId": f"partial_match_check_{i}",
"Description": f"Partial match description {i}",
},
muted=False,
)
existing_uids.append(finding.uid)
# Create a mute rule with both existing and non-existing UIDs
all_uids = existing_uids + [
"nonexistent_uid_1",
"nonexistent_uid_2",
]
mute_rule = MuteRule.objects.create(
tenant_id=tenant_id,
name="Test Partial Match Rule",
reason="Testing partial matching",
enabled=True,
created_by=test_user,
finding_uids=all_uids,
)
# Execute the muting function
result = mute_historical_findings(tenant_id, str(mute_rule.id))
# Verify only existing findings were muted
assert result["findings_muted"] == 3
assert result["rule_id"] == str(mute_rule.id)
# Verify the existing findings were muted
findings = Finding.objects.filter(tenant_id=tenant_id, uid__in=existing_uids)
assert findings.count() == 3
for finding in findings:
assert finding.muted is True
assert finding.muted_at == mute_rule.inserted_at
assert finding.muted_reason == mute_rule.reason
def test_mute_historical_findings_empty_uids(self, tenants_fixture, test_user):
"""
Test muting when the rule has an empty finding_uids array.
"""
tenant_id = str(tenants_fixture[0].id)
# Create a mute rule with empty finding_uids
mute_rule = MuteRule.objects.create(
tenant_id=tenant_id,
name="Test Empty UIDs Rule",
reason="Testing empty UIDs",
enabled=True,
created_by=test_user,
finding_uids=[],
)
# Execute the muting function
result = mute_historical_findings(tenant_id, str(mute_rule.id))
# Verify no findings were muted
assert result["findings_muted"] == 0
assert result["rule_id"] == str(mute_rule.id)
def test_mute_historical_findings_return_format(self, mute_rule_with_findings):
"""
Test that the return value has the correct format and fields.
"""
mute_rule = mute_rule_with_findings
tenant_id = str(mute_rule.tenant_id)
result = mute_historical_findings(tenant_id, str(mute_rule.id))
# Verify return value structure
assert isinstance(result, dict)
assert "findings_muted" in result
assert "rule_id" in result
assert isinstance(result["findings_muted"], int)
assert isinstance(result["rule_id"], str)
assert result["rule_id"] == str(mute_rule.id)
active_finding.refresh_from_db()
disabled_finding.refresh_from_db()
older_finding.refresh_from_db()
assert active_finding.muted is True
assert active_finding.muted_at == active_rule.inserted_at
assert disabled_finding.muted is False
assert older_finding.muted is False
assert result == {"findings_muted": 1, "scan_id": str(scan.id)}
+335 -66
View File
@@ -1,6 +1,5 @@
import csv
import json
import re
import uuid
from collections.abc import MutableMapping
from contextlib import contextmanager
@@ -10,6 +9,7 @@ from unittest.mock import MagicMock, patch
import pytest
from api.db_router import MainRouter
from api.db_utils import rls_transaction
from api.exceptions import ProviderConnectionError, ProviderDeletedException
from api.models import (
Finding,
@@ -2795,6 +2795,167 @@ class TestCreateComplianceRequirements:
assert count_after_first > 0
assert count_after_second == count_after_first
with rls_transaction(tenant_id):
row_versions = {
row_id.version
for row_id in ComplianceRequirementOverview.objects.filter(
scan_id=scan_id
).values_list("id", flat=True)
}
assert row_versions == {7}
def test_create_compliance_requirements_threatscore_counts_from_template(
self,
tenants_fixture,
scans_fixture,
aws_provider,
findings_fixture,
):
"""ThreatScore finding counts are derived from the template mapping,
not from each finding's stored ``compliance`` payload."""
from api.models import ComplianceRequirementOverview
with patch(
"tasks.jobs.scan.PROWLER_COMPLIANCE_OVERVIEW_TEMPLATE"
) as mock_compliance_template:
tenant_id = str(tenants_fixture[0].id)
scan_id = str(scans_fixture[0].id)
mock_compliance_template.__getitem__.return_value = {
"prowler_threatscore_aws": {
"framework": "ProwlerThreatScore",
"version": "1.0",
"requirements": {
"1.1.1": {
"description": "ThreatScore requirement",
"checks": {"test_check_id": None},
},
"1.1.2": {
"description": "Unrelated requirement",
"checks": {"other_check_id": None},
},
},
},
"other_framework": {
"framework": "Other",
"version": "2.0",
"requirements": {
"a": {
"description": "Same check, other framework",
"checks": {"test_check_id": None},
},
},
},
}
create_compliance_requirements(tenant_id, scan_id)
with rls_transaction(tenant_id):
counted = sum(
len(finding.resource_regions or [])
for finding in Finding.all_objects.filter(
scan_id=scan_id,
muted=False,
status__in=["PASS", "FAIL"],
check_id="test_check_id",
)
)
rows = list(
ComplianceRequirementOverview.objects.filter(
scan_id=scan_id
).values_list("compliance_id", "requirement_id", "total_findings")
)
assert counted > 0
assert (
sum(
total
for compliance_id, requirement_id, total in rows
if (compliance_id, requirement_id)
== ("prowler_threatscore_aws", "1.1.1")
)
== counted
)
assert all(
total == 0
for compliance_id, requirement_id, total in rows
if (compliance_id, requirement_id) == ("prowler_threatscore_aws", "1.1.2")
)
assert all(
total == 0
for compliance_id, _, total in rows
if compliance_id == "other_framework"
)
def test_create_compliance_requirements_rows_across_regions_and_frameworks(
self,
tenants_fixture,
scans_fixture,
aws_provider,
):
from api.models import ComplianceRequirementOverview
tenant_id = str(tenants_fixture[0].id)
scan_id = str(scans_fixture[0].id)
check_status_by_region = {
"us-east-1": {"check_a": "FAIL", "check_b": "PASS"},
"eu-west-1": {"check_a": "PASS"},
}
template = {
"fw_one": {
"framework": "One",
"version": "1",
"requirements": {
"r1": {"description": "a", "checks": {"check_a": None}},
"r2": {
"description": "a+b",
"checks": {"check_a": None, "check_b": None},
},
},
},
"fw_two": {
"framework": "Two",
"version": "2",
"requirements": {
"m1": {"description": "manual", "checks": {}},
},
},
}
with (
patch(
"tasks.jobs.scan.PROWLER_COMPLIANCE_OVERVIEW_TEMPLATE"
) as mock_compliance_template,
patch(
"tasks.jobs.scan._aggregate_findings_by_region",
return_value=(check_status_by_region, {}),
),
):
mock_compliance_template.__getitem__.return_value = template
result = create_compliance_requirements(tenant_id, scan_id)
assert result["requirements_created"] == 6
with rls_transaction(tenant_id):
rows = set(
ComplianceRequirementOverview.objects.filter(
scan_id=scan_id
).values_list(
"compliance_id",
"requirement_id",
"region",
"requirement_status",
"passed_checks",
"failed_checks",
"total_checks",
)
)
assert rows == {
("fw_one", "r1", "us-east-1", "FAIL", 0, 1, 1),
("fw_one", "r1", "eu-west-1", "PASS", 1, 0, 1),
("fw_one", "r2", "us-east-1", "FAIL", 1, 1, 2),
("fw_one", "r2", "eu-west-1", "PASS", 1, 0, 2),
("fw_two", "m1", "us-east-1", "MANUAL", 0, 0, 0),
("fw_two", "m1", "eu-west-1", "MANUAL", 0, 0, 0),
}
def test_create_compliance_requirements_kubernetes_provider(
self,
@@ -4723,17 +4884,11 @@ class TestAggregateFindingsByRegion:
"""Test function returns correct data structure."""
tenant_id = str(uuid.uuid4())
scan_id = str(uuid.uuid4())
modeled_threatscore_compliance_id = "ProwlerThreatScore-1.0"
normalized_id = "prowlerthreatscore10"
# (check_id, status, resource_regions, compliance) tuples
finding_rows = [
(
"check1",
"FAIL",
["us-east-1"],
{modeled_threatscore_compliance_id: ["req1", "req2"]},
)
]
# (check_id, status, resource_regions) tuples
finding_rows = [("check1", "FAIL", ["us-east-1"])]
threatscore_by_check = {"check1": ["req1", "req2"]}
mock_queryset = MagicMock()
mock_queryset.values_list.return_value = mock_queryset
@@ -4747,13 +4902,16 @@ class TestAggregateFindingsByRegion:
check_status_by_region, findings_count_by_compliance = (
_aggregate_findings_by_region(
tenant_id, scan_id, modeled_threatscore_compliance_id
tenant_id,
scan_id,
normalized_id,
threatscore_by_check,
)
)
# Streaming query contract: column-scoped values_list + iterator
mock_queryset.values_list.assert_called_once_with(
"check_id", "status", "resource_regions", "compliance"
"check_id", "status", "resource_regions"
)
mock_queryset.iterator.assert_called_once()
@@ -4774,13 +4932,14 @@ class TestAggregateFindingsByRegion:
"""Test that FAIL status takes priority over other statuses."""
tenant_id = str(uuid.uuid4())
scan_id = str(uuid.uuid4())
modeled_threatscore_compliance_id = "ProwlerThreatScore-1.0"
normalized_id = "prowlerthreatscore10"
# Same check/region: PASS first, then FAIL — FAIL must win
finding_rows = [
("check1", "PASS", ["us-east-1"], {}),
("check1", "FAIL", ["us-east-1"], {}),
("check1", "PASS", ["us-east-1"]),
("check1", "FAIL", ["us-east-1"]),
]
threatscore_by_check = {}
mock_queryset = MagicMock()
mock_queryset.values_list.return_value = mock_queryset
@@ -4793,12 +4952,15 @@ class TestAggregateFindingsByRegion:
mock_findings_filter.return_value = mock_queryset
check_status_by_region, _ = _aggregate_findings_by_region(
tenant_id, scan_id, modeled_threatscore_compliance_id
tenant_id,
scan_id,
normalized_id,
threatscore_by_check,
)
# Streaming query contract: column-scoped values_list + iterator
mock_queryset.values_list.assert_called_once_with(
"check_id", "status", "resource_regions", "compliance"
"check_id", "status", "resource_regions"
)
mock_queryset.iterator.assert_called_once()
@@ -4813,8 +4975,9 @@ class TestAggregateFindingsByRegion:
"""Test that muted findings are filtered out (muted=False in query)."""
tenant_id = str(uuid.uuid4())
scan_id = str(uuid.uuid4())
modeled_threatscore_compliance_id = "ProwlerThreatScore-1.0"
normalized_id = "prowlerthreatscore10"
threatscore_by_check = {}
mock_queryset = MagicMock()
mock_queryset.values_list.return_value = mock_queryset
mock_queryset.iterator.return_value = []
@@ -4826,12 +4989,15 @@ class TestAggregateFindingsByRegion:
mock_findings_filter.return_value = mock_queryset
_aggregate_findings_by_region(
tenant_id, scan_id, modeled_threatscore_compliance_id
tenant_id,
scan_id,
normalized_id,
threatscore_by_check,
)
# Streaming query contract: column-scoped values_list + iterator
mock_queryset.values_list.assert_called_once_with(
"check_id", "status", "resource_regions", "compliance"
"check_id", "status", "resource_regions"
)
mock_queryset.iterator.assert_called_once()
@@ -4851,23 +5017,14 @@ class TestAggregateFindingsByRegion:
"""Test that ThreatScore compliance counts are processed correctly."""
tenant_id = str(uuid.uuid4())
scan_id = str(uuid.uuid4())
modeled_threatscore_compliance_id = "ProwlerThreatScore-1.0"
normalized_id = "prowlerthreatscore10"
# PASS and FAIL findings mapped to the same ThreatScore requirement
finding_rows = [
(
"check1",
"PASS",
["us-east-1"],
{modeled_threatscore_compliance_id: ["req1"]},
),
(
"check2",
"FAIL",
["us-east-1"],
{modeled_threatscore_compliance_id: ["req1"]},
),
("check1", "PASS", ["us-east-1"]),
("check2", "FAIL", ["us-east-1"]),
]
threatscore_by_check = {"check1": ["req1"], "check2": ["req1"]}
mock_queryset = MagicMock()
mock_queryset.values_list.return_value = mock_queryset
@@ -4880,19 +5037,19 @@ class TestAggregateFindingsByRegion:
mock_findings_filter.return_value = mock_queryset
_, findings_count_by_compliance = _aggregate_findings_by_region(
tenant_id, scan_id, modeled_threatscore_compliance_id
tenant_id,
scan_id,
normalized_id,
threatscore_by_check,
)
# Streaming query contract: column-scoped values_list + iterator
mock_queryset.values_list.assert_called_once_with(
"check_id", "status", "resource_regions", "compliance"
"check_id", "status", "resource_regions"
)
mock_queryset.iterator.assert_called_once()
# Verify compliance counts
normalized_id = re.sub(
r"[^a-z0-9]", "", modeled_threatscore_compliance_id.lower()
)
assert "us-east-1" in findings_count_by_compliance
assert normalized_id in findings_count_by_compliance["us-east-1"]
assert "req1" in findings_count_by_compliance["us-east-1"][normalized_id]
@@ -4909,13 +5066,14 @@ class TestAggregateFindingsByRegion:
"""Test aggregation across multiple regions."""
tenant_id = str(uuid.uuid4())
scan_id = str(uuid.uuid4())
modeled_threatscore_compliance_id = "ProwlerThreatScore-1.0"
normalized_id = "prowlerthreatscore10"
# One finding per region
finding_rows = [
("check1", "FAIL", ["us-east-1"], {}),
("check1", "PASS", ["us-west-2"], {}),
("check1", "FAIL", ["us-east-1"]),
("check1", "PASS", ["us-west-2"]),
]
threatscore_by_check = {}
mock_queryset = MagicMock()
mock_queryset.values_list.return_value = mock_queryset
@@ -4928,12 +5086,15 @@ class TestAggregateFindingsByRegion:
mock_findings_filter.return_value = mock_queryset
check_status_by_region, _ = _aggregate_findings_by_region(
tenant_id, scan_id, modeled_threatscore_compliance_id
tenant_id,
scan_id,
normalized_id,
threatscore_by_check,
)
# Streaming query contract: column-scoped values_list + iterator
mock_queryset.values_list.assert_called_once_with(
"check_id", "status", "resource_regions", "compliance"
"check_id", "status", "resource_regions"
)
mock_queryset.iterator.assert_called_once()
@@ -4951,16 +5112,10 @@ class TestAggregateFindingsByRegion:
"""A finding with multiple resource_regions is tallied in every region."""
tenant_id = str(uuid.uuid4())
scan_id = str(uuid.uuid4())
modeled_threatscore_compliance_id = "ProwlerThreatScore-1.0"
normalized_id = "prowlerthreatscore10"
finding_rows = [
(
"check1",
"FAIL",
["us-east-1", "eu-west-1"],
{modeled_threatscore_compliance_id: ["req1"]},
)
]
finding_rows = [("check1", "FAIL", ["us-east-1", "eu-west-1"])]
threatscore_by_check = {"check1": ["req1"]}
mock_queryset = MagicMock()
mock_queryset.values_list.return_value = mock_queryset
@@ -4974,19 +5129,19 @@ class TestAggregateFindingsByRegion:
check_status_by_region, findings_count_by_compliance = (
_aggregate_findings_by_region(
tenant_id, scan_id, modeled_threatscore_compliance_id
tenant_id,
scan_id,
normalized_id,
threatscore_by_check,
)
)
# Streaming query contract: column-scoped values_list + iterator
mock_queryset.values_list.assert_called_once_with(
"check_id", "status", "resource_regions", "compliance"
"check_id", "status", "resource_regions"
)
mock_queryset.iterator.assert_called_once()
normalized_id = re.sub(
r"[^a-z0-9]", "", modeled_threatscore_compliance_id.lower()
)
for region in ("us-east-1", "eu-west-1"):
assert check_status_by_region[region]["check1"] == "FAIL"
req_stats = findings_count_by_compliance[region][normalized_id]["req1"]
@@ -5000,12 +5155,13 @@ class TestAggregateFindingsByRegion:
"""A finding with no denormalized regions contributes nothing."""
tenant_id = str(uuid.uuid4())
scan_id = str(uuid.uuid4())
modeled_threatscore_compliance_id = "ProwlerThreatScore-1.0"
normalized_id = "prowlerthreatscore10"
finding_rows = [
("check1", "FAIL", [], {modeled_threatscore_compliance_id: ["req1"]}),
("check2", "PASS", None, {}),
("check1", "FAIL", []),
("check2", "PASS", None),
]
threatscore_by_check = {"check1": ["req1"]}
mock_queryset = MagicMock()
mock_queryset.values_list.return_value = mock_queryset
@@ -5019,13 +5175,16 @@ class TestAggregateFindingsByRegion:
check_status_by_region, findings_count_by_compliance = (
_aggregate_findings_by_region(
tenant_id, scan_id, modeled_threatscore_compliance_id
tenant_id,
scan_id,
normalized_id,
threatscore_by_check,
)
)
# Streaming query contract: column-scoped values_list + iterator
mock_queryset.values_list.assert_called_once_with(
"check_id", "status", "resource_regions", "compliance"
"check_id", "status", "resource_regions"
)
mock_queryset.iterator.assert_called_once()
@@ -5040,8 +5199,9 @@ class TestAggregateFindingsByRegion:
"""Test with no findings - should return empty dicts."""
tenant_id = str(uuid.uuid4())
scan_id = str(uuid.uuid4())
modeled_threatscore_compliance_id = "ProwlerThreatScore-1.0"
normalized_id = "prowlerthreatscore10"
threatscore_by_check = {}
mock_queryset = MagicMock()
mock_queryset.values_list.return_value = mock_queryset
mock_queryset.iterator.return_value = []
@@ -5054,13 +5214,16 @@ class TestAggregateFindingsByRegion:
check_status_by_region, findings_count_by_compliance = (
_aggregate_findings_by_region(
tenant_id, scan_id, modeled_threatscore_compliance_id
tenant_id,
scan_id,
normalized_id,
threatscore_by_check,
)
)
# Streaming query contract: column-scoped values_list + iterator
mock_queryset.values_list.assert_called_once_with(
"check_id", "status", "resource_regions", "compliance"
"check_id", "status", "resource_regions"
)
mock_queryset.iterator.assert_called_once()
@@ -5804,6 +5967,112 @@ class TestResetEphemeralResourceFindingsCount:
resource2.refresh_from_db()
assert resource2.failed_findings_count == 5
def test_runs_when_newer_scan_is_not_full_scope(
self, tenants_fixture, scans_fixture, aws_provider, resources_fixture
):
"""A newer scoped scan must not block the full-scope scan's cleanup.
The race guard used to pick the newest COMPLETED scan of any scope
despite being named after full-scope ones, so a single scoped scan
landing after a complete one made the complete scan skip its cleanup
and leave ephemeral resources with a stale count permanently.
"""
from datetime import timedelta
tenant, *_ = tenants_fixture
scan1, *_ = scans_fixture
resource1, resource2, _ = resources_fixture
Resource.objects.filter(id=resource2.id).update(failed_findings_count=5)
self._make_scan_summary(tenant.id, scan1.id, resource1)
newer_completed_at = scan1.completed_at + timedelta(minutes=5)
Scan.objects.create(
name="Newer scoped scan",
provider=aws_provider,
trigger=Scan.TriggerChoices.MANUAL,
state=StateChoices.COMPLETED,
tenant_id=tenant.id,
started_at=newer_completed_at,
completed_at=newer_completed_at,
scanner_args={"checks": ["check1"]},
)
result = reset_ephemeral_resource_findings_count(
tenant_id=str(tenant.id), scan_id=str(scan1.id)
)
assert result["status"] == "completed"
resource2.refresh_from_db()
assert resource2.failed_findings_count == 0
def test_runs_when_many_newer_scans_are_not_full_scope(
self, tenants_fixture, scans_fixture, aws_provider, resources_fixture
):
"""The walk must not give up before it reaches the full-scope scan.
A fixed look-back window returned None once more newer scoped scans
had landed than it inspected, and skipped the cleanup exactly like the
original bug did with one.
"""
from datetime import timedelta
tenant, *_ = tenants_fixture
scan1, *_ = scans_fixture
resource1, resource2, _ = resources_fixture
Resource.objects.filter(id=resource2.id).update(failed_findings_count=5)
self._make_scan_summary(tenant.id, scan1.id, resource1)
for minutes in range(1, 41):
newer_completed_at = scan1.completed_at + timedelta(minutes=minutes)
Scan.objects.create(
name=f"Newer scoped scan {minutes}",
provider=aws_provider,
trigger=Scan.TriggerChoices.MANUAL,
state=StateChoices.COMPLETED,
tenant_id=tenant.id,
started_at=newer_completed_at,
completed_at=newer_completed_at,
scanner_args={"checks": ["check1"]},
)
result = reset_ephemeral_resource_findings_count(
tenant_id=str(tenant.id), scan_id=str(scan1.id)
)
assert result["status"] == "completed"
resource2.refresh_from_db()
assert resource2.failed_findings_count == 0
def test_runs_when_completed_at_is_null(
self, tenants_fixture, scans_fixture, aws_provider, resources_fixture
):
"""NULL `completed_at` used to make the reset never run at all.
The old guard filtered `completed_at__isnull=False`, so a provider
whose completed scans all had a NULL `completed_at` resolved the
"latest" scan to None, which never equals `scan.id`.
"""
tenant, *_ = tenants_fixture
scan1, *_ = scans_fixture
resource1, resource2, _ = resources_fixture
Scan.all_objects.filter(id=scan1.id).update(completed_at=None)
Resource.objects.filter(id=resource2.id).update(failed_findings_count=5)
self._make_scan_summary(tenant.id, scan1.id, resource1)
result = reset_ephemeral_resource_findings_count(
tenant_id=str(tenant.id), scan_id=str(scan1.id)
)
assert result["status"] == "completed"
resource2.refresh_from_db()
assert resource2.failed_findings_count == 0
def test_does_not_touch_other_providers_resources(
self, tenants_fixture, scans_fixture, aws_provider, resources_fixture
):
+148 -124
View File
@@ -1,6 +1,7 @@
import json
import uuid
from contextlib import contextmanager
from datetime import UTC, datetime, timedelta
from datetime import UTC, datetime
from unittest.mock import MagicMock, patch
import httpx
@@ -14,6 +15,7 @@ from api.models import (
StateChoices,
Task,
)
from api.v1.serializers import TaskSerializer
from botocore.exceptions import ClientError
from celery import states
from django_celery_beat.models import IntervalSchedule, PeriodicTask
@@ -32,11 +34,12 @@ from tasks.tasks import (
_scan_tmp_output_directory,
check_integrations_task,
check_lighthouse_provider_connection_task,
create_scan_task_record,
generate_outputs_task,
mute_findings_in_latest_scans_task,
perform_attack_paths_scan_task,
perform_scan_task,
perform_scheduled_scan_task,
reaggregate_all_finding_group_summaries_task,
refresh_lighthouse_provider_models_task,
s3_integration_task,
security_hub_integration_task,
@@ -2959,6 +2962,7 @@ class TestPerformScheduledScanTask:
with (
patch("tasks.tasks.perform_prowler_scan", side_effect=_complete_scan),
patch("tasks.tasks._perform_scan_complete_tasks"),
patch("tasks.tasks.reconcile_scan_mute_rules") as mock_reconcile,
self._override_task_request(perform_scheduled_scan_task, id=task_id),
):
perform_scheduled_scan_task.run(
@@ -2982,6 +2986,13 @@ class TestPerformScheduledScanTask:
).count()
== 1
)
completed_scan = Scan.objects.get(
tenant_id=tenant.id,
provider=provider,
trigger=Scan.TriggerChoices.SCHEDULED,
state=StateChoices.COMPLETED,
)
mock_reconcile.assert_called_once_with(str(tenant.id), str(completed_scan.id))
assert (
Scan.objects.filter(
tenant_id=tenant.id,
@@ -3176,7 +3187,10 @@ class TestPerformScanTask:
task=queued_task,
)
events = []
def _complete_scan(tenant_id, scan_id, provider_id, checks_to_execute=None):
events.append("scan")
scan_instance = Scan.objects.get(id=scan_id)
scan_instance.state = StateChoices.COMPLETED
scan_instance.save()
@@ -3184,7 +3198,14 @@ class TestPerformScanTask:
with (
patch("tasks.tasks.perform_prowler_scan", side_effect=_complete_scan),
patch("tasks.tasks._perform_scan_complete_tasks"),
patch(
"tasks.tasks.reconcile_scan_mute_rules",
side_effect=lambda *_args: events.append("reconcile"),
),
patch(
"tasks.tasks._perform_scan_complete_tasks",
side_effect=lambda *_args: events.append("summaries"),
),
patch("tasks.tasks.perform_scan_task.apply_async") as mock_apply_async,
):
with django_capture_on_commit_callbacks(execute=True):
@@ -3196,6 +3217,7 @@ class TestPerformScanTask:
queued_task_result.refresh_from_db()
assert result == {"status": "ok"}
assert events == ["scan", "reconcile", "summaries"]
assert queued_task_result.status == states.PENDING
mock_apply_async.assert_called_once_with(
kwargs={
@@ -3241,10 +3263,7 @@ class TestPerformScanTask:
@pytest.mark.django_db
class TestReaggregateAllFindingGroupSummaries:
def setup_method(self):
self.tenant_id = str(uuid.uuid4())
class TestMuteFindingsInLatestScansTask:
@patch("tasks.tasks.chain")
@patch("tasks.tasks.group")
@patch("tasks.tasks.aggregate_attack_surface_task")
@@ -3253,10 +3272,10 @@ class TestReaggregateAllFindingGroupSummaries:
@patch("tasks.tasks.aggregate_finding_group_summaries_task")
@patch("tasks.tasks.aggregate_daily_severity_task")
@patch("tasks.tasks.perform_scan_summary_task")
@patch("tasks.tasks.Scan.objects.filter")
def test_dispatches_subtasks_for_each_provider_per_day(
@patch("tasks.tasks.mute_findings_in_latest_scans")
def test_reaggregates_only_changed_scans(
self,
mock_scan_filter,
mock_mute_findings,
mock_scan_summary_task,
mock_daily_severity_task,
mock_finding_group_task,
@@ -3265,119 +3284,36 @@ class TestReaggregateAllFindingGroupSummaries:
mock_attack_surface_task,
mock_group,
mock_chain,
tenants_fixture,
):
provider_id_1 = uuid.uuid4()
provider_id_2 = uuid.uuid4()
scan_id_today_p1 = uuid.uuid4()
scan_id_yesterday_p1 = uuid.uuid4()
scan_id_today_p2 = uuid.uuid4()
today = datetime.now(tz=UTC)
yesterday = today - timedelta(days=1)
mock_outer_group_result = MagicMock()
# The first `group()` call wraps the inner parallel step; subsequent
# calls wrap the outer per-scan generator.
mock_group.side_effect = lambda *args, **kwargs: (
list(args[0]) if args and hasattr(args[0], "__iter__") else None,
mock_outer_group_result,
)[1]
mock_scan_filter.return_value.order_by.return_value.values.return_value = [
{
"id": scan_id_today_p1,
"completed_at": today,
"provider_id": provider_id_1,
},
{
"id": scan_id_today_p2,
"completed_at": today,
"provider_id": provider_id_2,
},
{
"id": scan_id_yesterday_p1,
"completed_at": yesterday,
"provider_id": provider_id_1,
},
]
result = reaggregate_all_finding_group_summaries_task(tenant_id=self.tenant_id)
assert result == {"scans_reaggregated": 3}
expected_scan_ids = {
str(scan_id_today_p1),
str(scan_id_today_p2),
str(scan_id_yesterday_p1),
tenant_id = str(tenants_fixture[0].id)
mute_rule_id = str(uuid.uuid4())
provider_ids = [str(uuid.uuid4()), str(uuid.uuid4())]
scan_ids = [str(uuid.uuid4()), str(uuid.uuid4())]
result = {
"findings_muted": 2,
"rule_id": mute_rule_id,
"scan_ids": scan_ids,
}
for task_mock in (
mock_scan_summary_task,
mock_daily_severity_task,
mock_finding_group_task,
mock_resource_group_task,
mock_category_task,
mock_attack_surface_task,
):
assert task_mock.si.call_count == 3
dispatched = {
call.kwargs["scan_id"] for call in task_mock.si.call_args_list
}
assert dispatched == expected_scan_ids
for call in task_mock.si.call_args_list:
assert call.kwargs["tenant_id"] == self.tenant_id
assert mock_chain.call_count == 3
mock_outer_group_result.apply_async.assert_called_once()
@patch("tasks.tasks.chain")
@patch("tasks.tasks.group")
@patch("tasks.tasks.aggregate_attack_surface_task")
@patch("tasks.tasks.aggregate_scan_category_summaries_task")
@patch("tasks.tasks.aggregate_scan_resource_group_summaries_task")
@patch("tasks.tasks.aggregate_finding_group_summaries_task")
@patch("tasks.tasks.aggregate_daily_severity_task")
@patch("tasks.tasks.perform_scan_summary_task")
@patch("tasks.tasks.Scan.objects.filter")
def test_dedupes_scans_to_latest_per_provider_per_day(
self,
mock_scan_filter,
mock_scan_summary_task,
mock_daily_severity_task,
mock_finding_group_task,
mock_resource_group_task,
mock_category_task,
mock_attack_surface_task,
mock_group,
mock_chain,
):
"""When several scans run on the same day for the same provider, only
the latest one is dispatched (matching the daily summary unique key)."""
provider_id = uuid.uuid4()
latest_scan_today = uuid.uuid4()
earlier_scan_today = uuid.uuid4()
today_late = datetime.now(tz=UTC)
today_early = today_late - timedelta(hours=4)
mock_mute_findings.return_value = result
mock_outer_group_result = MagicMock()
mock_group.side_effect = lambda *args, **kwargs: (
list(args[0]) if args and hasattr(args[0], "__iter__") else None,
mock_outer_group_result,
)[1]
# Returned ordered by `-completed_at`, so the most recent comes first.
mock_scan_filter.return_value.order_by.return_value.values.return_value = [
{
"id": latest_scan_today,
"completed_at": today_late,
"provider_id": provider_id,
},
{
"id": earlier_scan_today,
"completed_at": today_early,
"provider_id": provider_id,
},
]
task_result = mute_findings_in_latest_scans_task(
tenant_id=tenant_id,
mute_rule_id=mute_rule_id,
provider_ids=provider_ids,
)
result = reaggregate_all_finding_group_summaries_task(tenant_id=self.tenant_id)
assert result == {"scans_reaggregated": 1}
assert task_result == result
mock_mute_findings.assert_called_once_with(
tenant_id=tenant_id,
mute_rule_id=mute_rule_id,
provider_ids=provider_ids,
)
for task_mock in (
mock_scan_summary_task,
mock_daily_severity_task,
@@ -3386,23 +3322,35 @@ class TestReaggregateAllFindingGroupSummaries:
mock_category_task,
mock_attack_surface_task,
):
task_mock.si.assert_called_once_with(
tenant_id=self.tenant_id, scan_id=str(latest_scan_today)
)
mock_chain.assert_called_once()
assert task_mock.si.call_count == 2
assert {
call.kwargs["scan_id"] for call in task_mock.si.call_args_list
} == set(scan_ids)
assert mock_chain.call_count == 2
mock_outer_group_result.apply_async.assert_called_once()
@patch("tasks.tasks.chain")
@patch("tasks.tasks.group")
@patch("tasks.tasks.Scan.objects.filter")
def test_no_completed_scans_skips_dispatch(
self, mock_scan_filter, mock_group, mock_chain
@patch("tasks.tasks.mute_findings_in_latest_scans")
def test_skips_reaggregation_when_no_scan_changed(
self, mock_mute_findings, mock_group, mock_chain, tenants_fixture
):
mock_scan_filter.return_value.order_by.return_value.values.return_value = []
tenant_id = str(tenants_fixture[0].id)
mute_rule_id = str(uuid.uuid4())
result = {
"findings_muted": 0,
"rule_id": mute_rule_id,
"scan_ids": [],
}
mock_mute_findings.return_value = result
result = reaggregate_all_finding_group_summaries_task(tenant_id=self.tenant_id)
task_result = mute_findings_in_latest_scans_task(
tenant_id=tenant_id,
mute_rule_id=mute_rule_id,
provider_ids=[],
)
assert result == {"scans_reaggregated": 0}
assert task_result == result
mock_group.assert_not_called()
mock_chain.assert_not_called()
@@ -3439,3 +3387,79 @@ class TestTaskTimeLimits:
"lighthouse-provider-connection-check",
):
assert celery_app.tasks[name].time_limit < default
@pytest.mark.django_db
class TestCreateScanTaskRecord:
"""`task_kwargs` is what a response built before the publish can report."""
def _scan(self, tenant, provider):
"""A manual scan, like the one `POST /api/v1/scans` creates."""
return Scan.objects.create(
tenant_id=tenant.id,
provider=provider,
name="Manual scan",
trigger=Scan.TriggerChoices.MANUAL,
state=StateChoices.AVAILABLE,
)
def _publish_kwargs(self, tenant, scan):
"""What `enqueue_scan_execution_on_commit` publishes for this scan."""
return {
"tenant_id": str(tenant.id),
"scan_id": str(scan.id),
"provider_id": str(scan.provider_id),
}
def _task_args(self, task):
"""Read the record back the way `TaskSerializer` does."""
return TaskSerializer(task).data["task_args"]
def test_the_stored_kwargs_are_the_ones_the_publish_would_send(
self, tenants_fixture, aws_provider
):
"""The 202 reports what is stored here, so it has to be the dispatch kwargs."""
tenant = tenants_fixture[0]
scan = self._scan(tenant, aws_provider)
task = create_scan_task_record(
tenant_id=str(tenant.id),
task_id=str(uuid.uuid4()),
task_kwargs=self._publish_kwargs(tenant, scan),
)
assert self._task_args(task) == {
"scan_id": str(scan.id),
"provider_id": str(aws_provider.id),
}
def test_a_record_created_without_kwargs_reports_none(self, tenants_fixture):
"""The argument is optional, so the other callers keep their behaviour."""
task = create_scan_task_record(
tenant_id=str(tenants_fixture[0].id),
task_id=str(uuid.uuid4()),
)
assert self._task_args(task) == {}
def test_the_publish_can_overwrite_the_stored_kwargs(
self, tenants_fixture, aws_provider
):
"""django-celery-results stores a Python repr; both must decode alike."""
tenant = tenants_fixture[0]
scan = self._scan(tenant, aws_provider)
task_id = str(uuid.uuid4())
kwargs = self._publish_kwargs(tenant, scan)
task = create_scan_task_record(
tenant_id=str(tenant.id), task_id=task_id, task_kwargs=kwargs
)
before = self._task_args(task)
# What `before_task_publish` writes once the task reaches the broker.
task_result = TaskResult.objects.get(task_id=task_id)
task_result.task_kwargs = json.dumps(repr(kwargs))
task_result.save(update_fields=["task_kwargs"])
task.refresh_from_db()
assert self._task_args(task) == before
Generated
+5 -5
View File
@@ -53,7 +53,7 @@ constraints = [
{ name = "aliyun-log-fastpb", specifier = "==0.2.0" },
{ name = "amqp", specifier = "==5.3.1" },
{ name = "annotated-types", specifier = "==0.7.0" },
{ name = "anyio", specifier = "==4.12.1" },
{ name = "anyio", specifier = "==4.14.2" },
{ name = "applicationinsights", specifier = "==0.11.10" },
{ name = "apscheduler", specifier = "==3.11.2" },
{ name = "argcomplete", specifier = "==3.5.3" },
@@ -969,15 +969,15 @@ wheels = [
[[package]]
name = "anyio"
version = "4.12.1"
version = "4.14.2"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "idna" },
{ name = "typing-extensions" },
]
sdist = { url = "https://files.pythonhosted.org/packages/96/f0/5eb65b2bb0d09ac6776f2eb54adee6abe8228ea05b20a5ad0e4945de8aac/anyio-4.12.1.tar.gz", hash = "sha256:41cfcc3a4c85d3f05c932da7c26d0201ac36f72abd4435ba90d0464a3ffed703", size = 228685, upload-time = "2026-01-06T11:45:21.246Z" }
sdist = { url = "https://files.pythonhosted.org/packages/61/cc/a381afa6efea9f496eff839d4a6a1aed3bfafc7b3ab4b0d1b243a12573dd/anyio-4.14.2.tar.gz", hash = "sha256:cfa139f3ed1a23ee8f88a145ddb5ac7605b8bbfd8592baacd7ce3d8bb4313c7f", size = 260176, upload-time = "2026-07-12T20:29:07.082Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/38/0e/27be9fdef66e72d64c0cdc3cc2823101b80585f8119b5c112c2e8f5f7dab/anyio-4.12.1-py3-none-any.whl", hash = "sha256:d405828884fc140aa80a3c667b8beed277f1dfedec42ba031bd6ac3db606ab6c", size = 113592, upload-time = "2026-01-06T11:45:19.497Z" },
{ url = "https://files.pythonhosted.org/packages/da/35/f2287558c17e29fafc8ef3daf819bb9834061cfa43bff8014f7df7f63bdc/anyio-4.14.2-py3-none-any.whl", hash = "sha256:9f505dda5ac9f0c8309b5e8bd445a8c2bf7246f3ce950121e45ea15bc41d1494", size = 125813, upload-time = "2026-07-12T20:29:05.763Z" },
]
[[package]]
@@ -4938,7 +4938,7 @@ dependencies = [
[[package]]
name = "prowler-api"
version = "1.42.0"
version = "1.45.0"
source = { virtual = "." }
dependencies = [
{ name = "cartography" },
+8
View File
@@ -212,6 +212,14 @@ mainConfig:
# MEDIUM
ecr_repository_vulnerability_minimum_severity: "MEDIUM"
# AWS Inspector2
# aws.inspector2_coverage_recently_scanned
# Maximum days since Inspector2 last scanned an actively covered resource
inspector2_max_days_since_last_scan: 3
# aws.inspector2_active_findings_within_max_age
# Maximum days an Inspector2 finding can stay active since it was first observed
inspector2_active_finding_max_age_days: 192
# AWS Trusted Advisor
# aws.trustedadvisor_premium_support_plan_subscribed
verify_premium_support_plans: True
@@ -1,46 +0,0 @@
import warnings
from dashboard.common_methods import get_section_containers_cis
warnings.filterwarnings("ignore")
def get_table(data):
aux = data[
[
"REQUIREMENTS_ID",
"REQUIREMENTS_DESCRIPTION",
"REQUIREMENTS_ATTRIBUTES_SECTION",
"CHECKID",
"STATUS",
"REGION",
"ACCOUNTID",
"RESOURCEID",
]
].copy()
# Shorten the long FedRAMP KSI descriptions for better display
ksi_short_names = {
"A secure cloud service offering will protect user data, control access, and apply zero trust principles": "Identity and Access Management",
"A secure cloud service offering will use cloud native architecture and design principles to enforce and enhance the Confidentiality, Integrity and Availability of the system": "Cloud Native Architecture",
"A secure cloud service provider will ensure that all system changes are properly documented and configuration baselines are updated accordingly": "Change Management",
"A secure cloud service provider will continuously educate their employees on cybersecurity measures, testing them regularly": "Cybersecurity Education",
"A secure cloud service offering will document, report, and analyze security incidents to ensure regulatory compliance and continuous security improvement": "Incident Reporting",
"A secure cloud service offering will monitor, log, and audit all important events, activity, and changes": "Monitoring, Logging, and Auditing",
"A secure cloud service offering will have intentional, organized, universal guidance for how every information resource, including personnel, is secured": "Policy and Inventory",
"A secure cloud service offering will define, maintain, and test incident response plan(s) and recovery capabilities to ensure minimal service disruption and data loss": "Recovery Planning",
"A secure cloud service offering will follow FedRAMP encryption policies, continuously verify information resource integrity, and restrict access to third-party information resources": "Service Configuration",
"A secure cloud service offering will understand, monitor, and manage supply chain risks from third-party information resources": "Third-Party Information Resources",
}
# Replace long descriptions with short names - use contains for partial matching
if not aux.empty:
for long_desc, short_name in ksi_short_names.items():
mask = aux["REQUIREMENTS_DESCRIPTION"].str.contains(
long_desc, na=False, regex=False
)
aux.loc[mask, "REQUIREMENTS_DESCRIPTION"] = short_name
return get_section_containers_cis(
aux, "REQUIREMENTS_ID", "REQUIREMENTS_ATTRIBUTES_SECTION"
)
@@ -1,46 +0,0 @@
import warnings
from dashboard.common_methods import get_section_containers_cis
warnings.filterwarnings("ignore")
def get_table(data):
aux = data[
[
"REQUIREMENTS_ID",
"REQUIREMENTS_DESCRIPTION",
"REQUIREMENTS_ATTRIBUTES_SECTION",
"CHECKID",
"STATUS",
"REGION",
"ACCOUNTID",
"RESOURCEID",
]
].copy()
# Shorten the long FedRAMP KSI descriptions for better display
ksi_short_names = {
"A secure cloud service offering will protect user data, control access, and apply zero trust principles": "Identity and Access Management",
"A secure cloud service offering will use cloud native architecture and design principles to enforce and enhance the Confidentiality, Integrity and Availability of the system": "Cloud Native Architecture",
"A secure cloud service provider will ensure that all system changes are properly documented and configuration baselines are updated accordingly": "Change Management",
"A secure cloud service provider will continuously educate their employees on cybersecurity measures, testing them regularly": "Cybersecurity Education",
"A secure cloud service offering will document, report, and analyze security incidents to ensure regulatory compliance and continuous security improvement": "Incident Reporting",
"A secure cloud service offering will monitor, log, and audit all important events, activity, and changes": "Monitoring, Logging, and Auditing",
"A secure cloud service offering will have intentional, organized, universal guidance for how every information resource, including personnel, is secured": "Policy and Inventory",
"A secure cloud service offering will define, maintain, and test incident response plan(s) and recovery capabilities to ensure minimal service disruption and data loss": "Recovery Planning",
"A secure cloud service offering will follow FedRAMP encryption policies, continuously verify information resource integrity, and restrict access to third-party information resources": "Service Configuration",
"A secure cloud service offering will understand, monitor, and manage supply chain risks from third-party information resources": "Third-Party Information Resources",
}
# Replace long descriptions with short names - use contains for partial matching
if not aux.empty:
for long_desc, short_name in ksi_short_names.items():
mask = aux["REQUIREMENTS_DESCRIPTION"].str.contains(
long_desc, na=False, regex=False
)
aux.loc[mask, "REQUIREMENTS_DESCRIPTION"] = short_name
return get_section_containers_cis(
aux, "REQUIREMENTS_ID", "REQUIREMENTS_ATTRIBUTES_SECTION"
)
@@ -1,46 +0,0 @@
import warnings
from dashboard.common_methods import get_section_containers_cis
warnings.filterwarnings("ignore")
def get_table(data):
aux = data[
[
"REQUIREMENTS_ID",
"REQUIREMENTS_DESCRIPTION",
"REQUIREMENTS_ATTRIBUTES_SECTION",
"CHECKID",
"STATUS",
"REGION",
"ACCOUNTID",
"RESOURCEID",
]
].copy()
# Shorten the long FedRAMP KSI descriptions for better display
ksi_short_names = {
"A secure cloud service offering will protect user data, control access, and apply zero trust principles": "Identity and Access Management",
"A secure cloud service offering will use cloud native architecture and design principles to enforce and enhance the Confidentiality, Integrity and Availability of the system": "Cloud Native Architecture",
"A secure cloud service provider will ensure that all system changes are properly documented and configuration baselines are updated accordingly": "Change Management",
"A secure cloud service provider will continuously educate their employees on cybersecurity measures, testing them regularly": "Cybersecurity Education",
"A secure cloud service offering will document, report, and analyze security incidents to ensure regulatory compliance and continuous security improvement": "Incident Reporting",
"A secure cloud service offering will monitor, log, and audit all important events, activity, and changes": "Monitoring, Logging, and Auditing",
"A secure cloud service offering will have intentional, organized, universal guidance for how every information resource, including personnel, is secured": "Policy and Inventory",
"A secure cloud service offering will define, maintain, and test incident response plan(s) and recovery capabilities to ensure minimal service disruption and data loss": "Recovery Planning",
"A secure cloud service offering will follow FedRAMP encryption policies, continuously verify information resource integrity, and restrict access to third-party information resources": "Service Configuration",
"A secure cloud service offering will understand, monitor, and manage supply chain risks from third-party information resources": "Third-Party Information Resources",
}
# Replace long descriptions with short names - use contains for partial matching
if not aux.empty:
for long_desc, short_name in ksi_short_names.items():
mask = aux["REQUIREMENTS_DESCRIPTION"].str.contains(
long_desc, na=False, regex=False
)
aux.loc[mask, "REQUIREMENTS_DESCRIPTION"] = short_name
return get_section_containers_cis(
aux, "REQUIREMENTS_ID", "REQUIREMENTS_ATTRIBUTES_SECTION"
)
+8
View File
@@ -15,6 +15,7 @@ services:
api:
hostname: "prowler-api"
image: prowlercloud/prowler-api:${PROWLER_API_VERSION:-stable}
restart: unless-stopped
env_file:
- path: .env
required: false
@@ -44,6 +45,7 @@ services:
ui:
image: prowlercloud/prowler-ui:${PROWLER_UI_VERSION:-stable}
restart: unless-stopped
env_file:
- path: .env
required: false
@@ -61,6 +63,7 @@ services:
postgres:
image: postgres:16-alpine@sha256:57c72fd2a128e416c7fcc499958864df5301e940bca0a56f58fddf30ffc07777
restart: unless-stopped
hostname: "postgres-db"
volumes:
- ./_data/postgres:/var/lib/postgresql/data
@@ -81,6 +84,7 @@ services:
valkey:
image: valkey/valkey:8-alpine@sha256:a038175878d66b9d274fbf8be73c0305e93798b83917647f167e18cef3c71eec
restart: unless-stopped
hostname: "valkey"
volumes:
- ./_data/valkey:/data
@@ -97,6 +101,7 @@ services:
neo4j:
image: graphstack/dozerdb:5.26.27.0@sha256:9b54d6b3a98a76c00bd23e8e78d8c82081ff168162aebd47b25c234e092cb0a0
restart: unless-stopped
hostname: "neo4j"
volumes:
- ./_data/neo4j:/data
@@ -129,6 +134,7 @@ services:
worker:
image: prowlercloud/prowler-api:${PROWLER_API_VERSION:-stable}
restart: unless-stopped
# Give Celery soft shutdown time to drain/re-queue in-flight tasks on stop.
stop_grace_period: 120s
env_file:
@@ -149,6 +155,7 @@ services:
worker-beat:
image: prowlercloud/prowler-api:${PROWLER_API_VERSION:-stable}
restart: unless-stopped
env_file:
- path: ./.env
required: false
@@ -165,6 +172,7 @@ services:
mcp-server:
image: prowlercloud/prowler-mcp:${PROWLER_MCP_VERSION:-stable}
restart: unless-stopped
environment:
- PROWLER_MCP_TRANSPORT_MODE=http
env_file:
+212
View File
@@ -4,6 +4,218 @@ description: "New features and improvements in each Prowler release"
rss: true
---
<Update label="v5.43.0" description="September 21, 2026">
### 🏛️ Compliance — FedRAMP 20x Consolidated Rules 2026
The FedRAMP 20x Phase One pilot frameworks (`fedramp_20x_ksi_low_aws`, `fedramp_20x_ksi_low_azure` and `fedramp_20x_ksi_low_gcp`) are replaced by two universal frameworks built from the FedRAMP Consolidated Rules for 2026, each covering AWS, Azure, GCP, Kubernetes and Microsoft 365 from a single definition:
- **FedRAMP 20x KSI** (`fedramp_20x_ksi_2026`): the 46 Key Security Indicators across 10 themes. There is one indicator catalog for every class instead of a separate Low fork; each indicator carries its class applicability and NIST SP 800-53 controls.
- **FedRAMP 20x Class C FRR** (`fedramp_20x_frr_class_c_2026`): the 158 FedRAMP Rules of the Class C ruleset that bind cloud service providers. Most are program obligations (reports, notifications, certification package) and stay manual; checks are mapped only where they evidence part of the rule text. Configurable checks carry configuration requirements, so a relaxed `audit_config` cannot turn a requirement green.
Automation or stored results that reference the pilot framework IDs need to move to `fedramp_20x_ksi_2026`. In Prowler App, both frameworks show the per-provider breakdown in the cross-provider compliance view and can be downloaded as OCSF.
Read more in the [Compliance documentation](https://docs.prowler.com/user-guide/compliance/tutorials/compliance).
### 🔎 AWS — Inspector Coverage, CISA KEV and FIPS Checks
Seven new AWS checks back the vulnerability detection and cryptography rules of FedRAMP 20x Class C:
- `inspector2_coverage_scan_status_active` and `inspector2_coverage_recently_scanned` report resources Amazon Inspector is not scanning, or last scanned more than `inspector2_max_days_since_last_scan` days ago (default 3).
- `inspector2_active_findings_no_known_exploited_vulnerabilities` and `inspector2_active_findings_kev_within_due_date` report active findings whose CVE is in the CISA Known Exploited Vulnerabilities catalog, and those still open past the CISA due date. The KEV data comes from Inspector itself through `inspector2:BatchGetFindingDetails`, so no external feed is needed.
- `inspector2_active_findings_within_max_age` reports active findings first observed more than `inspector2_active_finding_max_age_days` days ago (default 192).
- `elbv2_listener_fips_tls_enabled` and `transfer_server_fips_security_policy_enabled` report HTTPS/TLS load balancer listeners and Transfer Family servers without a FIPS security policy.
`inspector2:BatchGetFindingDetails` is not part of `SecurityAudit`, so it is now included in the Prowler additions policy and the CloudFormation scan role. Without it, the KEV checks report `MANUAL` naming the missing permission instead of a false `FAIL`.
Explore all AWS checks at [Prowler Hub](https://hub.prowler.com/check?provider=aws).
### ☁️ AWS — Partition Bootstrap Failover
When `PROWLER_AWS_PARTITION` is set, the bootstrap STS calls (validating credentials, assuming a role and getting an MFA session token) now try up to two more regions of the partition if the first one cannot be reached. A GovCloud host whose configured region belongs to another partition was still sent to `us-gov-east-1`, and on a network that routes only to `us-gov-west-1` the connection check and the scan failed on perfectly valid credentials. Only connection errors and timeouts move on to the next region; credential errors are reported from the first one as before. Later STS calls reuse the region that answered, and nothing changes when `PROWLER_AWS_PARTITION` is unset.
Read more in the [AWS Regions and Partitions documentation](https://docs.prowler.com/user-guide/providers/aws/regions-and-partitions).
### 🌐 Azure — Sovereign Cloud Endpoints for Defender and Key Vault
Defender security contacts and Key Vault key rotation policies now use the endpoints of the cloud selected with `--azure-region` instead of the hardcoded `management.azure.com` and `vault.azure.net` hosts, so both work on `AzureUSGovernment` and `AzureChinaCloud`. Key Vault clients are built from the vault URI that Azure returns for each vault.
Read more in the [Azure non-default cloud documentation](https://docs.prowler.com/user-guide/providers/azure/use-non-default-cloud).
### ✉️ Invitations — Expired Invitations No Longer Block Re-Invites
A pending invitation past its expiry date is now reported as expired, and inviting the same email again marks it as expired and creates the new invitation instead of returning a generic error. The Invitations table disables Edit and Revoke on expired and revoked invitations, and `filter[state__in]` on the invitations endpoint no longer returns a server error.
In Prowler Cloud, new organizations are offered an **Invite your team** step once the first provider is connected, and Prowler Private Cloud deployments can set `UI_SELF_REGISTRATION_ENABLED=false` to make sign-up invitation-only.
Read more in the [Invitations documentation](https://docs.prowler.com/user-guide/tutorials/prowler-app-rbac#invitations).
### 📄 Reports — Downloads on Self-Hosted Storage
Report downloads no longer depend on the storage host being the same inside and outside the container network. `DJANGO_OUTPUT_S3_AWS_PUBLIC_ENDPOINT_URL` signs the download URL against a browser-reachable host, so a deployment whose object storage answers only on an internal address serves the file instead of a link the browser cannot open. Leaving `DJANGO_OUTPUT_S3_AWS_DEFAULT_REGION` unset, common on S3-compatible storage with no meaningful region, no longer makes the download fail with a server error.
### 🔍 Checks
- **Huawei Cloud:** new `smn_topic_subscriptions` check reports SMN topics without any subscription.
- **Cloudflare:** the API token links in the provider wizard request the SSL and Certificates, Bot Management and Zone WAF read permissions the checks need, so a token created from the wizard no longer produces failures on permissions it was never granted.
- **Microsoft 365:** five Defender malware, anti-phishing and inbound anti-spam checks no longer fail with `KeyError` on tenants that use the Standard or Strict preset security policies, which dropped every finding of those checks. Preset policies are covered by `defender_strict_preset_security_policy_enabled`.
- **Google Workspace:** `security_2sv_enforced` reports domain-wide 2-Step Verification failures as `FAIL` even when every failing setting is overridden for a group or organizational unit.
Explore all checks at [Prowler Hub](https://hub.prowler.com/check).
### 🔐 Security Updates
- `libsqlite3-0`, `gzip`, `perl-base`, `libssh2-1t64` and `libpcre2-8-0` upgraded in the SDK and API container images, patching high-severity Debian CVEs.
- PowerShell upgraded to 7.5.11 in the SDK and API container images, bundling .NET runtime 9.0.20 and patching CVE-2026-62901.
- `anyio` upgraded to 4.14.2 in the SDK, the API and the MCP Server, patching CVE-2026-63374.
See the [full release notes on GitHub](https://github.com/prowler-cloud/prowler/releases/tag/5.43.0) for the complete list of changes.
</Update>
<Update label="v5.42.0" description="September 11, 2026">
### ☁️ AWS — ISO Partitions
Prowler now resolves regions and services for the AWS ISO partitions (`aws-iso`, `aws-iso-b`, `aws-iso-e` and `aws-iso-f`) the same way it does for the commercial, China, European Sovereign Cloud and GovCloud partitions. The region matrix is filled from the endpoint metadata bundled with botocore, which needs no credentials or network access, so it covers partitions that are air-gapped from the internet. Scanning them no longer requires a hand-edited `aws_regions_by_service.json`: ISO regions such as `us-isob-east-1` are accepted by `--region` and `--excluded-region`.
Deployments that declare `PROWLER_AWS_PARTITION` also keep their bootstrap STS calls in the configured region when it belongs to that partition. An install in `us-gov-west-1` that reaches AWS only through its own VPC endpoints is no longer sent to `us-gov-east-1`, where the connection check and the scan used to time out.
Read more in the [AWS Regions and Partitions documentation](https://docs.prowler.com/user-guide/providers/aws/regions-and-partitions).
### ⏱️ AWS — Configurable Timeouts for Restricted Networks
Scans from networks with restricted egress (VPC endpoints for only some services, GovCloud or private deployments) could take hours: Boto3 waits 60 seconds to connect by default and retries connection errors, so every service without a reachable endpoint cost up to four 60-second attempts in every region. Prowler now lowers the default connect timeout to 10 seconds, keeps the read timeout at 60 seconds, and exposes both through `--aws-connect-timeout` and `--aws-read-timeout`, or through the `PROWLER_AWS_BOTO3_CONNECT_TIMEOUT` and `PROWLER_AWS_BOTO3_READ_TIMEOUT` environment variables for deployments without a CLI. `--aws-retries-max-attempts 0` now disables retries instead of silently falling back to three, leaving a single attempt per call.
Read more in the [Boto3 configuration documentation](https://docs.prowler.com/user-guide/providers/aws/boto3-configuration).
### 🐳 Image Provider — Reusable Vulnerability Database
The Image provider now honors `TRIVY_CACHE_DIR`. When the variable names a directory, Trivy keeps its vulnerability database there and Prowler leaves the directory in place after the scan, so the database is downloaded once instead of on every scan. Hosts without internet access can now scan images by pointing `TRIVY_CACHE_DIR` at a pre-populated database and setting `TRIVY_SKIP_DB_UPDATE=true`. Without the variable, the temporary cache is created and removed as before.
Read more in the [Image provider documentation](https://docs.prowler.com/user-guide/providers/image/getting-started-image#vulnerability-database-cache).
### 🎫 Jira Integration — Faster Connection Test
Testing a Jira integration no longer reports a false failure on accounts with many projects. The connection test fetched the issue types of every project one request at a time, which could outlast the wait in the UI even when the check was about to succeed. Issue types are now fetched concurrently, a project whose issue types the integration user cannot see is no longer logged as an error, and the Integrations page keeps following the connection test instead of giving up after about a minute.
Read more in the [Jira integration documentation](https://docs.prowler.com/user-guide/tutorials/prowler-app-jira-integration).
### 📚 Compliance — Catalog Integrity Fixes
A new integrity test runs over every compliance framework, asserting unique requirement IDs, no check listed twice within a requirement, and that every referenced check exists for its provider. The fixes it drove span 42 frameworks across AWS, Azure, GCP, GitHub, Kubernetes and Microsoft 365:
- **Duplicate requirement IDs:** identical copies are removed, and distinct requirements that shared an ID get their own, such as `1.10` in CIS AWS 5.0 and `rc_rp_1` for RC.RP-1 in NIST CSF 1.1. In Prowler ThreatScore for Azure, SQL auditing retention moves from `3.2.1` to `3.2.4`, and requirement `1.2.1` of Prowler ThreatScore for GCP now points to `iam_sa_no_user_managed_keys`.
- **Stale check references:** checks that no longer exist are replaced with their current name when there is a direct equivalent, or removed so the requirement reports as manual. Most of these were in the FedRAMP 20x KSI frameworks.
Renamed requirement IDs appear as new requirements for scans run after the upgrade.
The compliance overview task that runs after every scan is also faster: ThreatScore mappings are read once from the compliance template instead of from every finding, and rows are inserted with time-ordered `uuid7` IDs grouped by framework and requirement.
Read more in the [Compliance documentation](https://docs.prowler.com/user-guide/compliance/tutorials/compliance).
### 🔍 Checks
`rolesanywhere_profile_restricts_session_permissions`, `iam_role_service_trust_restricts_source_to_account` and `codebuild_project_uses_allowed_github_organizations` no longer crash with `TypeError` when the scanning role is denied `iam:ListRoles`, which dropped every finding of those checks for the account. Without the role inventory, an enabled IAM Roles Anywhere profile without session scoping reports `MANUAL`, and CodeBuild projects whose service role cannot be resolved are skipped.
Explore all AWS checks at [Prowler Hub](https://hub.prowler.com/check?provider=aws).
### 🔐 Security Updates
- `next` upgraded to 16.3.3 in the UI, patching unauthenticated remote code execution through AVIF image optimization ([GHSA-2xp9-vwfh-vxw4](https://github.com/advisories/GHSA-2xp9-vwfh-vxw4)) and on Windows-hosted servers ([GHSA-p293-qw3h-jr36](https://github.com/advisories/GHSA-p293-qw3h-jr36)).
- `sharp` upgraded to 0.35.4 in the UI, patching libheif image-decoding vulnerabilities ([GHSA-rgj7-g3m4-5g8c](https://github.com/advisories/GHSA-rgj7-g3m4-5g8c)).
- `nanoid`, `js-yaml` and `postcss`, plus eleven transitive UI dependencies, upgraded to patched versions, resolving 40 npm audit advisories (21 high, 15 moderate, 4 low).
- `libuuid` upgraded to 2.41.6-r1 in the MCP Server image, patching CVE-2026-53612, CVE-2026-53613, CVE-2026-53614, CVE-2026-76642, CVE-2026-78408 and CVE-2026-78410.
See the [full release notes on GitHub](https://github.com/prowler-cloud/prowler/releases/tag/5.42.0) for the complete list of changes.
</Update>
<Update label="v5.41.0" description="September 2, 2026">
### 📥 Scans — Import Findings from the Browser
<Note>
This feature is available exclusively in **Prowler Cloud** and **Prowler Private Cloud** with a [subscription](https://prowler.com/pricing).
</Note>
Findings produced outside the platform, by the Prowler CLI or a CI pipeline, can now be brought into the app without leaving the browser. The Scans page gains an "Import Findings" dialog that takes a Prowler `.ocsf.json` report by drag-and-drop or file picker, hands it to the ingestion API and tracks the job to completion, reporting how many records were processed and how many were invalid. Files that are not a `.ocsf.json` report, or are empty, are refused before any upload starts, and a rejected upload or a failed status poll can be retried in place. The dialog is available to roles holding the Manage Ingestions permission.
![Import Findings button on the Scans page](/images/prowler-app/import-findings/import-findings-button.png)
![Import findings dialog with the drag-and-drop area](/images/prowler-app/import-findings/import-findings-dialog.png)
Read more in the [Import Findings documentation](https://docs.prowler.com/user-guide/tutorials/prowler-import-findings#using-the-ui).
### 🎫 Jira Integration — Finding Reference in Every Issue
Every Jira issue created from a finding now carries a stable reference back to it. Issues are labeled `prowler`, `prowler-<provider>`, `prowler-<severity>`, `prowler-<check-id>` and `prowler-finding-<finding-uid>`, so they can be filtered, searched with JQL or matched by automation; labels are sanitized to Jira's limits so a long or unusual value never blocks issue creation. The issue also links back to the finding in Prowler, filtered by its UID so the link keeps working after later scans, and names the Prowler organization that sent it. Prowler Cloud always includes the link; Prowler Local Server enables it by setting `DJANGO_UI_BASE_URL` in the API environment.
Read more in the [Jira integration documentation](https://docs.prowler.com/user-guide/tutorials/prowler-app-jira-integration).
### 📚 Compliance — CIS Google Workspace Foundations Benchmark v1.4.0
Prowler now ships the CIS Google Workspace Foundations Benchmark v1.4.0. Alongside the new framework, the Google Workspace checks mapped to CIS were reworked to evaluate the benchmark's full audit procedure instead of a single condition, so Gmail spoofing actions, 2-Step Verification, password expiration and alert severity left on Google's defaults no longer pass. Expect new `FAIL` findings on domains that rely on those defaults. Three accuracy fixes also land:
- `security_2sv_enforced` and `security_2sv_hardware_keys_admins` report `MANUAL` instead of judging domain-wide values that a group or a sub-organizational unit overrides; a domain-wide failure is still reported as such, with the override noted.
- `rules_*_alert_configured` no longer passes a rule whose delivery to the alert center is disabled.
- `security_password_policy_strong` no longer fails a domain that never touched the password strength setting, since Google enforces strong passwords by default.
`security_login_challenges_configured` was unmapped from CIS Google Workspace requirement 4.1.4.1 (Post-SSO verification) and `security_2sv_enforced` from CISA SCuBA `GWS.COMMONCONTROLS.1.1` (phishing-resistant MFA), because neither check can prove what those requirements ask for.
Read more in the [Compliance documentation](https://docs.prowler.com/user-guide/compliance/tutorials/compliance).
### 🔍 Checks
Ten new AWS checks land in this release, eight of them contributed by @tamg-aws. Thank you!
#### Amazon Bedrock AgentCore
- `iam_policy_passrole_to_bedrock_agentcore_restricted` flags customer-managed IAM policies that allow `iam:PassRole` over every role where the passed role can reach Bedrock AgentCore, so any principal holding the policy could run agent code under any role in the account.
- `iam_policy_no_agentcore_workload_access_token_wildcard` flags customer-managed IAM policies that allow `bedrock-agentcore:GetWorkloadAccessToken`, `GetWorkloadAccessTokenForJWT` or `GetWorkloadAccessTokenForUserId` on resources reaching workload identities other than the caller's own.
- `cloudwatch_log_group_agentcore_data_protection_policy_enabled` verifies that Bedrock AgentCore log groups mask sensitive data with a CloudWatch Logs data protection policy. The log group prefixes are configurable through `agentcore_log_group_name_prefixes` in `config.yaml`.
#### Amazon GuardDuty
- `guardduty_runtime_monitoring_enabled` flags detectors without unified Runtime Monitoring, the only feature that covers Amazon EC2 instances and Amazon ECS on AWS Fargate tasks in addition to Amazon EKS.
- `guardduty_ai_protection_enabled` flags detectors without AI Protection, which analyzes CloudTrail data events from Amazon Bedrock, Amazon Bedrock AgentCore and Amazon SageMaker AI. A detector that does not report the feature is `MANUAL` rather than `FAIL`.
`guardduty_eks_runtime_monitoring_enabled` no longer reports `FAIL` for detectors that use unified Runtime Monitoring, which is mutually exclusive with `EKS_RUNTIME_MONITORING` and already covers Amazon EKS.
#### Amazon ECR and EKS
- `ecr_registry_enhanced_scanning_enabled` verifies that the ECR registry scan type is enhanced (Amazon Inspector, covering programming language packages and continuous rescanning) instead of basic, reporting `MANUAL` when the registry scanning configuration cannot be read.
- `eks_cluster_vpc_cni_network_policy_enforced` flags EKS clusters whose Amazon VPC CNI managed add-on does not enable Kubernetes network policy enforcement, reporting `MANUAL` where the EKS API cannot show the setting.
#### AWS IAM, Elastic Beanstalk and MemoryDB
- `iam_role_service_trust_restricts_source_to_account` flags IAM roles whose trust policy lets an AWS service principal assume the role without confining the request to a specific source account, including trust policies that `iam_role_cross_service_confused_deputy_prevention` does not evaluate.
- `elasticbeanstalk_environment_no_secrets_in_configuration` scans the option settings of every Elastic Beanstalk environment for hardcoded secrets. Thanks to @haneul-24!
- `memorydb_cluster_in_transit_encryption_enabled` verifies that MemoryDB clusters have in-transit encryption (TLS) enabled. Thanks to @UTKARSH698!
Explore all AWS checks at [Prowler Hub](https://hub.prowler.com/check?provider=aws).
### 🐳 Image Provider — On-Premises Registries
Scanning registries that live on private networks is now supported end to end. `PROWLER_IMAGE_PROVIDER_ALLOWED_PRIVATE_NETWORKS` takes a comma-separated list of IPs and CIDRs the provider may reach, while every other non-public address, including link-local and loopback, stays blocked by the SSRF guard. Authentication negotiation is also more resilient: the provider falls back to Basic when a registry such as Harbor rejects the negotiated bearer token, and switches to a bearer token when the server answers a Basic or anonymous request with a Bearer challenge. `--registry-insecure` now propagates to Trivy through `TRIVY_INSECURE`, so images behind self-signed certificates can be pulled and scanned, not just enumerated. The flag now disables certificate validation for the image pull too, so keep it for trusted internal registries only.
Registry scans also skip non-image OCI artifacts (Helm charts, cosign signatures, SBOM attestations), no longer abort the whole scan when Trivy fails on a single image, and enumerate repositories in parallel instead of one request at a time.
Read more in the [Image provider documentation](https://docs.prowler.com/user-guide/providers/image/getting-started-image#on-premises-registries-and-private-networks).
### 🛠️ Prowler MCP Server — Tool Failures Reported as Errors
Prowler Local Server tools now report a failure as an MCP tool execution error (`isError: true`, with the explanation in `content`) instead of a successful result carrying an `{"error": ...}` object, which clients and models read as a success. The Prowler Documentation and Prowler Hub tools follow the same rule: `prowler_docs_search` no longer reports a failed search as zero matches, `prowler_docs_get_document` no longer reports a failed fetch as a missing page, and `prowler_hub_get_check_code` and `prowler_hub_get_check_fixer` now name the provider a check ID actually belongs to instead of reporting it as nonexistent. `prowler_get_compliance_framework_state_details` also rejects a call that passes both `scan_id` and `provider_id` instead of silently ignoring the provider.
Read more in the [Prowler MCP documentation](https://docs.prowler.com/getting-started/products/prowler-mcp).
### 🙌 External Contributors
Thank you to our community contributors for this release!
- @tamg-aws: GuardDuty unified Runtime Monitoring and AI Protection checks ([#12564](https://github.com/prowler-cloud/prowler/pull/12564)), EKS VPC CNI network policy check ([#12661](https://github.com/prowler-cloud/prowler/pull/12661)), ECR enhanced scanning check ([#12660](https://github.com/prowler-cloud/prowler/pull/12660)), Bedrock AgentCore IAM and service trust checks ([#12664](https://github.com/prowler-cloud/prowler/pull/12664)), AgentCore log group data protection check ([#12662](https://github.com/prowler-cloud/prowler/pull/12662)), and fixes to ECR scan frequency ([#12560](https://github.com/prowler-cloud/prowler/pull/12560)), CloudWatch metric filters ([#12561](https://github.com/prowler-cloud/prowler/pull/12561)) and SageMaker direct internet access ([#12659](https://github.com/prowler-cloud/prowler/pull/12659))
- @haneul-24: AWS `elasticbeanstalk_environment_no_secrets_in_configuration` check ([#12378](https://github.com/prowler-cloud/prowler/pull/12378))
- @UTKARSH698: AWS `memorydb_cluster_in_transit_encryption_enabled` check ([#12246](https://github.com/prowler-cloud/prowler/pull/12246))
- @ye11oc4t: GitHub repository discovery pagination for unscoped scans ([#12460](https://github.com/prowler-cloud/prowler/pull/12460))
See the [full release notes on GitHub](https://github.com/prowler-cloud/prowler/releases/tag/5.41.0) for the complete list of changes.
</Update>
<Update label="v5.40.0" description="August 28, 2026">
### 💬 Slack Integration — Alert Channel Destinations
@@ -68,7 +68,7 @@ The generic service pattern is described in [service page](/developer-guide/serv
- Directly in the code, in location [`prowler/providers/alibabacloud/services/`](https://github.com/prowler-cloud/prowler/tree/master/prowler/providers/alibabacloud/services)
- In the [Prowler Hub](https://hub.prowler.com/) for a more human-readable view.
The best reference to understand how to implement a new service is following the [service implementation documentation](/developer-guide/services#adding-a-new-service) and taking other services already implemented as reference. In next subsection you can find a list of common patterns that are used across all Alibaba Cloud services.
The best reference to understand how to implement a new service is following the [service implementation documentation](/developer-guide/services#adding-a-new-service) and taking other services already implemented as reference. In the next subsection you can find a list of common patterns that are used across all Alibaba Cloud services.
### Alibaba Cloud Service Common Patterns
+5 -5
View File
@@ -8,7 +8,7 @@ By default, Prowler will audit just one account and organization settings per sc
## AWS Provider Classes Architecture
The AWS provider implementation follows the general [Provider structure](/developer-guide/provider). This section focuses on the AWS-specific implementation, highlighting how the generic provider concepts are realized for AWS in Prowler. For a full overview of the provider pattern, base classes, and extension guidelines, see [Provider documentation](/developer-guide/provider). In next subsection you can find a list of the main classes of the AWS provider.
The AWS provider implementation follows the general [Provider structure](/developer-guide/provider). This section focuses on the AWS-specific implementation, highlighting how the generic provider concepts are realized for AWS in Prowler. For a full overview of the provider pattern, base classes, and extension guidelines, see [Provider documentation](/developer-guide/provider). In the next subsection you can find a list of the main classes of the AWS provider.
### `AwsProvider` (Main Class)
@@ -59,7 +59,7 @@ The generic service pattern is described in [service page](/developer-guide/serv
- Directly in the code, in location [`prowler/providers/aws/services/`](https://github.com/prowler-cloud/prowler/tree/master/prowler/providers/aws/services)
- In the [Prowler Hub](https://hub.prowler.com/) for a more human-readable view.
The best reference to understand how to implement a new service is following the [service implementation documentation](/developer-guide/services#adding-a-new-service) and taking other services already implemented as reference. In next subsection you can find a list of common patterns that are used across all AWS services.
The best reference to understand how to implement a new service is following the [service implementation documentation](/developer-guide/services#adding-a-new-service) and taking other services already implemented as reference. In the next subsection you can find a list of common patterns that are used across all AWS services.
### AWS Service Common Patterns
@@ -67,8 +67,8 @@ The best reference to understand how to implement a new service is following the
- Every AWS service class inherits from `AWSService`, ensuring access to session, identity, configuration, and threading utilities.
- The constructor (`__init__`) always calls `super().__init__` with the service name and provider (e.g. `super().__init__(__class__.__name__, provider))`). Ensure that the service name in boto3 is the same that you use in the constructor. Usually is used the `__class__.__name__` to get the service name because it is the same as the class name.
- Resource containers **must** be initialized in the constructor. They should be dictionaries, with the key being the resource ARN or equivalent unique identifier and the value being the resource object.
- Resource discovery and attribute collection are parallelized using `self.__threading_call__`, typically by region or resource, for performance. The first parameter of the method is the iterator, if not provided, it will be the region; but if present indicate an array of the resources to be processed.
- Resource filtering is consistently enforced using `self.audit_resources` attribute and `is_resource_filtered` function, it is used to see if user has provided some resource that is not in the audit scope, so we can skip it in the service logic. Normally it is used before storing the resource in the service container as follows: `if not self.audit_resources or (is_resource_filtered(resource["arn"], self.audit_resources)):`.
- Resource discovery and attribute collection are parallelized using `self.__threading_call__`, typically by region or resource, for performance. The first parameter of the method is the iterator, if not provided, it will be the region; but if present it indicates an array of the resources to be processed.
- Resource filtering is consistently enforced using `self.audit_resources` attribute and `is_resource_filtered` function, it is used to see if the user has provided some resource that is not in the audit scope, so we can skip it in the service logic. Normally it is used before storing the resource in the service container as follows: `if not self.audit_resources or (is_resource_filtered(resource["arn"], self.audit_resources)):`.
- All AWS resources are represented as Pydantic `BaseModel` classes, providing type safety and structured access to resource attributes.
- AWS API calls are wrapped in try/except blocks, with specific handling for `ClientError` and generic exceptions, always logging errors.
- If ARN is not present for some resource, it can be constructed using string interpolation, always including partition, service, region, account, and resource ID.
@@ -162,7 +162,7 @@ When you instantiate `Check_Report_AWS`, you must provide the check metadata and
If the resource object does not contain the required attributes, you must set them manually in the check logic.
Other attributes are inherited from the `Check_Report` class, from those you **always** have to set the `status` and `status_extended` attributes in the check logic.
Other attributes are inherited from the `Check_Report` class, from which you **always** have to set the `status` and `status_extended` attributes in the check logic.
#### Example Usage
+3 -3
View File
@@ -8,7 +8,7 @@ By default, Prowler will audit all the subscriptions that it is able to list in
## Azure Provider Classes Architecture
The Azure provider implementation follows the general [Provider structure](/developer-guide/provider). This section focuses on the Azure-specific implementation, highlighting how the generic provider concepts are realized for Azure in Prowler. For a full overview of the provider pattern, base classes, and extension guidelines, see [Provider documentation](/developer-guide/provider). In next subsection you can find a list of the main classes of the Azure provider.
The Azure provider implementation follows the general [Provider structure](/developer-guide/provider). This section focuses on the Azure-specific implementation, highlighting how the generic provider concepts are realized for Azure in Prowler. For a full overview of the provider pattern, base classes, and extension guidelines, see [Provider documentation](/developer-guide/provider). In the next subsection you can find a list of the main classes of the Azure provider.
### `AzureProvider` (Main Class)
@@ -57,13 +57,13 @@ The generic service pattern is described in [service page](/developer-guide/serv
- Directly in the code, in location [`prowler/providers/azure/services/`](https://github.com/prowler-cloud/prowler/tree/master/prowler/providers/azure/services)
- In the [Prowler Hub](https://hub.prowler.com/) for a more human-readable view.
The best reference to understand how to implement a new service is following the [service implementation documentation](/developer-guide/services#adding-a-new-service) and taking other services already implemented as reference. In next subsection you can find a list of common patterns that are used across all Azure services.
The best reference to understand how to implement a new service is following the [service implementation documentation](/developer-guide/services#adding-a-new-service) and taking other services already implemented as reference. In the next subsection you can find a list of common patterns that are used across all Azure services.
### Azure Service Common Patterns
- Services communicate with Azure using the Azure Python SDK, mainly using the Azure Management Client (except for the Microsoft Entra ID service, that is using the Microsoft Graph API), you can find the documentation with all the management services [here](https://learn.microsoft.com/en-us/python/api/overview/azure/?view=azure-python).
- Every Azure service class inherits from `AzureService`, ensuring access to session, identity, configuration, and client utilities.
- The constructor (`__init__`) always calls `super().__init__` with the service Azure Management Client and Prowler provider object (e.g `super().__init__(WebSiteManagementClient, provider)`).
- The constructor (`__init__`) always calls `super().__init__` with the service Azure Management Client and Prowler provider object (e.g. `super().__init__(WebSiteManagementClient, provider)`).
- Resource containers **must** be initialized in the constructor, and they should be dictionaries, with the key being the subscription ID, the value being a dictionary with the resource ID as key and the resource object as value.
- All Azure resources are represented as Pydantic `BaseModel` classes, providing type safety and structured access to resource attributes. Some are represented as dataclasses due to legacy reasons, but new resources should be represented as Pydantic `BaseModel` classes.
- Azure SDK functions are wrapped in try/except blocks, with specific handling for errors, always logging errors. It is a best practice to create a custom function for every Azure SDK call, in that way we can handle the errors in a more specific way.
+34 -3
View File
@@ -103,7 +103,7 @@ class <check_name>(Check):
# Set required fields and implement check logic
report.status = "PASS"
report.status_extended = f"<Description about why the resource is compliant>"
# If some of the information needed for the report is not inside the resource, it can be set it manually here.
# If some of the information needed for the report is not inside the resource, it can be set manually here.
# This depends on the provider and the resource that is being audited.
# report.region = resource.region
# report.resource_tags = getattr(resource, "tags", [])
@@ -129,12 +129,42 @@ Each check **must** populate the `report.status` and `report.status_extended` fi
- Status field: `report.status`
- `PASS` – Assigned when the check confirms compliance with the configured value.
- `FAIL` – Assigned when the check detects non-compliance with the configured value.
- `MANUAL` – This status must not be used unless manual verification is necessary to determine whether the status (`report.status`) passes (`PASS`) or fails (`FAIL`).
- `MANUAL` – This status must not be used unless manual verification is necessary to determine whether the status (`report.status`) passes (`PASS`) or fails (`FAIL`). This includes the case where Prowler could not retrieve the data needed to evaluate the resource (see below).
- Status extended field: `report.status_extended`
- It **must** end with a period (`.`).
- It **must** include the audited service, the resource, and a concise explanation of the check result, for instance: `EC2 AMI ami-0123456789 is not public.`.
### Permission and Data-Availability Errors Are Not Findings
A `FAIL` must only be emitted when an insecure condition has actually been detected. A check **must never** report `FAIL` because the underlying API call failed: missing permissions or scopes on the scanning identity, an API that is not enabled, a feature that is not licensed, or data that could not be retrieved are scan-configuration problems, not security issues. Reporting them as `FAIL` surfaces a misleading (and often high-severity) finding to the user and skews compliance scores.
When the service layer cannot obtain the data a check depends on, the check must:
1. Emit a single `MANUAL` finding scoped to the widest affected resource (the tenant, account, project or subscription), not one finding per resource. For example, if user registration details cannot be read, emit one tenant-level `MANUAL` instead of one per user.
2. Explain in `status_extended` that the check could not be evaluated and what to fix, naming the permission, scope, API or license required, for instance: `Cannot evaluate credential exposure for privileged users: unable to query Microsoft Defender XDR Advanced Hunting. Verify that the ThreatHunting.Read.All permission is granted to the scanning application.`
3. Leave the check's severity untouched. Do not override `report.check_metadata.Severity` to hide the problem.
The service layer must make the distinction possible: log the error and expose it to checks in a way that cannot be confused with a legitimate empty result. Common patterns already used in Prowler are:
- Defaulting the attribute to `None` (data could not be read) instead of `[]`/`{}` (data was read and is empty), e.g. the `metric_filters is not None` guard in `prowler/providers/aws/services/cloudwatch/lib/metric_filters.py`.
- Keeping an availability flag raised on any denied listing, e.g. `logs_client.metric_filters_unavailable` consumed by the AWS CloudWatch metric filter checks.
- Keeping an error flag or message next to the data, e.g. `entra_client.user_registration_details_error` in M365 or `*_scan_errors` in AWS Bedrock.
- Keeping a set of resources whose lookup failed, e.g. `accessapproval_client.settings_lookup_failed` in GCP.
Make sure the error branch only captures real access errors. A `404`/not-found response frequently means the feature is simply not configured, which **is** a legitimate `FAIL`; a `403` or an unexpected exception is not. An "API not enabled" error is usually a scan-configuration problem too — **except** when the API's activation is itself the control being audited (e.g. GCP Access Approval: with `accessapproval.googleapis.com` disabled the feature provably cannot be enabled, so a definitive API-disabled state is a legitimate `FAIL`, while an undetermined state stays `MANUAL`).
```python
if <service>_client.<data> is None:
report = CheckReport<Provider>(metadata=self.metadata(), resource={})
report.resource_name = "<Tenant/Account-level resource>"
report.resource_id = "<stable-id>"
report.status = "MANUAL"
report.status_extended = "Cannot evaluate <requirement>: <data> could not be retrieved. Verify that <permission/API/license> is granted to the scanning identity."
findings.append(report)
return findings
```
### Prowler's Check Severity Levels
The severity of each check is defined in the metadata file using the `Severity` field. Severity values are always lowercase and must be one of the predefined categories below.
@@ -407,7 +437,7 @@ For the complete list of available categories, see [Categories Guidelines](/deve
#### DependsOn
List of check IDs of checks that if are compliant, this check will be a compliant too or it is not going to give any finding.
List of check IDs of checks that if they are compliant, this check will be compliant too or it is not going to give any finding.
#### RelatedTo
@@ -437,6 +467,7 @@ The metadata structure is enforced in code using a Pydantic model. For reference
- Use clear, actionable, and user-friendly language in `status_extended` to explain the result. Always provide information to identify the resource.
- Use helper functions/utilities for repeated logic to avoid code duplication. Save them in the `lib` folder of the service.
- Handle exceptions gracefully: catch errors per resource, log them, and continue processing other resources.
- Never report `FAIL` because data could not be retrieved (missing permissions, API not enabled, feature not licensed). Emit a single `MANUAL` finding explaining what is required instead; see [Permission and Data-Availability Errors Are Not Findings](#permission-and-data-availability-errors-are-not-findings).
- Document the check with a class and function level docstring explaining what it does, what it checks, and any caveats or provider-specific behaviors.
- Use type hints for the `execute()` method (e.g., `-> list[CheckReport<Provider>]`) for clarity and static analysis.
- Ensure checks are efficient; avoid excessive nested loops. If the complexity is high, consider refactoring the check.
@@ -154,6 +154,8 @@ Only fields with a numeric range, a fixed value set, or a length cap are listed.
| `max_days_secret_unused` | `7..365` days | |
| `max_days_secret_unrotated` | `1..180` days | NIST IA-5: rotate quarterly; CIS ≤90 |
| `min_kinesis_stream_retention_hours` | `24..8760` h | 1 day .. 1 year |
| `inspector2_max_days_since_last_scan` | `1..90` days | |
| `inspector2_active_finding_max_age_days` | `1..365` days | Default `192` matches the FedRAMP 20x rule that marks vulnerabilities still open after 192 days as accepted |
| `shodan_api_key` | ≤512 chars | |
### Azure
@@ -40,6 +40,18 @@ The former build-time variables map to the new runtime variables as follows:
`UI_CLOUD_ENABLED` is a plain runtime boolean flag that enables Prowler Cloud behavior when set to the exact string `"true"` and defaults to off; unlike the other renamed variables it has no legacy fallback, so `NEXT_PUBLIC_IS_CLOUD_ENV` is no longer read.
`UI_SELF_REGISTRATION_ENABLED` is a runtime opt-out flag that Prowler Local Server reads only when `UI_CLOUD_ENABLED` is `"true"`. It defaults to on and turns off when set to `"false"`, matched case-insensitively so the same value can be shared with a backend setting written `False`. When it is off, the sign-up page only opens with an invitation token, the sign-in page drops its "Sign up" link, and the profile hides "Create organization"; invited users can still complete their registration. Outside a Prowler Cloud deployment the flag is ignored and account creation stays open.
## Registry UI Rollout and Rollback
`UI_REGISTRY_ENABLED` is an optional runtime flag for Prowler Cloud and Private Cloud. Registry is eligible only when both `UI_CLOUD_ENABLED` and `UI_REGISTRY_ENABLED` are the exact string `"true"` and the current user has the backend-authorized `manage_registry` permission. Unset, `"false"`, or malformed values fail closed. The flag defaults to off and is not a replacement for backend authorization. Registry access is independent of billing; Private Cloud can use it with `CLOUD_BILLING_ENABLED=false`.
Roll out Registry only after the Registry backend dependency is deployed, intended roles have `manage_registry`, and acceptance with real credentials has exercised installation, provider account creation, credentials, connection, and scan launch. Deploy the UI with `UI_REGISTRY_ENABLED` unset or `"false"`; set it to `"true"` only in the prepared process environment, then restart or otherwise apply the environment update required by the platform. A Registry key must belong to the configured Registry environment; a production key does not authenticate against a development Registry.
The catalog displays all artifacts, including built-ins and packages containing only checks or compliance frameworks. Only external provider artifacts support Add. After confirmed installation, open Providers and select the option labeled Registry to configure an account. Creating accounts and running scans also require the corresponding provider and scan permissions. Removing an artifact keeps existing provider accounts, but future connections or scans can fail until the artifact is installed again.
To roll back, set `UI_REGISTRY_ENABLED=false` or remove it and apply the environment update. Proxy, page, and action checks deny on their next request. Navigation refreshes from server-authorized access when the page is requested again. Rollback does not delete Registry credentials, tenant artifact records, or provider accounts.
The build-time-only Sentry variables used for source-map upload — `SENTRY_ORG`, `SENTRY_PROJECT`, `SENTRY_AUTH_TOKEN`, and `SENTRY_RELEASE` — keep their names, as they are not part of Prowler Local Server's runtime configuration.
## Enabling Third-Party Integrations
+2 -2
View File
@@ -102,7 +102,7 @@ The generic service pattern is described in [service page](/developer-guide/serv
- Directly in the code, in location [`prowler/providers/gcp/services/`](https://github.com/prowler-cloud/prowler/tree/master/prowler/providers/gcp/services)
- In the [Prowler Hub](https://hub.prowler.com/) for a more human-readable view.
The best reference to understand how to implement a new service is following the [service implementation documentation](/developer-guide/services#adding-a-new-service) and taking other services already implemented as reference. In next subsection you can find a list of common patterns that are used across all GCP services.
The best reference to understand how to implement a new service is following the [service implementation documentation](/developer-guide/services#adding-a-new-service) and taking other services already implemented as reference. In the next subsection you can find a list of common patterns that are used across all GCP services.
### GCP Service Common Patterns
@@ -161,7 +161,7 @@ When you instantiate `Check_Report_GCP`, you must provide the check metadata and
- Defaults to "global" if none are available.
All these attributes can be overridden by passing the corresponding argument to the constructor. If the resource object does not contain the required attributes, you must set them manually.
Other attributes are inherited from the `Check_Report` class, from those you **always** have to set the `status` and `status_extended` attributes in the check logic.
Other attributes are inherited from the `Check_Report` class, from which you **always** have to set the `status` and `status_extended` attributes in the check logic.
#### Example Usage
+5 -5
View File
@@ -17,7 +17,7 @@ A provider is any platform or service that offers resources, data, or functional
For providers supported by Prowler, refer to [Prowler Hub](https://hub.prowler.com/).
<Warning>
There are some custom providers added by the community, like [NHN Cloud](https://www.nhncloud.com/), that are not maintained by the Prowler team, but can be used in the Prowler CLI. The main purpose of this documentation is to guide you through creating a new provider and integrating it not only in the CLI, but also in the API and UI. Non official providers can be checked directly at the [Prowler GitHub repository](https://github.com/prowler-cloud/prowler/tree/master/prowler/providers).
There are some custom providers added by the community, like [NHN Cloud](https://www.nhncloud.com/), that are not maintained by the Prowler team, but can be used in the Prowler CLI. The main purpose of this documentation is to guide you through creating a new provider and integrating it not only in the CLI, but also in the API and UI. Non-official providers can be checked directly at the [Prowler GitHub repository](https://github.com/prowler-cloud/prowler/tree/master/prowler/providers).
</Warning>
---
@@ -105,8 +105,8 @@ Once you have decided the provider you want or need to add to Prowler, the next
#### Implementation Complexity
- **SDK Providers**: Low complexity. You have mature examples like AWS, Azure, GCP, Kubernetes, etc. that you can leverage to implement your provider.
- **API Providers**: Medium complexity. You need to implement the authentication and session management, and the API calls to the provider. You now have NHN and MongoDB Atlas as example to follow.
- **Tool/Wrapper Providers**: High complexity. You need to implement the argument/output mapping to the provider and handle problems that the tool/wrapper may have. You now have IAC and the PowerShell wrapper as example to follow.
- **API Providers**: Medium complexity. You need to implement the authentication and session management, and the API calls to the provider. You now have NHN and MongoDB Atlas as examples to follow.
- **Tool/Wrapper Providers**: High complexity. You need to implement the argument/output mapping to the provider and handle problems that the tool/wrapper may have. You now have IAC and the PowerShell wrapper as examples to follow.
- **Hybrid Providers**: High complexity. You need to "customize" your provider, mixing the other types of providers to achieve the desired result. You have M365 (msgraph SDK + PowerShell wrapper) and GitHub (PyGithub SDK + graphql API requests) as examples.
### Determining Regional vs Non-Regional Architecture
@@ -1756,7 +1756,7 @@ Argument validation ensures that the API provider receives valid configuration p
**File:** `prowler/providers/<provider_name>/lib/arguments/arguments.py`
Arguments depends on the provider and not the type, so the pattern for this step is the same as the [SDK providers](#step-4-implement-arguments).
Arguments depend on the provider and not the type, so the pattern for this step is the same as the [SDK providers](#step-4-implement-arguments).
#### Step 5: Implement Mutelist
@@ -2397,7 +2397,7 @@ class ToolProvider(Provider):
# Build the tool command
tool_command = [
"your_tool_command",
# Add your tool-specific arguments here, this are just examples
# Add your tool-specific arguments here, these are just examples
"-d",
directory,
"-o",
+3 -3
View File
@@ -5,7 +5,7 @@ title: 'Prowler Services'
Here you can find how to create a new service, or to complement an existing one, for a [Prowler Provider](/developer-guide/provider).
<Note>
First ensure that the provider you want to add the service is already created. It can be checked [here](https://github.com/prowler-cloud/prowler/tree/master/prowler/providers). If the provider is not present, please refer to the [Provider](./provider.md) documentation to create it from scratch.
First ensure that the provider you want to add the service to is already created. It can be checked [here](https://github.com/prowler-cloud/prowler/tree/master/prowler/providers). If the provider is not present, please refer to the [Provider](/developer-guide/provider) documentation to create it from scratch.
</Note>
## Introduction
@@ -102,7 +102,7 @@ class <Service>(ServiceParentClass):
# A try-except block must be created in each function.
try:
# If pagination is supported by the provider, is always better to use it, call to the provider API to retrieve the desired data.
# If pagination is supported by the provider, it is always better to use it, call to the provider API to retrieve the desired data.
describe_<items>_paginator = regional_client.get_paginator("describe_<items>")
# Paginator to get every item.
@@ -133,7 +133,7 @@ class <Service>(ServiceParentClass):
# When handling exceptions, use the following approach to log errors appropriately based on the cloud provider being used:
except Exception as error:
# Depending on each provider we can must use different fields in the logger, e.g.: AWS: regional_client.region or self.region, GCP: project_id and location, Azure: subscription
# Depending on each provider we must use different fields in the logger, e.g.: AWS: regional_client.region or self.region, GCP: project_id and location, Azure: subscription
logger.error(
f"{<provider_specific_field>} -- {error.__class__.__name__}[{error.__traceback__.tb_lineno}]: {error}"
)
+3 -3
View File
@@ -284,7 +284,7 @@ class Test_iam_password_policy_uppercase:
assert len(results) == 1
assert result[0].status == "FAIL"
assert result[0].status_extended == "IAM password policy does not srequire at least one uppercase letter."
assert result[0].status_extended == "IAM password policy does not require at least one uppercase letter."
assert result[0].resource_arn == f"arn:aws:iam::{AWS_ACCOUNT_NUMBER}:root"
assert result[0].resource_id == AWS_ACCOUNT_NUMBER
assert result[0].resource_tags == []
@@ -882,7 +882,7 @@ from unittest import mock
from uuid import uuid4
# Import some constans values needed in almost every check
# Import some constant values needed in almost every check
from tests.providers.azure.azure_fixtures import (
AZURE_SUBSCRIPTION_ID,
@@ -1024,7 +1024,7 @@ from prowler.providers.azure.services.appinsights.appinsights_service import (
Component,
)
# Import some constans values needed in almost every check
# Import some constant values needed in almost every check
from tests.providers.azure.azure_fixtures import (
AZURE_SUBSCRIPTION_ID,
@@ -2,6 +2,8 @@
title: 'Basic Usage'
---
import { VersionBadge } from "/snippets/version-badge.mdx"
## Running Prowler
Running Prowler requires specifying the provider (e.g. `aws`, `gcp`, `azure`, `kubernetes`, `m365`, `github`, `iac` or `mongodbatlas`):
@@ -91,6 +93,18 @@ By default, `prowler` will scan all AWS regions.
</Note>
See more details about AWS Authentication in the [Authentication Section](/user-guide/providers/aws/authentication) section.
- **AWS Retrier and Timeout Configuration**
<VersionBadge version="5.42.0" />
Tune the Boto3 standard retrier and the endpoint timeouts when AWS throttles the scan or when some endpoints are unreachable from the network Prowler runs in:
```console
prowler aws --aws-retries-max-attempts 5 --aws-connect-timeout 5 --aws-read-timeout 30
```
See the [Boto3 configuration](/user-guide/providers/aws/boto3-configuration) page for defaults and environment variables.
## Azure
Azure requires specifying the auth method:
@@ -128,8 +128,8 @@ To update the environment file:
Edit the `.env` file and change version values:
```env
PROWLER_UI_VERSION="5.40.0"
PROWLER_API_VERSION="5.40.0"
PROWLER_UI_VERSION="5.43.0"
PROWLER_API_VERSION="5.43.0"
```
<Note>
@@ -6,7 +6,7 @@ This section contains the instructions to subscribe to **Prowler Cloud** through
## How to Subscribe
To get to the **Prowler Cloud** product listing in the AWS Marketplace, and click the `View purchase options` button:
Go to the **Prowler Cloud** product listing in the AWS Marketplace, and click the `View purchase options` button:
1. Use this link to be taken directly to the [Prowler Cloud Marketplace Listing](https://aws.amazon.com/marketplace/pp/prodview-6ochhig5kxpok):
@@ -24,7 +24,7 @@ After you have subscribed to the **Prowler Cloud** product, you will need to set
![](/images/aws-marketplace/marketplace-message.png)
2. You will be redirected to **Prowler Cloud Sign In** page. You can sign in with an existing account or sign up with a new account:
2. You will be redirected to the **Prowler Cloud Sign In** page. You can sign in with an existing account or sign up with a new account:
![](/images/aws-marketplace/marketplace-sign-up.png)
@@ -5,7 +5,7 @@ title: 'Overview'
import { SubscriptionBanner } from "/snippets/subscription-banner.mdx"
import { VersionBadge } from "/snippets/version-badge.mdx"
Prowler Cloud runs an enhanced version of Lighthouse AI in Open Source repository, the Agentic Cloud Defender that helps teams understand, prioritize, and remediate security findings across cloud environments.
Prowler Cloud runs an enhanced version of Lighthouse AI in the Open Source repository, the Agentic Cloud Defender that helps teams understand, prioritize, and remediate security findings across cloud environments.
<SubscriptionBanner />
@@ -10,7 +10,7 @@ Prowler Cloud automates scanning single or multiple accounts and has all of the
<Card title="Create your account here to see Prowler Cloud in action" href="https://cloud.prowler.com/sign-up" />
With 100% consistency across our open source policies and APIs. Prowler Cloud provides the following added benefits:
With 100% consistency across our open source policies and APIs, Prowler Cloud provides the following added benefits:
<ul>
<li> Immediate sign-up and account provisioning, including a trial period with zero billing details needed at registration. </li>
@@ -20,5 +20,5 @@ With 100% consistency across our open source policies and APIs. Prowler Cloud pr
<li> Zero touch third party notifications to Slack, Jira, and more. </li>
</ul>
The team who built [Prowler](https://github.com/prowler-cloud/prowler), has helped thousands of companies get Cloud Security under control, is now making it easier by taking
The team who built [Prowler](https://github.com/prowler-cloud/prowler), which has helped thousands of companies get Cloud Security under control, is now making it easier by taking
[Prowler](https://github.com/prowler-cloud/prowler) to the [Cloud](https://prowler.com)!
@@ -2,7 +2,7 @@
title: "Overview"
---
**Prowler Hub** is our growing public library of versioned checks, cloud service artifacts, and compliance frameworks with its mappings. It’s searchable, explainable, and built to serve the community.
**Prowler Hub** is our growing public library of versioned checks, cloud service artifacts, and compliance frameworks with their mappings. It’s searchable, explainable, and built to serve the community.
**Why this matters**: Every engineer has asked, “What does this check actually do?” Prowler Hub answers that question in one place, lets you pin to a specific version, and pulls definitions into your own tools or dashboards.
@@ -62,7 +62,7 @@ Prowler Lighthouse AI is powerful, but there are limitations:
- **Continuous improvement**: Please report any issues, as the feature may make mistakes or encounter errors, despite extensive testing.
- **Access limitations**: Lighthouse AI can only access data the logged-in user can view. If you can't see certain information, Lighthouse AI can't see it either.
- **NextJS session dependence**: If your Prowler application session expires or logs out, Lighthouse AI will error out. Refresh and log back in to continue.
- **Response quality**: The response quality depends on the selected LLM provider and model. Choose models with strong tool-calling capabilities for best results. We recommend `gpt-5` model from OpenAI.
- **Response quality**: The response quality depends on the selected LLM provider and model. Choose models with strong tool-calling capabilities for best results. We recommend the `gpt-5` model from OpenAI.
## Architecture
@@ -82,7 +82,7 @@ If you encounter issues with Prowler Lighthouse AI or have suggestions for impro
### What Data Is Shared to LLM Providers?
The following API endpoints are accessible to Prowler Lighthouse AI. Data from the following API endpoints could be shared with LLM provider depending on the scope of user's query:
The following API endpoints are accessible to Prowler Lighthouse AI. Data from the following API endpoints could be shared with the LLM provider depending on the scope of the user's query:
## FAQs
@@ -110,7 +110,7 @@ Lighthouse AI [automatically filters](https://github.com/prowler-cloud/prowler/b
**3. Is my security data shared with LLM providers?**
Minimal data is shared to generate useful responses. Agent can access security findings and remediation details when needed. Provider secrets are protected by design and cannot be read. The LLM provider credentials configured with Lighthouse AI are only accessible to the Next.js server and are never sent to the LLM providers. Resource metadata (names, tags, account/project IDs, etc.) may be shared with the configured LLM provider based on query requirements.
Minimal data is shared to generate useful responses. The agent can access security findings and remediation details when needed. Provider secrets are protected by design and cannot be read. The LLM provider credentials configured with Lighthouse AI are only accessible to the Next.js server and are never sent to the LLM providers. Resource metadata (names, tags, account/project IDs, etc.) may be shared with the configured LLM provider based on query requirements.
**4. Can the Lighthouse AI change my cloud environment?**
Binary file not shown.

Before

Width:  |  Height:  |  Size: 289 KiB

After

Width:  |  Height:  |  Size: 234 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 193 KiB

After

Width:  |  Height:  |  Size: 194 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 185 KiB

After

Width:  |  Height:  |  Size: 187 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 150 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 169 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 120 KiB

After

Width:  |  Height:  |  Size: 120 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 110 KiB

After

Width:  |  Height:  |  Size: 111 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 156 KiB

After

Width:  |  Height:  |  Size: 156 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 93 KiB

After

Width:  |  Height:  |  Size: 93 KiB

+1 -1
View File
@@ -71,7 +71,7 @@ Prowler supports a wide range of providers organized by category:
| [LLM](/user-guide/providers/llm/getting-started-llm) | Official | Models | CLI |
| [M365](/user-guide/providers/microsoft365/getting-started-m365) | Official | Tenants | UI, API, CLI |
| [MongoDB Atlas](/user-guide/providers/mongodbatlas/getting-started-mongodbatlas) | Official | Organizations | UI, API, CLI |
| [Okta](/user-guide/providers/okta/getting-started-okta) | Official | Organizations | CLI |
| [Okta](/user-guide/providers/okta/getting-started-okta) | Official | Organizations | UI, API, CLI |
| [Vercel](/user-guide/providers/vercel/getting-started-vercel) | Official | Teams / Projects | UI, API, CLI |
### Kubernetes
+1 -1
View File
@@ -14,7 +14,7 @@ In macOS Ventura, the default value for the `file descriptors` is `256`. With th
If you have a different OS and you are experiencing the same, please increase the value of your `file descriptors`. You can check it running `ulimit -a | grep "file descriptors"`.
This error is also related with a lack of system requirements. To improve performance, Prowler stores information in memory so it may need to be run in a system with more than 1GB of memory.
This error is also related to a lack of system requirements. To improve performance, Prowler stores information in memory so it may need to be run in a system with more than 1GB of memory.
See section [Logging](/user-guide/cli/tutorials/logging) for further information or [contact us](/contact).
+1 -1
View File
@@ -239,7 +239,7 @@ The **Code** tab in the Claude desktop app runs the same Claude Code as the CLI,
</Step>
<Step title="Verify in the Code tab">
Open a **Code** tab session and ask for a Prowler tool: "Do you have access to the Prowler MCP tools?", it should respond with a list of available tools or confirming that it has access.
Open a **Code** tab session and ask for a Prowler tool: "Do you have access to the Prowler MCP tools?", it should respond with a list of available tools or confirm that it has access.
</Step>
</Steps>
+1 -1
View File
@@ -31,7 +31,7 @@ For Prowler, the **global** scope is usually the right choice — your findings
<Steps>
<Step title="Open the MCP settings">
From Agent Window open **Customize** in the Cursor sidebar, then select the MCP section.
From the Agent Window open **Customize** in the Cursor sidebar, then select the MCP section.
On earlier versions, press `Cmd + Shift + J` (macOS) or `Ctrl + Shift + J` (Windows/Linux) to open Cursor Settings, then click **Tools & MCP** in the sidebar.
+1 -1
View File
@@ -16,7 +16,7 @@ Pick your agent below. Each guide is a full walkthrough with screenshots, verifi
The Chat tab, via a local bridge
</Card>
<Card title="Codex / ChatGPT Desktop App" icon="code" href="/user-guide/ai-agents/codex">
ChatGPT Desktop App, Codex CLI, the VS Code extension through same config file
ChatGPT Desktop App, Codex CLI, the VS Code extension through the same config file
</Card>
<Card title="Cursor" icon="arrow-pointer" href="/user-guide/ai-agents/cursor">
Agentic code editor. Global and project scopes
@@ -4,7 +4,7 @@ title: "Configuration File"
import { VersionBadge } from "/snippets/version-badge.mdx"
Several Prowler's checks have user configurable variables that can be modified in a common **configuration file**. This file can be found in the following [path](https://github.com/prowler-cloud/prowler/blob/master/prowler/config/config.yaml):
Several Prowler checks have user configurable variables that can be modified in a common **configuration file**. This file can be found in the following [path](https://github.com/prowler-cloud/prowler/blob/master/prowler/config/config.yaml):
```
prowler/config/config.yaml
@@ -51,6 +51,7 @@ The following list includes all the AWS checks with configurable variables that
| `cloudtrail_threat_detection_privilege_escalation` | `threat_detection_privilege_escalation_actions` | List of Strings | See `config.yaml` |
| `cloudtrail_threat_detection_privilege_escalation` | `threat_detection_privilege_escalation_minutes` | Integer | `1440` |
| `cloudtrail_threat_detection_privilege_escalation` | `threat_detection_privilege_escalation_threshold` | Float | `0.2` |
| `cloudwatch_log_group_agentcore_data_protection_policy_enabled` | `agentcore_log_group_name_prefixes` | List of Strings | See `config.yaml` |
| `cloudwatch_log_group_no_secrets_in_logs` | `secrets_ignore_patterns` | List of Strings | `[]` |
| `cloudwatch_log_group_retention_policy_specific_days_enabled` | `log_group_retention_days` | Integer | `365` |
| `codebuild_project_no_secrets_in_variables` | `excluded_sensitive_environment_variables` | List of Strings | `[]` |
@@ -90,6 +91,8 @@ The following list includes all the AWS checks with configurable variables that
| `iam_user_access_not_stale_to_sagemaker` | `max_unused_sagemaker_access_days` | Integer | `90` |
| `iam_user_accesskey_unused` | `max_unused_access_keys_days` | Integer | `45` |
| `iam_user_console_access_unused` | `max_console_access_days` | Integer | `45` |
| `inspector2_active_findings_within_max_age` | `inspector2_active_finding_max_age_days` | Integer | `192` |
| `inspector2_coverage_recently_scanned` | `inspector2_max_days_since_last_scan` | Integer | `3` |
| `kinesis_stream_data_retention_period` | `min_kinesis_stream_retention_hours` | Integer | `168` |
| `neptune_cluster_backup_enabled` | `minimum_backup_retention_period` | Integer | `7` |
| `opensearch_service_domains_not_publicly_accessible` | `trusted_ips` | List of Strings | `[]` |
@@ -350,7 +353,7 @@ This is the new Prowler configuration file format. The old one without provider
# AWS Configuration
aws:
# AWS Global Configuration
# aws.mute_non_default_regions --> Set to True to muted failed findings in non-default regions for AccessAnalyzer, GuardDuty, SecurityHub, DRS and Config
# aws.mute_non_default_regions --> Set to True to mute failed findings in non-default regions for AccessAnalyzer, GuardDuty, SecurityHub, DRS and Config
mute_non_default_regions: False
# AWS Resource Scan Limit Configuration
@@ -489,6 +492,14 @@ aws:
# MEDIUM
ecr_repository_vulnerability_minimum_severity: "MEDIUM"
# AWS Inspector2
# aws.inspector2_coverage_recently_scanned
# Maximum days since Inspector2 last scanned an actively covered resource
inspector2_max_days_since_last_scan: 3
# aws.inspector2_active_findings_within_max_age
# Maximum days an Inspector2 finding can stay active since it was first observed
inspector2_active_finding_max_age_days: 192
# AWS Trusted Advisor
# aws.trustedadvisor_premium_support_plan_subscribed
verify_premium_support_plans: True
@@ -10,7 +10,7 @@ You can use `--custom-checks-metadata-file` followed by the path to your custom
## Available Fields
The list of supported check's metadata fields that can be overridden are listed as follows:
The list of supported check's metadata fields that can be overridden is listed as follows:
- Severity
- CheckTitle
+1 -1
View File
@@ -65,7 +65,7 @@ This page shows all the info related to the compliance selected. Multiple filter
<img src="/images/cli/dashboard/dashboard-compliance.png" />
To add your own compliance to compliance page, add a file with the compliance name (using `_` instead of `.`) to the path `/dashboard/compliance`.
To add your own compliance to the compliance page, add a file with the compliance name (using `_` instead of `.`) to the path `/dashboard/compliance`.
In this file use the format present in the other compliance files to create the table. Example for CIS 2.0:
+1 -1
View File
@@ -163,7 +163,7 @@ Mutelist:
Resources:
- "test"
Tags:
- "environment=prod" # Will mute every resource except in account 123456789012 except the ones containing the string "test" and tag environment=prod
- "environment=prod" # Will mute every resource in account 123456789012 except the ones containing the string "test" and tag environment=prod
"*":
Checks:
@@ -26,4 +26,4 @@ By default, it extracts resources from all the regions, you could use `-f`/`--fi
## Objections
The inventorying process is carried out with `resourcegroupstaggingapi` calls, which means that only resources they have or have had tags will appear (except for the IAM and S3 resources which are done with Boto3 API calls).
The inventorying process is carried out with `resourcegroupstaggingapi` calls, which means that only resources that have or have had tags will appear (except for the IAM and S3 resources which are done with Boto3 API calls).
+2 -2
View File
@@ -114,7 +114,7 @@ The CSV format follows a standardized structure across all providers. The follow
#### CSV Headers Mapping
The following table shows the mapping between the CSV headers and the providers fields:
The following table shows the mapping between the CSV headers and the providers' fields:
| Open Source Consolidated| AWS| GCP| AZURE| KUBERNETES
|----------|----------|----------|----------|----------
@@ -134,7 +134,7 @@ The following table shows the mapping between the CSV headers and the providers
### JSON-OCSF
The JSON-OCSF output format implements the [Detection Finding](https://schema.ocsf.io/classes/detection_finding) from the [OCSF](https://schema.ocsf.io)
The JSON-OCSF output format implements the [Detection Finding](https://schema.ocsf.io/classes/detection_finding) from the [OCSF](https://schema.ocsf.io).
```json
[{
@@ -253,7 +253,7 @@ Standard results plus the framework breakdown are printed to the terminal. A ded
<img src="/images/cli/compliance/compliance-cis-sample1.png" />
<Note>
If Prowler cannot find a resource related with a check from a compliance requirement, that requirement is omitted from the output.
If Prowler cannot find a resource related to a check from a compliance requirement, that requirement is omitted from the output.
</Note>
### List Available Compliance Frameworks
@@ -167,7 +167,7 @@ The same `External ID` entered in the Prowler UI must match the `ExternalId` par
```
<Note>
Check the aws documentation [here](https://docs.aws.amazon.com/IAM/latest/UserGuide/sts_example_sts_GetSessionToken_section.html)
Check the AWS documentation [here](https://docs.aws.amazon.com/IAM/latest/UserGuide/sts_example_sts_GetSessionToken_section.html)
</Note>
2. Copy the output containing:
@@ -1,14 +1,55 @@
---
title: "Boto3 Retrier Configuration in Prowler"
title: "Boto3 Retrier and Timeout Configuration in Prowler"
---
import { VersionBadge } from "/snippets/version-badge.mdx"
Prowler's AWS Provider leverages Boto3's [Standard](https://boto3.amazonaws.com/v1/documentation/api/latest/guide/retries.html) retry mode to automatically retry client calls to AWS services when encountering errors or exceptions.
## Timeout Configuration
<VersionBadge version="5.42.0" />
Every AWS API call is bounded by two timeouts:
- Connect timeout: seconds to wait to establish a connection (TCP, proxy tunnel and TLS handshake) to the AWS endpoint. Prowler's default is 10 seconds, configurable via `--aws-connect-timeout 5`.
- Read timeout: seconds to wait for a response once connected. Prowler's default is 60 seconds, configurable via `--aws-read-timeout 30`.
Both timeouts can also be set through environment variables, which is the way to tune them in Prowler Cloud and other deployments without a CLI:
```console
export PROWLER_AWS_BOTO3_CONNECT_TIMEOUT=5
export PROWLER_AWS_BOTO3_READ_TIMEOUT=30
```
CLI flags take precedence over the environment variables. Prowler sets both timeouts explicitly, so `AWS_DEFAULTS_MODE` and a `connect_timeout` in `~/.aws/config` are ignored; use the flag or the environment variable instead.
<Note>
Boto3 defaults both timeouts to 60 seconds. In networks with restricted egress (for example VPC endpoints for a subset of services, GovCloud or private deployments), every AWS service without a reachable endpoint used to cost up to 4 attempts × 60 seconds (the first call plus the 3 retries) for each region. Prowler lowers the connect timeout to 10 seconds so unreachable endpoints fail fast; lower it further together with `--aws-retries-max-attempts 0`, which disables retries and leaves a single attempt per call, if a scan still spends most of its time waiting on unreachable services.
</Note>
## Retries Configuration
<VersionBadge version="5.44.0" />
The number of retries is set with `--aws-retries-max-attempts`, where `0` disables retries. It can also be set through an environment variable, which is the way to tune it in Prowler Cloud and other deployments without a CLI:
```console
export PROWLER_AWS_BOTO3_RETRIES_MAX_ATTEMPTS=0
```
The CLI flag takes precedence over the environment variable. The value must be a non-negative integer; when neither is set, Prowler uses 3 retries.
<Warning>
The environment variable is process-wide: it applies to every AWS provider built in the process where it is set, not only to a connection check. A scan started in that same process picks it up too. Boto3's Standard retry mode, which Prowler uses, also retries service-side throttling responses (see the errors listed below), so `0` disables retries for those as well. On a large account a scan can hit throttling under normal load, and with retries disabled that throttling becomes a hard failure instead of a retried call. Set the variable only on the processes that run connection checks; leave scan workers on the default, or raise their retry count instead of lowering it.
</Warning>
## Retry Behavior Overview
Boto3's Standard retry mode includes the following mechanisms:
- Maximum Retry Attempts: Default value set to 3, configurable via the `--aws-retries-max-attempts 5` argument.
- Maximum Retry Attempts: Default value set to 3, configurable via the `--aws-retries-max-attempts 5` argument. `0` disables retries.
- Expanded Error Handling: Retries occur for a comprehensive set of errors.
@@ -40,7 +81,7 @@ Boto3's Standard retry mode includes the following mechanisms:
- Nondescriptive Transient Error Codes: The retrier applies retry logic to standard HTTP status codes signaling transient errors: 500, 502, 503, 504.
- Exponential Backoff Strategy: Each retry attempt follows exponential backoff with a base factor of 2, ensuring progressive delay between retries. Maximum backoff time: 20 seconds
- Exponential Backoff Strategy: Each retry attempt follows exponential backoff with a base factor of 2, ensuring progressive delay between retries. Maximum backoff time: 20 seconds.
## Validating Retry Attempts
@@ -51,4 +92,4 @@ For testing or modifying Prowler's behavior, use the following steps to confirm
This approach follows the [AWS documentation](https://boto3.amazonaws.com/v1/documentation/api/latest/guide/retries.html#checking-retry-attempts-in-your-client-logs), which states that if a retry is performed, a message starting with "Retry needed" will be prompted.
It is possible to determine the total number of calls made using `grep -i 'Sending http request' debuglogs.txt | wc -l`
It is possible to determine the total number of calls made using `grep -i 'Sending http request' debuglogs.txt | wc -l`.
@@ -21,10 +21,31 @@ When scanning the China (`aws-cn`), European Sovereign Cloud (`aws-eusc`) or Gov
- Specify the regions to audit within that partition using the `-f/--region` flag.
- Declare the partition with the `PROWLER_AWS_PARTITION` environment variable, set to `aws`, `aws-cn`, `aws-eusc` or `aws-us-gov`.
<Note>
Refer to: https://boto3.amazonaws.com/v1/documentation/api/latest/guide/credentials.html#configuring-credentials for more information about the AWS credential configuration.
</Note>
### Declaring the Partition
`PROWLER_AWS_PARTITION` tells Prowler which partition the scan runs against, without relying on a region being configured:
```bash
export PROWLER_AWS_PARTITION="aws-us-gov"
```
It matters most where nothing else does. Resolving an identity means calling STS before anything is known about the credentials, and with no region configured Prowler would otherwise start from the commercial endpoints. Declaring the partition makes that first call go to the right place, which is the difference between a scan that starts and one that fails on an endpoint the credentials cannot use.
A region configured for the session still wins when it belongs to the declared partition, so a deployment in `us-gov-west-1` is not sent to `us-gov-east-1`. A region belonging to a different partition is ignored, since a partition that has been declared explicitly is the more deliberate statement of the two.
When no configured region says which one to prefer, the first region of the partition is tried, and up to two more follow if it cannot be reached. A network that routes to only one region of its partition therefore works without having to declare which one that is. Only a connection failure moves on to the next region: a credential error is reported from the first, since it would be the same everywhere. A region excluded from the scan is tried last, so it is avoided whenever another region of the partition answers.
<Note>
Set it wherever the scan runs. For deployments that scan from containers, that means the environment of the containers doing the scanning, not only the one accepting the request.
</Note>
### Scanning Specific Regions
To scan a particular AWS region with Prowler, use:

Some files were not shown because too many files have changed in this diff Show More