GCVE Workshop - 22 September 2026 (14:00-18:00), Luxembourg Before The Vulnopticon Conference - Registration

GHSA-WJV6-JCFJ-MF9R

Vulnerability from github – Published: 2026-07-28 21:50 – Updated: 2026-07-28 21:50
VLAI
Summary
`datamodel-code-generator` vulnerable to code injection via unescaped carriage return in `--extra-template-data` `comment` field
Details

Summary

datamodel-code-generator is vulnerable to code injection when a developer passes an --extra-template-data file whose comment value contains a literal \r (carriage return). The comment variable is rendered into a Python # comment in six built-in templates with no line-terminator escaping. Python's tokenizer treats a bare CR as a physical-line terminator (see Python language reference — Physical lines), so the comment ends at the \r and the text after it is parsed as Python, including, when the CR is followed by suitable indentation, as a statement within the class body that follows on the next template line.

Details

The vulnerable templates each contain # {{ comment }} with no escaping:

  • src/datamodel_code_generator/model/template/TypeAliasAnnotation.jinja2:12 and :19
  • src/datamodel_code_generator/model/template/TypeAliasType.jinja2:12 and :19
  • src/datamodel_code_generator/model/template/TypeStatement.jinja2:12 and :19
  • src/datamodel_code_generator/model/template/pydantic_v2/BaseModel.jinja2:4
  • src/datamodel_code_generator/model/template/pydantic_v2/RootModel.jinja2:19
  • src/datamodel_code_generator/model/template/pydantic_v2/RootModelTypeAlias.jinja2:13

The pydantic_v2/BaseModel.jinja2:4 site is representative:

class {{ class_name }}({{ base_class }}):{% if comment is defined %}  # {{ comment }}{% endif %}

When the developer-supplied extras file populates comment for a model, the value reaches the template via DataModel.extra_template_data (set in src/datamodel_code_generator/model/base.py:736-742) and Jinja2 interpolates it raw. None of the templates use comment_safe, escape_docstring, or any other line-terminator filter.

PoC

Complete self contained POC is available at my secret gist: https://gist.github.com/thegr1ffyn/8ad6b8cb3cc2be9d3a0144aeb6896a3f

Impact

  • Who's affected: any developer or CI pipeline that runs datamodel-codegen --extra-template-data <file> where the extras file is influenced by attacker-controlled input. Realistic scenarios include:
  • Extras file generated from a third-party schema-annotation system.
  • Extras file vendored from an upstream repository.
  • Extras file produced by a script that merges multiple comment sources.
  • Build pipelines that template the extras file from environment variables, ticket descriptions, or commit metadata.
  • What it gains: arbitrary Python code execution in the importer's process at import time.
  • What it does NOT need: the schema itself can be entirely benign; only the extras file needs to contain the malicious comment.
  • What does block it: not passing --extra-template-data, or rejecting extras files whose comment values contain \r, \x0b, or \x0c before invocation.

Resolution

The fix normalizes comment values from built-in --extra-template-data before template rendering. Inline comments now convert CRLF, bare CR, vertical tab, and form feed into LF and prefix continuation lines with #, so attacker-controlled text stays inside the generated Python comment block.

Remediation

Upgrade to datamodel-code-generator 0.60.2 or later.

This issue affects datamodel-code-generator versions >= 0.14.1, <= 0.60.1 and is fixed in 0.60.2.

Submitted by: Hamza Haroon (thegr1ffyn)

Show details on source website

{
  "affected": [
    {
      "database_specific": {
        "last_known_affected_version_range": "\u003c= 0.60.1"
      },
      "package": {
        "ecosystem": "PyPI",
        "name": "datamodel-code-generator"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "0.14.1"
            },
            {
              "fixed": "0.60.2"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    }
  ],
  "aliases": [
    "CVE-2026-54654"
  ],
  "database_specific": {
    "cwe_ids": [
      "CWE-1336",
      "CWE-94"
    ],
    "github_reviewed": true,
    "github_reviewed_at": "2026-07-28T21:50:09Z",
    "nvd_published_at": null,
    "severity": "HIGH"
  },
  "details": "### Summary\n\n`datamodel-code-generator` is vulnerable to code injection when a developer passes an `--extra-template-data` file whose `comment` value contains a literal `\\r` (carriage return). The `comment` variable is rendered into a Python `#` comment in six built-in templates with **no** line-terminator escaping. Python\u0027s tokenizer treats a bare CR as a physical-line terminator (see [Python language reference \u2014 Physical lines](https://docs.python.org/3/reference/lexical_analysis.html#physical-lines)), so the comment ends at the `\\r` and the text after it is parsed as Python, including, when the CR is followed by suitable indentation, as a statement within the class body that follows on the next template line.\n\n### Details\n\nThe vulnerable templates each contain `# {{ comment }}` with no escaping:\n\n- `src/datamodel_code_generator/model/template/TypeAliasAnnotation.jinja2:12` and `:19`\n- `src/datamodel_code_generator/model/template/TypeAliasType.jinja2:12` and `:19`\n- `src/datamodel_code_generator/model/template/TypeStatement.jinja2:12` and `:19`\n- `src/datamodel_code_generator/model/template/pydantic_v2/BaseModel.jinja2:4`\n- `src/datamodel_code_generator/model/template/pydantic_v2/RootModel.jinja2:19`\n- `src/datamodel_code_generator/model/template/pydantic_v2/RootModelTypeAlias.jinja2:13`\n\nThe `pydantic_v2/BaseModel.jinja2:4` site is representative:\n\n```jinja2\nclass {{ class_name }}({{ base_class }}):{% if comment is defined %}  # {{ comment }}{% endif %}\n```\n\nWhen the developer-supplied extras file populates `comment` for a model, the value reaches the template via `DataModel.extra_template_data` (set in `src/datamodel_code_generator/model/base.py:736-742`) and Jinja2 interpolates it raw. None of the templates use `comment_safe`, `escape_docstring`, or any other line-terminator filter.\n\n### PoC\nComplete self contained POC is available at my secret gist: https://gist.github.com/thegr1ffyn/8ad6b8cb3cc2be9d3a0144aeb6896a3f\n\n### Impact\n\n- **Who\u0027s affected**: any developer or CI pipeline that runs `datamodel-codegen --extra-template-data \u003cfile\u003e` where the extras file is influenced by attacker-controlled input. Realistic scenarios include:\n  - Extras file generated from a third-party schema-annotation system.\n  - Extras file vendored from an upstream repository.\n  - Extras file produced by a script that merges multiple `comment` sources.\n  - Build pipelines that template the extras file from environment variables, ticket descriptions, or commit metadata.\n- **What it gains**: arbitrary Python code execution in the importer\u0027s process at `import` time.\n- **What it does NOT need**: the schema itself can be entirely benign; only the extras file needs to contain the malicious `comment`.\n- **What does block it**: not passing `--extra-template-data`, or rejecting extras files whose `comment` values contain `\\r`, `\\x0b`, or `\\x0c` before invocation.\n\n### Resolution\n\nThe fix normalizes `comment` values from built-in `--extra-template-data` before template rendering. Inline comments now convert CRLF, bare CR, vertical tab, and form feed into LF and prefix continuation lines with `# `, so attacker-controlled text stays inside the generated Python comment block.\n\n### Remediation\n\nUpgrade to `datamodel-code-generator` `0.60.2` or later.\n\nThis issue affects `datamodel-code-generator` versions `\u003e= 0.14.1, \u003c= 0.60.1` and is fixed in `0.60.2`.\n\nSubmitted by: Hamza Haroon (thegr1ffyn)",
  "id": "GHSA-wjv6-jcfj-mf9r",
  "modified": "2026-07-28T21:50:09Z",
  "published": "2026-07-28T21:50:09Z",
  "references": [
    {
      "type": "WEB",
      "url": "https://github.com/koxudaxi/datamodel-code-generator/security/advisories/GHSA-wjv6-jcfj-mf9r"
    },
    {
      "type": "WEB",
      "url": "https://github.com/koxudaxi/datamodel-code-generator/commit/b73abb5cd703a50471b8950bbd3bd0b82ad71de7"
    },
    {
      "type": "PACKAGE",
      "url": "https://github.com/koxudaxi/datamodel-code-generator"
    },
    {
      "type": "WEB",
      "url": "https://github.com/koxudaxi/datamodel-code-generator/releases/tag/0.60.2"
    }
  ],
  "schema_version": "1.4.0",
  "severity": [
    {
      "score": "CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H",
      "type": "CVSS_V3"
    }
  ],
  "summary": "`datamodel-code-generator` vulnerable to code injection via unescaped carriage return in `--extra-template-data` `comment` field"
}



Log in or create an account to share your comment.




Tags
Taxonomy of the tags.


Loading…

Loading…

Loading…

Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.

Sightings

Author Source Type Date Other

Nomenclature

  • Seen: The vulnerability was mentioned, discussed, or observed by the user.
  • Confirmed: The vulnerability has been validated from an analyst's perspective.
  • Published Proof of Concept: A public proof of concept is available for this vulnerability.
  • Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
  • Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
  • Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
  • Not confirmed: The user expressed doubt about the validity of the vulnerability.
  • Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.

Loading…

Detection rules are retrieved from Rulezet.

Loading…

Loading…

Related by attack behaviour

Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.


Loading…