Statistics#
Vulnerability-Lookup answers eight questions about a corpus: how many records it holds per period, how they split by lifecycle state, how severity and CVSS version are distributed, how much of it is known to be exploited, how long that took, and which assigners and vendors it is dominated by.
Every one of them is asked per source, so ?source=cvelistv5 and ?source=euvd
are the same question about two corpora rather than two different figures. Six of
them accept any source the instance holds records of; the two rankings accept a fixed
list of four, given with each of them below.
This page is about what the answers mean. The request and response shapes are in the
interactive Swagger UI at /api/ of any instance, the keys behind them are in
Architecture, and the passes that prepare them are in
Prepared statistics figures.
One implementation, two namespaces#
Figure |
Any source |
EUVD |
|---|---|---|
Records per period |
|
|
Lifecycle status |
|
|
CVSS version distribution |
|
|
Severity bands |
|
|
Known-exploited inclusion |
|
|
Time to exploit |
|
|
Top assigners |
|
|
Top vendors |
|
|
The generic namespace takes ?source= and defaults to cvelistv5. The EUVD namespace
pins source=euvd, takes no source parameter, and implements nothing of its own: it
calls the same readers, behind the same cache window (STATS_CACHE_TIMEOUT, one hour
by default). It exists so a page has one stable set of URLs to hold. Because there is
one implementation of each figure, the two surfaces cannot report a number
differently.
/api/stats/vulnerability/count is the older single-period form and returns a bare
number for one year or month. Prefer /counts, which returns the whole series in one
request: twenty years of a chart is otherwise twenty round trips against counters that
are a single MGET.
Reading an answer#
available: false is not zero. A figure whose preparation has never run says so
instead of reporting empty buckets. The distinction is the point: a severity chart of
five empty bars claims the source holds nothing severe, while “no data” is the truth
about an instance that has not run the pass. Most sources prepare only some of the
eight, so this is the normal answer for a figure a given source does not maintain.
The two rankings are the exception. They predate this convention and answer an unprepared leaderboard with an empty list, so an empty ranking is not “no vendors”: it usually means the pass has not run for that source or that period.
Most figures are as old as the last pass, not as old as the request. Nothing here
walks the corpus while a visitor waits. Lifecycle status is read live from the
indexes. The known-exploited cardinalities are live for EUVD, which maintains its
membership sets during ingestion, and as old as the last euvd --stats-scan for every
other source. The counters, both CVSS distributions, the two rankings and the
time-to-exploit summary come from what a pass wrote.
time_to_exploit carries a computed_at timestamp for exactly this reason.
A year means a publication date, never the year embedded in an identifier. The two differ for anything published across a year boundary.
The eight figures#
Records per period#
GET /api/stats/vulnerability/counts?source=<name>[&granularity=month&period=YYYY]
Records of the source per publication year, or per month within one year. The years covered are derived from the source’s own published index, so a source holding only recent records does not report two decades of zeros.
state=reserved reads the reservation counters instead, keyed by reservation date.
The year axis still comes from the published index, so the two series of one source
line up.
period selects the year to break into months and is refused with a 400 when sent
without granularity=month, rather than ignored: a caller who sent one asked for
something, and the whole corpus returned as though it were one year is the single
answer they cannot tell apart from a correct one.
Lifecycle status#
GET /api/stats/vulnerability/status?source=<name>
How many records the source holds in total, how many carry a publication date, and how many carry a reservation date.
The states are not a partition and do not sum to the total. A record can carry both dates, and a source may hold records carrying neither. Charting them as slices of one pie is the easier presentation and the wrong one.
There is no rejected count: no index holds rejected records, so a third slice would have to be invented.
CVSS version distribution#
GET /api/stats/cvss/versions?source=<name>[&period=YYYY]
Which CVSS version the source’s records are scored under, bucketed as 4.0, 3.1,
3.0, 2.0, other and unscored, so the axis is the same six values whatever the
corpus holds.
Severity bands#
GET /api/stats/cvss/scores?source=<name>[&period=YYYY]
The CVSS v3.1 qualitative ratings (none, low, medium, high, critical) applied
to whichever version a record is scored under, plus an explicit unscored bucket. A
v4.0 and a v3.1 record with the same base score land in the same band, which is what a
severity mix means.
unscored is a real bucket rather than an omission: how much of a corpus carries no
usable metric is part of the answer, and on the CVE list it is the largest bucket.
The band comes from the record’s own metrics with the highest CVSS version winning,
which is the score its detail page displays. It is deliberately not derived from
index:<source>:cvss, which carries the maximum across a record’s metrics and is
absent for sources whose pass does not write it. A band computed the other way would
put a record in a severity its own page contradicts.
Both distributions cover published records only, so their total equals the sum of the per-year counts. Counting reserved records would report a lifecycle state as a severity.
Known-exploited inclusion#
GET /api/stats/kev?source=<name>
How many of the source’s records appear in a known-exploited catalogue, as a count, a ratio of the corpus, and a count per configured catalogue.
The per-catalogue counts overlap and do not partition exploited. A record listed
by two catalogues is counted by both, because that is what each catalogue asserts about
it, so charting them as slices of the total overshoots. They are independent
memberships.
Every configured catalogue is listed, including those matching no record: a catalogue
missing from the response would read as “not configured” rather than “nothing of ours
is in it”. unattributed counts exploited records whose catalogue this instance cannot
name, and a non-zero value means the local catalogue reference data is absent or stale.
EUVD maintains its membership sets as a by-product of ingestion. Every other source
gets them rebuilt by euvd --stats-scan --stats-source <name>, which matches KEV
entries against the records the source holds, so the figure is meaningful for CVE-keyed
sources.
Time to exploit#
GET /api/stats/time_to_exploit?source=<name>
Days between a record being published and its first appearance in a known-exploited catalogue, as a mean, a median, a 90th percentile, and the extremes.
The denominator is on the wire, because it is not the corpus. The figure rests on
records that were ever exploited, a small fraction of any source, so the response
carries count beside cves_exploited and entries_scanned, and reports the median
next to the mean: one 2007 CVE listed in 2024 moves a mean by thousands of days and
leaves a median untouched.
Intervals can legitimately be negative, when a catalogue listed a vulnerability before
the record was published. They stay in the distribution and are counted separately as
exploited_before_publication rather than clamped away, so min_days below zero is an
expected reading and not a defect.
unmatched_no_record and unmatched_no_publication_date report entries the scan could
not date, rather than dropping them.
Top assigners#
GET /api/stats/assigners/ranking?source=<name>&limit=10[&period=YYYY]
The assigners (CNAs) that account for most of the source, whole or for one publication period. Capped at ten.
source is one of cvelistv5, certfr_avis, certfr_alerte and euvd, and defaults
to cvelistv5. Any other value is refused with a 400, including sources the other six
figures accept.
For EUVD the assigner is taken from the linked CVE, never from the EUVD record’s own assigner field, which names the organisation that allocated the identifier and is the same for every record an instance mints.
Top vendors#
GET /api/stats/vendors/ranking?source=<name>&limit=10[&period=YYYY]
The vendors most affected, whole or for one publication period. Capped at ten. Vendors, like the assigner and the CVSS metrics, come from the linked CVE for sources whose records are aliases rather than advisories of their own.
source takes the same four values as the assigner ranking.
Two pairs that are easy to confuse#
Known exploited is not the sightings ratio. /api/stats/kev counts membership of a
known-exploited catalogue. /api/stats/sighting/exploitation_ratio reports the share of
records with an exploitation sighting, which is a community observation. Both figures
are legitimate and neither may be labelled as the other.
An EUVD corpus is not a second population. An EUVD record is normally an alias of a CVE, so the EUVD figures largely count the same vulnerabilities under a different identity. The per-source keys exist precisely so the two are never added together.
Preparing the figures#
Two idempotent commands, documented in full under Prepared statistics figures:
# counts per period, the two rankings and both CVSS distributions
$ poetry run index_vulnerabilities --source cvelistv5
# time to exploit, and the known-exploited sets for non-EUVD sources
$ poetry run euvd --stats-scan --stats-source cvelistv5
Run the reindex at least once before trusting any counter on a new, restored or upgraded instance. VL’s derived indexes are rebuilt from the stored records, and a stale index makes a status figure report a fraction of reality while looking entirely normal.
The assigner scorecard – how well each CNA and GNA fills the records it publishes – is a prepared figure too, with its own page: see Assigner scorecard.
The small, stable figures are also exported from /metrics when that endpoint is
enabled, where they acquire a history: see Metrics. The two rankings are
deliberately excluded from the scrape, since a leaderboard mints a new series every
time its membership changes.