# Statistics

Vulnerability-Lookup answers eight questions about a corpus: how many records it
holds per period, how they split by lifecycle state, how severity and CVSS version
are distributed, how much of it is known to be exploited, how long that took, and
which assigners and vendors it is dominated by.

Every one of them is asked *per source*, so `?source=cvelistv5` and `?source=euvd`
are the same question about two corpora rather than two different figures. Six of
them accept any source the instance holds records of; the two rankings accept a fixed
list of four, given with each of them below.

This page is about what the answers mean. The request and response shapes are in the
interactive Swagger UI at `/api/` of any instance, the keys behind them are in
[Architecture](architecture.md), and the passes that prepare them are in
[Prepared statistics figures](command-line-interface.md#prepared-statistics-figures).

## One implementation, two namespaces

| Figure | Any source | EUVD |
| --- | --- | --- |
| Records per period | `/api/stats/vulnerability/counts` | `/api/euvd/stats/count` |
| Lifecycle status | `/api/stats/vulnerability/status` | `/api/euvd/stats/status` |
| CVSS version distribution | `/api/stats/cvss/versions` | `/api/euvd/stats/cvss/versions` |
| Severity bands | `/api/stats/cvss/scores` | `/api/euvd/stats/cvss/scores` |
| Known-exploited inclusion | `/api/stats/kev` | `/api/euvd/stats/kev` |
| Time to exploit | `/api/stats/time_to_exploit` | `/api/euvd/stats/time_to_exploit` |
| Top assigners | `/api/stats/assigners/ranking` | `/api/euvd/stats/assigners/ranking` |
| Top vendors | `/api/stats/vendors/ranking` | `/api/euvd/stats/vendors/ranking` |

The generic namespace takes `?source=` and defaults to `cvelistv5`. The EUVD namespace
pins `source=euvd`, takes no `source` parameter, and implements nothing of its own: it
calls the same readers, behind the same cache window (`STATS_CACHE_TIMEOUT`, one hour
by default). It exists so a page has one stable set of URLs to hold. Because there is
one implementation of each figure, the two surfaces cannot report a number
differently.

`/api/stats/vulnerability/count` is the older single-period form and returns a bare
number for one year or month. Prefer `/counts`, which returns the whole series in one
request: twenty years of a chart is otherwise twenty round trips against counters that
are a single `MGET`.

## Reading an answer

**`available: false` is not zero.** A figure whose preparation has never run says so
instead of reporting empty buckets. The distinction is the point: a severity chart of
five empty bars claims the source holds nothing severe, while "no data" is the truth
about an instance that has not run the pass. Most sources prepare only some of the
eight, so this is the normal answer for a figure a given source does not maintain.

The two rankings are the exception. They predate this convention and answer an
unprepared leaderboard with an empty list, so an empty ranking is not "no vendors":
it usually means the pass has not run for that source or that period.

**Most figures are as old as the last pass, not as old as the request.** Nothing here
walks the corpus while a visitor waits. Lifecycle status is read live from the
indexes. The known-exploited cardinalities are live for EUVD, which maintains its
membership sets during ingestion, and as old as the last `euvd --stats-scan` for every
other source. The counters, both CVSS distributions, the two rankings and the
time-to-exploit summary come from what a pass wrote.
`time_to_exploit` carries a `computed_at` timestamp for exactly this reason.

**A year means a publication date**, never the year embedded in an identifier. The two
differ for anything published across a year boundary.

## The eight figures

### Records per period

`GET /api/stats/vulnerability/counts?source=<name>[&granularity=month&period=YYYY]`

Records of the source per publication year, or per month within one year. The years
covered are derived from the source's own published index, so a source holding only
recent records does not report two decades of zeros.

`state=reserved` reads the reservation counters instead, keyed by reservation date.
The year axis still comes from the published index, so the two series of one source
line up.

`period` selects the year to break into months and is refused with a 400 when sent
without `granularity=month`, rather than ignored: a caller who sent one asked for
something, and the whole corpus returned as though it were one year is the single
answer they cannot tell apart from a correct one.

### Lifecycle status

`GET /api/stats/vulnerability/status?source=<name>`

How many records the source holds in total, how many carry a publication date, and how
many carry a reservation date.

**The states are not a partition and do not sum to the total.** A record can carry both
dates, and a source may hold records carrying neither. Charting them as slices of one
pie is the easier presentation and the wrong one.

There is no rejected count: no index holds rejected records, so a third slice would
have to be invented.

### CVSS version distribution

`GET /api/stats/cvss/versions?source=<name>[&period=YYYY]`

Which CVSS version the source's records are scored under, bucketed as `4.0`, `3.1`,
`3.0`, `2.0`, `other` and `unscored`, so the axis is the same six values whatever the
corpus holds.

### Severity bands

`GET /api/stats/cvss/scores?source=<name>[&period=YYYY]`

The CVSS v3.1 qualitative ratings (`none`, `low`, `medium`, `high`, `critical`) applied
to whichever version a record is scored under, plus an explicit `unscored` bucket. A
v4.0 and a v3.1 record with the same base score land in the same band, which is what a
severity mix means.

`unscored` is a real bucket rather than an omission: how much of a corpus carries no
usable metric is part of the answer, and on the CVE list it is the largest bucket.

The band comes from the record's own metrics with the **highest CVSS version winning**,
which is the score its detail page displays. It is deliberately not derived from
`index:<source>:cvss`, which carries the maximum across a record's metrics and is
absent for sources whose pass does not write it. A band computed the other way would
put a record in a severity its own page contradicts.

Both distributions cover published records only, so their total equals the sum of the
per-year counts. Counting reserved records would report a lifecycle state as a
severity.

### Known-exploited inclusion

`GET /api/stats/kev?source=<name>`

How many of the source's records appear in a known-exploited catalogue, as a count, a
ratio of the corpus, and a count per configured catalogue.

**The per-catalogue counts overlap and do not partition `exploited`.** A record listed
by two catalogues is counted by both, because that is what each catalogue asserts about
it, so charting them as slices of the total overshoots. They are independent
memberships.

Every configured catalogue is listed, including those matching no record: a catalogue
missing from the response would read as "not configured" rather than "nothing of ours
is in it". `unattributed` counts exploited records whose catalogue this instance cannot
name, and a non-zero value means the local catalogue reference data is absent or stale.

EUVD maintains its membership sets as a by-product of ingestion. Every other source
gets them rebuilt by `euvd --stats-scan --stats-source <name>`, which matches KEV
entries against the records the source holds, so the figure is meaningful for CVE-keyed
sources.

### Time to exploit

`GET /api/stats/time_to_exploit?source=<name>`

Days between a record being published and its first appearance in a known-exploited
catalogue, as a mean, a median, a 90th percentile, and the extremes.

**The denominator is on the wire, because it is not the corpus.** The figure rests on
records that were *ever* exploited, a small fraction of any source, so the response
carries `count` beside `cves_exploited` and `entries_scanned`, and reports the median
next to the mean: one 2007 CVE listed in 2024 moves a mean by thousands of days and
leaves a median untouched.

Intervals can legitimately be negative, when a catalogue listed a vulnerability before
the record was published. They stay in the distribution and are counted separately as
`exploited_before_publication` rather than clamped away, so `min_days` below zero is an
expected reading and not a defect.

`unmatched_no_record` and `unmatched_no_publication_date` report entries the scan could
not date, rather than dropping them.

### Top assigners

`GET /api/stats/assigners/ranking?source=<name>&limit=10[&period=YYYY]`

The assigners (CNAs) that account for most of the source, whole or for one publication
period. Capped at ten.

`source` is one of `cvelistv5`, `certfr_avis`, `certfr_alerte` and `euvd`, and defaults
to `cvelistv5`. Any other value is refused with a 400, including sources the other six
figures accept.

For EUVD the assigner is taken from the **linked CVE**, never from the EUVD record's
own assigner field, which names the organisation that allocated the identifier and is
the same for every record an instance mints.

### Top vendors

`GET /api/stats/vendors/ranking?source=<name>&limit=10[&period=YYYY]`

The vendors most affected, whole or for one publication period. Capped at ten. Vendors,
like the assigner and the CVSS metrics, come from the linked CVE for sources whose
records are aliases rather than advisories of their own.

`source` takes the same four values as the assigner ranking.

## Two pairs that are easy to confuse

**Known exploited is not the sightings ratio.** `/api/stats/kev` counts membership of a
known-exploited catalogue. `/api/stats/sighting/exploitation_ratio` reports the share of
records with an exploitation *sighting*, which is a community observation. Both figures
are legitimate and neither may be labelled as the other.

**An EUVD corpus is not a second population.** An EUVD record is normally an alias of a
CVE, so the EUVD figures largely count the same vulnerabilities under a different
identity. The per-source keys exist precisely so the two are never added together.

## Preparing the figures

Two idempotent commands, documented in full under
[Prepared statistics figures](command-line-interface.md#prepared-statistics-figures):

```bash
# counts per period, the two rankings and both CVSS distributions
$ poetry run index_vulnerabilities --source cvelistv5

# time to exploit, and the known-exploited sets for non-EUVD sources
$ poetry run euvd --stats-scan --stats-source cvelistv5
```

Run the reindex at least once before trusting any counter on a new, restored or
upgraded instance. VL's derived indexes are rebuilt from the stored records, and a
stale index makes a status figure report a fraction of reality while looking entirely
normal.

The assigner scorecard -- how well each CNA and GNA fills the records it publishes --
is a prepared figure too, with its own page: see [Assigner scorecard](scorecard.md).

The small, stable figures are also exported from `/metrics` when that endpoint is
enabled, where they acquire a history: see [Metrics](metrics.md). The two rankings are
deliberately excluded from the scrape, since a leaderboard mints a new series every
time its membership changes.
