Command Line Interface#

This section describes the various commands available.

All commands below are run from the Vulnerability-Lookup project directory using poetry run.

Core#

Start all the services#

$ poetry run start

Stop all the services#

$ poetry run stop

Start/stop the backend (Valkey and Kvrocks)#

$ poetry run run_backend --start
$ poetry run run_backend --stop

Start/stop all enabled feeders#

$ poetry run feeders_manager --start
$ poetry run feeders_manager --stop

Start only the website#

$ poetry run start_website

Restart only the website#

$ poetry run restart_website

Dump a source in a JSON file#

$ poetry run dump --feed nvd

Bootstrap from Vulnerability-Lookup dumps#

Download the NDJSON files published by a trusted Vulnerability-Lookup instance (for example, the CIRCL dump directory), then import the directory into a fresh instance while Kvrocks is running:

$ poetry run import_dump /path/to/dumps
$ poetry run index_vulnerabilities

The filename (without .ndjson) becomes the source name. Existing vulnerability documents are never overwritten by default — a skipped document only gains metadata and source-index membership it does not have yet, so the command is safe to resume and re-running it never re-triggers notifications. Pass --replace to replace existing documents. Collaborative SQL exports (comments, bundles, sightings, and KEV entries) are ignored, documents whose shape is not recognized are counted and reported without aborting the file, and a failing file does not block the remaining dumps.

Dumps produced by the current dump command annotate each line with the physical storage key and index score (vulnerability-lookup:id and vulnerability-lookup:score), which the importer uses to rebuild storage faithfully. Older dumps without these annotations still import, with the identifier derived from the document content and the file’s modification time used as the index score — freshness-based features (recent listings, since= queries) will then treat those documents as modified at import time.

After bootstrapping, enable and run the normal feeders to receive regular updates. Dumps are intended for one-time initial bootstrap, not periodic synchronization: do not schedule recurring dump downloads and imports (a daily cron, for instance). Without --replace an import never updates a document that already exists, so a polled corpus silently goes stale — and repeatedly fetching full dumps shifts significant bandwidth costs onto the publisher, who may rate-limit or block clients seen doing it. The feeders and the API’s since= parameter exist precisely so nobody has to re-download the whole dataset.

Run an individual feeder manually#

$ poetry run nvd_importer
$ poetry run cvelist_importer
$ poetry run github_importer

See pyproject.toml under [project.scripts] for the full list of available feeder commands.

Update the documentation#

$ cd docs; poetry run make html

The build reads the API reference from docs/_static/files/swagger.json, which is not written by hand – regenerate it first with dump_openapi whenever the API changed.

Web service#

This section describes the main commands related to the web service. All Flask commands are prefixed with poetry run flask --app website.app.

Database#

Init the database:

$ poetry run flask --app website.app db_init

Backup the PostgreSQL database (also runs automatically during poetry run update):

$ poetry run flask --app website.app db_backup

Generate a database models diagram:

$ poetry run flask --app website.app db_diagram

User management#

Create a user:

$ poetry run flask --app website.app create_user --login <login> --email <email> --password <password>

Create an admin:

$ poetry run flask --app website.app create_admin --login <login> --email <email> --password <password>

List all users:

$ poetry run flask --app website.app user_list

Delete a user:

$ poetry run flask --app website.app user_delete --login <login>

Retrieve a user’s API key:

$ poetry run flask --app website.app user_get_api_key --login <login>

Data management#

Update MISP warning lists (the administrator will be prompted to run this during updates):

$ poetry run flask --app website.app update_warninglists

Update the local copy of the GCVE registry (can be scheduled via cron):

$ poetry run flask --app website.app update_gcve_registry

Update the list of EU CSIRT coordinated CNAs the public dashboard presents (can be scheduled via cron, and is also offered as a button on the administration maintenance page). It mirrors the Vulnerability Disclosure Policies table the CSIRTs Network publishes at enisaeu/CNW, taking the organisations whose CNA column reads Yes. An unreachable or unparseable source leaves the previously mirrored copy in place:

$ poetry run python tools/update_cnw_assigners.py

Import OSI-approved licenses:

$ poetry run flask --app website.app import_osi_approved_licenses

Import source code languages:

$ poetry run flask --app website.app import_languages

Clean sightings by source or pattern:

$ poetry run flask --app website.app clean_sightings

Backfill sightings from existing data:

$ poetry run flask --app website.app backfill_sightings

Creates sightings retroactively from all existing comments (seen), bundles (seen, one per vulnerability), and KEV entries (exploited). Each sighting inherits the creation_timestamp of its source object. The command is idempotent: running it multiple times will not create duplicates, as it skips any entry whose source URL already exists in the database.

Remove duplicate sightings of the same observation:

$ poetry run flask --app website.app dedup_sightings        # dry run: counts and largest groups
$ poetry run flask --app website.app dedup_sightings --yes  # delete them

A sighting whose source is a URL or a Telegram/… identifier records one observation. Until the #660 fix the API rejected a repeat of it only within the same UTC day, so collectors re-scanning an overlapping window stored it again after midnight UTC. For each group sharing vulnerability, source, type, author and origin, the command keeps the oldest sighting and deletes the others with their role assignments. The same source reported by different authors is kept: each is an independent assertion. Label sources such as CISA-KEV are left alone. Only sightings created on this instance are considered; --all-origins includes synced ones (the next sync may bring them back) and --author <login> narrows to one user.

Move the sightings of one instance to another:

$ poetry run flask --app website.app export_sightings sightings.ndjson              # on the source instance
$ poetry run flask --app website.app import_sightings sightings.ndjson --dry-run    # on the target: what would happen
$ poetry run flask --app website.app import_sightings sightings.ndjson --author admin

Both stream: the export writes one sighting per line as it walks the table and the import commits every thousand lines, so they handle volumes the browser export/import of /admin/maintenance times out on. - stands for stdout and stdin (export_sightings - | gzip). The line format is the one of the published dumps/sightings.ndjson, and the import also reads the JSON array the browser export produces. --local-only exports the sightings created on this instance and leaves the synced ones out.

The import keeps the uuid, the origin instance and the timestamp of every sighting, so the copy is the same assertion rather than a new one, and a sighting already present is skipped: the same file can be imported twice. The author is matched by uuid; a sighting whose author does not exist on the importing instance (the usual case between two instances) is owned by the --author user instead, admin by default. Records the model refuses (unknown type, malformed identifier) are counted and reported, not fatal.

Regenerate the OpenAPI specification#

The API reference published in the documentation is a snapshot of the specification, committed at docs/_static/files/swagger.json. This command rewrites it from the API object itself, so no instance has to be running and the result cannot be an older deployment’s:

$ poetry run flask --app website.app dump_openapi

Two things make this preferable to curling /api/swagger.json off a live instance. The output is stable: flask-restx serialises the methods of an endpoint out of a set, so two dumps of an unchanged API otherwise differ by well over a thousand lines of reordering, and the diff of a real change is unreadable. And the instance UUID that reaches the schema as the documented default of every vulnerability_lookup_origin field is replaced by a fixed example, so the published reference does not advertise whichever instance happened to produce it (--example-origin chooses another).

The specification only covers what the instance exposes: the namespaces behind user_accounts, cna and cpe_enabled are registered at import time, so an instance with one of them off has no routes to document for it. The shipped samples alone are three paths short (/cna_credential/, /cna_credential/{credential_id} and /cpe_editor/status/{vulnerability_id}). Rather than quietly publishing the smaller file, the command refuses to write and names the modules to turn on; pass --allow-partial if you really do want only what this instance exposes.

--check writes nothing and exits non-zero when the committed file is not what the current code would produce:

$ poetry run flask --app website.app dump_openapi --check

That is the form the OpenAPI workflow runs on every pull request, which is what keeps the snapshot honest. It needs no storage, cache or database – the specification comes from the API object, which touches none of them – only the samples with every optional module enabled, the profile .github/workflows/openapi.yml writes before checking. Regenerating from a full instance reproduces that byte for byte.

A pre-commit hook regenerates the snapshot for you when you touch website/web/api/ or bump the version in pyproject.toml (tools/regenerate_openapi_snapshot.py). It cannot be the gate – it runs on whoever is committing, with whatever modules they happen to have enabled, and pre-commit install is a choice each contributor makes – so it is built to step aside rather than get it wrong: it checks that the instance can produce the complete reference before doing anything, and skips with a one-line reason when it cannot. What it does not do is pass quietly when the generator is broken; that fails, with the error. When it does regenerate, the commit stops so you can stage the result and look at it.

Note

A relative --output is taken from VULNERABILITYLOOKUP_HOME: importing the application changes the working directory. The command prints the path it actually used.

Health and integrity checks#

Report whether the derived state indexes still account for the corpus. Nothing else surfaces a stale index: every read path that trusts one is quietly wrong while the UI, the API and the logs look normal.

$ poetry run flask --app website.app index_health

Count EUVD identifier invariant breaches — records carrying several CVEs or none, pointer disagreements, and published CVEs with no EUVD — and store the result for the /metrics endpoint to report. Exits non-zero when an invariant is broken, so it can be scheduled and left to speak up:

$ poetry run euvd --integrity-scan

It walks the whole identifier space, so run it on a schedule (daily is ample) rather than expecting the metrics endpoint to compute it per scrape. See Metrics.

Prepared statistics figures#

The per-source statistics endpoints (/api/stats/vulnerability/counts, /api/stats/cvss/versions, /api/stats/cvss/scores, /api/stats/kev and /api/stats/time_to_exploit, all taking ?source=<name>, cvelistv5 by default) and the pages built on them read figures that are prepared in advance rather than recomputed per visitor: at ~400k records, counting severities or scanning the known-exploited catalogue on request would not answer in a page load. A figure whose preparation has never run answers available: false rather than zero. /api/stats/vulnerability/status is the exception: it reads the recency indexes directly and is always live.

Two commands produce them, and both are meant to run on a schedule, once per source the statistics should cover:

# counts per year and status, top vendors, top assigners, and the CVSS
# distributions by band and by version -- all derived from the stored records
$ poetry run index_vulnerabilities --source cvelistv5
$ poetry run index_vulnerabilities --source euvd

# average and median time to exploit, and the known-exploited share -- both span
# PostgreSQL (KEV entries) and Kvrocks (the records), so they are computed as a job
$ poetry run euvd --stats-scan --stats-source cvelistv5
$ poetry run euvd --stats-scan                          # defaults to euvd

euvd and euvd_seed_counters are the same command; the flag lives on the EUVD maintenance script for historical reasons but is source-agnostic.

A third prepares the assigner scorecard, the ranking of every CNA and GNA by the quality of its publications:

# the CVE list and every GNA source with records, whole history; also run by
# every index_vulnerabilities pass
$ poetry run scorecard
# only the rolling window the ranking is built on: cheap enough to run hourly
$ poetry run scorecard --window-only

--stats-scan needs the KEV catalogues imported first, and refuses to overwrite the stored figures when it finds no exploited KEV entries at all: an empty catalogue table is indistinguishable from an import that has not run yet, and the previous figures are the better answer. It exits non-zero when it refused either figure.

Both commands are idempotent. A cron entry at the cadence the statistics page should refresh on is enough, for example:

# m h dom mon dow  command
30 3 * * *  cd /path/to/vulnerability-lookup && poetry run index_vulnerabilities --source cvelistv5
30 4 * * *  cd /path/to/vulnerability-lookup && poetry run euvd --stats-scan --stats-source cvelistv5
15 * * * *  cd /path/to/vulnerability-lookup && poetry run scorecard --window-only

Whatever the cadence, run the reindex at least once before trusting any counter on a new, restored or upgraded instance: VL’s derived indexes are rebuilt from the stored records, and a stale index makes a “by status” figure report a fraction of reality while looking entirely normal. An instance that ran index_vulnerabilities before the CVSS distributions existed has its per-year counts but not the distributions, so /api/stats/cvss/* stays available: false until the pass is run again.

Known-exploited share#

/api/stats/kev?source=<name> counts the membership sets index:<source>:kev and index:<source>:kev:<catalog>. They are kept two ways:

  • EUVD maintains its own as a by-product of ingestion, and on every KEV entry change, so euvd --stats-scan leaves them alone.

  • Every other source gets them from euvd --stats-scan --stats-source <name>, which rebuilds them from the KEV entries on each run. A CVE listed in a KEV catalogue counts when the source holds a record with that id, so the figure is meaningful for CVE-keyed sources (cvelistv5, nvd, …). Each run replaces the sets, so a retracted entry leaves them on the next run.

The rebuild is refused, keeping the previous sets, when no KEV entry is exploited (the catalogues are not imported yet) or when none of the exploited CVEs is a record of the source (its index is not rebuilt yet, or its records are not CVEs). Until a first successful run the endpoint answers available: false. An exploited entry whose origin maps to no configured catalogue still counts as exploited, and is reported under unattributed.

Upgrading from the EUVD-only figures#

These figures were EUVD’s before they were the platform’s, and were keyed euvd:stats:* accordingly. Making them per-source moved the source into the key — euvd:stats:cvss_band:high is now stats:euvd:cvss_band:high — so an instance that ran the earlier code holds a set of keys nothing reads any more, including a time-to-exploit snapshot that the statistics endpoints will report as available: false while it sits under its old name.

Run once, after upgrading:

$ poetry run euvd --migrate-stats-keys

It moves each figure to the name its readers use and drops any old key the new namespace already holds, so it is idempotent and can be run before or after the reindex. Without it the counters return on the next index_vulnerabilities --source euvd, but the snapshot does not — it is a cross-store job, and re-running it means re-deriving data that is already stored.

Background services#

Launch the email notification service:

$ poetry run flask --app website.app notify_users

Launch the synchronization service (see Synchronization service for details):

$ poetry run flask --app website.app sync