Skip to content

Instantly share code, notes, and snippets.

@ctron
Last active September 9, 2026 10:22
Show Gist options
  • Select an option

  • Save ctron/aee0f728b71ede2292170042f4b79821 to your computer and use it in GitHub Desktop.

Select an option

Save ctron/aee0f728b71ede2292170042f4b79821 to your computer and use it in GitHub Desktop.
Correlation Engine Redesign — Roadmap

Correlation Engine Redesign — Roadmap

What Correlation Is

Correlation is the core value proposition of Trustify: matching advisory data (CSAF, OSV, CVE, NVD) against SBOMs to answer "what vulnerabilities affect my software?"

The system must take a component identifier — a PURL, CPE, or file hash — and determine which advisories assert something about it, what status each advisory declares (affected, not_affected, fixed, under_investigation), and what the resolved answer is when multiple advisories disagree. Every step of this process must be explainable to the user.

Cornerstone Requirements

These are non-negotiable properties that define what "correct" means:

  • Spec compliance. Match using identifiers endorsed by CSAF v2.0, VEX, and OSV: PURL, CPE (2.2 and 2.3), and file hashes. Not ad-hoc heuristics like bare package name strings.

  • Correctness. The 18 test scenarios (S1–S18) are the acceptance gate. All must pass — zero false positives from flawed matching paths, zero false negatives from dropped statuses or aliases.

  • Component identity only. Match based on what the component is (its PURL, CPE, hash), not which product it belongs to. No product-scoping assumptions (describing CPE, product_status name matching). Product-level questions are the user's responsibility.

  • All VEX statuses honored. not_affected must suppress affected. fixed must suppress affected. under_investigation must surface. The current system only reads affected and ignores the rest.

  • Single source of truth. One matching implementation, not five divergent code paths with inconsistent behavior. A bug fixed in one place is fixed everywhere.

  • Transparency. Users must be able to understand correlation decisions. The engine collects evidence (all advisory assertions about a component), then resolves them into a result. Users can inspect:

    • Why did this package match that vulnerability? — which identifier matched, which version range applied
    • Why didn't it match? — no matching identifier, version outside range, different PURL type
    • Where is this score coming from? — which advisory, which CVSS vector
    • How were conflicting statements resolved? — advisory A says affected, advisory B says not_affected → result and rationale

Verification: The 18 Scenarios

The scenario test suite (etc/test-data/scenarios/S1–S18/) is the acceptance gate. Each scenario contains:

  • Advisory documents (CSAF, CVE, OSV) to ingest
  • SBOMs in both CycloneDX and SPDX formats
  • expected.json with ground-truth verdicts per CVE per SBOM

Each stage must pass its relevant subset. The full suite must pass by the end of Stage 0.

Capability Scenarios
PURL matching + version comparison S2, S5, S6, S8, S9, S14
CPE identity matching S7, S10
Status resolution (not_affected) S10, S12
Unbounded ranges (bare known_affected) S11
Non-CVE aliases (GHSA/RUSTSEC) S13
Cross-stream / cross-product scoping S1, S3, S4, S9
Container / advanced scenarios S15–S18

Stages

Stage 0 — Define and Demonstrate Correct Matching

Goal: Build the authoritative matching logic as a standalone, testable component — before any persistence, API, or worker infrastructure. Prove it produces correct results.

What this delivers:

  • A matching engine that takes a set of advisories and an identifier (PURL, CPE 2.2/2.3, or file hash) and returns:

    • Which vulnerabilities match
    • Which status each advisory asserts
    • How conflicts between advisories are resolved
    • A detailed decision trace explaining the process
  • The matching logic is the single source of truth — the same code is used by the test suite, the diagnostic endpoint, and (later) the background workers.

  • A temporary/ephemeral mode: advisories can be loaded into a scratch pool without full ingestion, identifiers tested against them interactively (API, CLI, or test harness). This allows exploring and demonstrating matching behavior before any infrastructure exists.

  • Decision trace output: for each identifier, a step-by-step trace of what was checked, what matched, what didn't, and why — suitable for display in UI or test assertions.

  • All 18 scenarios pass against this engine, validating correctness before infrastructure is built around it.

Matching dimensions:

  1. PURL (confidence 0.9): match by base PURL identity with type enforcement (an RPM version range cannot match a golang PURL) and version comparison using the scheme appropriate for the PURL type. An advisory assertion without a version range matches only SBOM entries without a version — if you want to match versioned packages, provide a version range.

  2. CPE (confidence 0.8): full CPE name matching per NIST IR 7696, comparing all 11 fields. See the NIST reference implementation for the authoritative compare() function. Supports CPE 2.2 and 2.3.

  3. Digest (confidence 1.0): exact hash-value lookup — the strongest possible match. No version comparison needed.

Verification: The scenario test suite calls this engine directly.

Transition: This stage produces the core matching library. No database tables, no background workers, no API endpoints yet. Subsequent stages wrap this engine in infrastructure.


Stage 1 — Persistence & v4 API

Goal: Wrap the Stage 0 matching engine in database persistence and a fresh v4 API. Materialize results so they don't need to be recomputed on every request.

What this delivers:

  • Database tables for materialized matches and evidence.

  • The matching engine runs against ingested data and writes results to the database.

  • Diagnostic "explain" endpoint: accepts a PURL, CPE, or hash and runs the Stage 0 engine live, returning the full decision trace. Real-time, not stored. Answers "why does/doesn't this package show as affected?"

  • v4 API with slim, opt-in responses. Core match data (vulnerability, status, dimension, confidence) by default. Evidence, scores, package details, and advisory metadata available via query parameters (e.g., ?include=evidence, ?include=scores).

API design principle: Responses are minimal by default. The common case (list of matches with status) is fast and small. Detail is opt-in, not mandatory.

Transition: v3 endpoints remain untouched. v4 exists alongside as a separate API version.


Stage 2 — Freshness & Lifecycle

Goal: Correlation results stay current automatically as documents are ingested, updated, or deprecated.

What this delivers:

  • Automatic re-evaluation when SBOMs or advisories change.

  • When an advisory is added, affected SBOMs are identified and re-evaluated.

  • When an advisory is deprecated (CSAF update), stale matches are removed by re-evaluating affected SBOMs against the updated advisory landscape.

  • Correlation status visible on SBOM/advisory API responses (pending, completed, stale).

  • Staleness detection — users can see whether results reflect the latest advisory landscape or lag behind recent changes.

Transition: After this stage, correlation results are always up-to-date. The UI can show correlation status indicators.


Stage 3 — Scalability & Operations

Goal: Handle large-scale deployments with many SBOMs, advisories, and concurrent users.

What this delivers:

  • Multi-worker concurrency — multiple processors can run in parallel across pods or hosts.

  • Read-only replica usage for matching reads, primary for writes.

  • Configurable batch processing and backpressure during bulk ingestion.

  • Per-ecosystem policy overrides — different matching rules for different package ecosystems.

Transition: Production-ready for large deployments.


Stage 4 — v3 Sunset

Goal: Migrate v3 consumers to the materialized data, deprecate the old query-time correlation code paths.

What this delivers:

  • v3 endpoints backed by materialized data — an adapter reads from the correlation results and returns v3-shaped responses.

  • Deprecation and removal of the five independent SQL correlation paths in modules/fundamental/.

Transition: v3 consumers get correctness and performance benefits without changing their integration. Migration at their pace.


Open Concerns

Items identified during design and review that need further exploration before staging:

  • Containers and layered products — non-RPM content where PURL namespace conventions must be agreed rather than derived (Go modules, Maven in container images).

  • Static linking and vendored dependencies — a single vendored-module vulnerability fans out across every binary that embedded it (Go stdlib). Presentation/rollup concern.

  • Vendor/NVD name mismatch — package names used by vendors differ from NVD names for the same component. May need an aliasing/mapping table.

  • Source-to-binary PURL gap — an advisory referencing a source RPM (pkg:rpm/redhat/kernel?arch=src) won't match the binary RPM (pkg:rpm/redhat/kernel-core). Requires relationship-based matching (SPDX GENERATED_FROM, build metadata).

  • SBOM trust signals — signed SBOMs should carry higher trust than unsigned ones. Policy model should accommodate this.

References

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment