Phase 1 · the hardest step

Finding every place you use cryptography.

You can't migrate what you can't see, and no single technique finds it all. This is a practical guide to layered cryptographic discovery — including how to mine the logging and SIEM tools you already run — plus a plan for an AI-based inventory proof-of-concept.

Why it's mandated: OMB M-23-02 requires U.S. federal agencies to submit a prioritized cryptographic inventory and update it annually through the migration. Inventory is step one of the CISA/NSA/NIST quantum-readiness process — and the foundation everything else builds on.
01

The discovery methods, at a glance

The NCCoE project uses a deliberate mix of active and passive techniques — static code analysis, passive network sniffing, and agent-based host scanning — because each sees a different slice of the estate. Here is the full landscape and, crucially, what each one misses.

MethodWhat it findsData sourceBlind spots
Passive networkTLS versions, cipher suites, curves, certificates on the wireNetwork taps/SPAN, Zeek/Corelight, packet brokers, JA3/JA4 fingerprintsData-at-rest, internal app crypto, key sizes of private keys
SIEM / log miningTLS & cert metadata already logged across the estateSplunk, Elastic, Sentinel, QRadar, Chronicle — Zeek ssl.log, proxy/WAF, IIS/Nginx, Schannel eventsOnly what devices already log; gaps where TLS logging is off
Active scanningCipher suites, protocol versions, cert chains, weak configtestssl.sh, sslyze, nmap ssl-enum-ciphers, internal scanners; CT logsRisky for OT/IoT; sees only reachable listening services
Certificate & PKICert key algorithms/sizes, signature algorithms, expiry, issuersCA databases, ACME, Venafi/Keyfactor, cloud cert managers, crt.shCerts not managed centrally; embedded/self-signed certs
Source / SASTCrypto API calls, algorithms, modes, key sizes → CBOMIBM/PQCA CBOMkit (Sonar), CodeQL + cryptobom-forge, SemgrepRuntime/config-driven choices; native OpenSSL C usage
Binary / firmwareCrypto constants, S-boxes, certs/keys in artifactscbomkit-theia, GhidraFindCrypt, YARA, binwalkStripped/obfuscated/table-free & hardware-accelerated crypto
Host / endpoint agentsKeystores, SSH keys, TLS libs & versions, configs, cert storesosquery (user_ssh_keys, certificates), Keyfactor/InfoSec Global, TychonAgentless legacy/OT; custom application logic
Cloud & KMS APIsManaged key specs, cert key algorithms, HSM inventoriesAWS KMS/ACM, Azure Key Vault, GCP KMS/CAS, K8s cert-manager, PKCS#11How apps use keys; client-side/embedded crypto
App & configDatabase TDE, broker TLS, middleware settingsDB DMVs (SQL Server, Oracle), Kafka/RabbitMQ configsColumn-level/app-driven crypto chosen in code
The core truth: no method is sufficient alone. Network & SIEM see traffic but not keys at rest; cloud APIs see managed keys but not how they're used; SAST sees code but not runtime choices; host agents see disks but not embedded firmware. A defensible inventory fuses them and normalizes into a CBOM.
02

Mining your logging & SIEM tools

Many organizations start with network-device logs, but the far bigger opportunity is the centralized logging you already run. If your SIEM ingests TLS/handshake telemetry, you have a partial cryptographic inventory sitting in it today — no new sensors required. This is often the fastest way to get a first, broad picture.

Where the crypto signal lives in logs

  • Network sensors — Zeek/Corelight ssl.log is the richest source: one record per TLS connection with version, cipher suite, elliptic curve, server name, and certificate details. JA3/JA4 fingerprints summarize the client/server handshake.
  • Web & proxy tiers — IIS, Apache/Nginx, load balancers, and forward/reverse proxies can log negotiated protocol and cipher per request.
  • Endpoint & OS — Windows Schannel/CAPI events and TLS logs; VPN and auth logs.
  • Firewalls & gateways — TLS-inspection appliances expose handshake metadata.

Turning logs into an inventory

The pattern is the same across platforms: extract the cipher-suite / protocol / cert fields, translate raw cipher IDs to names, classify each as quantum-vulnerable or not, then aggregate by service, host, or business unit.

Splunk — Zeek ssl.log

The Splunk Add-on for Zeek maps fields to CIM. Extract cipher IDs, join a lookup that translates ID → name and flags KEX type, then aggregate quantum-vulnerable usage by server.

sourcetype="zeek:ssl"
| lookup cipher_map cipher_id OUTPUT cipher_name kex_family
| eval pq_vulnerable=if(kex_family IN
    ("ECDHE","DHE","RSA"),"yes","no")
| stats count values(version) as tls_versions
    by server_name kex_family pq_vulnerable
| sort - count

Elastic / others

Zeek → Filebeat → Elastic maps to ECS tls.* fields. Aggregate on tls.cipher, tls.version, and tls.server.x509.public_key_algorithm. Microsoft Sentinel (KQL), IBM QRadar, Chronicle, Datadog, and Cribl follow the same shape — the trick is a good cipher-suite → PQC-vulnerability lookup table.

// Sentinel / KQL sketch
Zeek_SSL_CL
| extend pq = iff(KexFamily in
    ("ECDHE","DHE","RSA"), "vulnerable","ok")
| summarize count() by ServerName,
    Cipher, Version, pq
Realistic value & limits. SIEM mining is fast, safe, and broad — but it only sees what is already logged. Many internal TLS flows aren't captured, cipher IDs need translation, and logs reveal negotiated algorithms, not private-key sizes or data-at-rest. Treat it as a high-value first layer, then fill gaps with scanning, PKI, host, and cloud discovery.
03

A layered inventory methodology

Sequence the layers so you get breadth fast, then depth — reconciling everything into one CBOM.

LAYER 1 · DAYS

Mine what you have

Query SIEM/logs and cloud KMS/PKI APIs. Zero new sensors; produces a broad first-cut of TLS, certs, and managed keys.

LAYER 2 · WEEKS

Observe & scan

Add passive network capture (Zeek/JA4) and targeted active TLS/cert scans of reachable services — carefully, avoiding OT.

LAYER 3 · WEEKS

Look inside

Run SAST/CBOM on source and dependencies (CBOMkit, cryptobom-forge); scan container images and key artifacts.

LAYER 4 · ONGOING

Go to the hosts

Deploy host agents (osquery) for keystores, SSH keys, TLS-library versions, and config — the data-at-rest picture.

LAYER 5 · ONGOING

Reconcile

Normalize every source into a single CBOM, de-duplicate, and attribute each asset to a system and owner.

LAYER 6 · CONTINUOUS

Keep it live

Schedule re-discovery so the CBOM tracks drift and proves regressions haven't reintroduced vulnerable crypto.

04

The output: a CBOM (CycloneDX 1.6)

A Cryptography Bill of Materials is the machine-readable record of every cryptographic asset and its dependencies. CycloneDX added native cryptographic support in v1.6 (now ECMA-424). Capture at least these fields per asset:

FieldExample / notes
Asset typealgorithm · certificate · protocol · related-crypto-material (key)
Algorithm & parametersRSA-2048, ECDSA P-256, AES-256-GCM, ML-KEM-768; mode, curve, key size, OID
Primitive / functionencryption · signature · key-encapsulation · hash · MAC
Quantum-security levelNIST category 1–5, or 0 if not quantum-safe
Evidencefile path + line, host, log source, cert serial — where it was found
Context (add for migration)system/service, owner, data-confidentiality lifetime, exposure, dependencies

The last row isn't in the base CBOM spec but is what turns an inventory into a prioritizable migration backlog.

05

Tool landscape

Open source

  • PQCA / IBM CBOMkit — Sonar Cryptography plugin (source→CBOM), theia (containers/dirs), CBOM Viewer.
  • cryptobom-forge (Santander) — CodeQL/SARIF → CycloneDX CBOM.
  • Zeek + JA3/JA4, testssl.sh, sslyze, nmap — network & TLS discovery.
  • osquery — host keystores, SSH keys, certificates.
  • GhidraFindCrypt / YARA / binwalk — binary & firmware.

Commercial (representative)

  • Keyfactor / InfoSec Global AgileSec + CipherInsights (host + passive network).
  • IBM Quantum Safe Explorer/Advisor, SandboxAQ AQtive Guard.
  • CryptoNext COMPASS, Venafi/CyberArk, AppViewX, Tychon.
  • Encryption Consulting CBOM Secure, QryptoCyber, ISARA Advance.

Selection criteria that matter: coverage breadth (network + host + code + cloud), source-code crypto discovery, dependency/impact mapping, crypto-aware risk scoring, and open-standard CycloneDX CBOM export to avoid lock-in.

06

An AI-based inventory PoC

Existing tools produce a lot of raw signal. The gap they leave — and where AI adds the most value — is normalization, classification, correlation, and prioritization: turning heterogeneous scan output and logs into a clean, deduplicated, risk-ranked CBOM with human-readable findings. Here is a pragmatic proof-of-concept plan.

What the AI does (and doesn't) do

AI is well-suited to

  • Parsing messy, varied outputs (testssl JSON, Zeek logs, SIEM exports, certutil dumps) into one schema.
  • Classifying algorithms as quantum-vulnerable and assigning NIST security categories.
  • De-duplicating and correlating the same asset seen by multiple tools.
  • Drafting risk scores (Mosca-style) and plain-language remediation guidance.
  • Mapping findings to the right migration wave and owner.

Keep deterministic / human

  • The scanning/collection itself (use proven tools; don't let a model touch production).
  • Ground-truth cipher-ID → algorithm lookups (a table, not a guess).
  • Final risk sign-off and business-context inputs (data lifetime).
  • Anything acting on OT/embedded systems.

Reference architecture

A · Collect
Reuse existing tools/exports: SIEM queries, Zeek logs, testssl/sslyze/nmap, CBOMkit/CodeQL, osquery, cloud KMS/PKI API pulls. No new production risk.
B · Normalize
An LLM-assisted parser maps each source's fields into a common intermediate record. Deterministic lookups resolve cipher IDs and OIDs; the model handles the messy/unstructured remainder.
C · Classify & enrich
Tag each asset quantum-vulnerable/safe + NIST category, infer purpose, and correlate duplicates across sources into a single CBOM entity.
D · Prioritize
Combine algorithm risk with business-context inputs (data lifetime, exposure) to produce a Mosca-based score and a migration wave per asset.
E · Report
Emit a CycloneDX 1.6 CBOM plus a human-readable summary, per-owner backlog, and draft remediation notes.

Delivery model & guardrails

  • Bring-your-own-key, in-browser. A client-side page that calls the model with your API key keeps pasted data off any third-party server. Good for a fast PoC on redacted samples.
  • For production/sensitive data, run the same pipeline inside your own environment (a private endpoint or self-hosted deployment) so nothing leaves your boundary.
  • Redact before you paste. Strip hostnames/IPs/secrets from samples during evaluation; keep the algorithm/parameter fields that matter.
  • Human-in-the-loop. Treat AI output as a high-quality draft to verify, not an authoritative inventory.

Start your inventory

Take the CBOM-aligned template and the reference architecture above into your own environment, then work the results through the migration phases.