
Joel Calce
Senior Technical Program Manager

Christina DePinto
Senior Product Marketing Manager

James Shank
Senior Software Engineer

Mallory Mooney
Staff Technical Content Writer
A year ago, Datadog’s cloud environment had grown to support tens of thousands of users across the globe. As our host fleet expanded alongside that user growth, we saw that our original system for vulnerability scanning was becoming less effective at reaching and assessing every host.
Over the past year, we migrated host vulnerability scanning to the Datadog Agent to address that issue. Through our migration, Datadog Cloud Security became the system of record for host vulnerability findings. Our Agent-based workflow now maintains scan freshness above 99%, measured as the percentage of in-scope hosts whose latest vulnerability scan was completed within the previous 24 hours. That scanning is one part of our broader vulnerability management program, which supports our work under SOC 2 Type II, ISO 27001, PCI DSS, and FedRAMP® High.
Before cutover, we had to confirm that the Agent-based approach could meet our detection, reporting, and audit requirements at scale. In this post, we’ll explain why we reconsidered the previous architecture, how we compared the two approaches during a six-month parallel run, and how that evaluation informed our cutover decision.
Why we reconsidered remote host vulnerability scanning
We wanted our new vulnerability management tooling to let teams scan hosts within the same workflows they already used to provision, observe, and replace those hosts. By consolidating those workflows, we could more easily investigate whether repeated findings pointed to a shared source, such as a common image that included a vulnerable package. Our teams could then update the image and replace affected hosts through their typical release process.
But managing vulnerabilities together with infrastructure in this way required us to scan every in-scope host consistently, which became harder to do with remote scanning as our fleet grew. Assessing every host took more than 18 hours, while the scanner’s SSH credentials expired after 12 hours. Once those credentials expired, the scanner couldn’t authenticate to the remaining hosts. And without authentication, it couldn’t inspect the local software and package versions required for vulnerability detection.
We could have extended the credential lifetime or continued expanding the scanner fleet. However, either option would still require our Security team to maintain dedicated scanners and manage the credentials and network access those scanners needed to reach hosts. Because the Datadog Agent was already deployed across the fleet for observability, it gave us another way to collect the host data vulnerability detection required. When configured, the Agent could collect an inventory of installed packages and their versions and send it to Datadog through its existing data path. Cloud Security could then use that inventory to identify known vulnerabilities.
The existing Agent deployment gave us a practical alternative to extending the previous architecture, so we compared the two approaches based on cost and coverage.
How we evaluated Agent-based host vulnerability scanning
As Customer Zero, our Security team worked directly with the Cloud Security product team to evaluate host vulnerability scanning in Datadog Cloud Security against our operational and audit requirements. We focused on whether the replacement could keep host coverage current, support reporting, and provide evidence for compliance audits. We also evaluated whether teams could use production context to prioritize findings and remediate vulnerabilities through existing infrastructure workflows.
For six months, we ran our previous scanner and Cloud Security in parallel to determine when we could switch to Cloud Security without losing host coverage or vulnerability detection. We compared which hosts each system covered, how it identified packages, how many vulnerability findings it produced, and how it reported the results. We then evaluated the Agent-based replacement against four requirements.
Keep authenticated host coverage current
Our replacement needed to keep coverage current as hosts entered and left the fleet. A scanner can accurately assess every host it reaches and still leave a substantial gap if new or short-lived hosts never enter its scan schedule.
An Agent-based design fit that requirement because infrastructure teams already managed the Datadog Agent configuration through their workflows to provision and maintain hosts. Enabling host vulnerability scanning through those existing processes meant coverage could follow the host life cycle without requiring a separate scanner to authenticate to each host over SSH.
Prioritize findings with production context
As part of our cutover criteria, we evaluated whether we could combine Cloud Security findings with production context already available in Datadog. That context included whether the affected host was publicly accessible or supported a critical service.

We still kept the full set of findings available for compliance reporting and later review.
Fit vulnerability data into infrastructure releases
At our scale, teams usually remediate shared host vulnerabilities through the existing infrastructure release process. For example, suppose the same vulnerable OpenSSL version appeared across hundreds of API hosts built from one Amazon Machine Image (AMI). The responsible team could update OpenSSL in the image, publish a new AMI, and replace the affected hosts through its typical deployment process. The new scanning workflow needed to support that process without requiring our teams to patch each host or open a separate ticket for every CVE.
Preserve reporting and audit evidence
The replacement also needed to support internal risk reviews and provide evidence for compliance audits, including for SOC 2 Type II, ISO 27001, PCI DSS, and FedRAMP® High.

In the new workflow, we used the same vulnerability findings for both, filtering them by the owning team and by any applicable compliance requirements. Using current findings for reviews and audits kept reporting tied to the data that teams used for remediation. We also avoided maintaining separate, static audit reports that became outdated as findings changed.
Why the scanners produced different results
During the migration, we couldn’t treat differences between the existing scanner and the Agent-based workflow as evidence that one had missed a vulnerability. We first had to determine whether each difference came from how the tools identified installed packages, grouped vulnerabilities into findings, or calculated coverage. Before cutting over an environment, we investigated those differences to determine whether they represented a real gap in detection or a difference in how the two systems reported the same underlying state.
When our investigation exposed a product gap, our Security team worked directly with the Cloud Security product team to address it before cutover. That collaboration is part of how our Security team works as Customer Zero, using Datadog in production and bringing pain points back to the teams building the product.
The scanners identified installed software differently
Package identification sometimes caused the scanners to report different vulnerabilities for the same installed software. During the parallel run, for example, we found that our existing scanner and its proposed replacement, Cloud Security, identified a fork of containerd differently. Because the scanners classified that software differently, comparing their raw finding counts wouldn’t tell us whether either one had missed a vulnerability. We therefore compared the unique CVEs each scanner reported for the same asset.
Hosts with multiple installed kernels revealed another difference. Cloud Security initially reported vulnerabilities for every installed version, including kernels that were installed but not loaded, while the previous scanner excluded those versions. The parallel run surfaced this behavior as a product gap, and before cutover, the Cloud Security product team updated how Datadog reported vulnerabilities for installed kernel versions that weren’t loaded.
The scanners counted vulnerability findings differently
The total number of vulnerability findings sometimes differed even when both scanners identified the same vulnerabilities on the same assets. For example, our previous scanner grouped multiple CVEs under a single plugin, but Cloud Security reported each CVE separately. We decided to match unique CVEs by asset and investigated meaningful differences, including possible false positives. The asset-level comparison gave our Security team and the Cloud Security product team a repeatable way to compare detection results without relying solely on the total number of vulnerability findings.
The scanning methods differed in how quickly they assessed hosts
The previous scanner reported which hosts it had evaluated, but we also needed to verify that coverage stayed current. For hosts enrolled in the Agent-based workflow, we measured scan freshness as the percentage of in-scope hosts whose latest vulnerability scan was completed within the previous 24 hours. That percentage remained above 99%, which gave us a reliable way to identify when hosts were falling behind rather than relying on a point-in-time coverage count.
We separately used the Cloud Security coverage analysis API to identify hosts that hadn’t entered the workflow because the Agent was missing or vulnerability scanning wasn’t configured correctly.
What we learned from the migration
After cutover, our Security team retired the remote credentials and dedicated scanner infrastructure that the previous approach required. With host vulnerability management integrated into the Datadog workflows we already used to operate our infrastructure, teams could use vulnerability findings to easily identify and address shared infrastructure issues.
If you’re considering a similar migration, you can use the same steps to define and test the requirements for a replacement:
Document what your current scanning control provides across coverage, detection, reporting, and audit evidence.
Define which capabilities the replacement must preserve or improve, along with any new capabilities it should provide.
Run both approaches in parallel until you understand each meaningful difference and confirm that the replacement meets those requirements.
Read our documentation to learn more about Datadog Cloud Security and how to enable it on hosts, or you can sign up for a free 14-day Datadog trial.
