A Deep Dive into VirusTotal: Characterizing and Clustering a Massive File Feed

van Liebergen, Kevin; Caballero, Juan; Kotzias, Platon; Gates, Chris

Computer Science > Cryptography and Security

arXiv:2210.15973v2 (cs)

[Submitted on 28 Oct 2022 (v1), last revised 31 Oct 2022 (this version, v2)]

Title:A Deep Dive into VirusTotal: Characterizing and Clustering a Massive File Feed

Authors:Kevin van Liebergen (1), Juan Caballero (1), Platon Kotzias (2), Chris Gates (2) ((1) IMDEA Software Institute, (2) Norton Research Group)

View PDF

Abstract:Online scanners analyze user-submitted files with a large number of security tools and provide access to the analysis results. As the most popular online scanner, VirusTotal (VT) is often used for determining if samples are malicious, labeling samples with their family, hunting for new threats, and collecting malware samples. We analyze 328M VT reports for 235M samples collected for one year through the VT file feed. We use the reports to characterize the VT file feed in depth and compare it with the telemetry of a large security vendor. We answer questions such as How diverse is the feed? Does it allow building malware datasets for different filetypes? How fresh are the samples it provides? What is the distribution of malware families it sees? Does that distribution really represent malware on user devices?
We then explore how to perform threat hunting at scale by investigating scalable approaches that can produce high purity clusters on the 235M feed samples. We investigate three clustering approaches: hierarchical agglomerative clustering (HAC), a more scalable HAC variant for TLSH digests (HAC-T), and a simple feature value grou** (FVG). Our results show that HAC-T and FVG using selected features produce high precision clusters on ground truth datasets. However, only FVG scales to the daily influx of samples in the feed. Moreover, FVG takes 15 hours to cluster the whole dataset of 235M samples. Finally, we use the produced clusters for threat hunting, namely for detecting 190K samples thought to be benign (i.e., with zero detections) that may really be malicious because they belong to 29K clusters where most samples are detected as malicious.

Comments:	16 pages, 4 figures
Subjects:	Cryptography and Security (cs.CR)
Cite as:	arXiv:2210.15973 [cs.CR]
	(or arXiv:2210.15973v2 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.2210.15973

Submission history

From: Kevin van Liebergen [view email]
[v1] Fri, 28 Oct 2022 08:14:49 UTC (1,072 KB)
[v2] Mon, 31 Oct 2022 10:14:29 UTC (135 KB)

Computer Science > Cryptography and Security

Title:A Deep Dive into VirusTotal: Characterizing and Clustering a Massive File Feed

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:A Deep Dive into VirusTotal: Characterizing and Clustering a Massive File Feed

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators