The diamond is barely visible

More in Malware
- This is Siemens...
Recently, our colleagues from the Positive Industrial Expertise Center discovered a curious Windows sample on MalwareBazaar.…
- Anti-antivirus
Recently, we came across an APK with an intriguing and trust-inspiring name: «Антивирус ФСБ.apk». After installing…
- .exe .docm .xlsm
.exe .docm .xlsm Malicious files with these extensions are most often found in corporate network traffic.…
- Operation Chewbacca
At the end of June, the PT ESC team, during incident investigations, discovered a new group…
- Your Zimbra server is at risk
Recently, our PT ESC IR team encountered a new attack by ransomware groups on Zimbra mail…
The diamond is barely visible 💎
During the analysis of PT ESC IR dumps, we periodically encounter new malware families that are not detected by known indicators and YARA signatures and that mimic legitimate or system files well.
When a certain number of machine dumps with similar OSes containing file scan results are available, “fuzzy” hashes can be used for searching; their application is traditionally limited to the task of finding files relatively similar to previously identified malware samples.
🧐 The first screenshot shows the distribution of executable files of a nix-like system taking into account file sizes (the scale is “inverse-logarithmic”: large system files are located closer to the center, small ones on the periphery). To group files taking into account their size and content similarity (it must be kept in mind that these variables are not always independent — for example, when using the TLSH algorithm), it is necessary to perform a “clustering” procedure taking into account the matrix of “cross distances” between all (N) files of the system, which will have size (N^2).
Obviously, to reduce the size of this matrix, it is possible to introduce a partition of the file size range of one or several jointly analyzed systems — the entire size range can be represented as a set of non-overlapping intervals [x-ax;x+ax], where a<1, and x is the central point of the interval. Experience shows that such a division, as a rule, makes it possible to obtain fewer than a hundred size “bands” with a value of a=0.1. In further analysis within individual “bands,” the number of samples will be substantially smaller than the original total number.
The analysis of the “anomalousness” of samples within a separate “band” can be performed taking into account various factors — the average distance to the remaining samples, the number of samples similar to the given one within a specified threshold value, etc., except in cases where a very small number of files (for example, fewer than 10) falls within the “band” — in that case, all of them can be considered “conditionally anomalous.”
The final stage of such an analysis is the identification, within the “anomalous” groups of files obtained for each of the analyzed systems, of those that satisfy the following criteria:
1️⃣ have a small number of similar samples or a high average distance to the remaining samples (for formalization, one can set the upper half of the dynamic range);
2️⃣ do not have, within a single system, files with an identical name but a different path (which makes it possible to filter system files of nix-like systems);
3️⃣ have a small number of files with a similar path/name on the jointly analyzed systems (or have no analogues at all — that is, they are not mandatory for the functioning of the system).
👀 The result of such an analysis is the picture shown in the second screenshot: file sizes are again on an “inverse-logarithmic” scale, the “maximally different” files within a “size band” tend toward the angular coordinate of π radians, while the minimally different ones tend toward 0. The “conditionally anomalous,” i.e., practically unique in size, files have an angular coordinate of 3π/2.
The total number of identified “anomalies” ranges, for different systems, from 0.7% to 8% of the original number of analyzed executable files, which makes it possible in a number of cases to conduct further analysis simply “by eye” — out of the original thousands of files, about fifty remain.
The very first “substantially different” file in this case is indeed malware, and moreover, for an ensemble of 12 analyzed systems, a similar sample is confidently detected on one more machine, while further searching at a “TLSH distance” of no more than 70 units makes it possible to identify 5 more samples on different machines of the ensemble with identical paths — all of them belong to the same family and implement malware persistence through the use of system-generators.

#tip #ir #malware
@ptescalator
More in Malware
- This is Siemens...
Recently, our colleagues from the Positive Industrial Expertise Center discovered a curious Windows sample on MalwareBazaar.…
- Anti-antivirus
Recently, we came across an APK with an intriguing and trust-inspiring name: «Антивирус ФСБ.apk». After installing…
- .exe .docm .xlsm
.exe .docm .xlsm Malicious files with these extensions are most often found in corporate network traffic.…
- Operation Chewbacca
At the end of June, the PT ESC team, during incident investigations, discovered a new group…
- Your Zimbra server is at risk
Recently, our PT ESC IR team encountered a new attack by ransomware groups on Zimbra mail…






