[ << ALL_FEED ]

How to fix CFG

More in Reverse engineering

How to fix CFG 🔧

In the process of reverse engineering malware, we encounter cases where obfuscation hinders understanding the overall algorithm. One example is the creation of two consecutive opposite conditional jumps to a single point.

Schematically it looks like this:


start:
    jnX labelA
    jX labelA
labelA:
    <bytes>
Code language: YAML (yaml)


That is, when analyzing code, IDA first follows the False branch and creates code there, deferring the True branch for later. Upon encountering the jnX instruction, IDA creates code immediately after the current one.

Next, upon encountering the opposite conditional jump, it again follows the False branch and creates code at the next address, building a garbage instruction after which disassembly is no longer possible. Then IDA returns to the deferred queue and takes the address from there, but the trouble is that code has already been created there, which means analysis terminates (clearly visible in screenshot 1).

During live execution, regardless of the flag state, the jump to the operand will be taken, meaning execution will proceed as in screenshot 2 if we correct the control flow.

The difficulty in researching such shellcodes is that there can be many jumps and each time the analysis will stumble on them. You can try to fix it manually, but what if there are hundreds of such blocks?

✍️ To handle this, let’s write a simple IDAPython script that will fix the problem automatically. The task is to find these blocks and patch them.

First, let’s compile a list of opposite conditional jumps and put it into a function that will compare two consecutive instructions against the list:


def c_jumps(addr, n_addr):
    ops = [
        ("jz", "jnz"),
        ("jnz", "jz"),
        ("je", "jne"),
        ("jne", "je"),
        ...
    ]
    if (ida_ua.print_insn_mnem(addr), ida_ua.print_insn_mnem(n_addr)) in ops:
        return True
    return False
Code language: Python (python)


We will sequentially iterate through each instruction until we find the ones we need or hit the limit.


def deobf(start, limit=BADADDR):
    while addr != BADADDR:
        n_addr = ida_search.find_code(addr, ida_search.SEARCH_DOWN)
        if n_addr == BADADDR:
            break
 
        if not c_jumps(addr, n_addr):
            addr = n_addr
            continue
Code language: Python (python)


🧐 Using the find_code method from the ida_search module, we find the next address where code exists, and using the c_jumps function, we check whether the instructions at this and the next address are opposite jumps. Having found them, we need to check whether these jumps point to a single point (that is, whether their operands are equal):


o1 = get_operand_value(addr, 0)
o2 = get_operand_value(n_addr, 0)
 
if o1 != o2:
    addr = n_addr
    continue
 
insn = ida_ua.insn_t()
l1 = ida_ua.decode_insn(insn, addr)
l2 = ida_ua.decode_insn(insn, n_addr)
Code language: plaintext (plaintext)


Using get_operand_value, we obtain the operand values (for jX and jXX it is one) and check their equality. To determine the length of the segment that needs to be patched, using decode_insn from ida_ua, we find the instruction lengths.

Then we fix the conditional jumps and ask IDA to analyze the new code. This is necessary so that further obfuscation blocks can be found:


after_addr = n_addr + l2
ida_bytes.patch_bytes(addr, bytes([0x90] * (l1 + l2 + 1)))

ida_auto.auto_wait()
ida_bytes.del_items(after_addr, ida_bytes.DELIT_EXPAND)
ida_auto.auto_wait()

ida_ua.create_insn(o1)
addr = o1
ida_auto.auto_wait()
Code language: plaintext (plaintext)


🗂 Using the patch_bytes method from ida_bytes, we patch the instructions, with auto_wait from ida_auto we ask IDA to analyze the new code, then, using del_items, we delete the garbage instructions created during the initial analysis and analyze again. Using create_insn, we create a valid instruction where the conditional jumps pointed, and re-analyze one last time.

Thus, this script allows obtaining normal ASM code from obfuscated code, as in screenshot 3.

Sometimes after running it, isolated unrecognized bytes remain, but they are easy to fix. From this code, you can create a function, decompile it, and analyze it (as in screenshot 4).

Study IDAPython, write scripts. Happy reversing!


#tip #reverse #idapython
@ptescalator

More from global_author

More from global_author

More in Reverse engineering