Benchmarking DeepDiff Against BinDiff on Stripped and Cross-Compiler Binaries
Back to Blog

Benchmarking DeepDiff Against BinDiff on Stripped and Cross-Compiler Binaries

Sheng Yu, Research Scientist @Deepbits

Benchmarking DeepDiff Against BinDiff on Stripped and Cross-Compiler Binaries

DeepDiff matches related functions across binaries. This matching layer can support patch verification, version analysis, firmware comparison, malware research, and other work that requires code alignment.

In our earlier DeepDiff post, we introduced the product through a firmware patch-verification example. This post focuses on a narrower question: how accurately can DeepDiff match functions when two binaries no longer look alike?

All of these tasks start with the same question:

Which function in one binary matches a function in the other binary?

This sounds simple when both binaries come from the same build system. In that case, addresses may move, but much of the code keeps the same shape.

The problem becomes much harder when the compiler changes, symbols are removed, libraries are linked into the program, or code moves from one function to another. The two functions may come from the same source and still look very different at the binary level.

To answer that question, we compared DeepDiff with Google BinDiff across 21 binary pairs. The tests cover three common problems in binary matching:

  1. Different compilers can create very different control-flow graphs.
  2. A diffing tool cannot match functions that its disassembler never found.
  3. A version update can move code to a function with a different name.

This is a benchmark of DeepDiff's core matching layer. It is not a test of one specific workflow built on top of those matches.

What We Tested

We tested 21 binary pairs:

  • 6 Lua pairs built with GCC or Clang at O2 or O3
  • 15 minigzip pairs built with different zlib versions

Both tools received stripped binaries. To check their results, we used matching unstripped binaries with symbols.

We counted a match as correct when both addresses had the same symbol name. For example, if a tool produced this match:

0xf610 -> 0xee50

and both addresses had the symbol name luaD_call, we counted it as correct. Some tool output used rebased addresses, so we converted those addresses back to ELF virtual addresses before checking them.

We used two standard measures:

  • Precision: Of all matches returned by the tool, how many were correct?
  • Recall: Of all expected function matches, how many did the tool find?

Across all 21 pairs, the combined results were:

ToolPrecisionRecallCorrect matchesWrong matches
DeepDiff0.8920.8795,171626
Google BinDiff0.7570.6423,7801,215

DeepDiff found more correct matches and made fewer wrong matches in these tests. The reason was not one single feature. Each test exposed a different part of the matching problem.

Finding 1: The Same Source Can Produce a Different Graph

BinDiff uses structural signals such as function graphs, basic-block graphs, call graphs, and edge patterns. These signals work well when the two binaries have similar structures.

That condition often holds when both binaries use the same compiler. In our Lua test, BinDiff reached about 0.95 precision and recall when both sides were built with Clang.

The result changed when one side used GCC and the other used Clang. The table below shows the average result in both directions for each pair:

PairBinDiff precision / recallDeepDiff precision / recall
Clang O2 vs. Clang O30.948 / 0.9410.964 / 0.951
GCC 11 O2 vs. GCC 13 O30.810 / 0.7930.861 / 0.898
GCC O2 vs. Clang O20.589 / 0.5790.844 / 0.838
GCC O2 vs. Clang O30.590 / 0.5690.885 / 0.834
GCC O3 vs. Clang O20.598 / 0.5970.842 / 0.850
GCC O3 vs. Clang O30.635 / 0.6290.865 / 0.813

Across the GCC-to-Clang pairs, BinDiff fell to about 0.59. DeepDiff stayed between about 0.84 and 0.89.

One pair makes the difference clear. The lua-clang-O2 -> lua-O2 ground truth contained 659 functions:

ToolCorrect matchesWrong matchesPrecisionRecall
DeepDiff5521020.8440.838
BinDiff3802650.5890.577

For this pair, DeepDiff found 172 more correct matches and made 163 fewer wrong matches.

When Graph Size Changes

The Lua function luaH_getn shows how large the compiler difference can be:

Build of luaH_getnBasic blocksEdgesInstructions
Clang O286137298
GCC O24654137

Both versions came from the same source function. The Clang version had more than twice as many instructions as the GCC version.

BinDiff gave the correct luaH_getn -> luaH_getn match a similarity score of only 0.104. It matched Clang's luaH_getn to GCC's close_func instead.

DeepDiff found the correct match:

0x1d9e0 -> 0x1b280, similarity 0.826

When Graph Size Looks Similar

Large graph changes are not the only problem. The two versions of luaD_call were close in size:

  • Clang: 8 basic blocks and 42 instructions
  • GCC: 7 basic blocks and 40 instructions

BinDiff still matched luaD_call to dothecall. DeepDiff matched it correctly:

0xf610 -> 0xee50, similarity 0.937

The image below shows BinDiff's match. In the stripped binary, sub_F610 is luaD_call and sub_10D80 is dothecall.

BinDiff graph view showing luaD_call matched to dothecall

This is an important failure case. Many small functions have similar graph shapes. When the compiler changes code layout, graph structure can point to the wrong nearby function even when the number of blocks and instructions looks close.

More Examples from Lua

On this one Lua pair, DeepDiff correctly matched 220 functions that BinDiff matched incorrectly or did not match.

Here are several examples:

FunctionDeepDiff matchDeepDiff similarityBinDiff result
luaD_call0xf610 -> 0xee500.937matched to dothecall
luaD_precall0xf2d0 -> 0xeaf00.916matched to luaT_getvarargs
luaD_pcall0xfb60 -> 0xf2f00.934matched to luaD_call
luaH_getn0x1d9e0 -> 0x1b2800.826matched to close_func
luaH_next0x1bda0 -> 0x19f500.959matched to luaG_errormsg
luaH_size0x1c580 -> 0x1a8f00.997not matched
luaG_errormsg0xd680 -> 0xd4600.975matched to luaG_addinfo
luaT_callTM0x1e140 -> 0x1b7100.965matched to gmatch
luaS_resize0x1af00 -> 0x192e00.903matched to luaD_checkminstack
luaO_pushfstring0x15990 -> 0x142900.995matched to luaK_concat

Two more graph views show the same pattern.

BinDiff matched luaD_pcall to luaD_call. DeepDiff matched luaD_pcall correctly at 0xfb60 -> 0xf2f0.

BinDiff graph view showing luaD_pcall matched to luaD_call

BinDiff also matched luaG_errormsg to luaG_addinfo. DeepDiff found the correct luaG_errormsg match at 0xd680 -> 0xd460.

BinDiff graph view showing luaG_errormsg matched to luaG_addinfo

BinDiff labels stripped functions with names such as sub_XXXX. The function names above come from the symbol-based ground truth used only for checking the results.

Finding 2: You Cannot Match a Function You Never Found

The Lua test was mainly about matching quality. The minigzip test exposed a different problem: function coverage.

BinDiff works from functions identified by IDA. In a stripped, statically linked binary, IDA may not identify every library function. This is common for linked functions that the main program does not call.

For minigzip64-1.2.11:

  • The original binary with symbols had 133 function symbols.
  • Versions 1.2.11 and 1.2.12 had 130 functions in common.
  • IDA identified 138 function nodes in the stripped binary.
  • Only 76 of the 130 expected functions were present in IDA's function list.

The total count of IDA functions was not the important number. IDA also found functions outside our symbol-based ground truth. The key result was that 54 of the expected functions were missing from the input given to BinDiff.

For minigzip64-1.2.11 -> minigzip64-1.2.12, the results were:

ToolCorrect matchesWrong matchesPrecisionRecall
DeepDiff12190.9310.931
BinDiff7601.0000.585

BinDiff was correct for every function it matched. This is a strong result. But it found only 76 of the 130 expected matches because the other functions were not available to it.

DeepDiff found 47 of the 54 functions that BinDiff missed. Examples include:

FunctionSide 1 addressDeepDiff matched addressDeepDiff similarityBinDiff
gzopen0x22000x32101.000missed
gzbuffer0x22c00x32d01.000missed
gzseek640x23e00x33f01.000missed
deflateBound0x6da00x7d101.000missed
deflateSetHeader0x6ae00x7a201.000missed
inflateCopy0xc4a00xd4601.000missed
zlibCompileFlags0x101000x111101.000missed
zError0x101100x111200.985missed
adler32_combine0x108300x118401.000missed
main.cold0x143a0x243a1.000missed

Many of these functions were almost identical in the two zlib versions. The difficult part was not deciding whether the functions matched. The difficult part was finding the function boundaries in the first place.

gzopen is a simple example:

  1. The symbol table placed gzopen at 0x2200.
  2. IDA did not create a function at 0x2200.
  3. DeepDiff still found its match at 0x3210 in the next version.

This coverage problem explains most of the recall difference in this minigzip pair. It is separate from the GCC-to-Clang result, where both tools had the functions but differed in how well they matched them.

Finding 3: The Name Can Stay While the Code Moves

Most minigzip version pairs were close. The largest change appeared in zlib 1.3.2.

In version 1.3.2, several public API functions became small tail-call wrappers. Their main code moved into another function, often one with a *64, _z, or _gen64 name.

The symbol sizes make the change easy to see:

Function1.3.1 size1.3.2 sizeChange
gzseek4979became a wrapper
gzseek64497520now contains the main code
deflateBound4729became a wrapper
deflateBound_z-647now contains the main code
crc32_combine2479became a wrapper
crc32_combine_gen64167192contains shared code
gzgetc_1569became a wrapper
gzgetc156188now contains the main code

This creates a problem for any result checked only by symbol name. The old symbol can still exist, but it no longer contains the old code. A useful binary match should help the analyst follow the code to its new location, even when the name changes.

For older zlib versions compared with 1.3.2, the results were:

ToolCorrect matchesWrong matchesPrecisionRecall
DeepDiffabout 110-112about 11-200.85-0.91about 0.86
BinDiffabout 65-71about 6-110.855-0.922about 0.50

For 1.3.1 -> 1.3.2, the strict symbol-name check counted 11 wrong matches for each tool. The number was the same, but the errors were not the same kind:

ToolName-based wrong matchesMean similarityMatches with similarity >= 0.90
DeepDiff110.97811 / 11
BinDiff110.5371 / 11

Nine of DeepDiff's 11 name-based errors followed code to its new function:

  • gzseek in 1.3.1 matched gzseek64 in 1.3.2.
  • deflateBound in 1.3.1 matched deflateBound_z in 1.3.2.
  • crc32_combine in 1.3.1 matched crc32_combine_gen64 in 1.3.2.

The strict check marked these as wrong because the names differed. For an analyst, these are useful matches: they show where the code moved.

The other two DeepDiff matches were clear errors:

  • adler32_combine64 -> adler32_combine
  • deflatePending -> deflateUsed

BinDiff's wrong matches were mostly low-similarity matches between unrelated functions:

Function in 1.3.1BinDiff match in 1.3.2Similarity
gzputcgzgetc0.739
inflateInit2_gz_look0.418
gz_lookgz_error0.352
gzprintfgzvprintf0.289
gz_errordeflateReset0.171
gzgetc_gzprintf0.067

The zlib 1.3.2 case shows why binary matching is more than checking whether two symbols have the same name. Sometimes the most useful result is the function with a different name but the same code.

What This Means for Binary Analysis

The benchmark shows where DeepDiff's matching layer is most useful.

When the compiler changes, DeepDiff can still match functions even when their control-flow graphs have very different sizes and shapes.

When a stripped static binary contains functions that IDA did not identify, DeepDiff can find many of those function boundaries and include them in the comparison.

When an update moves a function's main code behind a wrapper, DeepDiff can follow that code to a new symbol instead of stopping at the old name.

Release-to-release comparison is one use of this alignment, but it does not define the product. The same matching layer can support any analysis that needs to connect related code across two binaries. The analyst's question decides how those matches are used.

What This Test Does Not Show

There are a few limits worth stating clearly.

  • This is not a test of every binary type. The test set contains Lua and minigzip binaries. Other programs, compilers, platforms, and forms of obfuscation may produce different results.
  • The symbol check is strict. It treats different symbol names as a wrong match, even when code moved and the result is useful. We reviewed those cases separately instead of changing the scoring rule after the test.
  • The minigzip gap includes function coverage. Much of DeepDiff's recall gain came from finding functions that were missing from IDA's function list. This is different from the cross-compiler Lua result, which tested matching quality.
  • BinDiff is strong when its input is complete and the binaries are close. It reached 1.000 precision on the detected functions in unchanged zlib pairs. Our result is not that BinDiff performs poorly in every setting.

The narrower result is more useful:

In these tests, DeepDiff was more reliable when the binaries were built with different compilers, when stripped static code was missing from IDA's function list, and when an update moved code between functions.

These cases are common in real binary analysis. The results give measured evidence for DeepDiff's core purpose: finding related code across binaries even when compiler choices, missing function boundaries, or code movement make the match difficult.

Found this useful? Share it.