changliu8541/Assemblage_LinuxELF
Assemblage Linux Dataset Please note, the Assemblage code is published under the MIT license, while the dataset specify each binary's source code repository license, please obey the original repository's license. This reposotory holds the Linux public dataset for Assemblage, and you can find the paper here. We also provide all the licensed source code (w/ all Git info) for these binaris available upon request, as the compressed file is too large (~1.5T after at highest compression level)… See the full description on the dataset page: https://huggingface.co/datasets/changliu8541/Assemblage_LinuxELF.
Assemblage Linux Dataset Please note, the Assemblage code is published under the MIT license, while the dataset specify each binary's source code repository license, please obey the original repository's license.
This reposotory holds the Linux public dataset for Assemblage, and you can find the paper here.
We also provide all the licensed source code (w/ all Git info) for these binaris available upon request, as the compressed file is too large (~1.5T after at highest compression level), please contact us and we can setup sftp or some netdisk for data transfer.
Files
linux_licensed.duckdb.zst— DuckDB database (zstd-compressed). Decompress withzstd -d linux_licensed.duckdb.zst→ 127 GiB DuckDB file. Open with the duckdb Python/CLI client. Schema:binaries,functions,rvas,lines,pdbs(same column layout as the previous SQLite release).binaries.tar.xz— raw ELF binaries.
Changelog
2024 June 28th: Update binaries on license info.
2025 Nov 13th: Added GCC O2 binaries for the Oz misflagged binaries. Please refer to the dataset docs.
2026 Apr 5th: Added new binaries with source codes.
2026 May 7th: Replaced linux_licensed.sqlite.tar.xz with linux_licensed.duckdb.zst. The DWARF extraction pipeline was rewritten end-to-end to fix several bugs in the previous SQLite release; the data was re-extracted from scratch on every binary in the corpus. Concrete fixes:
- Function names — preferred
DW_AT_linkage_name, fall back toDW_AT_name, then followDW_AT_specification/DW_AT_abstract_origin. The previous extractor stored 16-char-truncated names for ~5.9% of rows; now <0.7% are 16 chars (and those are real 16-char names). - RVA
end— derived fromDW_AT_high_pcform-aware (DW_FORM_addr→ absolute,DW_FORM_data*→ offset fromlow_pc);DW_AT_rangesproperly handled including DWARF v5BaseAddressEntry+ offset-pair lists. The previous extractor'sint(str(size), 16)bug inflated function sizes; ranges in v5 binaries with split text sections are now correct. - Line program — skip
state.line == 0synthetic entries; per-CU bisect-walk handles overlapping subprogram + inlined-subroutine ranges; line→function assignment is exhaustive. - Source-code text — populated from on-disk source, byte-exact with the source file at the recorded line. Path heuristics resolve
/tmp/projects/...and<32-hex>/...build-tmpdir prefixes. - Section pseudo-symbols (
.text,.bss, etc.) and zero-length RVAs are filtered. binaries.optimizationis the literal compiler flag (-O0..-O3,-Oz).
Re-extraction was validated bit-exact against an independent DWARF + source walk on a 160-binary stratified sample (zero discrepancies across ~280K functions, ~1.4M RVAs, ~2M lines, ~50K source-code text strings) plus full-corpus integrity checks (4.7B rows, zero PK/FK/range/null violations).
Row counts: 249,121 binaries · 613,573,055 functions · 685,044,264 rvas · 4,028,246,246 lines.
