UnsafeChecker: A Bug Detector for Rust's Safe Abstractions
A compiler-integrated analyzer reconstructs the ownership, lifecycle, and layout facts that raw pointers erase — and finds 114 real soundness bugs reachable from safe Rust.
Xizhe Yin, Yaokun Zhang, Yang Feng, Baowen Xu · 2026 · arXiv preprint | arXiv:2609.09641 · PDF
Rust makes a strong promise: a program that compiles without ever writing the word unsafe will not corrupt memory. However, that promise leans on the libraries underneath it — and those libraries are full of unsafe code. When one of them gets a single detail wrong, an ordinary program that never wrote a line of unsafe can still read freed memory, run off the end of a buffer, or be exploited. A recent paper from Nanjing University — Xizhe Yin, Yaokun Zhang, Yang Feng, and Baowen Xu — builds a tool called UnsafeChecker to hunt for exactly these hidden mistakes, and reports 114 previously unknown ones in real, widely used crates. In this post we walk through what it does, how it works, and — just as carefully — where it falls short.
In one sentence
UnsafeChecker is a compiler-integrated static analyzer that reconstructs the three kinds of information Rust’s raw pointers throw away — who owns a value, whether it is still alive, and its physical layout (where it sits and how big it is) — carries all three in one shared, path-by-path abstract state over the compiler’s own intermediate representation, and uses them to flag unsafe library code that would let ordinary safe code trigger undefined behavior.
How safe code can still break
Start with the puzzle in the title. If Rust is safe, how can safe code be buggy?
Rust keeps memory safe through ownership and borrowing — compile-time rules about who is responsible for each value and who may look at it — with no garbage collector. Every value has exactly one owner; no reference outlives the data it points at; every access stays in bounds. However, for low-level work — talking to hardware, squeezing out performance — those rules are too strict, so Rust provides an escape hatch: the unsafe keyword, which relaxes the compiler’s checks and permits raw pointers, values that are just bare memory addresses and carry none of the ownership, lifetime, or layout guarantees the rest of the language relies on.
Library authors do not want to push that danger onto their callers, so they hide it. They wrap the unsafe code behind an ordinary, safe-to-call API — a safe abstraction (think of Vec, Rust’s growable array, or Mutex). The arrangement is a contract: inside the wrapper, the unsafe code must uphold by hand the very invariants the compiler would otherwise enforce. Keep the contract and callers are safe. Break it and the abstraction is unsound — meaning safe callers who never wrote a line of unsafe can nonetheless trigger undefined behavior (UB). The Rust Reference states this literally: unsafe code that can be misused by safe code to exhibit UB is unsound. And because the compiler never validates the inside of an unsafe block, a single implementation slip can compromise an entire library.
Unsoundness comes in two flavors, and keeping them apart matters for the rest of the story:
- Direct UB: the
unsafecode itself executes an illegal operation on safe inputs — an out-of-bounds read, a dereference of freed or uninitialized memory. - Leaked invalid state: the API returns or stores something broken — a dangling pointer, or two owners of the same resource — and the UB happens later, in the caller, even though the function itself never crashed.
The core idea. Soundness rests on invariants across three interconnected dimensions — ownership, lifecycle, and layout — and those are exactly the dimensions raw pointers erase. This is also where existing tools fall short: analyzers built for C and C++ work on a flat memory model and ignore Rust’s ownership contract entirely, while current Rust tools mostly pattern-match on risky unsafe constructs without a semantic model of what those constructs do to ownership, liveness, or layout. Ordinary dataflow analysis treats a pointer as a plain value: it can say which address a pointer holds, but not that a ptr::read just created a second owner of a heap buffer, or that a map is handing back a reference into memory it already freed. UnsafeChecker’s move is to reconstruct those erased facts as an abstract state it propagates along the program, and to let each safety check pull out only the slice of facts its corresponding Rust obligation needs. The three dimensions are not independent — dropping an owner invalidates every alias of the underlying resource — so they live in one shared state rather than three separate analyses.
An analogy that maps onto the mechanism. Picture an automated warehouse. Every item on the shelves carries three facts in a manifest:
- who holds the claim ticket — ownership, and there is supposed to be exactly one ticket per item;
- its current status stamp — on the shelf, an empty slot not yet stocked, shipped out, incinerated — this is validity, or lifecycle;
- its physical footprint — which shelf address it starts at, how many slots wide it is, and which slot a worker is reaching for — this is layout.
Normally Rust’s compiler maintains this manifest flawlessly. But a raw pointer is a bare shelf coordinate scribbled on a sticky note — “aisle 5, shelf 3” — with none of the three facts attached. The instant code goes unsafe, it is working from sticky notes. UnsafeChecker is the auditor who walks the floor and rebuilds the manifest from those bare coordinates: for every pointer it re-derives the owner, the status, and the footprint, then raises a flag whenever an operation would break a rule — reaching past an item’s footprint, reading something already incinerated, or photocopying a claim ticket so that two parts of the program each believe they must dispose of the same item. The mapping is exact: claim ticket → ownership, status stamp → lifecycle, footprint → layout, sticky note → raw pointer, auditor → the tool.
unsafe behind a normal API. The interior must uphold Rust's invariants by hand; if it doesn't, safe callers who never wrote unsafe can still trigger undefined behavior. UnsafeChecker guards this boundary.How UnsafeChecker works
UnsafeChecker is a compiler plugin, not a source-text scanner. It is built on Rust’s own compiler libraries (rustc_middle and rustc_mir_dataflow), so it sees code the way the compiler does. Its symbolic memory model draws on MIRAI (an abstract interpreter for Rust); it uses the Apron library for the numerical abstract domains behind the arithmetic and the Z3 solver for satisfiability-modulo-theories (SMT) constraint solving. The pipeline runs in five stages.
1. Lower to MIR, then sort every location into three kinds
The tool works on MIR (Mid-level Intermediate Representation), the desugared, simplified form the compiler already produces on its way to machine code, where tangled source becomes a flat list of small, explicit steps. In MIR a place is a path to a memory location — a local x, a field x.f, a dereference *p. UnsafeChecker sorts these places into three disjoint kinds of abstract location, because it treats each differently:
- Stack owners: non-
Copyroots of ownership, such as aVecorBox, which live on the stack but may manage heap resources. - Pointer locations: references and raw pointers (
&T,*mut T). Raw pointers are technicallyCopy, but the tool deliberately keeps them here so it can track aliasing precisely. - Copyable data: other
Copylocals such asusizeandbool, whose concrete values feed the layout domain’s arithmetic.
Two more modeling choices keep the analysis both precise and finite. First, it is field-sensitive: x.f is a distinct abstract location from its parent x, so partial ownership and independent borrows of separate fields are tracked. Dereferences *p are not given their own location; they are resolved through the points-to relation to whatever p may target. Second, dynamic allocations use allocation-site abstraction: every allocation produced by the same MIR instruction collapses to one canonical heap object, which is what keeps the set of tracked objects finite. The trackable resources are then the stack owners together with the heap objects.
2. Carry one shared state with three domains
The state propagated through the program is a triple — an ownership domain, an object-validity (lifecycle) domain, and a layout domain. Crucially it is one bundle: the transfer rules update only the parts they touch, and the checks read only the parts they need, but a single object or pointer identity is shared across all three. That sharing is what lets a composite bug — a drop in the lifecycle domain invalidating an alias in the ownership domain — be caught at all.
Ownership domain. This reconstructs the heap’s shape as three relations, ⟨Pt, Own, Alias⟩:
- Points-to (
Pt) — a may-points-to map from each pointer to the set of objects it might reference. Dereferencingpis treated as touching every objectpmay point at; at branch joins the sets are unioned. - Ownership hierarchy (
Own) — the vertical structure: which heap resources an object is currently responsible for (aVecowning its buffer). This is what tells the analysis which resources die when an object is dropped. - Ownership aliasing (
Alias) — the horizontal relation: which other objects claim the same underlying resource. Safe Rust guarantees each resource has exactly one owner (|Alias(o)| = 1); anunsafeoperation such asptr::readcan grow that set past one, which is the structural precursor to a double-free. The join is set union, so if a duplication exists on any incoming path it survives to be checked later.
Lifecycle domain (object validity). This tracks each object’s runtime status. (“Lifecycle” here means runtime object state — deliberately distinct from Rust’s compile-time lifetimes.) The states are a small fixed set: Live (initialized and valid), Uninit (allocated but not yet initialized; reading it is UB), Moved (ownership transferred out; later access is use-after-move), Dropped (destructor has run; the pointer now dangles, and use is use-after-free), ManuallyDropped (still valid storage but with automatic destruction suppressed, e.g. behind ManuallyDrop), and Forgotten (consumed by mem::forget — semantically moved, resources leaked rather than freed). The clever part is the shape of the domain: it is a power-set lattice, so an object can hold a set of possible states and the join is set union. If a value is Live on one branch and Moved on another, its merged state is exactly {Live, Moved} — not a vague collapsed “unknown.” That precision is what lets the tool report a conditional finding like “potential use-after-free on some paths” instead of drowning in false alarms.
Layout domain. This gives each pointer a byte-level tuple:
(base, size, offset, align, null, mut)
base is the target memory block, size is the minimum guaranteed buffer capacity, align is the alignment guarantee, and null/mut are the nullability and mutability capabilities. The key element is that offset is an interval [l, u], not a single number — so the tool can reason about “somewhere between byte 8 and byte 56” when an index is not statically known. Everything is tracked in bytes, so pointer casts (reinterpreting *mut u64 as *mut u8) work.
The ordering is designed so that larger elements are weaker guarantees. At a branch join the tool keeps only what holds on all incoming paths: size becomes the minimum of the two, offset intervals are unioned, and capabilities downgrade (to immutable, or to maybe-null) unless they are guaranteed everywhere; there is an explicit invalid-pointer bottom element for a pointer known to be unusable. One subtlety the authors are careful about: size is allocated capacity, not initialized length — so reading uninitialized bytes that are still inside bounds is fine as far as layout is concerned, yet is UB, and that case is caught separately by the lifecycle domain’s Uninit state. The two domains divide the labor cleanly.
3. Propagate the state to a fixpoint
The analysis is a worklist fixpoint over each function’s control-flow graph. At every instruction it applies a transfer function that updates the shared state; where branches rejoin it applies the join — the conservative over-approximation above — so it never quietly assumes a path is safe. The transfer rules are written as inference rules; a representative set:
-
T-Alloccreates a fresh heap object and initializes all three domains at once (lifecycleLive, ownership edges, points-to, and a layout tuple marked mutable). -
T-Movetransfers the ownership topology and marks the sourceMoved. -
T-ReadBitsmodelsptr::read: it creates a new owner and merges alias sets, which is precisely how ownership duplication gets recorded. -
T-Drop,T-Forget, andT-ManDropupdate the lifecycle state and fire the matching topology changes (note thatManuallyDroppedis treated as still valid, so access through the wrapper is not reported). -
T-PtrWritemodelsptr::write(overwrite without dropping the old value);T-AddrOf,T-Cast, andT-UsePtrhandle address-taking, casts (which preserve points-to and layout except mutability), and plain pointer propagation. -
T-Offsetmodels pointer arithmeticp.offset(i): it advances the offset interval byi · sizeof(T)and, tellingly, sets the new alignment togcd(align, sizeof(T))— a conservative lower bound on the alignment that survives the step.
One refinement matters for precision: a write to a single, unambiguous location gets a strong update that replaces the old fact, while a write through a merged or allocation-site target — where the tool cannot be sure which concrete object is meant — gets a weaker update that can only add possibilities. That is the precision price of having collapsed many allocations into one canonical object.
These rules do their bookkeeping through a handful of auxiliary operations that keep the ownership graph and lifecycle consistent:
- Transfer — move semantics: structurally replace the old owner with the new one everywhere.
- AliasMerge — fuse two ownership cliques, as
ptr::readdoes. - Isolate — cut a variable out of all alias sets before it is overwritten.
- RecursiveDrop — cascade the
Droppedstate to an object and all its aliases; this is what makes a drop through one alias detectable as a use-after-free through another. - Clear — freshen a variable’s identity before reassignment.
- Consume — move a variable and its fields to
Moved. - Forget — push
Forgottento a variable and all its fields, so that even an&x.ftaken earlier is invalidated.
Why it terminates. The ownership and lifecycle domains range over finite sets — the locations in a function body are finite, and heap objects are finite thanks to allocation-site abstraction — and the transfer functions are monotone, so those chains stabilize. The layout domain’s interval component, though, has infinite height: a loop could keep growing an offset forever. So the tool applies widening at loop headers when an interval grows, with optional narrowing on simple loop guards (like i < len) to win back some precision. That combination guarantees the analysis finishes.
4. Set up realistic entry conditions, and handle calls
To analyze a library without ever seeing its callers, UnsafeChecker builds a symbolic initial state at each public-API entry point that encodes Rust’s application-binary-interface (ABI) guarantees: mutable references do not alias, and owned objects form disjoint trees (each ownership node starts as a singleton alias set); each pointer argument points to a fresh symbolic object; arguments start Live while locals start Uninit, to catch use-before-initialization; and a container argument’s internal buffer is given a symbolic capacity and the element type’s alignment. That symbolic capacity is what lets the tool later check a constraint like “index < capacity” without ever knowing the concrete size.
Calls to internal functions are handled by full inlining — a recursive, depth-first, context-sensitive traversal that initializes the callee from the caller’s current state, runs it to a fixpoint, and maps the return value’s abstract state back. This buys full context sensitivity without pre-building an inter-procedural graph. To stay terminating and tractable it imposes an inline depth limit, default k = 3; when the limit is hit, or a callee’s body is unavailable — a foreign-function-interface call (FFI) into C, say — it substitutes a conservative summary that may mutate any reachable state of the mutable arguments and returns an unconstrained value. This is a deliberate trade: pessimism here guards against false confidence, but it is also a source of missed bugs, as the evaluation shows.
5. Check at two granularities
This two-level checking is what separates UnsafeChecker from a pattern-matcher.
Instruction-level checks fire during propagation, as a dangerous operation is about to execute, reading the relevant domains at that moment:
-
E-Access— use-after-free, use-of-uninitialized, or use-after-move, keyed off the lifecycle state and a sharedInvalidSet = {Dropped, Moved, Forgotten, Uninit}. -
E-DoubleFree— a second drop of a resource whose owners were duplicated. -
E-RefCreation— a reference must be built from a valid, aligned, non-null pointer. -
E-OOB/MA— out-of-bounds or misaligned access, keyed off the layout tuple: flag it if the offset can start before the allocation or ifu + sizeof(T) > size, or ifalign < align_of(T). -
E-NullDeref,E-SizeMismatch,E-Proj(field projection in bounds),E-PtrArith(pointer arithmetic in bounds),E-SliceInvalid(e.g. aslice::from_rawlength exceeding the allocation), andE-MutateImmut(a write through an immutable pointer).
Boundary-level checks fire at function exit and look only at the values that escape to the caller — return values, out-parameters, exposed references. These catch the “leaked invalid state” flavor of unsoundness, where nothing looked wrong locally but the shipping dock still hands out something broken:
-
E-AliasRet— two returned values owning the same resource (double ownership). -
E-MutAlias— two returned references aliasing the same location, with at least one mutable. -
E-Dangling— an escaping pointer that targets an object in an invalid lifecycle state. -
E-Layout— an assembled container (such as one built withVec::from_raw_parts) whose pointer, length, and capacity do not satisfy the type’s invariants, including theisize::MAXoverflow limit. -
E-Leak— a heap object still live at exit but unreachable from the caller.
What is in scope, and what is not
This bounds every claim below, so it is worth stating plainly. UnsafeChecker targets the memory-safety UB categories of the Rust Reference — [pointer-access], [place-projection], [alias], [immutable], and the pointer- and allocation-related parts of [invalid] — plus resource leaks, plus three lightweight auxiliary checks used in the evaluation (integer overflow in index or size computations, division by zero, and invalid UTF-8 passed to unchecked string constructors such as str::from_utf8_unchecked). It explicitly excludes [race], [intrinsic], [target-feature], [call], [asm], and [runtime]. The authors are careful to frame the result as scoped bug-finding, not a theorem of full Rust soundness: the claim is limited to the MIR operations, library and API models, and memory-safety categories actually modeled, and it does not cover concurrency, opaque callees, unmodeled library behavior, the inline-depth cutoff, FFI, or pointer→integer→pointer round-trips where object identity is lost.
A worked example
Here is the whole problem in about twenty lines — this exact snippet compiles cleanly under rustc (edition 2021). Imagine a tiny Buffer type that wraps a data: Vec<u8> of four bytes and offers a get(i) method:
/// A tiny "safe abstraction": callers never write `unsafe`.
pub struct Buffer {
data: Vec<u8>,
}
impl Buffer {
pub fn new() -> Self {
Buffer { data: vec![0u8; 4] }
}
/// SAFE signature, but the body forgot the bounds check.
/// For `i >= 4` this reads past the allocation: undefined behavior,
/// reachable from 100% safe code.
pub fn get(&self, i: usize) -> u8 {
unsafe { *self.data.as_ptr().add(i) }
}
}
fn main() {
let b = Buffer::new();
// No `unsafe` here -- yet this call is unsound.
let _leaked = b.get(999);
}
Look at the signature: get is a completely ordinary, safe method — no caller ever types unsafe. But the body reaches into an unsafe block and offsets a raw pointer by i without checking that i is in range. Call buffer.get(999) on a four-byte buffer from 100% safe code and the read lands far past the allocation — undefined behavior. And the counterintuitive part: this compiles, because the danger is sealed inside unsafe.
Trace it through the layout domain the way UnsafeChecker would:
-
Buffer::newallocatesdata: Vec<u8>of length 4. Its internal buffer pointer gets the layout tuple(base = heap object, size = 4, offset = [0,0], align = 1, non-null, mutable). - In
get,self.data.as_ptr()produces a*const u8into that buffer:(base, size = 4, offset = [0,0], align = 1, non-null, immutable). -
iis a public-API argument, so it enters symbolically as the fullusizerange..add(i)triggersT-Offset, which advances the offset interval byi · sizeof(u8) = i · 1, giving offset[0, huge]. - The dereference is an access.
E-OOB/MAchecks whetheru + sizeof(T) > size— herehuge + 1 > 4— and fires an out-of-bounds alarm.
Notice that this bug needed only the layout domain — a domain-local check. The complementary case is the one raw pointers are infamous for, and it needs the domains to cooperate. It walks through in the same kind of steps:
- A
ptr::readon a non-Copyvalue runsT-ReadBits, which callsAliasMergeand grows the object’s alias set past one. Two owners now claim the same heap resource. - The first owner is dropped.
RecursiveDropmarks the resourceDroppedacross all its aliases — including the second owner. - The second owner is dropped in turn. Its target is already
Dropped, so this drop is a detectable double-free, andE-DoubleFreefires.
The first example is ownership-blind and lives entirely in layout; the second is ownership and lifecycle working together through one shared object. The two are exactly the two rule families — instruction-level checks and the multi-domain boundary checks — in miniature. The double-free case is worth drawing out, because it is where the “one shared state” design earns its keep:
ptr::read (the paper's motivating pattern). A Vec owns a heap buffer; ptr::read (rule T-ReadBits → AliasMerge) mints a second owner of the very same resource. Safe Rust keeps exactly one owner; now there are two, so when both are dropped the buffer is freed twice. Catching it needs the ownership and lifecycle domains together — not pattern matching.How it was evaluated, and what it found
The authors built two datasets and asked four questions: does it beat existing tools (effectiveness), what are its precision and recall trade-offs, is it fast enough, and can it find genuinely new bugs at scale. Everything ran on one machine — an Intel Core i7 with 64 GB of RAM on Ubuntu 22.04 LTS — with every baseline at its latest stable version and default configuration. Throughout, “detected” means the tool raised at least one true-positive alert matching the vulnerability, confirmed by manual audit.
The setup
Dataset A (known answers). 46 CVEs drawn from the RustSec Advisory Database — soundness issues caused by unsafe memory operations that fall inside the tool’s sequential scope — which together contain 53 distinct ground-truth bugs. Because the baseline tools often need older compilers that cannot build the original crates, the vulnerable code was extracted into minimized libraries that faithfully reproduce each bug, then verified by hand. This is the ground truth for recall and precision.
Dataset B (the wild). A scan of the crates.io registry. Exhaustively auditing every result is impossible, so the authors manually inspected the reports from roughly 10,000 crates, and — this is the strict part — counted a warning as a real bug only if they could write a proof-of-concept that actually triggered UB in Miri (Rust’s interpreter for detecting undefined behavior) through the crate’s safe public API. That yielded the 83 crates that make up Dataset B.
Baselines. Three tools spanning the main Rust-analysis paradigms: MirChecker (MIR-level numerical and symbolic analysis for generic errors like integer overflow, but no ownership tracking), Rudra (a pattern-matcher that flags risky unsafe constructs but has no semantic model), and SafeDrop (deallocation-focused — good at double-free and use-after-free, but blind to alignment, out-of-bounds, and nested ownership).
RQ1 — Effectiveness
On Dataset A, UnsafeChecker detected 32 of 46 CVEs (69.6% CVE-level recall) and covered 36 of 53 bugs (67.9% bug-level recall). These two recall figures are genuinely different quantities — one counts vulnerabilities, the other counts individual bug patterns — and are worth keeping apart. The comparison with the baselines is lopsided:
| Tool | Detected (Uninit/OOB/Own) | Bugs /53 | CVEs /46 | Bug recall | TP alerts | Total alerts | Precision |
|---|---|---|---|---|---|---|---|
| MirChecker | 0 / 0 / 0 | 0 | 0 | 0.0% | 0 | 45 | 0.0% |
| SafeDrop | 0 / 0 / 1 | 1 | 1 | 1.9% | 1 | 4 | 25.0% |
| Rudra | 11 / 0 / 0 | 11 | 11 | 20.9% | 11 | 11 | 100.0% |
| UnsafeChecker | 15 / 19 / 2 | 36 | 32 | 67.9% | 64 | 124 | 51.6% |
Laid side by side, the gap is hard to miss: on both CVEs detected and bugs covered, UnsafeChecker stands well above all three baselines, while MirChecker sits flat at zero. The best baseline, Rudra, caught 11 CVEs — all of a single kind, uninitialized memory — which is exactly why its precision is a perfect 100% and its recall is low: it fires only on the patterns it is certain about. UnsafeChecker found 21 more CVEs than Rudra, and more than 20 vulnerabilities that every baseline missed.
For reference, the 53 ground-truth bugs break down as 11 uninitialized-memory, 13 length- or state-consistency, 25 out-of-bounds or pointer-boundary, 1 use-after-free, 1 double-ownership, and 2 dangling-pointer. These roll up into the table’s three columns as follows: the 11 + 13 become Uninit (24), the 25 become OOB (25), and the 1 + 1 + 2 become Own (4). Cases that only UnsafeChecker caught include RUSTSEC-2023-0078 (the tracing crate: mem::forget(self) followed by reads of field pointers, which trips a Forgotten-state alarm), RUSTSEC-2018-0022 (std::mem::uninitialized used to seed entropy — an Uninit read), and RUSTSEC-2025-0072 (wrflib: an unchecked ptr.add(offset), caught as out-of-bounds by the layout domain).
RQ2 — Precision, and the unflattering half
That recall is not free. Of UnsafeChecker’s 124 alerts, 64 were true positives — 51.6% alert-level precision. Roughly one flag in two is a real bug; the rest are false alarms a human must triage. It also missed 14 CVEs (30.4%) and 17 bugs (32.1%), about a third of the benchmark.
The 60 false positives break down by domain: layout, 51.7% (31/60) — the single largest source, driven by the conservative lower-bound layout facts and the interval and widening arithmetic; lifecycle, 35.0% (21/60) — from control-flow merges and opaque calls; and ownership, 13.3% (8/60) — from path-insensitive aliasing. As shares of the 60, layout alone is more than half, lifecycle roughly a third, and ownership the small remainder. The 14 misses split into 7 that need unmodeled library or API semantics (e.g. Cell, alignment-related allocation APIs, debug_assert; RUSTSEC-2023-0017 needs Layout::from_size_align, which is not modeled), 3 that need finer numeric, loop, or iterator reasoning, 2 that need per-element array or init-state tracking, and 2 that depend on external, cross-boundary logic.
For scale context — and this is explicitly not a closed-world precision figure — on the ~10,000 inspected crates the tool flagged 378 crates and produced 2,639 raw warnings, de-duplicated to 1,736 unique, which triage narrowed to the 114 Miri-validated bugs. On the subset Rudra could compile (2,892 crates, where Rudra flagged 119 with 326 warnings), UnsafeChecker flagged 141 crates with 1,296 raw, 784 unique warnings. These are workload and discovery-yield numbers, not “X% of warnings were bugs.”
RQ3 — Efficiency
Across a 9-crate efficiency study the median end-to-end time was 11.4 s and the median memory 475.2 MB; the largest case (meilisearch, 244.6k LOC) took 264.0 s and peaked at about 4.3 GB. Small crates (<10k LOC) finish within roughly 10 seconds and large ones (>100k LOC) within 5 minutes — and time does not track LOC strictly (diesel, at 126.8k LOC, took 113.5 s, while nushell, at 287.6k LOC, took only 94.6 s), because the cost is driven by unsafe complexity, not raw size. On Dataset A the whole run — compile plus analyze — took 35.7 s, about 0.78 s per case, versus Rudra’s 179.7 s (5× slower), MirChecker’s 781.9 s (more than 20× slower), and SafeDrop’s 36.1 s (similar speed, but it made only one detection). At ecosystem scale the tool analyzed more than 100,000 crates in two days. Roughly speaking, that is fast enough for routine, per-crate use.
RQ4 — Discovery in the wild
This is the result that lands. UnsafeChecker uncovered 114 previously unknown, Miri-validated bugs across 83 crates. Maintainers and RustSec experts acknowledged 45, 27 have already been fixed, and 11 received public vulnerability identifiers. The affected crates have more than 30 million downloads combined and span 34 categories — most often data-structures, no-std, and memory-management libraries, i.e. foundational, heavily-depended-on code.
By bug type (detected / confirmed / fixed): out-of-bounds 60/20/13, ownership-duplication 17/10/6, misaligned-access 11/4/0, integer-overflow 3/2/1, invalid-pointer 3/2/1, dangling-pointer 7/2/1, division-by-zero 1/1/1, invalid-UTF-8 1/1/1, resource-leak 5/1/1, null-pointer 2/1/1, uninitialized-memory 1/1/1, and other 3/0/0 — totaling 114/45/27. Layout-related bugs dominate in the wild: out-of-bounds alone is 52.6% (60/114) and misaligned access another 9.6% (11/114).
This is the same layout reasoning that produced the most false alarms in RQ2, and the coincidence is worth stating plainly: the conservative interval arithmetic that finds the most real out-of-bounds bugs is also the single largest generator of noise. The tool’s biggest strength and its biggest cost share one root.
One more finding cuts against a common assumption. The notorious ptr::read and mem::forget patterns account for only 5 and 1 of the 53 Dataset-A bugs, and 19 and 4 of the 114 wild bugs; the validated wild bugs in fact touch 129 unsafe-API call sites across 21 distinct APIs. In other words, a tool that only pattern-matches those two calls would miss most of the real problem.
Four representative cases give the texture:
- a dangling-pointer/double-free in
multiversx_chain_vm(132k downloads) — aptr::readbehind a&mut Tcombined with panic unwinding; - a resource leak in
bytes-kman(Box::leakplusptr::read, caught byE-Leak); - a double-free in
doubly’spop_back(ptr::readduplication followed by aBoxdrop); - an out-of-bounds in
auto_vec((&children[0]).sub(1)moving a pointer before the start of the allocation, caught byE-OOB/MA).
From raw signal to landed fix, the funnel narrows steeply: 1,736 unique warnings collapse to 114 validated bugs — roughly 15 to 1 — and from there to 45 confirmed, 27 fixed, and 11 with public identifiers.
Questions worth asking
If Rust is safe, how can safe code be buggy? Because “safe” means one’s own code delegates the dangerous parts to libraries that use unsafe internally. That safety is only as strong as those libraries’ contracts. When a library’s unsafe code is wrong, the breakage surfaces in ordinary unsafe-free code that calls it.
What is a “safe abstraction,” concretely? Any type whose public API can be used without writing unsafe, but which uses unsafe inside. Vec and Mutex are the canonical examples. The entire ecosystem is built on the assumption that these wrappers keep their promises.
Why three domains instead of one general analysis? Because the three invariants raw pointers erase are qualitatively different — who owns a value, whether it is alive, where it physically sits — and different bugs need different subsets: an out-of-bounds read is pure layout, a double-free is ownership plus lifecycle. Keeping them in one shared state, rather than three separate tools, is what lets a single event like a drop ripple correctly across all of them.
Why inline calls, and why only to depth 3? Full inlining gives full context sensitivity — the callee is analyzed in the exact state of its caller — which matters because unsafe code tends to sit in shallow, statically-dispatched call chains. The depth-3 cap (and the pessimistic summary beyond it, or for FFI) is the price of guaranteed termination; it is also a known source of missed bugs.
Where do the false positives come from? Mostly from layout reasoning. Of the 60 false alarms on the benchmark, 51.7% traced to the layout domain’s conservative, range-based arithmetic, 35.0% to lifecycle reasoning at branch merges and opaque calls, and 13.3% to ownership aliasing.
How was a bug confirmed in the wild, so that these are not just more warnings? By construction. A wild warning counted only if the authors could write a safe-API proof-of-concept that made Miri — Rust’s UB-detecting interpreter — actually report undefined behavior. That is a much higher bar than “the analyzer flagged it.”
Is this a tool for everyday development? It is fast enough to be (median ~11 s per crate), but it is really aimed at the authors of unsafe-heavy libraries and at auditors. For most application developers the benefit is indirect: the 27 fixes it has already prompted make the crates they depend on a little safer.
Limitations and threats to validity
- It is a bug-finder, not a soundness proof. A clean run does not certify a library as sound; it means only that these particular checks found nothing, within the modeled scope.
- About half its alerts are false alarms (51.6% precision), so triage remains a human job, and the layout domain’s conservative pointer-range arithmetic is the dominant noise source.
- It misses roughly a third of bugs (17 of 53 on the benchmark). The blind spots are library APIs it does not model, FFI and deeply recursive code it approximates pessimistically (the depth-3 inline cutoff), per-element reasoning inside arrays, and non-linear pointer math — bitwise tricks an interval cannot capture precisely.
- Sequential memory safety only. There is no concurrency: it cannot detect data races or improper
Sync/Senduse, and it does not target application-level logic bugs. - It models key types, not the whole language.
Vec,Box, and raw-pointer operations are covered; the full standard library, specialized APIs such as custom allocators, and deep interior mutability (RefCell) are not. - Threats to validity. Internally, accuracy depends on the ground truth — Dataset A uses established CVEs; Dataset B has no ground truth, mitigated by the strict Miri-PoC criterion. Externally, Dataset A’s strict inclusion criteria give an imbalanced bug-type mix that may not match the natural wild distribution, and results may not generalize to non-public codebases. And the wild numbers (378 flagged, 2,639 raw, 1,736 unique) are triage workload and open-world discovery yield, not a closed-world precision measurement — the remaining warnings were never exhaustively labeled.
Open questions
These are prompts for reflection, not findings from the paper.
- Could the layout domain trade some conservatism for precision — shrinking that ~52% false-positive share — without letting real out-of-bounds bugs slip through?
- The 114 wild bugs came from a ~10,000-crate hand-inspected sample. What is sitting in the other ~90,000 crates the tool scanned but nobody checked?
- How much of the “half the alerts are noise” problem is fundamental to over-approximation, and how much would simply disappear with better models of common library APIs — given that 7 of the 14 benchmark misses were exactly that?
- Could the same three-domain, boundary-aware idea extend to the parts deliberately left out — concurrency and
Sync/Sendsoundness — or does that need a fundamentally different analysis? - If a tool like this became a routine step in publishing an
unsafe-heavy crate, how many of these bugs would never ship in the first place?
This write-up was produced by a multi-agent pipeline I built with Claude: a paper is drafted, independently fact-checked, edited for writing, and rewritten to match my own voice — and then reviewed by me before publishing. The plots are my own, generated from the paper's reported numbers (not copied figures). Every number and claim is traced back to the source paper. Read the original: arXiv:2609.09641.