Changelog

Included verbatim from CHANGELOG.md in the repository root. The format follows Keep a Changelog and the project aims to follow Semantic Versioning.

Changelog

All notable changes to asm-test are documented here. The format follows Keep a Changelog, and the project aims to follow Semantic Versioning.

Unreleased

Added

  • The Session strip desktop view — the whole session as per-thread lanes with run seams, a kernel rail, address bands, and run-length density, simplified to the top lanes by default (the rest counted, never vanished) — and its 3D companion, the session flow scene: the strip’s channels as smooth, depth-stacked ribbon rates over stream order, with seam walls at the strip’s run boundaries.

  • Desktop live sessions: live union weave (the 3D pane accumulates every capture the session makes, so a fresh Start adds to the scene) and stable plane layout (regions keep their first-seen slot), both ON by default; a GPU 3D rendering settings toggle (OFF takes the honest degraded 2D branch); and a two-step clear previous sessions affordance.

  • asmspy --serve emits the target’s address-space map as a vmmap event, re-sent only when the map changes; the desktop’s 3D plane names its regions from it (pick detail, region roster, provenance chip).

  • asmspy --dataflow --auto --sampler=ptrace — a third, perf-free sampler rung (ptrace residency probes → call-target expansion → int3 arrival confirmation), so --auto keeps working where perf_event_open is locked down entirely; the chain advances to it on a perf refusal.

  • The 3D plane’s simplified posture: the top-8 worldlines with the five per-event spike layers withheld and counted on a placard; the HUD’s detail toggle restores the full population. Plus occupancy and permissions plane layers (default off).

  • The live dataflow producer reads ymm16–31 (the EVEX registers glibc’s EVEX code paths keep vectors in), gated on XCR0 rather than CPUID alone.

  • asmspy --info <pid> — an attach-free process snapshot (identity, runtime, per-thread wchan + current syscall, symbol/JIT surface, and which tracing modes will work on the target). --json emits it as a one-event .asmtrace recording carrying the new procinfo kind.

  • Four standalone 3D scenes, on a substrate that is not the address plane (docs/internal/gui/59-standalone-scenes.md). The 3D pane gains a scene selector; every kind states what each of its axes is and what it is not, because the axis label is part of the host contract rather than each scene’s choice:

    • Divergence worldline — two recordings, one fused tube up their shared prefix, then A and B ribbons with a rib per step of architectural-state disagreement. A failed identity/basis/arch check renders a refusal card, not an empty scene; a truncated side caps the tube TORN (“we stopped looking” ≠ “they agreed to the end”); an uncomputed step is a hollow rib, never a zero-width one.

    • Invocation stack — one slab per invocation of a routine, stacked on a discrete invocation # axis that is never scrubbed as time. A block absent from an invocation is a hole in that slab, not a zero column; an unterminated invocation is a labelled prefix; a slab is tinted by the codeimage version in force when it started. Drill-in uses the dedicated dt_link::invocation field and leaves dt_link::step untouched.

    • Module excursion ribbon — one depth-vs-call-order sheet per thread, banded by module, with the boundary crossings brightened. A depth-capped capture marks its lane floor CAPPED (and the cap survives the drill-in); an unresolved module gets its own hue rather than being blanked or merged; a single-threaded recording renders the honest 2D icicle form instead.

    • SIMD lane prism — what happens inside one wide vector register across its writes, built from the already-decoded ValRec::wide/bytes. A stable value→hue hash makes a shuffle’s permutation visible. Element width is not recorded, so the default 16 byte-lanes is labelled as a default and only an unambiguous mnemonic subdivides; uncaptured bytes render a [wide] wireframe rather than zeros.

  • Inspect before you leap: hover readout and pickable overlays in the 3D scene (docs/internal/gui/47-scene-inspect-and-pickable-overlays.md). The 3D pane now answers “what is this, and where would a click send me?” before any navigation happens:

    • A throttled hover pick (SceneHost::pick, at most once per actual mouse-move pixel, zero readbacks during a drag/orbit/pan) resolves to a PickHint via the new resolve_pick_hint, which shares one classification helper with resolve_pick so the preview can never disagree with what a click does.

    • Convergence arcs and access spurs — previously undrawn in the pick pass — are now pickable via new id bands past the vertex band, decoded through an explicit PickBands so band sizes are never inferred; a convergence arc resolves to whichever tid is nearer the current playhead and states which side was chosen.

    • A hover tooltip shows identity/location/quantity/fidelity and the click destination (or “click → nothing here”); the HUD legend gains swatches for both overlay classes, a convergence layer toggle, and a persistent “hover to inspect, click to open the flat reader” hint.

  • Go there: camera pan, recentre, address/region goto, and discoverable controls in the 3D scene (docs/internal/gui/48-scene-navigation-and-goto.md). The plane is a reproducible address layout; now it is navigable:

    • Camera gains pan/frame plus four keyboard CamKey pan values, bound to middle-drag or shift+left-drag in the viewport (plain left-drag stays orbit).

    • Double-click recentres on whatever is under the cursor — a cell or a placed vertex’s address — without reorienting; a background/padding double-click is a stated no-op, not silent nothing.

    • A “go to” row in the HUD accepts a hex address or a region from Projection.regions; an address the recording does not map refuses with a stated reason rather than snapping to the nearest cell, and a region target frames its real cell footprint, never a bounding box. The global find bar offers the same resolver as “show in 3D” for any hit carrying an absolute address.

    • SceneView computes a stable “home” landmark once per weave — the code region’s centroid, not the plane’s fixed centre — so “reset view” returns to a place that does not drift as a live capture grows; “default view” keeps the literal fixed-centre preset as a separate button. A “you are here” readout names the region and address under the camera target.

    • The HUD’s collapsible “controls” block is generated from the CamKey enum so an unadvertised key fails a test, not a silent drift.

  • A GL-free flat terrain surface for the 3D pane (docs/internal/gui/ 52-flat-terrain-surface.md). The 3D overview’s spatial channel was unavailable wherever GL was absent (headless tests, the null backend, a driver whose shader will not build) and had no perspective-free reading mode even with GL, despite the pane’s own rule that precise reading is a 2D job:

    • views/scene2d.{h,cpp}’s cell_paint() mirrors kTerrainFrag’s branch chain exactly (kind hue via space::region_style() -> height mix -> churn -> stat -> unknown -> torn), cited back from embedded.h as the keep-in-sync C++ mirror so the two renderers cannot silently disagree about a cell’s fidelity state.

    • build_scene2d_plan() folds a terrain slice into down-sampled blocks (max height, OR’d flags — a torn or statistical cell is never dropped to binning, and the factor is stated on screen), breaks trajectory strips at unplaced vertices instead of interpolating across them, and dims — never drops — path points past the terrain-time playhead, the same rule 49’s worldline clipping uses.

    • The three previously GL-only degraded branches (no scene_host, not ready(), no frame texture) now draw the flat surface beside their existing placard; a new “flat surface” toggle swaps it in for the viewport on a working GL driver. Hover and click route through scene2d_pick_cell into the SAME scene3d::resolve_pick / resolve_pick_hint the 3D pick path uses, so a flat pick and a 3D pick of the same cell can never disagree — proven directly in the new test_scene2d.cpp.

    • Region-transition boundaries and labels, a compacted-domain-edge mark (padding is drawn as nothing, never a dark cell), and a scale strip (scene3d::height_scale_note) round out the surface as a real reading instrument, not just a fallback.

  • Launch a process and trace it from birth, and target a running one by window-pick (docs/internal/archive/gui/45-launch-and-window-target.md). Two new ways onto a live target, alongside attach-by-PID:

    • asmspy --serve’s wire protocol gains a launch command ({"cmd":"launch","mode":...,"argv":[...],"cwd":...}) that forks the target itself, PTRACE_TRACEMEs it, and hands it straight to the requested mode’s engine flagged as already-traced — the recorded session starts at the target’s true first instruction, with no detach/reattach gap. asmspy --launch <mode> -- <cmd> [args...] is the headless CLI equivalent. Scoped to the whole-process modes plus trace/watch for v1; dataflow/auto (a different, deeper seize subsystem) are refused with a stated reason rather than silently degraded.

    • The desktop Home rail gains a fifth entry, “Launch a new process”: a form (command/arguments/working directory/mode) landing on the same live-capture workflow every attach uses (LiveSession::send_launch, inspect_launch_full_detail, the Launch pane).

    • A crosshair affordance next to “Capture a live process”: drag it onto any window on screen (even outside the app’s own) to attach to that window’s owning process by PID, Spy++-style — X11-only (desktop/src/platform/window_picker.{h,cpp}, new), engine-free (D4), honestly unavailable under pure Wayland or over an ssh-remote capture target. libx11-dev + xvfb are pinned in Dockerfile.desktop, and make desktop-test-xvfb proves window resolution against a live (virtual) X11 display and a real second window, not a mock.

  • The faithful city, Phase A — the MVP terrain reskin (docs/internal/archive/gui/44-faithful-city-phase-a-mvp-terrain-reskin.md). The 3D spacetime overview now reads as a place, not an abstract density field: the terrain is zoned by memory-region kind (Scene::set_zoning, a new uKind R8UI texture blended into kTerrainFrag), an in-domain but never-touched cell reads as a sunken fog-of-war pit (TerrainFlag::TF_UNKNOWN), the recording’s dominant fidelity tier drives a damped weather sky (scene3d/atmosphere.h, byte-identical in colour source to the 2D fidelity banner), IBS survey residency renders as a physically separate, stippled “ghost district” terrain that never scrubs with the playhead (Scene::set_stat_terrain), and a followed “citizen” vehicle glyph with a comet tail rides the placed PC vertex at a new, independently-ticking follow_step playhead (SceneView::follow_play) that deliberately never fuses with the terrain-time or execution-step clocks. All five new SceneLayers toggles (zoning/weather/ghost_fog/vehicle, plus the pre-existing five) default on, so the “city” preset is the out-of-the-box view; no .asmtrace schema change.

  • Desktop GUI packaging: a Debian package and an AppImage for asmtest-desktop (packaging/debian-desktop/, packaging/appimage/). The full app links the GPL-2.0 Unicorn emulator and the Keystone/Capstone engines directly, none of which is a distro package on the pinned base images, so both artifacts vendor the three privately (rpath $ORIGIN) via a new scripts/package-native.sh --app mode; GLFW/GL/FreeType stay real host dependencies (declared Depends: for the .deb, checked at launch by packaging/appimage/AppRun for the AppImage — GL in particular must come from the host to match its driver). Both build straight from a checkout (no released tarball needed) and are verified end to end in CI via two new Docker lanes, make docker-syspkg-deb-desktop and make docker-syspkg-appimage (folded into docker-syspkg and a new syspkg-desktop CI job). The AppImage pack step pins its own tooling (scripts/fetch-appimagetool.sh, scripts/fetch-appimage-runtime.sh) rather than letting appimagetool fetch its runtime stub from a rolling tag on every pack. See installation.md.

  • Live blame + statediff from the asmspy serve/dataflow leg (docs/internal/archive/gui/41-live-blame-statediff-serve-leg.md). The backward-attribution blame cone and the step-to-step statediff register delta were producible only by the emulator recorder; a live attach could not emit them. Both are now emitted by the live single-step leg as pure projections over data it already captures — blame is a backward slice over the def-use graph the sink already builds (seeded at the penultimate step), statediff is a delta of the per-step register ring — spelled with the same wire builders the recorder uses, so a live artifact and a golden one are byte-identical. Opt in with serve blame:true / statediff:true or CLI --blame / --statediff (the latter self-arms the register ring); off by default. No capture-engine, wire, or schema change (both kinds were already defined). The blame cone and state-diff views already worked live client-side; this adds the reproducible, deep-linkable precomputed artifact.

  • Per-invocation dataflow views over a continuous capture (docs/internal/archive/gui/40-segment-dataflow-by-invocation.md). A continuous dataflow/auto capture appends many invocation passes into one recording, each restarting df_step at 0 (delimited by a df_invocation marker). The desktop’s decode_streams indexed those steps in one flat space, so the passes aliased — offsets were last-write-wins and operands/edges merged across passes that merely shared a step number — the last dataflow consumer that still conflated passes after the register Scrubber learned to segment its ring. build_segmented_dataflow now splits them (one decoded stream per pass, bucketed exactly as the Scrubber’s build_segmented_step_index buckets regstate), Streams::df resolves to the latest pass (the live default), and the Slice / Timeline / Loom panes carry a per-pass invocation pager — following the latest by default, pinnable to an earlier pass. A one-shot recording stays a single pass, byte-identical to before. Pure desktop decode + a selector: no producer, wire, or schema change.

  • df_step states its region base on the wire (rbase) (docs/internal/archive/gui/37-region-tag-on-df-step.md). 36 anchors a routine-relative df_step offset by deriving the base from a recording’s single codeimage span, and must refuse whenever a recording carries zero or ≥2 spans — which a live auto candidate walk produces routinely. The producer already knows the base as it writes the offset, so it now states it: df_step gains an optional rbase (emitted only when nonzero, so a rbase == 0 step is byte-identical to pre-37). All four producers state it (live ptrace, both corpus recorders, the Author VM). The desktop reader places a tagged step from rbase + off — a stated fact — with per-event precedence over 36’s recording-wide anchor, and grades HOW it was placed (wire / single-span / mixed) in the HUD, so a multi-span capture that 36 alone must refuse now places every vertex; 36’s single-span anchor remains the permanent fallback for pre-37 recordings and for rel trace. The terrain churn walk is redeemed as a sound “region as-of this step” resolver — it counts df_step offsets as steps (a dataflow recording’s churn no longer pins at step 0) and keys the churn join on each step’s own rbase. Additive optional field on a known kind — no envelope bump, no break.

Fixed

  • asmspy attach to a multi-process tree no longer kills a followed child: the teardown drained a queued single-step trap against the wrong address space, which was fatal to the child moments after detach.

  • asmspy hardens its JIT perf-map/debug file opens against planted /tmp and /proc paths (symlink/ownership checks before trusting a path).

  • auto reliably captures — one shot is not a policy (docs/internal/archive/gui/39-auto-capture-reliability.md). The complaint “start and arm a process, it starts then stops, and the pane says refused: no session is running” was two independent bugs. The region picker: the candidate walk (arm a ranked pick, and on “not seen entering” try the next) lived inline in two identical, untestable copies, and that duplication is exactly how the strong AMD-IBS/entry sampler ended up with no walk at all — it returned a single pick, so ncand was 0 and the guard never fired (on an AMD box the better sampler gave the less resilient capture). The walk is now one pure, unit-tested asmspy_autoregion_walk, and auto_pick hands back the ranked list like its sw-clock sibling, so BOTH samplers walk; an empty sample window is a retry, not a verdict (an idle target yields nothing in one 400 ms window but may in the next), and the window is settable end to end (--window=<ms>, a wire ms param, a capture-pane input); a continuous capture now survives a quiet region instead of ending at the first lull, surfacing the quiet window as a 0-step df_invocation marker. The session lifecycle: a capture that ends on its own is now announced from the tracer tail (an idle client learns without sending a command), the desktop frees the ptrace jack it used to hold forever (reconciling active against the terminal-event count), a stop for a self-ended session is acked rather than refused for a precondition the host’s own reap just removed, and a stale refused: banner no longer haunts the next healthy capture. Every policy change is proven by pure tests (no CI lane has AMD silicon), per CLAUDE.md.

  • The 3D overview places a live routine-relative path, or says why not (docs/internal/archive/gui/36-anchor-the-3d-plane.md). A live dataflow/auto capture (df_step + codeimage, no trace) is basis:"rel", so its 3D-overview tab used to open onto an empty, unlabelled plane — three defects stacked: the promised “rel: routine-relative” HUD chip never fired (it keyed on the empty canvas basis), every routine-relative vertex silently failed to project (so a total placement failure was indistinguishable from success), and the value producer emits no block coverage (a flat plane even when anchored). Now a routine-relative offset is anchored to the recording’s single codeimage code span (base + off is the true address — a derivation from a stated fact, not a guess): the PC path is placed on the plane (TRAJ_ANCHORED alongside TRAJ_RELATIVE_BASIS; the wire basis stays rel), the terrain grows a labelled single-step-residency height rung (df_step, never block coverage), placement is counted (an off-plane vertex is named, not dropped — the 4096-byte codeimage clamp), and a recording with zero or ≥2 code spans refuses louder — no geometry and a stated reason (HEIGHTS NOT PLACED / PATH NOT PLACED), never the old silent empty plane. Convergence detection admits an anchored rel path (a trace recording with tid) while still refusing an unanchored one. Pure space/ + HUD — no engine, GL, or wire change.

Added

  • scene-df-loop golden (docs/internal/archive/gui/36-anchor-the-3d-plane.md T4). A 3D-overview golden in the live dataflow shape no existing golden carried — absolute codeimage + region-relative df_step, no trace — so the anchor fallback and the terrain’s df_step residency rung are exercised end to end by a committed artifact (record_scene_df, regenerated byte-stable under docker-cli).

  • Author-mode arm64 run/trace (docs/internal/archive/gui/32-per-guest-value-producer.md R5 T3). The Author door’s Run button now dispatches assembled AArch64 code through the per-guest emulator L0 value-fabric producer (asmtest_dataflow_emu_run_arch, src/dataflow_emu.c) instead of refusing it — never through the x86-64-only emu_call_traced/emu_result_t path, and never touching src/emu.c’s register ring. The result is a genuinely distinct shape (def-use edges + per-step operand values, no register file, no fault kind/address — the door says so explicitly rather than showing zeros) that materialises, on Save, into trace/df_step/df_edge events through the existing recording writer, so an arm64 Author run opens in the Loom / Slice / Timeline like any other recording. The arch-gating table’s AArch64 row is now runnable, the shared refusal label now names only the arches still unsupported (ARM32 / RISC-V), and the capability panel states which arches Author mode runs straight from that same table.

  • arm64 register-ring time-travel for the Scrubber (docs/internal/gui/ 32-per-guest-value-producer.md R5 T2, the regstate half). The emulator’s per-step register ring (src/emu.c) is arch-parameterized: emu_arm64_t gains its own opt-in drop-oldest ring (emu_arm64_step_capture/_clear/ _count/_dropped/_at), mirroring the x86-64 ring’s shape over emu_arm64_regs_t instead of a union grafted onto the x86-64 handle — zero changes to emu_t, emu_x86_regs_t, or emu_snapshot/emu_restore. A new emu_arm64_regs_t@aarch64/aapcs64 regstate descriptor (docs/internal/gui/asmtrace-schema.md) names the AArch64 GP file (x0..x30, sp, pc, nzcv); the arm64-df-chain golden now also carries a regstate ring (steps_cap = 8, the arm64 analogue of add_signed’s worked example), and a new arm64-regstate-truncated golden proves the ring’s faithful truncation (D7). No desktop/reader change: the Scrubber’s existing generic “unnamed integer field” fallback already renders the new register names (in a plain sorted rather than hand-curated order, a cosmetic follow-on).

  • A command palette (Ctrl+Shift+P / Ctrl+P) over the dt_nav_go router (docs/internal/archive/gui/21-spine-navigation.md T1). A modal fuzzy finder that makes the whole spine reachable by typing: view-switch, go-to step/offset/link, open-recent (the open workspace), attach-a-process, run-walkthrough, reset-layout, and jump-to-a-recorded-offset — every command dispatching only through dt_nav_go (or the exact want_view/show_* intent the keymap uses), plus a discoverable hint for every advertised accelerator, enumerated from dt_nav_bindings() so the palette can never advertise a key the app does not honour. Filtered with the “showing N of M” idiom (app-only ImSearch relevance, degrading to an unranked list under the null backend). No new dependency.

  • A persistent wayfinding breadcrumb (recording ▸ session ▸ view + step/filter/process scope) with same-basename disambiguation (docs/internal/archive/gui/21-spine-navigation.md T2). Sourced from nav.current and drawn in the outer shell (the docked menu bar / the windowed top strip) so it is visible from every pane; two same-basename .asmtrace files are now distinguishable in both the band and the recording tab titles (basename + the shortest distinguishing parent-dir segment, else a short path hash). Builds on the doc-18 back/forward affordance; the empty state prompts rather than blanks.

  • An always-visible overview/minimap on the timeline and the Loom (whole trace, current viewport marked, click-to-jump) (docs/internal/gui/21-spine- navigation.md T3). A compressed projection of each recording’s own rows — never a fabricated or padded layout (docs 04/08 ban; a sparse trace yields a sparse strip, tested) — with the click routed through dt_nav_go so a minimap click and a typed go-to land identically. Completes the doc-14 T5 overview-strip / timeline-windowing follow-on (the ImZoomSlider now drives a real timeline window). No new dependency (ImPlot, ImZoomSlider already vendored).

  • Keyboard camera for the 3D overview, Tab-focusable panes and 3D viewport, and keyboard-operable slice cones (docs/internal/archive/gui/22-selection-and-search.md T2, F18). The 3D spacetime overview was a mouse-only island; because Dear ImGui exposes no OS screen-reader tree, keyboard operability is the only accessibility substitute. Arrows now orbit, +/- (and =) dolly, R resets and T is the faithful top-down 2D-ish fallback — routed through the SAME Camera methods the mouse drag uses (a pure camera_key), so keyboard and mouse are one code path. The 3D viewport gains a Tab-reachable focus target that exists even under the null backend (no GL), which is what makes the keyboard camera headlessly testable; the arrow keys defer to the camera only when a 3D pane holds focus. The slice cones (b/f/Enter) were already keyboard-operable.

  • Global find (Ctrl+F): highlight-all, match count, aggregate cost, and Enter/Shift+Enter cycling; type-to-narrow extended to disasm and hot-edges (docs/internal/archive/gui/22-selection-and-search.md T3, F17). Find is a measurement, not just a jump: it highlights EVERY hit in the timeline (with a clean seam for the doc-21 minimap), reports the match count AND the aggregate cost (summed hot-edge samples across sites — “this symbol retires N samples across M sites”), and cycles matches through the ONE navigation spine. It searches the timeline, disasm, order-preserving syscall stream and hot-edges in stream order (never relevance-ranked), and never hides a row. The doc-16 client-side “showing N of M” narrowing now also covers the disasm and hot-edges lists. The call tree is deliberately untouched — it stays engine-filtered, so surviving depths never lie (D7).

  • App-level undo/redo (Ctrl+Z/Ctrl+Y) over filter / cone / selection / take-set state — distinct from the Author editor’s text undo; the Loom takes gutter gains per-take remove and clear-forks (docs/internal/gui/22-selection- and-search.md T4, F12). Applying an aggressive filter, lighting the wrong cone or forking exploratory Loom takes is now reversible: an app-level command stack reverses and replays those view-model changes, deliberately DISJOINT from the Author editor’s own text undo (both guard on text-input focus and own separate state) and from the router’s back/forward history. The Loom takes gutter — which held no take set before — now accumulates forks with a working per-take remove and a clear forks action, both reversible; clearing removes a whole take node with its loud refusal intact, never silently dropping a failure (D7).

  • Derivable severity fidelity-chrome tier in the .asmtrace schema (docs/internal/archive/gui/23-graded-truth-layer.md T1). An optional provenance.severity string (neutral | caution | integrity) names the tier a reader renders a recording’s fidelity chrome at. It is derivable from the fidelity fields the schema already carries (torn / truncated / trust / redacted / skip / basis / drops), so a producer need not emit it and every old recording still grades; additive, no new envelope major, no field on any existing kind. Landed under doc 01’s append-only rule with 01-owner sign-off recorded in the schema doc (the Phase-3-freeze checkpoint, D5). It grades loudness; it gates no truth off (D7).

  • A uniform elapsed-time + Cancel busy signal on long operations, and a degrade-to-coarse 3D scrub (docs/internal/archive/gui/23-graded-truth-layer.md T4). progress.h grows a pure LongOp (elapsed clock + faithful determinate-vs- indeterminate mode + observable cancel flag) and a draw_progress half, so an op that can exceed a frame shows a spinner + elapsed + Cancel instead of a silent UI-thread stall an expert would mistake for a hang. The 3D scrub degrades to the labelled coarse terrain plane (a pure should_degrade decision + coarse_slice) while a full re-slice would exceed the frame budget, then swaps to the full slice — the coarse plane is the same labelled rung the terrain shows normally, so it hides nothing. The no-fabricated-total fidelity rule is unchanged.

  • A Queue path on the live patch bay (docs/internal/archive/gui/23-graded-truth-layer.md T3). A blocked capture can be queued as a visible, cancellable chip; it starts automatically the moment the jack frees — and only then (never an auto-swap, so the one-ptrace-jack invariant is never bypassed).

  • The desktop remembers your workspace, and a Settings pane (docs/internal/archive/gui/20-workspace-and-settings.md T3/T4/T5). Open recordings, the active view and each pane’s selection now restore across launches (persisted as asmtrace-links in build/desktop-workspace.json beside the dock .ini), with an MRU recents list on the home rail and a File ▸ Open Recent menu — each entry reopening to its exact prior position — plus drag-drop to open a file. A recording whose file has vanished is kept in recents with its load error, never silently dropped. Named dock perspectives and named saved filter presets persist in the same store. A new Settings pane adds a user text-scale (0.8×–2.0× via FontGlobalScale with a DPI-aware atlas re-bake on content-scale change), a remembered window size (retiring the hardcoded 1280×720), and a light theme that keeps the fidelity-chrome (warn/refuse) contrast — and states transparently that text-scale is the only in-app accessibility lever, because Dear ImGui exposes no OS screen-reader tree.

  • In-app term registry, per-view “?” caveats, domain-term-first headings, and a searchable Terms pane (docs/internal/archive/gui/24-one-visual-language.md T3). The coined GUI lexicon (Loom, fabric, patch-bay, hollow span, born-untraced, patient-zero, hot-edges, knot, jack, worldline, Reweave, dim/hot take, terrane) is now defined in the ONE Sphinx glossary (docs/project/glossary.md) with each term’s plain-language definition and its expert synonym; scripts/gen-terms.py generates the app’s registry from that one file (no hand-copied second list), so a tooltip, a legend and the docs cannot drift. Every coined surface now leads with its canonical domain term and the metaphor as a subtitle (“Data-flow lineage (Loom)”, “First divergence (patient zero)”, “Edge execution counts (hot-edges)”), carries a per-view “?” with the verbatim metric caveat (“hot-edges are edge counts, not a call stack”), and offers a searchable Terms pane (the doc-16 ImSearch idiom, guarded so the null backend degrades to a plain list). The registry/lookup/heading metadata are asserted headlessly against the same glossary the build parses (one source).

  • First-open primer + legend on the Loom and the 3D overview (docs/internal/archive/gui/24-one-visual-language.md T5). The two heaviest surfaces now open a dismissible, in-canvas primer (“what this is / how to read it” + the shared legend) instead of dropping a learner into a raw fabric/terrain; it shows once per session, never nags again, and is re-openable from the per-view “?”. Until it is acknowledged the view holds its lean default (nothing heavier is front-loaded). State transitions are asserted headlessly.

  • Convention-alignment keyboard shortcuts, accurately advertised (docs/internal/archive/gui/18-breach-stops.md T1). Shift+F fits / frames the current selection, ,/. step to the previous / next sibling, F10/F11 are step / step-back aliases of j/k, W/S/A/D drive camera zoom/pan when a spatial pane (timeline / 3D) holds focus — a labelled context that also keeps d’s diff meaning outside it — and Ctrl+C copies a deep link (an alias of y). The help overlay is now generated from a wired flag, so it greys any not-yet-mapped binding “planned” and can never again advertise a dead key as live. Every advertised binding has a passing desktop-ui-test.

  • Back/forward navigation history over the deep-link router (docs/internal/archive/gui/18-breach-stops.md T6): Alt+Left/Alt+Right walk a bounded, serialisable stack of asmtrace-links and land identically to a fresh navigation, with a minimal breadcrumb affordance by the status line. A new jump clears the forward branch (browser-history discipline).

  • A reset-layout keybinding (Ctrl+Shift+R) and a pre-commit confirm before arming a perturbing single-step capture (docs/internal/gui/18-breach- stops.md T2/T5): the confirm states the page-dirty / timing cost and, on arm64, the blocking-syscall termination that detach cannot undo; the capture picker now defaults to the least-perturbing substrate the host supports (AMD IBS where available, else the lightest ptrace mode), and arm64 single-step modes are greyed and annotated with the stated hazard.

  • Author output can be saved, and an unsaved-work marker in the tab title (docs/internal/archive/gui/18-breach-stops.md T3): an Author run materialises into a Recording and writes to a .asmtrace file through the same confirm-overwrite dialog live captures use; a dirty (authored + unsaved) tab shows a trailing *.

  • Pan/zoom, fit-graph and selection for the graph views (docs/internal/gui/15- plotting-and-graph-nav.md T3, via imgui-node-editor). The process topology, the live/replay call tree and the frozen hot-edge snapshot now draw on a real pan/zoom canvas: drag to pan, wheel to zoom, “Fit graph” frames the whole graph (NavigateToContent), and double-clicking a node navigates through the same deep-link router the buttons use (a process’s syscalls, a call’s or edge’s address in the region view). The node positions are the app’s own deterministic layout, fed every frame — node-editor’s settings file is disabled so it can neither persist a dragged position nor invent one on load: force-directed layout stays banned (docs 04/08), the fidelity guardrail is that the layout lives in a pure, tested builder, never in the drawing library. Large graphs (a 10k-routine trace) cull off-viewport nodes before emitting them. imgui-node-editor is vendored (MIT, master 021aa0ea — the last release fails to compile on the pinned imgui) with its bundled crude_json; its four TUs compile into both desktop binaries and it is in the imgui-repin compile-probe. The layout is a pure model, so the null-backend tests assert node positions, the disabled settings file, click-routing and culling without pixels.

  • The keyboard shortcuts now all work, and are tested. The desktop help overlay advertised twelve keybindings but only [/] were wired; the other ten are now implemented (docs/internal/archive/gui/17-interaction-testing-and-editor.md T1): 1/2/3/4 switch view (canvas / timeline / slice / diff), j/k (and the arrows) and PgDn/PgUp step the selection, Enter opens the slice explorer at the selection, b/f light the backward / forward dependence cone and c clears it, d attaches or detaches a second recording for the diff, x swaps A and B, n/p walk the divergences of an attached pair, y copies a deep link to the current position, and Ctrl+G opens a go-to box that accepts a step number or a full asmtrace-link. They are decided in one place, so the help overlay and the behaviour cannot drift.

  • A headless interaction-test lane for the desktop app (make desktop-ui-test), built on the Dear ImGui Test Engine. It drives the real UI through simulated clicks and keypresses on the null backend — the layer the golden-text tests cannot reach — and writes JUnit XML for CI. Every one of the twelve keybindings has a passing test (which is what proves they all work). The engine is fetched at build and compiled into the test binary only, never into the shipped asmtest-desktop/asmtest-viewer, so those stay MIT.

  • Non-modal toasts for live-session events (docs/internal/gui/16-live- feedback-and-filtering.md T1, via ImGuiNotify). Events that were previously silent — a capture refusal, a session ending, a fatal, a skip (success-with-nothing-to-report), a save completing or failing — now raise a transient toast in the corner. An exact saved capture’s toast carries an “Open in Loom” button that opens it through the same door the panes use; a statistical capture’s does not (there is no Loom for it). Toasts supplement, never replace, the in-pane refusal banners (the banner is still the record). The decision of which toast to raise is a pure, unit-tested function over the live status and save outcome, so it is driven as a model by the null-backend tests. ImGuiNotify is vendored (MIT) with a minimal Font Awesome range merged into the UI font for its type icons. Bundled in both the full app and the render-only viewer.

  • Client-side filtering for the long lists (docs/internal/gui/16-live- feedback-and-filtering.md T2). The Learn-door walkthrough catalog narrows as you type (via ImSearch). The syscall stream gains a name filter that hides non-matching rows while keeping each row’s original number and execution order and showing “showing N of M”, so a filtered view never reads as if the trace only made those calls; it matches syscall names only, never the redacted payload. (Filtering is applied only where it can stay faithful: the time-ordered syscall stream uses an order-preserving filter rather than a relevance re-rank, and the address-ordered disassembly keeps its existing time/range controls.)

  • The register scrubber and the ABI x-ray are now hosted in visible shell panes (docs/internal/archive/gui/09-teaching-producers.md T3/T4 — the integration surfacing pass). The register time-travel scrubber is a per-recording Scrubber tab: its regstate seek index (analysis/stepindex.h) is built once when the recording opens, parallel to the workspace like the decoded streams and Observer decks, and its playhead is persisted per recording so switching tabs holds each recording’s place. The ABI x-ray is a tab that locks the active recording (the SysV leg) against the attached B (the Win64 leg, the d binding) — reusing the Diff tab’s A/B mechanism — with the walkthrough rail rebuilt only when the pair changes; no B attached shows an “attach the Win64 leg” placard, the same shape Diff shows with no second recording. Both degrade to their own fidelity placards (absent producer, unaligned pair, torn ring) exactly as their standalone draws do. Both views are backend-free, so they ship in the render-only viewer too, and desktop/test/test_shell.cpp pins the wiring end to end: the seek index is built and present() for a regstate recording, a make_pair pair feeds a present and aligned x-ray through the very calls the tab makes, and the per-recording parallel vectors survive a close.

  • The 3D spacetime overview is now hosted in a visible shell pane (docs/internal/archive/gui/10-spacetime-3d-overview.md — the integration surfacing pass). It is a per-recording 3D overview tab. The GL boundary is a new abstract SceneHost (desktop/src/ui/scene_host.h): the shell weaves the pure, engine-free space/ models once per recording (a SceneView), draws the HUD (draw_scene_hud) and resolves picks (resolve_pick) itself — all GL-free, so ui/shell.o stays out of the viewer’s engine closure and the null backend drives the whole model + HUD + placard path — and reaches the GL scene only through ShellState::scene_host. main.cpp injects the one concrete host (ui/gl_scene_host.cpp: a scene3d::Scene over an offscreen RGBA8+depth FBO whose colour texture the shell blits with ImGui::Image, re-uploading the terrain/trajectory only when the recording or playhead changes); the headless tests inject none and the pane shows a placard where the viewport would be. A left-drag orbits, the wheel dollies, and a click that did not drag picks → decode_pick → resolve_pick → 04’s deep-link router (3D to find, 2D to read). Because the scene links no engine it ships in the render-only viewer too (D4). desktop/test/test_shell.cpp drives draw_scene_overview under the null backend (models woven, code regions placed, terrain heat, the exact trajectory, the terrain slice tracking the HUD playhead, and the faithful no-regions placard for a codeimage-less recording); the GL scene stays pinned by the test_scene_fbo smoke and the pick router by test_drillin. This completes the deliberately-outstanding integration item — all ten GUI docs’ views are now surfaced in the shell.

  • Golden scenes and a CI-runnable GL lane for the 3D spacetime overview (docs/internal/archive/gui/10-spacetime-3d-overview.md T7 — the view is now complete). Three committed scenes pin the whole stack end to end. Two are generated by make asmtrace-golden from one byte literal whose listing sits beside the bytes, so the terrain’s heat is countable on paper: tests/golden-asmtrace/scene-abs-loop.asmtrace (the coarse rung — codeimage + absolute-basis trace from a 3-iteration loop, giving one hot cell over cold ones: coarse terrain plus one exact trajectory) and scene-abs-loop-truncated.asmtrace (the same bytes with the trace buffer holding 6 of the 13 executed steps, so the producer flips truncated, drops.lost counts the rest, and every populated cell is TF_TORN) — both byte-stable under make asmtrace-golden-check. The third, the rich mem scene, cannot be generated at all (mem is a reserved schema kind with no producer) and is hand-authored under tests/golden-asmtrace/scenes/, marked schema-unfrozen and deliberately inert but present: its accesses fall outside every code region, so the coarse-rung projection gains no data cell and the scene renders pixel-identically to its coarse twin. desktop/test/test_scene_fbo.cpp now builds and renders all three under surfaceless EGL + software Mesa in the make docker-desktop lane, asserting terrain heat and the exact tube for the abs scene, the red gash for the truncated one, and — at the pixel level — that a survey-only scene draws its statistical layer and nothing on the exact one; the statistical scene reuses the committed obs-survey-ibs fixture rather than adding a second survey saying the same thing. Each scene’s model facts are asserted in the test’s pure half too, so they hold on a host with no GL. desktop/README.md gains a 3D spacetime overview section (controls, the coarse-vs-rich staging, and the “3D to find, 2D to read” rule).

  • Drill-in router + fidelity invariants for the 3D spacetime scene — every pick reaches the right 2D view, and truncation/statistical survive it (docs/internal/archive/gui/10-spacetime-3d-overview.md T6). The scene is an overview, not a reading surface: scene3d/pick.h’s resolve_pick (pure, GL-free) now routes every pickable kind through 04’s deep-link router to the flat 2D view that actually reads it — an exact code cell → the trace canvas at that offset (or the codeimage-versioned disasm pane, 08-T7, when the region churned); a data cell (rich mem rung) → the slice explorer at the step whose access last hit it; an exact PC vertex → the operand timeline at that step. Two fidelity invariants are pinned as tested behaviour: truncation survives the drill-in — opening a TF_TORN cell lands on a 2D view that still carries 04/08’s truncation banner, so the 3D tear is never the only signal; and statistical is never exact — a survey-only recording draws no exact trajectory tube, its statistical provenance is surfaced, and a statistical pick (a TF_STAT cell or a TRAJ_STATISTICAL vertex) opens the hot-edge view (08-T4), never the exact slice explorer. New desktop/test/test_drillin.cpp (headless, no GL) asserts each pickable-kind mapping and both invariants against the truncated and obs-survey-{ibs,sw} fixtures; builds into both binaries and stays engine-free (D4).

  • Live-observer overlay for the 3D spacetime scene — N per-thread trajectories and cross-thread convergence hints (docs/internal/archive/gui/10-spacetime-3d-overview.md T5). A live 07 LiveSession now feeds the 3D overview: each thread is one coloured trajectory growing over the shared terrain in real time, fed incrementally by re-running the trajectory builder on the session’s growing recording — no second capture path, the overlay consumes the session the app already owns and opens no ptrace of its own (D6/D9). A new engine- and GL-free detector desktop/src/space/converge.{h,cpp} surfaces where two threads meet: when two different tids place a PC vertex in the same projection cell within a sliding step window, it emits one convergence mark per thread-pair-and-cell (the closest crossing), which the scene draws as a bright magenta arc bowing above the plane on its own toggleable layer. It is explicitly a hint, never a proof — per-thread step indices are not a global clock, so it shows co-location, not a proven race or ordering — and a statistical or region-relative (single-step) path never converges, because a shared cell over a sampled or unplaced path would be a false address-space claim. Threads are placed by colour, not by physical core/NUMA (the repo has no scheduling feed — the affinity layer stays gated, recorded not faked). Builds into both binaries and stays engine-free (D4); covered by test_converge (the detector + the incremental feed against a two-thread fake-serve fixture) and an added arc case in test_scene_fbo.

  • ABI x-ray in the desktop GUI — SysV vs Win64 argument marshalling, side by side (docs/internal/archive/gui/09-teaching-producers.md T4). The classroom flagship: two register scrubbers (System V | Microsoft x64) LOCKED to one playhead, driven by an authored walkthrough (06’s note stops). asmtrace_record grew a Win64 recording path — it now records the same corpus routine twice, once through emu_call_traced and once through emu_call_win64_traced, into paired goldens (abixray-make_pair-{sysv,win64}, abixray-sum3-{sysv,win64}) that carry trace + per-step regstate + the walkthrough stops. Advancing a stop seeks both panes together; the per-step register deltas do the animation, and a cross-pane diff marks the registers whose SysV and Win64 values disagree — so a0 → rdi (SysV) vs rcx (Win64), the callee-saved rsi/rdi role reversal and struct-return eightbyte classification are the register diff, made mechanical. Faithful (D7): a pane with no regstate producer refuses and names --steps, two runs of different lengths are flagged not aligned (they cannot share a playhead), a dropped step renders UNKNOWN per pane, and a stop past the recorded window is refused not clamped. Builds into both binaries and stays engine-free (D4); covered by test_abixray (builder) and test_abixray_draw (the ImGui half).

  • Save a live capture to a .asmtrace file from the Inspect door — and open it in the Loom (docs/internal/archive/gui/07-serve-live-host.md). The live host kept its recordings in memory only; the Inspect door now has a save capture section that writes the growing (or last completed) recording to disk in the same NDJSON a --record run produces, via a new model-native serializer recording_to_asmtrace / save_recording_file (desktop/src/doc/recording.h) that round-trips through the loader: header, every event in stream order, and the end footer only when the recording had one — so a torn capture reloads torn and a truncated one stays truncated. A saved exact recording offers an Open in Loom button that loads the file into the Workspace and jumps the tab strip to its Loom (a statistical capture says why it cannot be woven instead of sending the user to the Loom’s refusal). Covered by new test_recording round-trip cases (string and file, clean and truncated).

  • Register time-travel scrubber in the desktop GUI (docs/internal/archive/gui/09-teaching-producers.md T3). A playhead over a recording’s steps showing the full register file at each step, seeked O(1) through a new shared regstate index (desktop/src/analysis/stepindex.h, which the Loom’s now-column reads too). Registers that changed vs the previous held step are highlighted. Truthful about the ring’s limits (D7): when the producer dropped a prefix of steps the scrubber renders that region as a torn edge — seeking into it shows UNKNOWN, never zeros — and a recording with no regstate events states the producer is absent and deep-links the docs (the re-run-with-larger-max_insns fallback is deliberately not offered). Builds into both binaries and stays engine-free (D4); covered by test_scrubber (builder: seek, diff-highlight, tear, absent message) and test_scrubber_draw (the ImGui half under the null backend).

  • The 3D spacetime overview’s GL scene — camera, terrain mesh, trajectory tubes, colour-ID picking (docs/internal/archive/gui/10-spacetime-3d-overview.md T4). The desktop now draws the address-space projection, a terrain slice and the execution trajectories as a 3D scene under the ImGui HUD, in the same GL context the shell already stands up. A 2^order × 2^order grid VBO is displaced by a per-cell height texture and coloured by a flags texture (a TORN tear is a red gash, statistical residency is dimmed, JIT churn is tinted); exact paths draw as opaque lines while statistical residency is translucent and stippled and is never joined into an exact tube. An orbit camera (drag to orbit, wheel to dolly, a “reset view” and a faithful “top-down 2D-ish” preset) frames it, and every pick — resolved through an offscreen R32UI colour-ID framebuffer — leaves the 3D overview for the flat 2D view that reads it: a terrain cell opens the trace canvas at that code offset, a trajectory vertex the slice explorer at that step (3D to find, 2D to read). The scene links no engine — only OpenGL and the pure space/ models — so it ships in the render-only asmtest-viewer as in the full app; the ImGui HUD and the pick/router logic are separate TUs that link no GL. Camera math is pinned by test_camera (no display), and the terrain + picking by test_scene_fbo, a gated GL smoke that renders offscreen via surfaceless EGL on software Mesa. The math comes from a newly pinned single header, linmath.h (WTFPL), fetched and digest-verified like imgui/json.

  • Opt-in per-step register-capture ring in the emulator (docs/internal/archive/gui/09-teaching-producers.md T1). emu_step_capture(e, cap) arms one more UC_HOOK_CODE on the x86-64 guest that snapshots the full emu_x86_regs_t before every executed instruction into a bounded, caller-sized ring — the per-step producer a register time-travel scrubber and the ABI x-ray replay from. It is drop-accounted (when a run exceeds cap the earliest entries are evicted and emu_step_dropped() counts them, so a torn timeline is genuine data, not a silent gap), read back with emu_step_count() / emu_step_at(), and never armed by default: the arming is handle-level, persists across emu_call_* until emu_step_capture_clear(), and survives emu_snapshot / emu_restore (the Track F arming discipline).

  • The live Observer views — seven of them, plus a codeimage kind and a PT-replay slice (docs/internal/archive/gui/08-observer-views.md T1–T8). The desktop now renders what a live asmspy --serve session produces: the syscall stream (payloads redacted by default, revealed per row, session-wide reveal behind a second confirmation), the watchpoint timeline, the process topology, the statistical hot-edge table, the call tree with the engine-side filter panel the TUI never exposed, the region trace as discrete invocation snapshots, and a disassembly pane that resolves bytes as of trace time.

    The load-bearing property is that these are not “live views” at all: each is a pure function of the recording document model, so the Inspect door and a replayed .asmtrace tab draw the same deck from the same code — which is how an AMD IBS survey, an arm64 watchpoint refusal and a JIT code image are all asserted in CI on hardware that has none of them.

    Each view carries the fidelity rule that view can most easily lose. Redaction distinguishes hidden here from withheld at record time (and revealing cannot conjure bytes that are not in the file). A watchpoint’s direction has three values, the third being “the trap fired and the instruction did not decode”, and a value that was never read back is never rendered as 0. A refused watchpoint arm is a successful session with nothing to report, carrying asmspy_hwdebug_reason()’s measured string verbatim. The topology view states that the procs engine holds the ptrace jack for the whole descendant tree, and never shows inv without saying whether it counts syscalls or calls. Hot edges are edges, not stacks — there is no flame graph, because nothing in an IBS sample observed a call stack — and IBS entry evidence is labelled differently from software-clock residency. The region view pages between invocations and never scrubs, because between two invocations the target ran unobserved for an unknown time.

  • codeimage — captured code bytes at a version (schema; produced by asmspy --serve). A JIT patches, frees and reuses code addresses, so bytes read after the fact are not the bytes that ran. A region-scoped serve session now tracks its region through asmtest_codeimage and streams versioned snapshots, and the desktop resolves an address at trace time t to the version with the greatest when ≤ t — never the newest, and unknown rather than the next one along, since that would be the succeeding method’s code. Where the recorder is unavailable (soft-dirty / PAGEMAP_SCAN, Linux ≥ 6.7) the session emits the measured reason as a note and captures without it; the pane then falls back to the recorded disasm strings and labels them as the weaker source. make cli-smoke asserts the events are well-formed, that the tracer’s own entry int3 never appears in them, and that a static region does not accumulate byte-identical versions.

  • The PT-replay def-use slice (08 T8). On a PT host a value slice can be produced with zero single-steps of the target: the hardware records the path, and the F5 producer replays it through the emulator against the recorded code image to reconstruct the values. The result is an ordinary def-use stream, so the existing slice explorer and Loom draw it unchanged. Capture is hardware-gated with the library’s own reason; replay is not — a path decoded at capture time (stitch) and the bytes it was decoded against (codeimage) replay anywhere Unicorn and Capstone are present, which is what desktop-test exercises on hosts with no Intel PT.

  • asmspy --serve[=<socket>] — the live-session control loop (docs/internal/archive/gui/07-serve-live-host.md T1/T2/T6). asmspy can now be driven as a capture host rather than a one-shot command: it reads NDJSON commands (start / pause / stop / quit) on stdin or a unix(7) socket and streams back the events of whichever engine is running. This is how the desktop GUI captures — it spawns asmspy --serve as a subprocess, locally or via ssh <host> asmspy --serve, so remote capture is the same code path as local and the viewer links no tracer at all.

    It is a thin wrapper over libasmspy, and deliberately so: every mode drives one engine through the public header with the same sinks and the same writer TU --record uses. There is no serve-specific event body anywhere, so a session’s events are exactly a recording’s events — slice [header … end] out of the stream, drop the control lines, and any reader in this tree parses it. The protocol is specified normatively (Serve protocol in docs/internal/gui/asmtrace-schema.md) rather than defined by the C code, including the three lifecycle kinds (session / cmd / err) that were reserved in the kind registry and are now defined.

    Two properties carry over rather than being re-argued. One session at a time, because a target has one ptrace jack: a second start is refused with an err naming the rule. And a session only ever ends through the engines’ two-phase detach — stop flag, SIGALRM to unblock a pending waitpid, join — so the target survives, which make cli-smoke now asserts end to end (start log → stop → start stream → quit, victim alive after). The flag matrix is the argument parser’s, verbatim, so the two front ends cannot disagree about what is legal.

    Fidelity is preserved where it would have been easiest to lose: pause suspends emission but not tracing, so the events it swallows are counted, the recording is marked truncated, and the terminal event carries paused_dropped. And a skip reports the measuring source’s reason — which surfaced a real defect: asmspy_strerror had no case for ASMSPY_SAMPLE_UNAVAIL, so that positive skip code rendered as the default "attach failed" — wrong twice over, since the IBS sampler is out of band and attaches nothing. Fixed.

  • libasmspy — the asmspy tracer engine is now a linkable library (docs/internal/archive/gui/07-serve-live-host.md T0). The ptrace engines (cli/asmspy_engine.c) and the /proc/ELF/JIT resolvers (cli/asmspy_proc.c) were loose objects compiled straight into the asmspy binary with no public header, so anything else that wanted them — the forthcoming --serve control loop, a future language binding — would have had to re-declare or re-implement them. They now ship as build/libasmspy.a plus make shared-asmspy (build/libasmspy.so), behind one public header cli/libasmspy.h, exactly as every other tier already ships (libasmtest_emu / _dataflow / _hwtrace).

    This is packaging, not a rewrite: the engines, sinks, and the one-tracer-thread / two-phase-detach contracts are byte-for-byte the code that was already there, and cli/asmspy.h now includes the public header and re-exports it, so every existing includer compiles unchanged. cli/asmspy.h keeps only the front end’s own entry point.

    What the split buys is a boundary that can be tested: the new cli/test_libasmspy.c includes only libasmspy.h and links only libasmspy.a plus the framework tier objects — not asmspy.o and not ncurses — then attaches to a live victim, streams syscalls through a real sink, detaches and asserts the target survived. A hidden dependency on the CLI front end, or a TUI dependency leaking into the engine, now fails there and nowhere else. It runs in make cli-smoke.

    cli/asmspy_autoregion.h became self-contained in the same change (it used asmspy_sample_edge_t without including anything that declared it, forcing consumers to pin an include order by hand — desktop/src/vm_compat.cpp had that pinned against clang-format). The desktop still links nothing of this (D9): it reaches the engines only through the asmspy --serve subprocess.

  • make desktop-setup — one command from a bare host to a runnable GUI. Nothing bootstrapped the desktop app: make deps covered the engines but knew nothing about GLFW or GL, so the app backends were reachable only by copying the apt line out of the dependency-gate’s guidance text. The new target runs install-deps.sh --desktop, then the pinned Capstone and Keystone source builds — not optional extras, since no Linux package manager ships either engine, so a package-manager-only setup would leave make desktop still gated — and finally builds both binaries. make desktop-setup-render does the viewer half: app backends only, no engines, no source builds (D4).

    The build step is a recursive $(MAKE), which is load-bearing rather than stylistic: DESKTOP_MISSING/DESKTOP_ENGINE_MISSING are $(shell pkg-config) probes expanded when make reads mk/desktop.mk, so a setup target that installed the dependencies and then merely depended on desktop would be judged against the pre-install answers and print the guidance text it had just made obsolete. Every step is idempotent, so re-running on a set-up host is a plain incremental build.

  • install-deps.sh learned the dependencies it was missing: --desktop (glfw + GL + unicorn + capstone + keystone + pkg-config + buildtools), --desktop-render (the app backends alone), and --buildtools (git + cmake — what the pinned source builds need in order to run, which nothing previously installed even though --asm/--emu both point at them). Package names cover apt/dnf/yum/pacman/zypper/apk/brew; gl_pkg is empty on brew because macOS OpenGL is an Xcode framework, not a package. --all and the no-flag default pick up the new dependencies; --emu/--asm/--nasm/--tidy are unchanged.

  • The three doors: Learn, Author, and a capability panel — plus runner record mode (desktop GUI plan, Phase 2; docs/internal/archive/gui/06-doors-and-learning.md). The GUI’s first-run promise is “no blank IDE”. The Learn door plays bundled walkthroughs, and a walkthrough is not a document beside a recording — it IS a recording, with ordered stop:true notes in it, so a story cannot drift away from the run it narrates. Four ship (square, demo-fail, ct_eq, and a truncated low-fidelity fixture), regenerated byte-identically by make asmtrace-walkthroughs; the truncated one’s last stop points past the recorded window and the player must say so rather than clamp. The Author door (full app only) assembles what you type, runs it, and renders faults as data — the assembler’s own diagnostic survives verbatim, because it is the one sentence that names the AT&T-under-Intel trap. The capability panel shows what this host can do and why not, straight from asmtest_trace_resolve / asmtest_hwtrace_status / both IBS reasons: a greyed row always carries its measured reason, and ticking “native only” into an empty cascade says the library returns EUNAVAIL rather than silently downgrading to the emulator.

  • ct_eq ships as a real suite (examples/ct_eq.s + test_ct_eq.c, make ct-eq-test), replacing the illustrative docs snippet: block-coverage union across secret-differing inputs, with leaky_eq as a negative control that asserts the union does grow — without it, “no new blocks” would also pass for a routine that never ran.

  • Runner record mode: --record-dir=DIR (also ASMTEST_RECORD_DIR). Every test gets a recording path, and a failing test’s TAP block and JUnit <failure> text carry recording: and step: — additive keys, so existing report parsers keep working. The runner links no engine and records nothing itself: it arms a directory and carries a producer’s note (new asmtest_record_path / asmtest_note_recording, with asmtest_rec_emu() as the emulator-tier glue), so a suite with no producer accepts the flag, writes nothing, and emits no recording: key. A directory that cannot be created is a hard exit 2 — a run asked to record that silently recorded nothing is the exact outcome the flag exists to prevent. A recording noted by a test that then crashes still reaches the report; a child that died before reporting names none.

  • The Loom — a recorded run as a spacetime fabric (desktop GUI plan, Phase 2 flagship; docs/internal/archive/gui/05-loom-day-one.md). New desktop/src/loom/ turns a value trace and its def-use graph into lanes, worldline spans, hops and knots: a register deck in the producer’s own fixed order, memory coalesced into bands on first touch, and a pure zoom-aware draw plan (spans under 3px collapse into a live-count density ribbon; a band at >=12px per byte explodes into per-byte rows). Clicking a worldline lights its whole thread-tree, [/] walk it a generation at a time, a biography narrates how the value was born, every hop it takes and where it escapes to memory, and a zeroization audit answers “who still holds a descendant at time T” — with its own title stating that “clear” means not overwritten within the traced window. The lane annex joins what companion recordings saw about the same place under a deliberately closed two-verdict enum: corroborates or unconfirmed, never “contradicts”, because a statistical feed’s silence proves nothing. Forks change exactly one fact — an entry argument or the routine’s source — and re-run from entry, rendering an interventional dim/hot/neutral verdict per step, where neutral is what you get whenever either side never captured the value. Five fidelity rules are structural and tested verbatim: a statistical producer is REFUSED with a reason and no partial fabric, an uncaptured value stays hollow, a value alive at the last recorded step gets a fade-out and never a death cap, a worldline whose first record is a read is marked born-of-untraced-state, and an unassemblable fork patch fails loudly rather than weaving code the user did not write. Four committed golden looms (including a generated truncation fixture) and ten new desktop-test binaries cover it; test_loom_parity pins the generation walk’s closure to src/dataflow.c’s slicer over 200 pseudo-random graphs.

  • Desktop replay views — trace canvas, operand timeline, slice explorer, recording diff, deep links (desktop GUI plan, Phase 2; docs/internal/archive/gui/04-replay-views.md). The desktop app now renders what a recording contains rather than just its summary: per-offset heat with a block-granular coverage gutter, the per-step operand-value timeline, the def-use slice explorer (click a step; see everything that produced the value and everything it affects), and a two-recording diff. Slicing is computed client-side from recorded df_edge events, so the engine-free asmtest-viewer slices with no engine linked — and test_slice_diff pins the viewer’s closure to src/dataflow.c’s slicer over 200 pseudo-random graphs, so a divergence between the GUI’s answer and the TUI’s fails the build. The slice layout is layered by step index and fully deterministic, never force-directed. Value annotations call the TUI’s own cli/asmspy_dataview.h helpers, so both frontends speak one dialect. Every position is addressable as asmtrace-link:v=slice&rec=...&step=4, round-trip byte-stable, and every view and keyboard binding routes through one router. Fidelity is enforced structurally: a recording mixing region-relative and absolute events draws no rows (a placard instead), a truncated recording carries a non-collapsible banner naming how much is missing, cones over a truncated stream are labelled lower bounds, a dropped step renders as unknown rather than as offset 0, statistical hot-edge data is never merged into exact heat, a refused diff produces a reason instead of plausible numbers, and a diff bounded by truncation says “no divergence observed within the recorded window” — never “identical”. make desktop-test grew ten binaries covering all of it.

  • Backend-completeness panel and its data readers (docs/internal/archive/gui/02-exporters-and-readers.md T5/T6). New desktop/src/data/ readers for the three shapes the existing producers emit (a live asmfeatures sweep, a committed benchmarks/boxes/<box>/ record, a full asmtest-bench-report/v1) plus each box’s append-only perf-history.jsonl, whose torn final line is counted rather than fatal. “Not measured” survives as std::nullopt through both spellings the producers use (JSON null and an omitted key) and renders as an em dash, never as 0. The panel shows tier x backend x arch with skip_reason rendered verbatim and truncation (trace_insns < insns_truth, or complete:false) made loud — both pinned by byte-compared golden renders.

  • asmtrace_record now emits trace events, so the golden corpus feeds the trace canvas with real recorded data. It deliberately emits no coverage event: the L0 value producer measures executed steps, not basic blocks, and block starts cannot be recovered from an offset stream without instruction lengths.

  • .asmtrace exporters — recordings open in speedscope, Perfetto, genhtml and Graphviz (desktop GUI plan, Phase 1; docs/internal/archive/gui/02-exporters-and-readers.md). New tools/asmtrace_export.c (make asmtrace-export) reads a recording and writes a speedscope evented profile (--speedscope), Chrome Trace Event JSON for Perfetto (--chrome), a block-offset lcov record (--lcov) or the --tree call graph as Graphviz DOT (--dot-tree) — so a capture made in CI or a container can be rendered later, anywhere, without re-running the traced program. One TU, libc only: no engine objects, no Capstone, no JSON library. Fidelity is enforced rather than documented: statistical survey events are never exported as stacks (exit 2, naming the reason), truncation and dropped samples surface in every mode, the time axis is labelled as the event ordinal it is — no producer records timestamps — and a mixed address basis, a newer format major or a compressed container is refused by name instead of best-efforted. make asmtrace-export-test byte-compares every mode against committed expected files and pins each refusal by exit code and by the reason it prints.

  • Desktop GUI skeleton — a Dear ImGui shell over the .asmtrace document model (desktop GUI plan, Phase 2; docs/internal/archive/gui/03-desktop-shell.md). New desktop/ tree building two binaries: asmtest-desktop, the full app, which links the Author-tier engines and so is GPL-2.0 as a whole; and asmtest-viewer, a render-only viewer with zero engine dependencies that stays permissively distributable. Dear ImGui 1.91.9 and nlohmann/json 3.11.3 are fetched pinned + digest-verified the way the native engines are (scripts/fetch-imgui.sh, scripts/fetch-json.sh), and mk/desktop.mk adds make desktop / desktop-render / desktop-test plus the docker-desktop lane. The .asmtrace loader groups events by kind, enforces the schema’s forward-compat and fidelity rules (a stream with no provenance is refused, a newer major is refused by name, and truncation / drops / redaction / a torn file each survive into the model), and the headless desktop-test drives ImGui through its null backend — no display, no GL, no engines — so it runs on any host with a C++17 compiler and opens every committed golden recording.

  • .asmtrace recordings — one NDJSON format for every headless asmspy mode (desktop GUI plan, Phase 1). --record=<file> on --log --trace --dataflow --stream --graph --tree --procs --sample --watch writes a .asmtrace recording beside the existing text/JSON output, and --json on --log / --stream streams that same format to stdout — so `asmspy –log –json

    x.asmtrace*is* a recording. Every stream carries mandatory provenance   (which backend, exact vs statistical, the measured skip reason), and   truncation, drops, throttling and redaction are fields rather than renderer   discipline: a run that SKIPS still produces a closed recording naming the   gate, and a producer killed mid-record leaves a visibly torn file. Syscall   recordings split the line from its payload — the recorded line keeps the   syscall name, fds, flag words, counts and return value while every decoded   buffer, path, sockaddr and fd-backing path becomes a placeholder — so a reader   can default-redact content without losing the call. The draft schema is  docs/internal/gui/asmtrace-schema.md; make asmtrace-goldenregenerates the   committedtests/golden-asmtrace/corpus (deterministic emulator L0   recordings, plus hand-authored truncated/dropped/redacted/torn fixtures) and  make asmtrace-golden-checkgates it byte-for-byte.vec512_tjoins the  asmtest_abi.json` manifest, closing the named-descriptor precondition for AVX-512 register state.

  • Standing self-hosted Intel PT runner — unattended nightly coverage (self-hosted-ci-runners.md T5). New scripts/runner-jit-loop.sh <owner/repo> <lane> <runner-dir>: the runbook’s production JIT/ephemeral loop as a script — mint a fresh JIT config, run exactly one job, re-register (60 s backoff on mint failure) — deployed on the bare-metal i7-8559U PT box as a systemd user unit with linger, HW_RUNNER_INTEL_PT left at 1 so the nightly hw.yml schedule runs unattended. Re-registration proven by two consecutive green dispatches on freshly minted ephemeral runner identities (run 29999081537, run 29999251602). The runbook (docs/internal/ci/runners.md) gains the unit template, the deployment record, and the standing form’s recorded posture tradeoffs; the remaining HW_RUNNER_CORESIGHT/_MACOS_TART/_KVM variables now exist at 0 per the settings checklist.

  • Self-hosted Intel PT CI lane first green run (self-hosted-ci-runners.md T5). hw.yml’s hwtrace-pt-baremetal executed live for the first time (run 29997961188, 2026-07-23) on an ephemeral pinned v2.335.1 runner registered --labels intel-pt on the bare-metal i7-8559U PT box: the broad hardware-capture tier under CAP_PERFMON (# 649 passed, 0 failed, PT tier live) plus the fail-not-skip make docker-hwtrace-pt-live (ASMTEST_REQUIRE_PT=1, 1..644, # 644 passed, 0 failed). One-shot ephemeral per the power-down rule (runner de-registered, HW_RUNNER_INTEL_PT back to 0); unattended nightly coverage still needs a standing runner. See docs/internal/ci/runners.md.

  • macOS clean-room Track D lane validated green — first shakedown complete (macos-cleanroom-lanes.md T6). make docker-osx-bindings now runs a vanilla x86-64 macOS 13.7.8 (Ventura) guest under KVM on a bare-metal Linux host, SSHes in, and runs the Track-A clean-room install test: on 2026-07-23 it exited rc=0 (stable ×2) with clean-room-test: OK on darwin-x86_64 — ruby PASS (the freshly-installed gem resolved its bundled native/darwin-x86_64/libasmtest_emu.dylib, proving no dev-build//Homebrew/ /usr/local leak) and python/node/lua/java/dotnet/hdr SKIP (toolchain-free guest, the expected shape). The one-time Ventura install was driven headless over QEMU VNC (no X11), producing a reusable prebuilt disk (build/osx/mac_hdd_ng.img, DOCKER_OSX_DISK). Getting there hardened the lane script: QEMU -display none (gtk default is fatal in an X-less container), -i on docker run (a closed stdin EOFs the -monitor stdio monitor and quits QEMU), DOCKER_OSX_CPU / DOCKER_OSX_SMP / DOCKER_OSX_VNC / DOCKER_OSX_CPUSET passthroughs (the image’s Penryn default spins newer macOS userlands — use Haswell-noTSX-IBRS), and a tree-copy exclude so the lane never tars the guest’s own disk into the guest. Hard host gate: the box’s current_clocksource must read tsc — an earlier 2026-07-22/23 attempt (Ryzen 9 4900HS) froze four times on a warped-TSC boot (clocksource demoted to hpet; macOS is TSC-only and livelocks under load there, at any vCPU count); the green run was on a healthy-TSC Ryzen 9 9950X/Zen 5 box. Two corrections vs the earlier runbook: enable Remote Login via System Settings → Sharing → Remote Login (Ventura refuses systemsetup -setremotelogin on without Full Disk Access), and each fresh-container run re-downloads the ~850 MB recovery before QEMU starts. Full evidence + runbook: docs/internal/docker-osx-linux-host.md.

  • Bare-metal Intel PT self-hosted lane + a fail-not-skip docker target, and the dark CoreSight placeholder (self-hosted-ci-runners.md T5). hw.yml gains hwtrace-pt-baremetal (runs-on: [self-hosted, linux, x64, intel-pt]) and hwtrace-coresight-board, both carrying the workflow’s required guard pair — an HW_RUNNER_* variable that is 0/absent by default and the github.actor == github.repository_owner actor guard — so they land and stay green with zero runners. The PT job’s non-vacuity check is not a skip-string grep (the PT skip text has a known cosmetic misreport, so a grep would assert on a string that lies): new make docker-hwtrace-pt-live runs the require-mode target hwtrace-pt-live (ASMTEST_REQUIRE_PT=1) in the hwtrace image under --cap-add=PERFMON with default seccomp, turning the PT tier’s availability self-skip into a hard failure — verified on a non-PT host to fail for the right reason (661 passed, 1 failed, non-zero exit, not a GenuineIntel x86-64 host). It also prints /sys/bus/event_source/devices/intel_pt/type from inside the container each run: that shakedown answer was already measured on the bare-metal i7-8559U box (PMU visible in-container; CAP_PERFMON bypasses perf_event_paranoid=4, so no --privileged and no host-native fallback) and is now recorded in the runner runbook. The CoreSight job is deliberately dark — flipping HW_RUNNER_CORESIGHT is the acceptance step of the OpenCSD decode work, not something to do early.

  • asmspy names why a hardware watchpoint/breakpoint could not arm, and the AArch64 cli CI leg is now gating (asmspy-aarch64-support.md T7). “Hardware watchpoint unavailable” was three different host facts behind one guessed message (“qemu / seccomp / permission”). Measured on the hosted ubuntu-24.04-arm runner: NT_ARM_HW_BREAK reports 6 slots and NT_ARM_HW_WATCH 4 (debug_arch=8), and PTRACE_SETREGSET on either returns ENOSPC — slots exposed, reservation refused, so nothing can arm and nothing can fire. asmspy_hwdebug_reason() now records what actually happened at the arm site, so --watch skips with “host reports 4 watchpoint slots but refused to reserve one: No space left on device” and --trace --tid adds the same note. The smoke’s --trace --tid= block takes the host-capability split the --watch block already had (named skip on AArch64, strict on x86-64) and still asserts the property that survives it: an unarmable entry is reported as an attach failure, never as a false “never executed”. A --watch success line that printed unconditionally — including on the skip path — now says which happened. With those the arm64 cli leg is green end to end and continue-on-error is removed.

  • Fixed: asmspy --stream / --graph / --tree killed every multi-threaded AArch64 target they traced (asmspy-aarch64-support.md T2). On AArch64 a single step armed on a thread parked in a blocking syscall survives PTRACE_DETACH — arm64’s ptrace_disable() sets SPSR.SS and clears only TIF_SINGLESTEP — and fires as a fatal SIGTRAP when that syscall returns, 200-400 ms after asmspy has exited. Whole-process tracing therefore left the target dying: measured on Neoverse-N2, a second trace of the same process found every thread already dead with termsig=5. The whole-process engines now resume a thread poised on an svc with PTRACE_SYSCALL instead of a step (step_resume), so nothing is armed across a call that may block and the thread is still stepped again from the syscall-exit stop; PTRACE_O_TRACESYSGOOD tags those stops on AArch64 only, leaving the x86 stop stream unchanged. The teardown drain gained the matching guard (/proc/<tid>/syscall, since a thread at a syscall-entry stop has its pc past the svc — stepping it hung asmspy unkillably) and now keeps stepping until it consumes a real trap rather than a queued PTRACE_EVENT_STOP. The detach-survival smoke assertion, which checked kill -0 immediately and so was blind to a delayed kill, now polls for ~2 s.

  • The macOS drgate gate now runs in CI — CPython signal chaining under attach is regression-protected on an independent host (macos-dynamorio-signal-chaining.md ✅ closure). The nightly drtrace-macos job (macos-15-intel) had run only the C harnesses (drtrace-test-macos + test_drtrace), so the very case whose wedge motivated the fork’s macOS sigreturn fix — test_drgate.py::test_signal_chaining, a signal delivered to a live CPython interpreter while DynamoRIO is attached — was validated only on the dev box that landed the fix. The job now installs pytest and runs make drtrace-python-test against the freshly built pinned fork (tests/test_drtrace.py 3 cases + tests/test_drgate.py 4 cases, including takeover scope, signal chaining, tracing under the managed host, and start/stop bracketing), with the lane’s standing self-skip-is-failure posture: a DR-not-found skip, a pytest skip, or a short collection fails the job rather than silently retiring the managed-host gate. gate (aarch64-sve-capture.md T8).** The SVE execution sign-off had been recorded as hardware-gated (“no SVE host in this environment”); the hosted ubuntu-24.04-arm runner is Azure Cobalt 100 / Neoverse-N2, which carries sve and sve2 in HWCAP at sve_default_vector_length = 16 bytes on Linux 6.17.0-1020-azure — so asm_call_capture_sve and its ptrue/fadd corpus body had been executing on real silicon in every CI run, not only under qemu-user TCG (ok 5 - simd.sve_adds_doubles_at_any_vl, 0 skipped; make check 57 passed / 0 failed). The test job’s arm64 leg gains an SVE silicon sign-off step that prints the silicon facts (kernel, CPU part, VL) and fails if that test ever self-skips there, so a runner-fleet change or a regressed HWCAP_SVE probe cannot silently retire the only real-silicon validation the SVE path has. Still gated, and reported as such: a native VL other than 16 B (Graviton3 at 32 B, A64FX at 64 B) — the sweep lane’s 48/128/256 B legs remain qemu-TCG emulation.

  • Self-hosted AMD Zen CI lane ran green live on real silicon (self-hosted-ci-runners.md T3). The hwtrace-privileged-zen job in .github/workflows/hw.yml runs make docker-hwtrace-privileged on a registered AMD Zen 4/5 runner and asserts the exact LbrExtV2 branch-stack + live-IBS paths RAN (a self-skip there is a hard failure). It was exercised end to end on the Ryzen 9 9950X (Zen 5) box on 2026-07-22: a pinned v2.335.1 ephemeral runner (tarball SHA-256 verified against GitHub’s published digest) picked up a workflow_dispatch and went green — # 667 passed, 0 failed, zero AMD-LBR/IBS self-skips, call_auto escalating off the LBR window (insns=77 truncated=0), assert step passing. hw.yml was additionally hardened with an explicit github.actor == github.repository_owner actor guard on the self-hosted job (on top of the existing no-push/no-pull_request posture), so the lane runs only for a trigger the repo owner initiated. The hosted hwtrace-privileged bitrot-gate comment in ci.yml was corrected: the call_auto non-escalation finding is FIXED (5d8e0d2) rather than open, and the self-hosted counterpart is no longer “future” — it now exists in hw.yml. See the runbook docs/internal/ci/runners.md (standing unattended-nightly coverage needs a persistent runner via the JIT/ephemeral loop).

  • macOS x86-64 DynamoRIO native-trace tier — M0 is GO (macos-dynamorio-fork-build.md FB1–FB3 driving macos-dynamorio-port.md T3–T5). DynamoRIO publishes no macOS release asset (0 across all 455 releases), so the macOS tier now builds the runtime from a git-commit-pinned source fork: scripts/build-dynamorio-macos.sh + make dynamorio-macos produce lib64/release/libdynamorio.dylib (plus the drmgr/drreg/drx extension set the drclient sub-build links), pinned in scripts/third-party-digests.txt with the license vendored, Darwin-x86-64-gated (clean skip everywhere else). Three root-caused fixes in the fork (wilvk/dynamorio, branch asmtest/macos-fixes) make the dr_app_* embedding functional on modern macOS: the dr_app_setup startup fault (exec path read from KERN_PROCARGS2 instead of the above-envp walk; baseline was 10/10 SIGSEGV), the invisible environment in dylib embeddings (dyld passes no envp to initializers, so DYNAMORIO_OPTIONS was never read and no client ever loaded; now captured live via _NSGetEnviron, the STATIC_LIBRARY approach), and dr_get_proc_address returning NULL on every modern binary (LC_DYLD_EXPORTS_TRIE unparsed, plus a trie rebase that was only correct for preferred-base-0 dylibs — never the main executable). On the asm-test side, libdynamorio resolution is dylib-aware (DR_LIBNAME), the DR make tier resolves .dylib client names on Darwin, and the new make drtrace-test-macos M0 harness traces a normally-compiled __TEXT function end to end — attach, Mach-O marker resolution, coverage accumulation, symbol mode, clean detach — 13/13 twice in a row on the macOS-14.7.5/Intel host, against the pinned script-built home. M1a (macos-dynamorio-port.md T6): the same target now also runs the generated-bytes test_drtrace harness on Darwin — the PROT_NONE → RW → RX exec_alloc W^X path, coverage accumulation, the exact instruction-mode offset stream, and the truncation bit all pass on native Intel silicon (18/18 on the same host; Rosetta stays must-verify, gated on Apple Silicon hardware). M1b groundwork (T7/T8): executable-memory allocation now sits behind a platform seam whose arm64-macOS arm uses MAP_JIT + per-thread pthread_jit_write_protect_np (byte-identical elsewhere; asmtest_asm_exec_native returns ENOSYS on non-x86-64 hosts), and on Darwin the test_drtrace harness is ad-hoc signed with the new drtrace.entitlements (com.apple.security.cs.allow-jit; never --options runtime, whose one-MAP_JIT-region limit the harness’s three allocations would trip). arm64 acceptance itself stays gated on Apple Silicon hardware plus an arm64 DR runtime (upstream i#5383); the guide’s new “macOS arm64” section records the hardened-interpreter limitation as a property of the OS model. M2 bindings (T9): the binding lanes resolve platform-correct library names on Darwin (drtrace_env → .dylib + DYLD_LIBRARY_PATH), with drtrace-cpp-test, drtrace-ruby-test, and drtrace-python-test green on macOS x86-64 — compiled-function/symbol mode is the documented macOS binding path. Signal chaining under an attached trace now works on macOS (a fourth fork fix): a signal raised while attached is delivered to the host runtime’s handler and the process resumes cleanly, so the Python test_drgate.py::test_signal_chaining gate runs on Darwin (previously Darwin-skipped). Root cause (confirmed single-threaded, not the multi-thread i#58 first suspected): the app handler is delivered fine, but its return through libsystem _sigtramp → macOS 3-arg sigreturn(uctx, infostyle, token) needs a per-delivery kernel token DR cannot forge for the frames it synthesizes, so the real sigreturn was rejected and the thread resumed into DR gencode → ud2 → SIGILL → terminate. The fork (pin b8785a5d8) now restores the app context in handle_sigreturn and skips the real sigreturn on macOS x86-64, exactly as the VMX86 path does. Fork api.startstop/api.detach run 10/10 crash-free (their multi-thread takeover assertions are upstream-NYI on macOS, i#58, and upstream macOS CI never runs them — they are outside the OSX ctest label set). arm64 stays gated on the upstream arm64 port (i#5383). A nightly/dispatch-only drtrace-macos CI job on macos-15-intel builds the pinned fork from source (commit-stamp cached) and runs the M0 harness, outside the test matrix so a DR failure never blocks the emulator tier — with a self-skip-is-failure guard, since the runner that just built the runtime must actually test it.

  • AArch64 out-of-process single-step stream validated live on real silicon (aarch64-ptrace-single-step-validation.md T1–T6). The out-of-process ptrace tracer’s AArch64 arm — written and decode/execute-validated under qemu, but whose live capture had never run — now runs LIVE on GitHub’s ubuntu-24.04-arm hosted runners (Azure Cobalt 100 / Neoverse-N2 VMs) via a gating hwtrace-arm64 CI job: 205 ok assertions, the whole ptrace tier (trace_call / trace_attached / run_to software-brk plant / call-out step-over) exercised on real hardware, with an anti-vacuity step that fails the build if the tier self-skips. A companion hwtrace-bindings-arm64 job runs the Go and Java wrappers’ AArch64 ptrace fixtures live (the other wrappers gate their whole hwtrace suite on the x86-only in-process single-step backend and self-skip), plus a live Python host step. The first arm64 benchmarks/boxes/ record (arm-linux-arm64-gha) is committed with a live native-oop capability row. Measured answer to the previously-undetermined hardware-breakpoint question: NT_ARM_HW_BREAK is armed-but-silent on this hypervisor (slots reported, arm accepted, the debug exception withheld), so the forced-hardware run_to self-skips there with that named reason and bare-metal arm64 hw-breakpoint firing moves to self-hosted-runner territory. qemu-user keeps self-skipping transparently. Also brought the asmspy CLI toward AArch64: fixed the SYS_stat/dup2 legacy-syscall gaps (arm64 uses newfstatat/dup3 — the latter now decoded), the host-arch disassembly of traced code, and R_AARCH64_JUMP_SLOT / 2-slot-PLT0 stub resolution (asmspy-aarch64-support.md T2/T7).

  • F5 live foreign-pid PT replay wired to the landed capture (dataflow-pt-replay-tier.md T4). The out-of-band data-flow value tier’s live case now CONSUMES the asmtest_hwtrace_pt_attach_* foreign-pid capture that landed with intel-pt-attach-foreign-pid: examples/test_dataflow_pt.c forks a deterministic victim, captures ONE in-region invocation over Intel PT with zero single-steps of the target, replays the decoded offset path through F5, and asserts the value trace matches both the emulator L0 and the force_singlestep block-step oracle’s executed path + result. It was previously a dead-code stub behind a never-defined macro; it is now a runtime-probed body (asmtest_hwtrace_available(ASMTEST_HWTRACE_INTEL_PT)) that self-skips off Intel PT and, under ASMTEST_REQUIRE_PT=1 (make dataflow-pt-live), fails rather than skips. Gated on bare-metal Intel PT silicon for the live oracle match; make docker-dataflow-pt runs the synthetic decode→replay bridge + both block-step-layout guards green (19/19) with the live half self-skipping.

  • Java binding publishable to Maven Central (distribution-packaging.md T6). make java-package now runs a real mvn package against bindings/java/pom.xml (the same POM Central publishes) instead of raw javac + jar cf, emitting the binding jar plus matching -sources/-javadoc jars. A new tag-gated, secret-guarded maven job in release.yml mvn deploys them — GPG-signed — to the Central Portal staging, no-opping unless both MAVEN_CENTRAL_TOKEN and MAVEN_GPG_KEY are set (mirroring the crates/npm publishes). Maven is a version-pinned Apache tarball in the java docker image; make docker-java-package proves the whole build locally with no credentials, and docs/reference/releasing.md carries the Central + LuaRocks runbooks.

  • Cross-system benchmark legs on real Windows and Intel macOS. The deterministic golden gate now runs on every OS the framework targets: a per-push benchmarks-windows leg (windows-latest, mingw/MSYS2) builds the PE benchmark producers and runs make win64-bench-check + win64-bench-report on a genuine Windows kernel, and a nightly benchmarks-macos-x86 leg (macos-15-intel) produces the Intel-macOS report. benchmarks-compare merges all five OS × arch reports.

  • Nightly auto-commit of per-box benchmark records. On the nightly schedule (and manual dispatch) each benchmark leg records its per-box history and a new benchmarks-record job commits benchmarks/boxes/gh-** back to main as github-actions[bot] (a GITHUB_TOKEN push, so it never re-triggers CI). Golden emu counts stay human-reviewed — never auto-committed — so a real count drift fails a leg’s bench-check instead of being laundered into history. New make win64-bench-record persists the Windows box record.

  • Ambient stitched operations (.NET, opt-in, Intel PT). AsmAmbientStitchedTrace follows one logical operation across await / thread hops with zero calls in the body — an AsyncLocal value-changed handler opens a per-thread intel_pt slice when the flow lands on a thread and decodes-at-disable when it leaves, then stitches the slices at close (op.Hops / op.Path). PT-only by construction; off Intel PT it self-skips and runs the body uninstrumented. Built on the new per-tid PT hop capture primitive (asmtest_hwtrace_pt_hop_open/_hop_close) on the one shared intel_pt perf-AUX arm. The default whole-window capture is unchanged (it stays the faithful per-thread window that flags truncated on a cross-thread hop — never auto-stitch).

  • Runtime-enabled jitdump byte recovery on an already-running process — turn jitdump emission on in a live foreign runtime and recover a JIT method’s recorded bytes with asmtest_jitdump_find, with no launch flag on the target. For CoreCLR (make docker-hwtrace-jit-dotnet-attach-jitdump) the harness sends EnablePerfMap(All) over the runtime’s diagnostics IPC socket (the documented DOTNET_IPC_V1 wire, hand-rolled in C — no NuGet), so an already-JITted Program::Add is rundown-emitted into /tmp/jit-<pid>.dump even though the victim was launched without DOTNET_PerfMapEnabled. For HotSpot (make docker-hwtrace-jit-java-attach-jitdump) it loads a new in-tree attach-capable JVMTI jitdump agent (examples/jvmti_jitdump_agent.c, test-support only, never shipped) into the running JVM with jcmd JVMTI.agent_load; on attach the agent replays every already-compiled method via GenerateEvents(COMPILED_METHOD_LOAD) into the dump. Correction: the linux-tools libperf-jvmti.so exports only Agent_OnLoad (no Agent_OnAttach), so HotSpot refuses to load it via jcmd — it is -agentpath-only and cannot serve the attach case; hence the bespoke agent. V8/Node has no runtime-enable path (--perf-prof is wired once at isolate init), so there is deliberately no Node attach lane. Both lanes run on any host with the runtime — no Intel PT / hardware gate.

  • Opt-in safe-managed whole-window policy. ASMTEST_WHOLEWINDOW_SAFE_MANAGED=1 makes asmtest_hwtrace_begin_window refuse an in-process EFLAGS.TF whole-window arm when a managed runtime (CoreCLR / JVM / Mono) lives in the process, returning the new distinct ASMTEST_HW_EMANAGED status instead of single-stepping code whose SIGTRAP disposition the runtime’s signal layer owns. The .NET empty-ctor new AsmTrace() builds on it to route a managed window to Intel PT where the silicon exists, else the §D3 out-of-process stepper, else a transparent self-skip — never in-process TF (ww.Route reports the chosen route). Default (env unset / no safeManaged) is byte-identical to today; asmtest_hwtrace_managed_runtime_present() exposes the probe.

  • Native trace-point → IL / bytecode / source-line attribution. A captured native offset of a managed method now resolves to a source line, a .NET IL offset, or a JVM bytecode index — not just a method name — from feeds the runtimes already emit. asmtest_jitdump_debug_find (include/asmtest_ptrace.h) recovers the per-address (line, column, file) table from a jitdump’s JIT_CODE_DEBUG_INFO records the byte reader skips, and asmtest_jitdump_debug_line_map bridges it into the shipped emu_line_map_t (works today for V8 node --perf-prof). A widened, backend-neutral schema (asmtest_srcmap_*, asmtest_srcreg_*, include/asmtest_trace.h) carries offset → {kind, value, file, column} with enclosing-point lookup and a version-keyed registry stamped on the code-image capture sequence, so a method re-JIT’d at a reused address resolves against the body live when the trace ran. For CoreCLR, a new IlToNativeMap EventListener (bindings/dotnet/hwtrace/HwTrace.cs) subscribes the MethodILToNativeMap JIT keyword (0x20000) in-process — no launch knob, no Intel PT — and resolves an address to (method, nativeOffset, ilOffset). For HotSpot, a new make docker-hwtrace-jit-java-bci lane loads an in-tree JVMTI agent (examples/jvmti_bci_agent.c, test-support only, never shipped) that captures CompiledMethodLoad’s address→bytecode-index map into an asmtest_srcreg and proves a native address resolves to a real bytecode index. The V8/HotSpot/CoreCLR jitdump debug lanes (make docker-hwtrace-jit-jitdump, -jit-java-jitdump, -jit-dotnet-jitdump) print the per-method attribution table and CHECK the reader against each real encoder. Attribution covers JIT-compiled code only: interpreted code keeps its bytecode index in VM state the native PC stream cannot see — see il-bytecode-attribution.md.

  • libdft64 differential taint oracle (make docker-taint-oracle). The shipped DynamoRIO in-band taint client had only an offline emulator/Capstone forward-slice as an independent check. This lane cross-validates it against a second, live, independently-implemented byte-level taint engine — libdft64 on Intel Pin — by running the same seed/sink fixtures (examples/taint_fixtures.h) through both and asserting byte-for-byte sink agreement on the general-purpose / integer-memory subset both cover (branch-condition, call-argument, and mem-copy-length sinks), with passing negative controls. A pinned, digest-gated Pin 3.20 kit (libdft64’s only tested pin) + git-commit-pinned libdft64 are fetched at build time and are test/oracle-only — never linked into libasmtest or any shipped binding. libdft64’s documented blind spots (SIMD is basic SSE/AVX, rules unverified; no eflags; no ZMM; no implicit flow; no x87/ternary) are enumerated as named skips, never a blanket pass — see data-flow-capture.md (even libdft punts on SIMD). The lane is CI-gated (the taint-oracle job) and self-skips only on non-x86 hosts (Pin is x86-64 gcc-linux).

  • Intel SDE future/absent-ISA test lane (make docker-sde / make sde-test SDE_HOME=$(scripts/fetch-sde.sh)). Assembly that uses an ISA extension the host CPU lacks — APX’s r16-r31, AVX10.2, AMX, or AVX-512 on an AVX2-only box — was untestable by this framework (the DynamoRIO tier runs on real silicon; the Unicorn tier’s vendored QEMU 5.0.1 predates AVX TCG). A pinned, digest-gated Intel SDE 10.8.0 (scripts/fetch-sde.sh + Dockerfile.sde, with APX-capable GAS from binutils 2.46.1 and NASM 3.02, all SHA-256-pinned) emulates those extensions for the whole process, so an unmodified suite binary runs under sde64 -future and gets the full register/flag/memory/ABI assertion battery on any x86-64 host, including CI runners. The lane proves SDE is byte-for-byte transparent to correct baseline code (native vs SDE TAP identical), adds an APX fixture suite (examples/apx_basic.s + test_apx_basic, gated on a new asmtest_cpu_has_apx() CPUID probe) that skips on real pre-APX silicon and runs green under emulation, asserts the AVX-512-on-AVX2 un-skip (an existing test_simd capability skip becomes a real execution under -future), cross-checks SDE against the Unicorn tier on overlapping baseline ISA, and offers an optional -mix instruction-mix report (make sde-mix) folded into the canonical asmtest_trace_t shape. SDE is proprietary freeware, fetched + digest-verified at build/test time and never bundled into a shipped artifact — test-lane only. Documented in the SDE testing guide and the implementation doc.

  • Live whole-window compose lane (.NET). The hwtrace-dotnet self-suite gains checks proving the zero-config compose seam (MethodLoadVerbose → codeimage_track → close-time versioned decode): over the in-process WEAK single-step tier (new AsmTrace()) against a managed method whose first JIT happens inside the window (genuinely-compiled-in-window, not a pre-warmed body); over the crash-proof §D3 out-of-process inline using-scope (new AsmTrace(outOfProcess: true)) against a resident method — the first suite coverage of that bare inline OOP whole-window ctor; and (self-skipping off Intel-PT silicon) the STRONG PT tier — plus a mid-window re-tier decode-at-version check. Runs as the named make docker-hwtrace-dotnet-unwarmed lane. Tasks T1–T5 of the managed-wholewindow-compose implementation doc. Two re-verification findings refine the doc’s original premise: (1) forcing a stop-the-world GC.Collect(0) on the single-stepped thread inside the window is intermittently fatal — the in-process EFLAGS.TF window dies (SIGTRAP) if the runtime spawns a thread in-window, per the repo’s own degradation note — so it is omitted (the unwarmed JIT itself supplies the live-window instruction noise); (2) the inline OOP ctor’s region-free window_stop stepper single-steps everything, so a first-call JIT inside that window steps the whole compiler and aborts CoreCLR (exit 134) — the unwarmed mid-window-JIT compose is therefore proven on the range-based AsmTrace.Window factory (whose stepper runs the JIT at native speed), while the inline OOP ctor is for resident (warm) code, matching the crashproof-showdown example.

  • Data-flow F5: an out-of-band PT + code-image + Unicorn-replay value producer (make dataflow-pt-test / make docker-dataflow-pt). The least-perturbing L0 value tier: it reconstructs an Intel PT trace’s executed instruction stream (captured with zero single-steps), supplies the bytes live at trace time from the code-image recorder, and replays that exact path through Unicorn to derive per-instruction values into the same asmtest_valtrace_t the shared def-use (L1) and slice (L2) analysis consume — byte-identical to the emulator L0 oracle on a deterministic region. src/dataflow_pt.c opens no perf event (it consumes a captured AUX blob + code-image); it reuses the block-step tier’s purity/replayability verdicts and truncates faithfully on an impure, VEX/EVEX, or nondeterministic region (no single-step fallback), a per-step path cross-check catching a divergence. The synthetic-AUX decode→rebase→materialize→replay bridge is validated in CI with no PT hardware (libipt’s own encoder, libipt-dev added to Dockerfile.dataflow-attach); live foreign-pid capture is silicon-gated (bare-metal Intel PT + the intel-pt-attach-foreign-pid capture arm) — wiring-complete, hardware-unvalidated, with a make dataflow-pt-live fail-not-skip target for a runner that claims PT. Documented in the native-tracing guide and the F5 implementation doc.

  • Whole-window in-process guards for the zero-config scope (deny regions / instruction budget / wall-clock watchdog). The using (new AsmTrace()) region-free whole-window scope, powered by the in-process EFLAGS.TF single-step tier, took the descent tier’s “step into everything” semantics without any of its safety guards: a window over code that reached a blocking libc call (a read on an empty pipe, a poll) stepped the runtime forever, bounded only by the capture ring’s memory, never by time. The three out-of-process descent guards are now ported onto that in-process path — a per-frame instruction budget (default 4x the ring cap), a process-global deny-region table with an opt-in blocking-libc default set (a stepped RIP inside a denied region ends the capture and the denied call then runs at native speed), and a repeating ITIMER_REAL/SIGALRM watchdog (default 10 s; breaks a blocked syscall via EINTR) — all malloc/lock-free inside the SIGTRAP handler. Plain asmtest_hwtrace_begin_window gets safe defaults (budget + watchdog on, denylist off), so a hung zero-config window is always bounded; the new additive asmtest_hwtrace_begin_window_ex (with the F27/F36 struct_size idiom and an asmtest_hwtrace_window_guards_t config) configures them per window, and asmtest_hwtrace_window_guard reports which guard fired (ASMTEST_HW_GUARD_*) render-on-close style. C-level knobs (parity-exempted); the zero-config defaults flow through the begin_window every binding already wraps.

  • Zero-config whole-window scope docs. The hardware-tracing guide gains a “zero-config whole-window scope (region-free)” section — the C begin_window / end_window / render_window surface, the .NET using (new AsmTrace()) form, the WEAK / STRONG / CEILING tier ladder (single-step / Intel PT / AMD LBR Zen 4+), and how the STRONG-tier PT decode is validated on a synthetic PT-packet fixture with no silicon. The troubleshooting reference gains a SkipReason table and a one-time-provisioning table (perf_event_paranoid / setcap cap_perfmon / --cap-add, ptrace CAP_SYS_PTRACE, eBPF CAP_BPF), and both it and portability record the whole-window facility’s Linux-only floor.

  • asmspy runs on AArch64 Linux. The out-of-process tracer’s register / single-step / detach reads are lifted behind an architecture shim (cli/asmspy_arch.h: PC / return / SP / LR / syscall-number accessors over PTRACE_GETREGS on x86-64 and PTRACE_GETREGSET(NT_PRSTATUS) on AArch64), so every engine — --stream, --graph, --tree, --region, --log, --procs — runs on both arches. The single-step teardown honours AArch64’s kernel-owned step model (no user trap flag; svc #0 syscall-instruction guard), the call-graph / call-tree frame logic uses AArch64 bl-writes-LR semantics (frame identity (entry_lr, sp)), and --log decodes the AArch64 syscall ABI (number in x8, args in x0-x5, *at-only name table). --watch gains an AArch64 arm over the NT_ARM_HW_WATCH regset (DBGWCR/DBGWVR/BAS encoding, pinned by a pure cli/test_arch.c unit test on every host); it self-skips where the host exposes no watchpoint slots (qemu-user, some hypervisors). Validated on the native ubuntu-24.04-arm CI runner (a real VM, not qemu) alongside the existing x86-64 leg; see the asmspy guide.

  • libFuzzer / AFL++ external-engine fuzzing shim (make docker-fuzz). Drive an x86-64 guest routine under the emulator with an industrial fuzzer, feeding the emulator’s basic-block coverage into the engine’s feedback channel without compiler-instrumenting the guest bytes (they run under Unicorn). A new tested seam emu_cover_hits (src/fuzz.c) reports one input’s distinct executed block offsets; examples/fuzz_libfuzzer.c registers them as SanitizerCoverage 8-bit counters (+ the PC table clang-18’s libFuzzer requires), and examples/fuzz_afl.c (native persistent-mode forkserver) plus an aflpp_driver reuse of the libFuzzer harness write them into AFL++’s shared-memory bitmap via a plain-compiled helper (examples/fuzz_afl_map.c). Dockerfile.fuzz (clang 18 + afl++ 4.09c on the bindings base) runs make fuzz-shim-test, which fails unless both engines steer to a planted crash — a real test, never a self-skip. Node (per-block) coverage; documented in the fuzzing-shim guide. src/capture.s; ret stubs elsewhere, incl. the NASM twin) marshals scalable-vector z0..z7 and predicate p0..p3 arguments per AAPCS64 and captures the whole z0..z31 / p0..p15 file into two new max-size containers — svec_t (256-byte VLmax) and spred_t (32-byte PLmax), of which only the low asmtest_sve_vl() (resp. /8) live bytes are written. ASM_SVCALL_1/_2 call and self-skip without SVE, and ASSERT_SVEC_EQ/ASSERT_SPRED_EQ compare exactly the live VL. A HWCAP_SVE runtime probe (asmtest_cpu_has_sve / asmtest_sve_vl) returns 0 everywhere SVE is absent (x86-64, and macOS arm64 — Apple silicon has no non-streaming SVE); the ABI manifest pins both containers, and the corpus routine sve_addd exercises the path. The new make docker-sve-sweep lane runs the SIMD suite under qemu-user at several vector lengths (VQ 1/3/8/16 → VL 16/48/128/256 bytes, including a non-power-of-two) to flush out VL-assumption bugs before any SVE silicon is available; execution sign-off on real SVE hardware (Graviton3/Grace/A64FX-class) remains pending — see aarch64-sve-capture.md.

  • XED-decoded Intel Pin trace lane (make docker-pintool / make pintool-test). A pinned, digest-gated Intel Pin 4.2 kit (scripts/fetch-pin.sh, SHA-256 pinned in scripts/third-party-digests.txt) drives a Pintool (pintool/asmtest_pintool.cpp) that fills the shared asmtest_trace_t offset model over POSIX shared memory. The lane asserts byte-for-byte instruction/block offset parity with both the in-process single-step backend and the DynamoRIO backend (Pin ≡ DynamoRIO ≡ single-step), and carries an Intel APX (EGPR/REX2) fixture whose bytes Pin’s XED decodes on any x86-64 host while the pinned DynamoRIO decoder rejects them — the decoder-currency gap the tier exists to close (DR #6226 is open; the APX execution halves are gated on APX silicon). Test-lane only: Pin is digest-verified at build/test time and never bundled into a shipped package.

  • Intel Pin probe-mode argument/return capture lane (make pin-probe-test, in make docker-pintool). A Pin probe-mode tool (pintool/probe_capture.cpp) splices a jump at a named routine’s entry/exit and records the SysV integer/FP argument registers, the return register(s) + flags, and up to a 4 KiB cap of a pointed-to buffer — at native speed (no code cache) — into at_val_rec_t records over a POSIX shm channel (include/asmtest_valtrace_shm.h). A pointer is validated against the target’s mapped ranges and the read clamped to the cap and the mapping end, so an invalid pointer is refused, never faulted; a routine too short or non-relocatable to probe is reported as an explicit per-target skip with a reason. The capture is proven by diffing it against the independent out-of-process ptrace stepper on the same routine (two producers agree on the arg + return registers). Captured buffers may contain secrets — a sensitive artifact. x86-64 Linux, test/oracle only (Pin is proprietary freeware, digest-verified at test time and never bundled), the same handling DynamoRIO gets; documented in the data-flow tracing guide.

  • Real object identity for managed memory def-use on the live-attach tier. A heap snapshot of {Address, Size, TypeID} nodes from the runtime’s GCBulkNode / GCBulkEdge / GCBulkType events, joined with the MovedReferences2 move feed, keys each captured memory record on (object, offset) where the snapshot has evidence and degrades to the landed address identity where it does not (asmtest_objid_canonicalize, src/dataflow_objid.c; unit suite test_dataflow_objid). On a live attach (make docker-gccanon-attach, new alias phase) the false def-use edge that address identity forges when a GC slides a live object onto a dead object’s vacated slot is reproduced under address identity and then eliminated by object identity.

  • Weighted cross-ISA cost proxy BM_MODEL_COST in the cross-system benchmark. emu-bench now emits a model_cost row beside each deterministic insns row: each executed instruction is classified with Capstone (new asmtest_disas_class → OTHER/MEM/BRANCH/MULDIV) and summed against a fixed weight table (1/3/2/8), a faithful cross-architecture cost model — comparable by construction, not silicon cycles. bench-compare renders it as its own Model cost matrix (never mixed with raw counts or real cycles). The metric needs Capstone: without it the bench emits the insns rows alone and every gate still passes. Model values depend on the Capstone version, so they are kept out of the golden file — bench-golden-check filters model_cost rows (and fails loudly if any insns row is missing its model sibling). See cross-system benchmarking.

  • Native RISC-V (rv64) host tier — the capture framework now runs on a RISC-V machine, not just as an emulator guest. A regs_t branch and trampolines (src/capture.s) for the RV64GC / LP64D psABI: a0/a1 return pair, integer callee-saved s0–s11, and FP callee-saved fs0–fs11 (checked by ASSERT_ABI_PRESERVED / ASSERT_ABI_PRESERVED_VEC after an _fp/_fp_n capture, since rv64gc has no vector file), plus rv64 bodies for every example suite and the framework self-tests. Two ISA facts are surfaced faithfully rather than faked: RISC-V has no condition-flags register, so ASMTEST_NO_FLAGS is set and ASSERT_FLAG_* is a compile error on rv64 (flag-only suites self-skip with a printed reason); and there is no 128-bit vector capture (ASM_VCALL* self-skips via asmtest_cpu_has_vec128() — RVV is a possible future arm). A make docker-riscv64 lane builds a linux/riscv64 image and runs the core suites + self-tests under QEMU binfmt (make binfmt-riscv64, pinned tonistiigi/binfmt), wired into CI as the test-riscv64 job. The tracing tiers stay x86-64/AArch64. See riscv-native-tier.md.

  • Live Intel PT whole-window smoke + a make hwtrace-pt-live lane that FAILS rather than skips where PT is claimed. test_pt_live_selfjit (examples/test_hwtrace.c) self-JITs the canonical routine, arms the region-free PT capture, decodes the REAL AUX stream through asmtest_pt_decode_window, and exercises PERF_EVENT_IOC_SET_FILTER (via the new asmtest_hwtrace_pt_set_filter knob), the anonymous-JIT decode-time fallback, and AUX-ring truncation on a 4 KiB ring. Off the intel_pt PMU (AMD/VMs/ containers) it self-skips with the specific reason — one of the two legitimate hardware gates — while make hwtrace-pt-live sets ASMTEST_REQUIRE_PT=1 to convert that skip into a build failure on a runner that is supposed to expose bare-metal Intel PT. The live capture is silicon-gated (no intel_pt on the reachable dev boxes); in make docker-hwtrace the test prints a clean # SKIP pt live: … everywhere. See intel-pt-whole-window-substrate.md.

  • System-package specs for the C core — Homebrew, Debian, AUR, vcpkg and Conan — each built, installed and consumed in a Docker CI lane. packaging/ now holds a Homebrew formula, a Debian libasmtest-dev source package, an AUR PKGBUILD + .SRCINFO, a vcpkg overlay port and a Conan 2 recipe for the MIT static core (lib

    • headers + asmtest.pc; the GPL engines stay in the dlopen binding packages only, never here). make docker-syspkg runs all five lanes and an additive syspkg CI job runs them as a matrix — each builds the package, runs its native linter (brew audit/style, lintian, namcap, vcpkg post-build validation), installs it, and compiles a pkg-config/CMake consumer against it. The lanes are hermetic on the reproducible make package-source tarball; per-manager index publication is a maintainer step (runbook in releasing.md). See distribution-packaging.md.

  • .NET inline using (new AsmTrace(HwBackend.IntelPt)) now arms the STRONG whole-window Intel PT capture (bindings/dotnet/hwtrace/HwTrace.cs), replacing the reserved "forward-look (not wired)" self-skip. The backend-keyed ctor gates on HwTrace.Available(IntelPt) (self-skip names the PT gate off bare-metal Intel PT), sets up the JitMethodMap + perf-map rundown before arming, and drives the native asmtest_hwtrace_pt_begin_window/_end_window pair via a finalizable PtWindowCtx (a leaked scope’s fd + AUX mappings are reclaimed drain-less by the finalizer) and a new Kind.PtWindow Dispose that decodes on close against the map’s code-image, filling Addresses with ABSOLUTE addresses and IsStatistical == false. hwtrace-dotnet-test carries the invariant-envelope case on any host; live capture is silicon-gated (make hwtrace-pt-live + make hwtrace-dotnet-test on a bare-metal Intel PT box). Only HwBackend.CoreSight keeps the forward-look self-skip. See intel-pt-whole-window-substrate.md.

  • Whole-window Intel PT STRONG tier wired behind the empty-ctor scope, with a runtime WEAK/STRONG decode-trust ladder. asmtest_hwtrace_begin_window/_end_window (src/hwtrace.c) now arm and drain a real region-free Intel PT capture on an inited INTEL_PT tier — the ONE shared perf-AUX arm (pt_aux_open) that also serves the region path, exposed as the native asmtest_hwtrace_pt_begin_window/_end_window pair the .NET inline ctor uses. asmtest_hwtrace_window_auto auto-selects STRONG only when the intel_pt PMU is present and asmtest_hwtrace_pt_window_trusted() proves the whole-window decode on the §Z2 synthetic fixture at runtime; else the WEAK single-step tier. The CEILING AMD LBR tier is deliberately never auto-selected for the exact whole-window contract (a sampled branch survey cannot meet it; live floor Zen 4+) — the quiet sampled complement stays explicit (new AsmTrace(HwBackend.AmdLbr)). The .NET empty-ctor using (new AsmTrace()) consults the ladder (AutoInitWindowBackend), and DegradationNote() names the PT probe outcome (present-but-untrusted vs no PMU). Live PT capture is silicon-gated; the ladder, the runtime trust probe, and the synthetic-fixture decode all run in make docker-hwtrace (green on this AMD host — no intel_pt PMU, so the ladder resolves to WEAK and the native PT pair self-skips). See intel-pt-whole-window-substrate.md.

  • Reproducible source tarball (make package-source) attached to every release. git archive of HEAD piped through gzip -n emits build/dist/asm-test-<version>.tar.gz + SHA256SUMS, byte-identical for a given commit across machines, so its digest is known ahead of the tag. The release.yml corresponding-source job builds it, uploads it as a dry-run artifact, and attaches both files to the tagged GitHub release — the digest-pinned source the system-package specs (Homebrew/Debian/AUR/vcpkg/conan) and Debian’s orig-tarball flow consume. See distribution-packaging.md.

  • Socket-syscall sockaddr contents are decoded (asmspy --log). connect/bind/ sendto render their struct sockaddr * as {AF_INET, 127.0.0.1:8080} / {AF_INET6, [::1]:80} / {AF_UNIX, "/path"} (abstract sockets as "@name"), and accept/accept4/recvfrom decode their OUT pointer on success (raw pointer on failure); socket()’s domain renders as AF_INET/AF_UNIX/… An unknown family prints {family=N, len=M}, never a guessed name. make docker-cli cli-smoke PASS. See asmspy-cli-enhancements.md.

  • ioctl requests and fcntl commands are named (asmspy --log). ioctl renders TIOCGWINSZ-style names, or a faithful _IOC(dir, type, nr, size) decomposition for an unknown request (never a guessed name); fcntl renders F_GETFL/F_SETFD/… with correct conditional arity (an arg-less command such as F_GETFL shows no third slot). make docker-cli cli-smoke PASS. See asmspy-cli-enhancements.md.

  • futex operations are named (asmspy --log). The op renders as FUTEX_WAIT/FUTEX_WAKE_PRIVATE/… with FUTEX_PRIVATE_FLAG and FUTEX_CLOCK_REALTIME masked off before naming and re-rendered as suffixes, never silently dropped; an unknown op keeps its number plus those suffixes. make docker-cli cli-smoke PASS. See asmspy-cli-enhancements.md.

  • stat/statx result buffers are decoded (asmspy --log). fstat/stat/lstat/newfstatat render {st_mode=S_IFREG|0644, st_size=18} on success (a raw pointer on failure), and statx renders its mask-honoring {stx_mode=…, stx_size=…} (a field the kernel did not fill is omitted, not invented). The path decode these calls already had is preserved. make docker-cli cli-smoke PASS. See asmspy-cli-enhancements.md.

  • Hot-edges → data-flow drill-in (asmspy TUI, mode 7). In the frozen hot-edges view, arrows select an edge and Enter opens a data-flow capture (mode 9) of the function containing the edge’s to_addr (falling back to from_addr), reusing the call-graph drill-in idiom. The pure decision (asmspy_edge_drill) requires a sized function and accepts a mid-function landing (drill ≠ rank); it is unit-tested in test_autoregion (6 checks) so it is covered on every host, not just an AMD IBS box. The ncurses wiring is pty-driven (manual-only); the decision logic runs in CI. See asmspy-cli-enhancements.md.

  • Block-step replay record-and-inject for rdtsc/rdtscp/rdrand/rdseed/cpuid, gated per block rather than per region. src/dataflow_blockstep.c’s step_block now injects each site’s recorded post-state (read from the T5 DR exec-breakpoint boundary) into the Unicorn replay and terminates the block there — the same record-and-inject shape as syscall/int 0x80, minus the producer-local write-set synthesis (Capstone already reports the complete architectural write set for all five mnemonics). region_scan’s injectable verdict now admits regions whose only impurities are syscall/int80/HWREC (subject to the existing 4-slot DR0-3 cap — a 5th+ distinct site still falls back to single-step, reason hwrec-overflow), and a region no longer forfeits the replay’s perturbation win for a hwrec site its real run never reaches (hw_hits/injected stay 0 for an unexecuted site while the region still replays). New opts test hook no_hw_record skips arming the DR breakpoints, reproducing the pre-injection fail-closed truncation on demand. make dataflow-blockstep-test 191/191 (was 186/186, +5), stable across 5 consecutive runs on this Zen 2 host. See dataflow-producer-correctness.md.

  • Extents-driven block-step region scan: a caller-vouched list of real instruction extents lets region_scan skip an embedded constant-pool island instead of desyncing on it (BSVS-2). New asmtest_blockstep_extent_t ({off, len}, blob-absolute) plus opts fields extents/nextents (src/dataflow_blockstep.c) — NULL/0 keeps today’s whole-region sweep. region_scan is split into a per-extent inner sweep (region_scan_extent, reused for both the implicit whole-region case and each real extent) whose verdicts aggregate across extents; bytes outside every extent are never decoded, so a data island sitting between two extents costs nothing. run() validates extents are sorted, non-overlapping, and fully inside [region_off, code_len) before any tracee is spawned (DF_BLOCKSTEP_EINVAL otherwise); the public is_pure/is_replayable/ is_injectable classifiers stay whole-blob (extents are a run()-only capability). New fixture island_sse (the same constant-pool-island shape as the existing island fixture, with a legacy-SSE paddq in place of island’s VEX-128 vpaddq so extents can actually recover it into the replay path) proves the positive case byte-identical to the single-step oracle with stops cut, while desyncing exactly like island without extents — the negative control. make dataflow-blockstep-test 199/199 (was 191/191, +8), stable across 5 consecutive runs; make docker-dataflow-attach 520/520 across all 8 suites, 0 skips; make docker-docs clean. See dataflow-producer-correctness.md.

  • Def-use graph and forward/backward slice surface in the Ruby, Lua, Zig, Rust, Go, Java and .NET data-flow bindings (previously producer-only), via a by-pointer slice-seed entry point (asmtest_slice_forward_seed/ _backward_seed, include/asmtest_valtrace.h — a 72-byte at_val_rec_t is SysV MEMORY-class and several of these FFIs, Ruby Fiddle chief among them, cannot pass it by value). Each binding’s ValueTrace now exposes defuse()/forward_slice(step)/backward_slice(step) alongside the existing live-attach producer, so all ten language bindings (with the prior Python/C++/Node) share one surface. Round-trip-tested with a hand-built r10→r11→r12 register-move chain (forward_slice(0)/backward_slice(2) both {0,1,2}) and, over the shared live-attach df_chain fixture, the memory def-use edge these seven could never reach before — backward_slice(4) and forward_slice(0) both equal {0,1,2,3,4} (the store at step 1 reached through the load at step 2), excluding the trailing ret. All seven docker-dataflow-<lang> lanes green at their new 40/40 (was 36/36), 0 skips. See dataflow-bindings-slice-codeimage.md.

  • Code-image recorder wrapper and versioned (time-correct) operand decode in all ten data-flow language bindings (asmtest_codeimage.h’s new/track/ now/bytes_at/free, plus the new asmtest_dataflow_ptrace_attach_pid_versioned entry point). Each binding gains a CodeImage wrapper (Python/Ruby/Lua’s available()/skip_reason()/track()/now()/bytes_at(), C++’s RAII CodeImage, Node’s mirroring the existing hwtrace-binding class, Zig/Rust/Go/ Java/.NET’s thin function wrappers) and a ValueTrace.attach_pid_versioned method; attach_jit no longer unconditionally passes NULL/null/nil for the versioned-decode img argument — a caller with a recorder now gets time-correct operand decode across a mid-capture JIT patch/free/reuse instead of NULL/live-snapshot bytes. Verified live: each binding tracks a recorder over a real victim’s published region and decodes an attach_pid_versioned capture through it (result and step-count assertions tied to that run’s own arguments, so a stubbed capture cannot pass). All ten docker-dataflow-<lang> lanes green with the new assertions, 0 skips. See dataflow-bindings-slice-codeimage.md.

  • Block-step pre-cover: an IBS covered-block table that memoizes the ptrace block-step reconstructors’ decode. asmtest_bs_precover_build/_free (include/asmtest_blockstep_internal.h, internal — no new public ABI symbol) pre-walks each asmtest_ibs_normalize_blocks-covered leader’s straight-line run ONCE and caches the per-instruction facts classify_branch would otherwise recompute on every #DB stop; blockstep_reconstruct (shared by the region and attached block-step drivers) resolves a cache hit by replaying asmtest_bs_scan_terminator’s exact decision procedure over the cached run — same bytes, same primitives, zero Capstone calls — so a hit is provably identical to a fresh scan, and a miss (including a leader that is not a real instruction boundary — the hostile case) falls back to the shipped path unconditionally. IBS stays statistical: pre-cover only memoizes tracer-side decode, never lets coverage skip recording anything (the exact parity contract in asmtest_ibs.h’s INVARIANT stands). A differential over the LOOP_X86 fixture (precover NULL vs. covering the loop head) proves byte-identical insns[]/blocks[]/truncated/result while cutting branch- probe decode calls; asmtest_bs_stats/_reset are test hooks that count cumulative probe calls and cache hits. See ptrace-blockstep-tracer-correctness.md.

  • ASMTEST_TRACE_IBS_PRECOVER: an opt-in policy bit that wires the block-step pre-cover table above into the cross-tier auto cascade (asmtest_trace_call_auto, include/asmtest_trace_auto.h). When set and IBS-Op is available, the block-step rung forks an isolated warm-up child that re-runs the routine for a bounded ~30ms while the parent surveys it out of band with asmtest_ibs_survey_process (no ptrace, no perturbation), builds a pre-cover table from the resulting live histogram, and installs it around the one asmtest_ptrace_trace_call_blockstep call the rung already made — so a trace produced with the bit set is byte-for-byte identical to one produced without it; any survey/build failure degrades silently to the plain rung. On a live Zen 2 host a 25-iteration loop’s block-step branch-probe decode calls dropped from 101 to 0 (a real live warm-up survey covering both of the routine’s basic blocks). Off AMD, or wherever the auto-cascade’s fast in-process single-step backend already completes the capture before reaching block-step, the bit is a proven no-op — never a behavior change. See ptrace-blockstep-tracer-correctness.md.

  • In-process, branch-granular single-step (W3): asmtest_ss_btf_available / asmtest_ss_btf_trace (src/ss_btf.c), the missing third single-step form. asm-test already had branch-granular stepping out of process (PTRACE_SINGLEBLOCK) and per-instruction stepping in-process (EFLAGS.TF, ss_backend.c); this arms DEBUGCTL.BTF alongside EFLAGS.TF over the same thread-pinned /dev/cpu/N/msr route asmtest_amd_msr_trace uses, so a taken-branch retiring — not every instruction — is what traps, with no 16-entry ceiling on the reconstructed stream (unlike AMD LBR). Gated by a hang-proof functional probe (some hypervisors silently mask DEBUGCTL.BTF and degrade to per-instruction stepping — the probe catches this, a build check cannot); deliberately scoped to a pinned leaf-routine envelope with per-trap re-arm (BTF is a hardware one-shot the CPU clears on every #DB) and faithful truncation on any observed context switch (Linux does not preserve a user-written BTF across one) — the general, context-switch-proof case stays owned by the shipped PTRACE_SINGLEBLOCK trio. Rides the existing docker-hwtrace-msr --privileged lane, no new capability. Live- verified on a Zen 2 host: the shared ROUTINE fixture reproduces the single-step baseline’s exact [0,3,6,c,11] stream and {0,0x11} block partition byte-for-byte, and a 20-trip loop (19 taken back-edges, past any 16-entry LBR window) reconstructs all 62 instructions, complete, 10/10 stable runs. See inproc-btf-block-step.md.

  • macOS out-of-process single-step tracer (asmtest_mach_*), completing the W2 foreign-process story on macOS. asmtest_mach_trace_call / _trace_attached / _run_to (asmtest_mach.h, src/mach_backend.c) mirror the Linux ptrace out-of-process tracer’s exact shape and offsets, but through task_for_pid + a Mach EXC_MASK_BREAKPOINT exception port + thread_set_state instead — macOS ptrace cannot edit RIP/RFLAGS at all. run_to’s breakpoint arm falls back from a software int3 to a DR0/DR7 hardware execution breakpoint on W^X code, same as the Linux tracer. x86-64 only for now. make mach-stepper-test (needs scripts/codesign-debugger.sh’s ad-hoc self-sign, or root) runs the lane live; self-skips (ASMTEST_MACH_EPERM) without either.

  • Pure tests for the IBS sample-period rounding/clamp and the additive-ABI struct_size guard (first coverage). The /16 rounding, the <16 → default clamp, and the period_jitter tail-guard had never executed against non-default values anywhere; internal asmtest_ibs_effective_period / asmtest_ibs_effective_jitter seams (shared with the live attr fill, so tested == shipped) now pin them.

  • Intel Pin vs. DynamoRIO analysis + a four-track umbrella plan — what Pin makes possible that the shipped DynamoRIO tier cannot. 2026-07-17-intel-pin-vs-dynamorio.md and intel-pin-capabilities-plan.md. Most “Pin advantages” are maturity, not impossibility (DR has drrun -attach, supports Windows, and the taint ground is already held in-tree), and the note separates those out. Four items survive as separable plan tracks: PIN-1 an Intel SDE lane that runs the existing TEST() suites under sde64 -future so APX / AVX10.2 / AMX / AVX-512 assembly gets full register/flag/memory/ABI assertions on any x86-64 host including CI — a true impossibility for both DR (executes on real silicon) and the Unicorn tier (QEMU 5.0.1 predates AVX TCG; vaddps ymm → UC_ERR_INSN_INVALID and VEX-128 is silently mis-run as SSE), converting CLAUDE.md’s “specific CPU generation” hardware self-skip into an installable, pinnable dependency; PIN-2 an XED-decoded Pin trace tier for the newest extensions DR’s own decoder rejects (APX is open — DR #6226; the once-broken AVX-512 VNNI is fixed — DR #5440 closed 2022-04-25, and the pinned DR post-dates it — so the case rests on APX alone); PIN-3 Pin probe-mode arg/return capture (original code runs native, no code cache — the capture-args-returns.md middle tier DR has no equivalent for); and PIN-4 libdft64 as an independent taint oracle diffed byte-for-byte against the shipped DR taint client — the ASSERT_MATCHES_REF cross-validation idiom. The reverse gaps are recorded so Pin is not mistaken for a superset (x86-only — no AArch64; no in-process no-IPC model; proprietary freeware). Every track is fetched-and-pinned, test/oracle-only, never shipped — DR’s exact handling — so none adds a bindings-parity obligation.

  • AMD hardware review + follow-up plan — an adversarial pass over the AMD tiers in which four of nine candidate findings were REFUTED and recorded as such. 2026-07-17-amd-hardware-review.md and amd-review-followup-plan.md. Every claim resting on kernel or silicon behaviour was checked against fetched primary source at pinned tags (v6.10/v6.12/v6.14), not recall — one verifier caught master shifting line numbers mid-review and re-pinned. Confirmed: Zen 3 BRS cannot open (the probe/capture use PERF_COUNT_HW_BRANCH_INSTRUCTIONS → 0x00c2, but amd_brs_hw_config demands raw 0xc4 and sample_period > lbr_nr(16); EINVAL is then reported as AMD_NOHW, so a real Zen 3 owner is told “no AMD branch records”) — code fix hardware-blocked per the house rule, but three docs assert the working path, including a Phase 0 marked landed that specifies the exact missing arm; branchsnap’s synthetic boundary edge inflates the depth check so a complete use == 15 window is spuriously truncated → a real re-execution (n_dec = use + 1 vs a check counting hardware slots only); IBS_MAX_RECORD is pre-callchain (112B vs ~1032-1184B — it landed in 68b53850, callchain in a266b91 two days later and never touched it), leaving the loss heuristic ~10× short where PERF_RECORD_LOST provably cannot cover the gap → lost==0 && throttled==0 silent loss in a fidelity-first lane; and asmtest_amd_freeze_available() is dead (nm: one definition, zero undefined refs) with a flatly false string printed to a human. The organising finding is process, not code: no CI lane exercises AMD silicon — the one AMD-targeted job is named hwtrace-privileged (PERFMON; AMD-exact self-skips off Zen) and runs on ubuntu-latest — so the only gate is a manual checklist with exactly one commit, timestamped identically to the fix for the bug it still calls open, which now instructs treating a truncated=0 regression as a known issue and gates on two signals this project measured false five days later. Refuted and recorded so they are not re-raised (the Matrix 3 convention): the MSR TOS-rotation claim (transplants Intel architecture — LbrExtV2 pins From[0]/To[0] by register renaming and hard-codes hw_idx = 0 because rotation is impossible; the linear read is correct), “IBS opts are unreachable” (NULL is the designed contract; SYSTEM_WIDE is tested live under --cap-add=PERFMON), the RipInvalidChk impact (the affected silicon population is empty — Family 10h hardwires BrnTrgt = 0, so the existing gate rejects it first), and the unread-IBS-regs “gap” (a dated non-goal in data-flow-tracing-plan.md:94; the original grep was a false negative on “address sampling”).

  • asmspy --dataflow --auto works off AMD now: a portable software-clock sampler (--sampler=ibs|sw), with the residency hazard owned out loud. Auto-targeting was AMD-IBS-only; asmtest_swclock_survey_process (new in the IBS lane’s backend: PERF_COUNT_SW_TASK_CLOCK + PERF_SAMPLE_IP, no PMU at all) samples any-vendor hosts and VMs under the same out-of-band, unprivileged envelope, with availability probed BY DOING and an errno-carrying reason (asmtest_swclock_unavail_reason) from birth — the --sample lesson, applied rather than re-learned. The pure ranking half (asmspy_autoregion_rank_ip) is transparent about being the WEAKER rule: an IP histogram measures residency, and residency’s winner is often the entered-once-never-returns shape whose entry breakpoint can never fire — test_autoregion #15 pins the two rules DISAGREEING on the same behaviour. So the sw path ranks up to 3 candidates and cmd_dataflow WALKS them: a winner never seen entering is refused at the bounded entry wait and the next-ranked candidate is tried, each refusal reported. Proven live in docker-cli-ibs on the built-to-disagree auto_victim: sw picks grind_forever, the wait refuses it, the walk captures entered_often. 9 new pure checks (31 total); both cli lanes PASS.

  • cli/asmspy_ghash.h + test_ghash.c — the graph’s hash index gets the collision test its faithful-gap note demanded. The --graph engine’s open-addressed index shipped with a measured blind spot: a probe loop that trusts the hash and skips the key compare emits byte-identical smoke output, because ≤7-node graphs in a 128-slot table never collide (on a larger graph that mutant silently over-merged an edge). The mechanism now lives in a pure header (asmspy_gh_find’s eq callback owns the key compare; the engine’s node/edge lookups supply it) and a unit test brute-forces three keys into one slot, so exactly that mutant fails. Mutation-proven 3/3 (measured): accept-first-slot → 5 FAILs, idx-not-idx+1 slot encoding → 3 FAILs, grow-without-rehash → 1 FAIL. Engine behavior unchanged (make docker-cli PASS, --graph uniqueness e2e included).

  • asmspy TUI symbol picker (modes 2 and 9) gains a Tab-cycle sort: address -> hot edges -> name. The picker used to be a flat, address-ordered list with no way to find “what’s actually running” short of guessing a name to filter by. Hot edges reuses the exact entry-arrival rule --dataflow --auto picks with (asmspy_autoregion_rank, one AMD IBS-Op window) rather than minting a second definition of “hot” — so the picker’s ranking and --auto’s pick always agree. An IBS-less host (or one where perf is locked down) falls back to address order with an on-screen reason rather than silently doing nothing. Name sorts case-insensitively. The row permutation is kept separate from the symtab’s own address-sorted storage (asmspy_symtab_at’s binary search depends on it), mirroring the process picker’s existing order[] indirection.

  • asmspy --dataflow <pid> --auto [--module=<m>] — trace what a process is DOING, no symbol needed. Auto-targeting samples the target OUT OF BAND (AMD IBS-Op, no ptrace, no perturbation) for 400 ms, ranks the hottest entry edge — an edge whose target is a function’s start is a direct observation of the exact event the data-flow producer blocks on, where the intuitive rules fail: the hottest raw edge is usually a mid-function loop back-edge no entry breakpoint can catch, and a PC histogram picks the functions entered once and never again (main, every event loop), which HANGS the producer — then hands the winner’s (base, len) to the existing --dataflow engine. Entry-edge counts are quantitative (measured: they reproduce a victim’s true 8× loop trip count and 2× call-site ratio). The ranking is a pure header (cli/asmspy_autoregion.h, 21-check unit test) so its correctness is covered on every host while the AMD-hardware live leg runs in the docker-cli-ibs lane. The resolver layers the JIT map over the ELF symtab, so a hot JIT’d/managed method (perf-map or jitdump, real size) can win the pick and names the capture — proven live against a perf-map victim whose only entry arrival is an anonymous-mapping function the ELF symtab cannot see. An idle target gets a truthful refusal (“no function was observed being ENTERED”), not a guess; zero-size symbols cannot win on their exact-start-only resolution technicality; --auto --tid is a usage error (the sampler carries no tid, so pinning could only arm a breakpoint on a thread that never arrives); --module= scopes the pick with the same substring rule as --tree --module=.

  • make docker-cli-ibs — the lane that actually tests --sample, and accurate skip reasons everywhere. The plain docker-cli lane runs under Docker’s default seccomp, which blocks perf_event_open — so every --sample assertion had self-skipped since the view landed (a green gate over an untested view, on hosts where IBS works fine). The new lane reruns the same image and smoke with --cap-add=PERFMON under the default profile (CAP_PERFMON bypasses perf_event_paranoid; no sysctl change needed), so the --sample and --auto blocks execute their else-branches for real on an AMD IBS host. Alongside it, a new public API asmtest_ibs_unavail_reason() carries the real perf_event_open errno out of the backend: # SKIP --sample: used to print an empty string precisely when the substrate probe passed but perf was blocked — the one case an operator on their own AMD box needed the reason most. EACCES now says paranoid/CAP_PERFMON, EPERM names seccomp, instead of one indistinguishable silence.

  • asmspy region view samples WORKER threads (--trace, TUI mode 2) + a new --tid=<t> filter. asmspy_engine_region used to PTRACE_ATTACH the thread-group LEADER and run it to the region, so a function that runs on a worker thread was never observed and the view reported ASMSPY_REGION_NEVER_RAN — structurally blind to exactly the code asmspy exists to show, since a managed method almost never runs on the leader. It now SEIZEs every thread (PTRACE_O_TRACECLONE, so a thread spawned mid-run can win a later round) and races them all to the entry, sampling whichever arrives first. --trace gains --tid=<t> to pin one thread, matching --stream/--graph/--tree/--dataflow. The design reuses the data-flow tier’s oracle-validated race (dfp_seize_all/dfp_run_to_multi) rather than inventing a second one, reimplemented over asmspy’s own thread table per the standing precedent that an engine stays in cli/ and leaves src/ptrace_backend.c untouched. --tid pins via a per-thread hardware execution breakpoint, not the shared int3: a shared int3 traps every thread, and stepping hot non-target threads back over it was measured not to converge. cli_smoke.sh asserts the worker sample, both --tid directions, and target survival past a settle; make docker-cli → cli-smoke PASS. Known gaps, unchanged: the any-thread entry breakpoint is POKETEXT-only (a W^X JIT page self-skips; no DR0 fallback), and the TUI has no thread picker.

  • CI gate for the DynamoRIO attach tier (taint-attach). Increments 1 and 3-5 — cooperative attach, the marker-less interactive nudge, external drrun -attach into a running native process with taint capture, and K-round attach/capture/detach cycling — had docker lanes but no CI job, so five increments of capability had no regression gate while the launch-only taint tier next door had four. The external-attach image needs --cap-add=SYS_PTRACE and its comment still described it as the Increment-2 research probe (“a manual diagnostic, not in the main gate”) — true when written, but Increments 4-5 were later added to the same image, so landed capability inherited a probe’s CI posture. Both lanes now run on every push. The two MANAGED probes stay out by design: they record a reproducible NO-GO, not capability.

  • Data-flow tracing Phase 6 Increment 1 — libasmtest_dataflow shared lib + Python binding. make shared-dataflow builds the pure analysis pipeline (L0 value sink + L1 def-use + L2 slice + method identity + GC-move canonicalization + runtime-helper summaries; the emu/ptrace/DR producers stay separate tiers) into a dlopen-able shared library — the packaging target the language bindings consume. First bindings: Python (asmtest.dataflow, ctypes), C++ (bindings/cpp/asmtest_dataflow.hpp, header-only), and Node (koffi, bindings/node/dataflow.js) all wrap the pure GC-move canonicalizer asmtest_gcmove_canon and the tiered-re-JIT-aware method resolver asmtest_method_resolve_pc; make dataflow-{python,cpp,node}-test build and run their TAP suites (8 / 16 / 16 checks, mirroring the C test_dataflow_gcmove / test_dataflow_method semantics). Each self-skips cleanly when the lib is not built, so none reddens a general binding job. A new dataflow CI job builds the lib and runs the Python + C++ bindings on every push (the Node binding runs in the docker bindings lane). All three host bindings (Python, C++, Node) additionally wrap the full L0->L1->L2 pipeline (ValueTrace: value-trace build -> def-use -> forward/backward slice), round-trip-validated against the C semantics (register move chains, load-after-store through memory, no spurious cross-links — which also validate the 13-field at_val_rec_t marshalling). The remaining seven language bindings are later increments.

  • Live GC-move detection feed for the data-flow tier (GcMoveMap, .NET). An in-proc EventListener on the CoreCLR runtime provider that enables the GCHeapSurvivalAndMovement keyword and captures GCBulkMovedObjectRanges from a compacting GC — the live source the pure GC-move canonicalizer (asmtest_gcmove_canonicalize) was built to consume. Validated via make docker-hwtrace-dotnet: an induced compacting gen2 collection is captured as 11 events / 20474 moved ranges (suite 169 → 177). Truthful scope: in-proc EventListener surfaces the reliable scalar range count but not the manifest struct-array Values payload, so the concrete {old_base, new_base, len} triples that drive the canonicalizer end-to-end are deferred to a raw EventPipe/nettrace path — the keyword, compacting-GC inducement, and listener wiring (the uncertain parts) are proven.

  • Data-flow tracing Phase 5 Increment 1 — DynamoRIO in-band L0 value producer (src/dataflow_dr.c, src/dataflow_dr_client.c). The in-band, whole-process analog of the scoped ptrace L0 producer: a DynamoRIO client instruments a target under real DR and captures per-instruction operand values into the same asmtest_valtrace_t sink, so the L1 def-use builder and L2 slicer work unchanged on an in-band capture. Cross-validated against the emulator L0 oracle on a shared fixture (in-band def-use edges and forward/backward slices equal the oracle’s), validated live via make docker-drtrace (dr-valtrace-test 14/14; DR is a software DBI engine, so the lane needs no privilege or special hardware). Self-skips (exit 0) without DYNAMORIO_HOME. Store values and RIP-relative/segmented/VSIB memory EAs are deferred to a later increment (the current fixture avoids them); DR-side taint shadowing and whole-process breadth are also later increments.

  • Data-flow tracing Phase 4 Increment 3 — runtime-helper summary edges (src/dataflow_helpers.c). asmtest_defuse_build_summarized recognizes a .NET runtime-helper call in an L0 value trace (via the Increment-1 method resolver → a helper table matched by name/prefix) and collapses the helper run into a summary node — emitting only its declared input reads and output writes and dropping the body — so caller dataflow connects across the helper (arg def → summary → return use) without instrumenting CoreCLR internals. Supports reg→reg helpers (allocation, generic-dict lookup) and a MEM_AT_REG write-barrier output. Conservative by construction: an unrecognized call is descended normally, never given a fabricated edge. Pure, host-independent suite test_dataflow_helpers (36 checks).

  • Data-flow tracing Phase 4 Increment 2 — GC-move canonicalization (src/dataflow_gcmove.c). A pure, host-independent transform (asmtest_gcmove_t, the shape of EventPipe GCBulkMovedObjectRanges: old_base/new_base/len) that remaps memory addresses across a heap compaction to a stable canonical identity, so a managed value’s def-use survives the move without pre/post-move false aliasing — the plan’s Phase-4 exit criterion. Synthetic suite test_dataflow_gcmove (26 checks) proves the exact trap: a def at the old address and a use at the new address unify into one object, while an unrelated object that later reuses the freed old address does not alias it. The live EventPipe feed is a later increment.

  • IBS-Fetch front-end coverage lane (AMD IBS statistical lane Phase 7). A second AMD IBS producer beside the retired-op edge sampler: ibs_fetch (PMU type 10) samples fetch addresses (front-end / i-cache / ITLB view). A pure, host-independent decoder turns one PERF_SAMPLE_RAW fetch record into {fetch_addr, valid, complete, icache_miss, itlb_miss, latency} (unit-tested with synthetic records on every CI host), plus an availability probe and a headless fetch-coverage survey with faithful throttled/lost provenance, self-skipping off IBS/permission. Kept fully internal (src/ibs_backend.h) — no public asmtest_ibs_* surface added, so no binding flag day. Verified live on Zen 5.

  • Data-flow tracing Phase 4 Increment 1 — PC→method-identity+version resolver (src/dataflow_method.c). A pure, host-independent resolver that labels each step of an L0 value trace with its owning method + version from a jitdump/perf-map-shaped method-map, correctly handling tiered re-JIT (newest code_index wins for an address; a re-JIT to a new address is a new version) — the managed-taint prerequisite. Synthetic suite test_dataflow_method (29 checks, incl. the moved-re-JIT version distinction); the hard GC-move canonicalization is deferred to a later increment.

  • hwtrace-privileged CI job + AMD hardware-validation doc. A CI job exercises make docker-hwtrace-privileged so the --cap-add=PERFMON lane can’t bitrot (the AMD-exact tests self-skip on GitHub’s non-AMD runners — transparent by design; it lights up on a future AMD runner). docs/internal/amd-hardware-validation.md documents the manual pre-release validation on real Zen 3+/Zen 5 silicon — closing the gap that let the call_auto LBR truncation bug hide (the exact AMD paths never ran in CI).

  • Size-negotiated hwtrace options ABI + machine-readable status surface + escalation mechanism, across all ten bindings (the AMD-followup API flag day — Phases 1, 3, and F22/F26/F37). asmtest_hwtrace_options_t now leads with a size_t struct_size the caller sets (the INIT_OPTS idiom, or explicitly after a zero-fill); asmtest_hwtrace_init copies min(struct_size, sizeof) and zero-fills the tail, so an older/newer caller is never read out of bounds — and a caller that fails to self-describe (struct_size == 0 or too small to reach backend) is rejected with EINVAL rather than having a set field silently dropped. New asmtest_hwtrace_status() (available / code / stage / probe errno / perf_event_paranoid / reason) and asmtest_hwtrace_perf_event_paranoid() distinguish ASMTEST_HW_EPERM (substrate present, permission denied — e.g. AMD LBR on an unprivileged paranoid > 2 host) from EUNAVAIL (missing silicon), backed by one shared classifier so status() and skip_reason() cannot drift. asmtest_trace_choice_t grew a mechanism field (HW_BRANCH / TF_STEP / MSR_LBR / BLOCKSTEP / PER_INSN / DBI / EMULATOR / STATISTICAL) plus ASMTEST_FIDELITY_STATISTICAL, so trace_call_auto reports which rung actually won and a statistical result is structurally unmistakable for an exact one. All ten wrappers mirror the new layouts and wrap the new calls; the parity gate passes with zero allow-list changes. Suite 358 → 383 (ABI guard, status incl. the live-EPERM assertion, mechanism).

  • Data-flow tracing gains a live scoped ptrace L0 producer (Phase 3 — real values, out of band). src/dataflow_ptrace.c single-steps a routine (fork+PTRACE_TRACEME, or PTRACE_SEIZE attach to a live victim that survives detach) and emits the same asmtest_valtrace_t stream the Phase-0/1 analyzers consume, so def-use + slicing work unchanged on live captures — reading each step’s registers (GETREGS, GETFPREGS for XMM, NT_X86_XSTATE for 256-bit YMM) and the memory its operands touch. Cross-validated edge-for-edge and value-for-value against the emulator L0 oracle; RIP-relative effective addresses resolve against the next instruction (a bug an adversarial verify caught before merge), gs-based and wide-vector operands captured. dataflow-test 26 → 36.

  • Data-flow tracing tier, Phases 0–2 (include/asmtest_valtrace.h, make dataflow-test) — the CI-runnable milestone of the data-flow plan. Phase 0: the shared L0 value-trace sink (asmtest_valtrace_t: caller-owned buffers, append/stash-wide/truncate discipline mirroring asmtest_trace_t) plus the Capstone operand read/write-set enumerator (explicit register + memory operands with base/index/scale/disp/segment, and the implicit ones — eflags writes, rsp read+write on push/pop — via cs_regs_access). Phase 1: the L1 def-use graph over a recorded value trace (register moves and load-after-store memory edges) and the L2 forward/backward slicer on top of it. Phase 2: the emulator (Unicorn) L0 producer — replay a routine under uc_hook instrumentation and emit the value trace the pure phases analyze; validated live (per-step values, def-use edges, and both slice directions asserted against a hand-traced fixture). Pure phases run on every host; the emulator cases self-skip without Unicorn. New mk/dataflow.mk; suites test_dataflow / test_operands / test_dataflow_emu (53 checks). The known raw-address aliasing false positive (pre-GC-canonicalization) is asserted AS a false positive, per the plan’s Phase 4 note.

  • asmspy reads binary jitdump files — the bytes-accurate, tiered-recompile-aware JIT symbol source (asmspy plan Theme A). A jit-<pid>.dump reader (LE header + JIT_CODE_LOAD / JIT_CODE_MOVE records; unknown record types skipped via total_size; a truncated in-flight tail ends the parse keeping what’s whole) is now tier 1 of the JIT resolve chain — discovered the way perf does (a mapped marker in /proc/<pid>/maps, then /tmp and the target’s cwd), parsed ahead of the text perf-map (tier 2, the LCD), re-JIT and code-motion aware (newest code_index wins, CODE_MOVE relocates). Same rate-limited refresh-on-miss discipline as the perf-map path. Hardened against hostile files (bounded name reads, zero/short total_size rejection — fuzzed under ASan/UBSan). Also new: --tree --json/--dot exports mirroring the --graph exporters, the extracted call-graph sort comparator (cli/asmspy_graphsort.h) with its ordering/tiebreak unit test, and a jitdump_victim end-to-end smoke.

  • .NET AsmTrace.Window captures methods JIT’d mid-window — the sibling-thread live JIT publish (extensions plan E3), closing the deep-BCL gap. JitMethodMap.SetPublishChannel now starts a dedicated, never-stepped publisher thread: the MethodLoadVerbose callback only enqueues (base,len) onto a lock-free queue (publishing inline could fire on the single-stepped thread and re-enter the runtime under step — the observed SIGABRT that kept this OFF), and the sibling drains it and P/Invokes each record into the shared address channel while the window runs. Stop joins the publisher before the channel is freed (no use-after-free window); the §E1 hybrid keeps live publish off by design (it must capture only the surveyed hot slice). New WholeWindowScope.LiveJitPublished counter; suite grows 161 → 169 checks including a ptrace-free mechanism test (native ring-head readback) and a mid-window-JIT integration case (52 records live-published in the docker lane).

  • Env-gated debug logging for the hwtrace/AMD tier (src/debug.{c,h}, followup Phase 4 / F32). ASMTEST_HWTRACE_DEBUG=1 (or ASMTEST_AMD_DEBUG=1) turns on stderr tier diagnostics; unset costs one cached getenv per process. Covered by two suite checks (silent when off, emits when on).

  • Host-independent synthetic-ring tests for the AMD branch-stack parse (review F43/F44). The hwtrace_end_amd ring-parse now has an internal, linkable entry (asmtest_amd_ring_parse_decode) driven by crafted PERF_RECORD_SAMPLE buffers — the nr-clamp / LOST / Tier-A-vs-Tier-B logic and amd_span_decodable’s dropped-jmp follow finally run on every CI host, AMD or not (+368 lines in examples/test_hwtrace.c, suite 341 → 358).

  • AMD LBR window-reach tuning guide (docs/guides/tracing/amd-lbr-tuning.md, review F47 / followup Phase 10) — what bounds a 16-deep LbrExtV2 window, the sizing/splitting levers, what truncated means and how it is reported, when the statistical IBS lane is the better tool, and the privileged-vs-unprivileged lanes including perf_event_paranoid.

  • asmspy --sample + TUI mode 7 — a live statistical hot-edge view, out of band (IBS lane Phases 2–3, the flagship deliverable). Built on the new asmtest_ibs_survey_process(pid, ms, opts, out): whole-process IBS-Op coverage that opens one perf event + ring per thread of the target (enumerating /proc/<pid>/task, with one mid-window rescan for threads spawned after start; the residual born-and-died-in-window race and the privileged system-wide remedy are documented, not hidden) and merges everything into one hot-edge histogram. asmspy_engine_sample resolves both endpoints of each edge through the existing ELF-symtab → JIT-perf-map chain, so managed Node/.NET/Java frames are named. Headless asmspy --sample <pid> [ms] [--json] prints the histogram — count  from -> to with [misp N%]/[ret] tags and faithful branch/total samples / throttled provenance — or machine-readable JSON; TUI menu item “7) Hot edges (sample)” shows the same table live, pausable + scrollable + Tab-sortable (count / mispredicts). Unlike the stream/graph/tree views this never attaches ptrace and never single-steps — the target runs at full speed — making it the only rich view that is safe on a live JIT, exactly the targets single-stepping can crash. Self-skips (# SKIP, exit 0) off IBS; new busy victim cli/sample_victim.c + a --sample smoke in cli/cli_smoke.sh; the TUI view is driven end-to-end through a pty harness. Verified live on Zen 2: both surfaces name the victim’s hot back-edge without perturbing it.

  • IBS-Op fallback for the AMD whole-window statistical survey (IBS lane Phase 4; fixes the AMD review’s F6). On Zen 2 the branch stack (BRS / LbrExtV2) does not exist, so asmtest_hwtrace_sample_window_amd (and its begin/end split) returned EUNAVAIL on the one AMD host class that most needs a crash-proof survey. New internal window primitives (asmtest_ibs_window_begin/_end, src/ibs_backend.h) arm IBS-Op on the calling thread around the caller’s window body, reusing the channel/drain/edge-hash machinery; on branch-stack perf_open failure — or when ASMTEST_FORCE_IBS_SURVEY is set, for cross-validation on Zen 3+/CI — the survey delegates to IBS and flattens each sampled edge’s target into the ips[] endpoint histogram weighted by count, so the caller’s bucket-by-method hotness view is unchanged in shape. Purely STATISTICAL: a separate producer that never feeds the exact insns[]/blocks[] parity cascade; the branch-stack path is byte-identical when the stack is present and the env unset. Covered by test_amd_sample_window_ibs (self-skips off IBS); verified live on Zen 2 (~468/468 endpoints in the hot loop, full hwtrace-test 341/341).

  • asmspy gained three whole-process structure views and a per-thread lens since the entry below. A call graph (TUI mode 4, headless --graph <pid> [n] [--sort=invocations|fanout]): every call attributed caller→callee across all threads, aggregated per function, with --json export (nodes and {caller,callee,count} edges, addresses as 0x strings) and --dot emitting a Graphviz digraph (kind-coloured nodes, count-labelled edges) ready for dot -Tsvg. A call tree (TUI mode 5, headless --tree <pid> [n]): the same feed with nesting/order preserved, indented by depth, in a two-pane TUI. A process/thread topology view (headless --procs, plus an F2 flat-list ↔ tree toggle in the process picker): the process forest drawn with ├─/└─/│ box glyphs, threads then child processes nested under each process. A --tid=<t> filter for --stream/--tree/--graph seizes and steps only that thread, leaving the rest of the process at full speed. The call-graph and region (assembly & funcs) TUI views now pause + scroll like the log views (space freezes a stable snapshot, arrows/PgUp/PgDn/Home/End move, Tab switches pane focus in the region view). And all single-step engines resolve JIT frames through the runtime’s perf map (/tmp/perf-<pid>.map — Node/V8 --perf-basic-prof, .NET DOTNET_PerfMapEnabled=1, OpenJDK perf-map-agent), refreshed rate-limited on miss so a compiling JIT keeps getting named: managed frames render name [jit] in the stream/tree and [JIT]-tagged internal nodes in the graph.

  • Statistical AMD IBS-Op tracing lane (asmtest_ibs.h, src/ibs_backend.c) — Phases 0–1. A new, self-contained statistical trace producer for AMD hosts where every branch-stack facility is absent (Zen 2 has no BRS / LbrExtV2, so every exact hwtrace backend self-skips and the machine falls back to ~1000×-slower single-stepping). IBS-Op (Instruction-Based Sampling) is the one branch-tracing facility this silicon has: it tags a retired op per NMI window and, for taken branches, delivers both the source (IbsOpRip) and target (IbsBrTarget) — a statistical from → to control-flow edge, sampled out of band, against a running thread, unprivileged (the kernel swfilt bit makes user-only sampling open at perf_event_paranoid=2) and without perturbing the target — exactly the case the single-step views are dangerous on (a live JIT / managed runtime). It needs no external library (raw perf_event_open + a pure decoder). Public surface: asmtest_ibs_available() / asmtest_ibs_skip_reason() (the full AMD/IBS/BrnTrgt/swfilt detect-and-skip chain), asmtest_ibs_decode_op() (a pure, host-independent decode of one IBS-Op PERF_SAMPLE_RAW record into an edge — unit-tested with synthetic records on every CI host, AMD or not), and asmtest_ibs_survey_pid() (attach IBS-Op to one thread, drain for N ms, return an aggregated hot-edge histogram sorted by count with faithful provenance — samples / branch_samples / lost / throttled). INVARIANT: statistical only — it can prove a block was seen, never that one was not, so it never feeds the exact insns[]/blocks[] parity contract; it is a separate diagnostic producer, not a member of the exact-trace cascade. Built into libasmtest_hwtrace; validated by make ibs-test (also folded into make hwtrace-test) and the containerized make docker-hwtrace-ibs; probe binary examples/ibs_probe.c. The live path is validated on an AMD Ryzen 9 4900HS (Zen 2, kernel 6.14): the test captures a spin loop’s back-edge out of band from a separate thread. Plan: docs/internal/plans/zen2-ibs-tracing-plan.md. Phases 2–4 landed subsequently — the whole-process survey, the asmspy --sample view (headless + TUI mode 7), and the statistical survey fallback; see their own entries above.

  • asmspy — an interactive process tracer (new cli/ subsystem, Linux x86-64) — a small ncurses front-end over the out-of-process (ptrace) tracer: attach to any running process and watch it live and out of band. Three live views: syscalls with data (a mini strace; every syscall named from a table generated against the host’s own <sys/syscall.h>, read/ write buffers and path arguments decoded, read/write file descriptors resolved to their path/socket/pipe via /proc/<pid>/fd like strace -y, decoded strings split into their own pane; the syscall stream and the whole-process instruction stream follow every thread of the target — PTRACE_SEIZE of all tasks plus PTRACE_O_TRACECLONE for threads spawned later, each line tagged [tid] when more than one is followed, and (for syscalls) entry/exit read from PTRACE_GET_SYSCALL_INFO so seizing a thread mid-syscall never desyncs), a chosen function’s assembly with per-instruction execution heat counts plus its callees ranked by call count (resampled each time the target calls it), and a whole-process live instruction stream (every instruction as it executes, resolved to its function). The two log feeds (syscall log, live stream) pause + scroll back through their history (space to freeze, ↑/↓/PgUp/PgDn/Home/End to move, End/space to resume the tail), and scrollback survives target exit. The process picker filters as you type and sorts by pid, recent CPU activity, or string-scan density (Tab cycles, r rescans, b navigates back). Every view is also a headless subcommand for scripts and CI: --list [active|scan], --syms <pid> [filter], --log <pid> [n], --trace <pid> <sym|0xADDR[:LEN]> [n] (an explicit 0xADDR:LEN range reaches stripped code or a JIT region no symbol covers), --stream <pid> [n]; a negative n runs until the target exits, and malformed arguments are rejected up front. Built by make cli (needs libncurses + Capstone; self-skips with guidance) or containerized via make docker-cli (Dockerfile.cli); carries its own /proc lister and ELF .symtab/.dynsym function resolver. End-to-end headless smoke (cli/cli_smoke.sh, make cli-smoke) drives all five subcommands against the example victims and is gated in CI (cli job). Guide: docs/guides/tracing/asmspy.md.

  • asmtest_trace_call_auto — the auto-escalating, call-owning cross-tier trace, now in all ten bindings. A single entry point that owns the invocation and traces a native routine under the fastest exact tier, then automatically escalates to a ceiling-free tier when the trace comes back truncated — walking the ladder fast HWTRACE backend → MSR-direct AMD-LBR rung → BTF block-step → per-instruction single-step until the capture is complete (or the tiers are exhausted, transparently flagged). This closes the “arm → detect truncation → re-resolve → re-run” loop that was previously only a documented idiom. Landed C-first (src/trace_auto.c, the MSR rung folded in later), then wrapped in every binding (python/cpp/rust/zig/node/java/dotnet/ruby/lua/go), removing its former ALL parity exemption. *used reports the tier that produced the final trace, so a caller can see whether escalation fired. Covered by test_call_auto* in the hwtrace suite.

  • asmtest_hwtrace_arm_tid wrapped in the remaining seven bindings (go, java, lua, node, ruby, rust, zig) — the §0.2 thread-scope assert accessor (the OS thread id that armed the active hardware-trace capture, -1 if none) was python/cpp/dotnet-only; it is now surfaced in all ten bindings with an idiomatic accessor (HwTrace.armTid() / arm_tid / HwTraceArmTid() per language), closing its seven per-binding parity exemptions. Each wrapper was built and its hwtrace test suite run green in the per-language docker lanes.

  • Whole-window attribution, version-aware render, and async-hop merge in the Node and Java bindings — dotnet-parity Phase 2, the remaining CI-runnable clusters. Wraps six .NET-lead C symbols across both bindings:

    • CodeImage.renderVersioned(when, trace) (asmtest_hwtrace_render_versioned) — disassemble a trace’s absolute addresses against a code-image timeline AS OF a capture sequence, not live memory. Version-aware (unlike render_window): tracking a region as add then rewriting it to sub renders add at the old sequence and sub at the new. Plus NativeTrace.appendInsn (wraps trace_append_insn, a non-tier symbol) to build such an absolute-address trace.

    • HwTrace.regionName / symbolizeBuckets / attributeWindow (asmtest_hwtrace_region_name / _symbolize_bucket / _attribute_window) — whole-window noise attribution: reverse-resolve an address to its mapped-region name, bucket a list of IPs by JIT symbol (perf-map) or region, and attribute a live whole-window capture’s absolute addresses to caller-named regions first (so two identical-byte leaves in distinct mappings split into separate buckets — what symbol/disasm attribution cannot do). AddrChannel-free; range classification, no Capstone.

    • HwTrace.stitchHandles(hops, …) (asmtest_hwtrace_stitch_handles) — the §D0.4 async-hop merge: order N already-captured hop traces by seq and concatenate into one logical trace with per-hop slice bounds. Host-independent (pure merge — runs on every lane, arm64 included); the hops must outlive the call (shallow-copy, not duplicated). asmtest_hwtrace_stitch (the C core) stays binding-internal.

    • Struct marshalling is pinned to the exact SysV layouts (bucket_t 136 B, slice_bound_t 32 B, named_region_t 80 B) and cross-checked against the dotnet [StructLayout]s. Validated in the docker-hwtrace-node / -java lanes against the C oracles (test_render_versioned, test_symbolize_bucket, test_wholewindow_buckets, test_stitch_slices). All six ALL allow-list lines stay (seven-eight bindings still don’t wrap them); trace_append_insn is a non-tier symbol (ungated).

  • Crash-proof WHOLE-WINDOW out-of-process capture in the Node and Java bindings — dotnet-parity Phase 2, increment 3, the out-of-process analog of the in-process window() form (which single-steps the calling thread and is fatal for arbitrary managed code). Wraps asmtest_ptrace_trace_window_call (Ptrace.windowCall / HwTrace.ptraceTraceWindowCall — fork-internal: a forked child runs the window frame and is stepped, so it asserts unconditionally on any ptrace lane) and asmtest_hwtrace_stealth_trace_windowed (HwTrace.stealthWindow — a helper child reverse-attaches and steps the calling thread’s window body out of band, mirroring dotnet’s AsmTrace.Window; self-skips on a refused reverse-attach), plus the five asmtest_addr_channel_* FFI shims behind a new AddrChannel class. Pre-publish the code regions the window frame calls into (its leaves/methods) on the channel; the capture records the frame plus every published region as ABSOLUTE addresses (classify by range — no Capstone), stepping over everything else. Validated in the docker-hwtrace-node / -java lanes against the C oracle’s driver-blob ceremony (a 35-byte frame calling two 7-byte leaves): result m2(7,3)==4, driver + both leaves recorded in call order, complete. The ALL allow-list lines for both windowed symbols stay (seven bindings still don’t wrap them); the addr_channel shims live in a non-tier header (ungated).

  • Crash-proof out-of-process stealth capture (stealthTrace) in the Node and Java bindings — dotnet-parity Phase 2, increment 2. HwTrace.stealthTrace(code, a, b) (Node) / HwTrace.stealthTrace(NativeCode, long...) (Java) wrap asmtest_hwtrace_stealth_trace: a helper child reverse-attaches (PR_SET_PTRACER + PTRACE_SEIZE) and single-steps the native leaf out of band, so no EFLAGS.TF is ever armed on the runtime’s own (V8 / JVM) thread — the crash-proof counterpart to the in-process callScoped/window forms, mirroring dotnet’s AsmTrace.Method(..., outOfProcess: true). The result is exact (the helper reads the caller’s RAX at the ret); the instruction stream is best-effort over a live runtime (its async signals can truncate the per-instruction walk — faithfully reported via truncated), so the tests assert the exact [0,3,6,c,11] stream only when not truncated. Validated in the docker-hwtrace-node / -java lanes; self-skips cleanly where a Yama ptrace_scope refuses the reverse-attach.

  • AMD LBR Zen 4/5 coverage: slot-efficient branch filtering (#2B), period-spaced stitch validation (#2A), and single-exit snapshot-by-default (#3). Three improvements that stretch how much of a routine each 16-deep AMD branch-record window reconstructs, all respecting the silicon ceiling and the “never emit corrupt as complete” rule.

    • #2B slot-efficient branch filtering (opt-in, SCOPE-SAFE). New asmtest_hwtrace_options_t.branch_filter (default 0 = PERF_SAMPLE_BRANCH_ANY, unchanged). Nonzero requests a reduced HW filter (COND | IND_JUMP | ANY_CALL | ANY_RETURN) that drops only the direct unconditional jmp — its target is statically decodable, so it need not consume a scarce LBR slot — and the reconstructor follows it from the region bytes for a byte-identical trace over a longer window. Dropping direct call too was deliberately rejected (an out-of-region-callee return strands the pre-call in-region code — a silent-corruption risk). The decoder is unified/no-flag: amd_replay follows a dropped jmp only when one appears mid-straight-line-walk, which under the default full filter can never happen (a taken jmp is the recorded from), so the follow path is provably dead code on the tested default. New primitive asmtest_disas_is_uncond_jump; the capture retries the full filter on EOPNOTSUPP/EINVAL so the tier stays available. Applies to both the sampled and the deterministic-snapshot paths (the statistical WindowHot survey keeps BRANCH_ANY). Host-independently validated (test_amd_reduced_filter F1–F5: dropped-jmp equivalence, back-edge-cycle termination, region-exit truncation, chained follow); two independent adversarial reviews confirmed the classify/follow logic exhaustive over every x86-64 CTI. Live-validated + reach-measured on Zen 5 (Ryzen 9 9950X, test_branchsnap): the deterministic snapshot with branch_filter=1 follows a dropped jmp to its target block on real LbrExtV2, and reconstructs 1.86× more executed instructions per 16-deep window (65 vs 35) on a loop whose body has a direct jmp plus a conditional back-edge.

    • #2A period-spaced Tier-B stitching — host-independent validation + documented caveat. test_amd_stitch_period_spaced proves the landed lbr_period path stitches period-spaced (P=4) windows of a distinct-edge path back to the exact sequence, and asserts the flip side: a self-similar loop silently undercounts under period>1 (the smallest-overlap heuristic can’t tell 1 iteration from P) — which is why the default stays lbr_period=0 (period=1, universally exact). Live-measured on Zen 5 (test_amd_reach_period): confirms the finding on real hardware — period=4 reconstructs fewer instructions than period=1 (231 vs 297) on a loop, since every loop is edge-self-similar, so period-spacing’s reach benefit is confined to (inherently short) distinct-edge paths, not loops.

    • #3 deterministic snapshot by default for single-exit regions. hwtrace_begin_amd now selects the Phase-3 boundary snapshot by default on the supporting substrate (amd_lbr_v2 + perfmon_v2 + Linux ≥ 6.10), but only when the region has a lone ret (amd_last_ret_off now counts rets) — the one exit breakpoint is then guaranteed hit, so the common small routine gets deterministic capture with no richest-window guessing. Multi-exit routines (which an earlier ret could make the breakpoint miss) keep the sampled path; explicit opts.snapshot is honored for any region and every arm failure falls through to sampling. Validated: docker-hwtrace-amd (328 decoder checks green) and docker-hwtrace-codeimage (branchsnap marker path green on the Ryzen 9 9950X). Only bindings/dotnet mirrors the new branch_filter field (matching the shipped lbr_period posture); the field is an ABI-safe tail append (struct stays 48 bytes).

  • Whole-window scope (begin_window/end_window/render_window) in the Node and Java bindings — the region-free, empty-ctor using (new AsmTrace()) §Z1 substrate (Phase 2, increment 1 of the dotnet-parity roadmap). HwTrace.window(fn) (Node) / HwTrace.window(Runnable) (Java) arm a single-step capture on the calling thread with NO registered region, run the body, disarm, and render the executed absolute addresses from live memory — returning {path, truncated, insns[]}. It is FAITHFUL-BUT-NOISY by design: single-stepping the managed runtime records everything between begin and end (the FFI dispatch + runtime), so the traced routine’s own addresses appear as a subset. A single V8-dispatched call runs ~100k instructions (captured cleanly, subset verified); a HotSpot + FFM call exceeds the single-step whole-window’s internal SS_WINDOW_CAP (1<<20), so the Java capture faithfully reports truncated (best-effort). Validated in docker-hwtrace-node / -java. The begin_window/end_window/render_window ALL exemptions stay in scripts/bindings-parity-allow.txt, now consumed by the seven bindings that don’t wrap them.

  • call_scoped — a registry-free traced native call — now in ALL TEN bindings. The Python/Ruby/Node/Java bindings shipped it first; the remaining five (C++, Rust, Zig, Lua, Go) now wrap it too. Each wraps asmtest_hwtrace_call_scoped_ex + asmtest_hwtrace_render_scope: arm, call the native leaf, and disarm entirely in native code — a tighter window than the scope form (whose FFI dispatch of code.call is stepped, though region-filtered) — returning the call’s result, the executed body’s disassembly, and the truncation bit in one step. Registry-free, so it is safe in a tight loop (no MAX_REGIONS exhaustion). HwTrace.call_scoped(code, *args) (Python/Ruby), HwTrace.callScoped(code, …args) (Node), HwTrace.callScoped(code, long…) (Java), HwTrace::callScoped(code, args…) (C++), HwTrace::call_scoped(&code, &[args]) (Rust), HwTrace.callScoped(&code, args) (Zig), HwTrace.call_scoped(code, ...) (Lua), CallScoped(code, args…) (Go); each returns {result, path, truncated, rc} and each is validated in its Docker lane (result 42, body renders to ret in 5 insn lines, a 40-call loop with no exhaustion). Struct-by-value for the 8-byte asmtest_hwtrace_scope_t handle is native in the five new bindings (C++ POD, Rust #[repr(C)], Zig callconv(.C), LuaJIT FFI, cgo) — no packing, unlike the Ruby/Java bridges. With all ten now wrapping the pair, the ALL exemptions for call_scoped_ex/render_scope leave scripts/bindings-parity-allow.txt.

  • §D0.4 async-hop stitching now has a LIVE producer — AsmStitchedTrace (.NET). The shipped asmtest_hwtrace_stitch merge core previously had no live producer (only synthetic-slice host tests). New asmtest_hwtrace_stitch_handles(traces[], scope_ids, seqs, tids, versions, n, out, bounds, nbounds) (src/hwtrace.c) is the binding-facing bridge — it merges N already-captured trace handles (the slice struct embeds heap pointers a binding can’t marshal by value). On top of it, the .NET AsmStitchedTrace carries an AsyncLocal<scopeId> across await/thread hops and feeds each hop’s managed-safe lazy-arm capture to the core, so one logical operation traced across a real Task.Run thread hop stitches its per-thread slices in seq order. Each hop uses the new registry-free asmtest_hwtrace_call_scoped_ex ([base,len) direct, no MAX_REGIONS slot) so a long-running operation with many hops cannot exhaust the fixed 32-slot region table process-wide. Validated on the single-step tier (no Intel PT needed): host test_stitch_handles / test_call_scoped_ex (incl. a 64-call no-exhaustion check) and the .NET lane (scope id flows across the hop; two different-thread hops merge with correct bounds; 40 operations all capture).

  • FP shim family for the lazy-arm scope — (double…)->double methods trace in-process. asmtest_hwtrace_call_scoped_fp (src/ss_backend.c / src/hwtrace.c) dispatches a homogeneous double signature through the SysV FP ABI (xmm0..7 args, 0-8 arity). The .NET AsmTrace.Method(...).Invoke now tries the integer (long…)->long shim, then the FP family, before falling back out-of-process — so a double-signature method is captured in-process instead of degrading. Host-tested (test_call_scoped_fp) and on the .NET lane. See managed-singlestep-lazy-arm-plan.md.

  • Managed single-step is now safe by construction — AsmTrace.Method() lazy-arms only the method body. New asmtest_hwtrace_call_scoped(name, fn, args, nargs, result, out) (src/ss_backend.c / src/hwtrace.c) arms the single-step window, calls the target through the SysV integer ABI, and disarms — all in native code — so the region filter keeps only the body’s offsets and NONE of the caller’s or a managed runtime’s machinery is ever under EFLAGS.TF. The .NET Invoke no longer steps DynamicInvoke in-process (the crash surface where an in-window pthread_create that blocks SIGTRAP force-killed the process on slow hosts): it marshals through a (long…)->long shim table and, for signatures the shims can’t express, auto-falls back to the out-of-process stepper with a loud SkipReason — never a silent miss. HwTrace.DegradationNote() gains the faithful managed-window warning. Host-tested (test_call_scoped, byte-for-byte parity with begin/end) and validated on the .NET lane; see managed-singlestep-lazy-arm-plan.md.

  • Slow-host crash-avoidance stress lane (make hwtrace-dotnet-stress, CI: docker-hwtrace-dotnet-stress in the hwtrace-bindings job) — the lazy-arm plan’s “Sharpening 1”. The ONE lane that runs with CoreCLR’s tiering worker unpinned (no DOTNET_TC_BackgroundWorkerTimeoutMs): it parks past the worker’s idle-exit, churns tier-up enqueues on the invoking thread (fresh DynamicMethods driven past the call-count threshold), and interleaves lazy-arm Invokes — recreating on the loaded CI runner the exact environment where the old stepped-DynamicInvoke path died with exit 133. Surviving with every capture intact is the pass signal.

  • The zig toolchain tarball is now integrity-pinned — the one third-party fetch P2’s supply-chain pass left unverified. DOCKER_SETUP_zig verifies a per-arch sha256 (ZIG_SHA256_x86_64/_aarch64 in mk/docker.mk) before extracting, the anchors are recorded in scripts/third-party-digests.txt, and check-thirdparty-versions.sh now asserts both anchors exist for the declared ZIG_VERSION — so a version bump that forgets the digests fails loudly.

  • Example suites are now auto-discovered. Every examples/test_foo.c + examples/foo.s pair (foo.asm under ASM_SYNTAX=nasm) links through a test_% pattern rule — drop the two files in and make test picks the suite up, with no Makefile edit. Legacy pairs whose routine object doesn’t match the test name (test_arith → add.o, test_capture → flags.o, test_struct → structs.o) keep explicit link rules, and SUITE_EXCLUDES lists the test_*.c files owned by other targets (bench, usecases, demos, the emulator/trace tiers). This makes the long-standing docs claim in writing-tests.md true instead of correcting it downward.

  • asm_call_capture_vec256_win64 and asm_call_capture_vec512_win64 are now declared in asmtest.h (under -DASMTEST_ABI_WIN64), completing the Win64 mirror of the System V capture surface. Both existed in src/capture_win64.asm and were exercised by the Win64 suite, but a consumer following the win64 guide had to hand-declare the prototypes; the guide’s entry-point table now lists _vec512_win64 too.

  • Wide-arity, mixed-FP, and struct-return capture reachable from all ten bindings (N4 of the 2026-07-04 review — previously the array-form C entry points existed but no binding referenced them). Three FFI-friendly shims join asmtest_capture6/_fp2/_vec_f32 in the opaque-handle layer: asmtest_capture_args (stack-spilling wide arity), asmtest_capture_mix (integer + FP register files together), and asmtest_capture_sret (hidden-pointer struct return). The struct-layout bindings (C++/Rust/Zig/Python) call the array forms directly per their existing idiom; the opaque-handle bindings (Node/Java/.NET/Ruby/Lua/Go) wrap the shims. Every binding gained wide-arity (sum8), mixed (mix_scale), and struct-return (make_big) conformance tests — all ten docker lanes green; fixtures registered in the corpus name table (no repeat of N7); NASM counterpart included.

  • Docs: Teaching with asm-test (the in-repo scope of P5 from the 2026-07-04 review). The instructor recipe the primitives always supported but nothing documented: a three-file assignment layout (student .s, rubric grade.c, grade-time Makefile), a complete GitHub Classroom autograding config (one scored step per rubric item via --filter + --fail-if-no-tests, so a deleted rubric test is a scored zero, not a free pass), and instructor notes on hidden tests, timeouts, and grading non-x86 courses through the emulator tier. The separate “Use this template” assignment repository remains a maintainer action.

  • .NET examples roadmap — the full remaining tail (11 items from dotnet-examples-roadmap.md, all instruction-count-faithful, all green in the docker lane). Five new reports: flatprofile (perf-report parity: self / Overhead % / cumulative %), amplification (user vs BCL vs native-runtime split + the WEAK-tier factor), runtimegaps (largest RuntimeBefore bursts by the method they precede), footprint (code working-set pages + jump-distance locality), and runtimebuckets (the ~1M-insn runtime lump named by module — resolved per 4 KB page, not per address, so ~hundreds of /proc lookups instead of ~1M). Six new example projects: instructionmix, perfannotate, loops (backedge trip counts), descent (native call-descent tree with self/inclusive counts), descent_dotnet (out-of-process call descent into a live CoreCLR — descends Program::Leaf twice as nested frames; jit_dotnet gained an additive chain mode), and codeimage (one address, two code bodies over logical time). The binding gained HwTrace.SymbolizeBuckets over the already-exported asmtest_hwtrace_symbolize_bucket (.NET suite 123 → 126).

  • Consumer-facing CI integration (P2 of the 2026-07-04 review). A composite GitHub Action at the repo root (action.yml, “Setup asm-test”: POSIX-sh steps; inputs version/prefix/optional-tiers/test-command; exports PKG_CONFIG_PATH and library paths), an includable GitLab CI template (ci/asmtest.gitlab-ci.yml, .asmtest-install + a documented consumer job with JUnit wiring), and a CI integration guide covering both plus the raw make install fallback. The wrapped install recipe is proven end-to-end locally; the Action’s uses: path needs a real Actions run.

  • macOS clean-room plan — Track E finished, Tracks C/D written. release.yml’s seven smoke blocks now source scripts/clean-env.sh instead of ad-hoc cd /tmp && env -u scrubbing (behavior-preserving; interpreters resolved to absolute paths before the PATH scrub), and the methodology is documented in docs/clean-room-testing.md. Tracks C (scripts/osx-vm.sh + make osx-vm-test, tart VM) and D (scripts/docker-osx-bindings.sh + make docker-osx-bindings, Docker-OSX/KVM) are written per the plan’s spec and clearly banner-marked UNVALIDATED — they need Apple-Silicon-tart / bare-metal-KVM hosts this environment lacks.

  • AMD tracing plan Phase 2 & 3 follow-ups — attached block-step + snapshot marker routing. Completes the two sub-items the earlier block-step / snapshot commits left open:

    • asmtest_ptrace_trace_attached_blockstep — the third public block-step symbol. Block-steps a SEPARATE, externally-attached process (one debug exception per taken branch, intra-block instructions reconstructed with Capstone), reading foreign bytes via process_vm_readv and leaving the target stopped past the region for the caller — the rootless managed-runtime completeness fallback. Wrapped in all ten bindings; a new test_ptrace_attach_blockstep asserts the stream is byte-identical to the per-instruction attached tracer over a true external attach.

    • opts.snapshot begin/end routing on AMD — the deterministic boundary LBR snapshot (bpf_get_branch_snapshot at a region-exit hardware breakpoint) is now reachable through the ordinary begin/end markers, not just the standalone asmtest_amd_snapshot_trace. The capture split into asmtest_amd_snapshot_begin/_end (armed single-slot); the AMD marker path derives the exit from the region’s last ret and falls back to the sample_period=1 sampled path when the BPF toolchain/caps/LbrExtV2 substrate is absent.

  • AMD hardware-trace improvements — Phases 0, 4, 5 of the AMD tracing plan. Completes the P0/P1 near-term work on the AMD LBR backend, all validated live on the Zen 5 dev box (Ryzen 9 9950X, amd_lbr_v2) via make docker-hwtrace-amd:

    • Phase 4 — LbrExtV2 speculation-bit filtering. amd_replay now drops a perf_branch_entry whose spec == PERF_BR_SPEC_WRONG_PATH (a speculative, never-retired phantom edge) before reconstruction; dropping it is expected, so it does not set truncated. The spec field (Linux ≥ 6.1) is gated behind a -fsyntax-only struct-member build probe (ASMTEST_HAVE_PERF_BR_SPEC), so the filter compiles out cleanly on older headers / Zen 3 BRS. amd_edge_eq (the stitcher’s from+to overlap key) is untouched.

    • Phase 5 — Tier-B stitch hardening. asmtest_amd_stitch gained a decodable-distance guard: a smallest-overlap match is accepted only if the adjacency it splices is real straight-line code, so a dropped/throttled-sample mis-stitch becomes a faithful gap instead of a silently-wrong trace. The AMD data ring default grew 64 KB → 256 KB to extend gapless stitch reach before the kernel drops samples; the data_size header comment now documents both backend defaults.

    • Phase 0 — runtime branch-stack depth. asmtest_amd_lbr_depth() reads the true depth from CPUID 0x80000022 EBX[9:4] (lbr_v2_stack_sz), replacing the hardcoded 16 in the Tier-A/Tier-B split, stitch bound, and LOST heuristic. A no-op today (every shipping Zen reports 16) that removes the assumption.

    • Phases 6 (Zen 3 BRS period-adjust) and 7 (IBS-Op coverage) remain forward-look — they require Zen 3 / Zen 2 silicon the dev box lacks, and the project does not ship hardware code it cannot self-validate.

  • Docker-OSX clean-room lane (Track D): containerized sshpass, repointed at surviving upstream tags. New Dockerfile.sshpass + make docker-sshpass build a small asmtest-sshpass image; scripts/docker-osx-bindings.sh now runs every ssh/scp-equivalent call through it instead of requiring a host sshpass install (and the sudo that would need), per CLAUDE.md’s “add it where the work runs” rule. Separately, sickcodes/docker-osx deleted every tag but :latest/:master from Docker Hub in 2024 (:ventura and friends now 404) — DOCKER_OSX_IMAGE defaults to :latest, and the script gained DOCKER_OSX_DISK support (-v <disk>:/image -e IMAGE_PATH=/image) plus a one-time-install recipe in its header, since a virgin :latest boots the macOS installer rather than a headless system. See macos-cleanroom-lanes.md.

Changed

  • Selection is now one shared brushing-and-linking model — a pick in any pane cross-highlights the same entity in detail/disasm/Loom/3D at once (docs/internal/archive/gui/22-selection-and-search.md T1, F7). Every view used to hold its own selection, so an analyst re-found the same address by hand in each pane. Selection is now ONE Workspace/shell-level entity ({rec, step, offset, lane} plus a bumped epoch), held distinctly from navigation (nav.current points a view; the selection brushes an entity): a pick in any pane — the timeline, the slice explorer, the Loom, a 3D drill — cross-highlights that same entity in every pane it appears in, and only there. A pane that cannot show the entity shows nothing selected rather than a fabricated row (D7), and cross-highlighting brushes in place without yanking every pane’s viewport.

  • Fidelity chrome is now a graded 3-tier system over a derivable severity field (docs/internal/archive/gui/23-graded-truth-layer.md T1, F5). The proliferating fidelity forms — a redaction placard, a statistical chip, a coarse chip, a bounded-window note and a torn banner, all equally loud and non-collapsible — collapse into ONE vocabulary: one banner, one inline chip, one glyph set, with mandated placement (a banner is pane-top, a chip is on a header row, enforced by the component API). Loudness follows the schema’s own severity gradient: neutral (skip / statistical / redacted / coarse / bounded) is a quiet chip; caution (truncated-but-usable / paused gap) is an amber banner that collapses to its chip after first read; integrity (torn / mixed-basis refusal / a drop on an exact capture) is a loud red banner that never collapses. This restructures the fidelity layer and removes no truth — every field still renders, graded against the committed low-fidelity fixtures (D7).

  • The live session-end state is now a persistent, cause-distinguished in-pane placard with an inline fix (docs/internal/archive/gui/23-graded-truth-layer.md T2, F20). The single collapsed “ended” is fanned into its cause — stopped-clean, torn (host crashed), torn (EOF), or a PROTOCOL-MISMATCH (a stale build/asmspy that printed a usage banner and exited 0) — each with the trust of the on-screen data, and the protocol-mismatch placard carries the verbatim one-line fix (make cli; Disconnect + reconnect). It persists across frames until the next Connect/Start; toasts (doc 16) supplement it, never replace it.

  • “Paused” is split into an operator pause and a budget block (docs/internal/archive/gui/23-graded-truth-layer.md T3, F23). The bare word named two states with disjoint recoveries; now an operator pause reads “PAUSED (you) — Resume” and a budget preemption reads “BLOCKED — jack held by <session> on <target>” with explicit Swap (a named two-step confirm), Queue (a cancellable chip that starts when the jack frees), and Cancel. No path auto-swaps without confirmation.

  • Desktop view tabs are now data-driven, and first run is a task rail (docs/internal/archive/gui/20-workspace-and-settings.md T1/T2). Only the views a recording can actually fill are shown — a bare-log recording no longer presents empty Loom / ABI-x-ray / 3D / Scrubber tabs; the views it cannot fill collapse into ONE faithful “unavailable views (N)” affordance that still names each absent view and its verbatim machine reason (D7 — restructured, never removed), and the set is scoped by the active mode. First run replaces the “choose a door” chooser with a persistent task rail (Learn how assembly runs / Open a trace / Capture a live process / Author a routine); an empty workspace auto-lands in Learn (the dependency-free path), and the chosen mode drives its dock perspective so the label and the layout agree. Resolves the plan-says-3-doors vs build-renders-4 drift (F13): there are four task modes, named as tasks, and the “door” vocabulary is retired.

  • CVD-safe categorical palette + a second channel on every colour-coded distinction (docs/internal/archive/gui/24-one-visual-language.md T2). Every categorical distinction now also carries a NON-colour channel so ~5% of users who cannot read the hue still read the axis: cone direction gains a shape glyph beside each node (◄ inflow / ► outflow / ● selection / · off-cone), the Loom take axis reads by pattern (solid = hot / hollow = dim / dashed = unaligned) with a named inline legend, and the src×dst hot-edge heatmap uses a CVD-safe, perceptually-uniform ImPlot colormap (Viridis) whose labelled ColormapScale is the magnitude channel. A pure ui/cvd.{h,cpp} simulates protan/deuter/tritan and computes WCAG contrast; the palette is verified in test (text ≥4.5:1, fills/borders ≥3:1 at the smallest font), and the caution amber (dt_warn) is marked large-text-only. The shared legend renders from ONE encoding table, so the legend is itself the proof no distinction rides on colour alone.

  • One filter affordance + one time-position widget across the desktop views (docs/internal/archive/gui/24-one-visual-language.md T4). A single type-to-narrow filter with a “showing N of M” count (ui/filter.h) replaces the ad-hoc client-side idioms, free ImGui column-sort landed on the hot-edge table (reordering the view, never the recorded model order), and ONE time-position widget (ui/timepos.h) now carries two faithful variants — a continuous scrub where a real total exists (the Loom/3D playheads) and a discrete step where it does not (Invocations, the disassembly logical-time control), the discrete case VISIBLY MARKED as an intentional fidelity choice with its verbatim reason. The counts, sort order and discrete-reason registry are asserted headlessly.

  • The capability panel leads with what the host can do (docs/internal/archive/gui/18-breach-stops.md T4). It opened with a wall of red errno rows; it now leads with a one-line positive summary (the available backends) and the standing “Learn and Author work here — no root, hardware, or attach needed” floor, and demotes the unavailable backends under an expandable “why can’t I capture X?” that keeps every verbatim machine reason and adds the shared attach_verdict remedy (paranoid / Yama / i386 / CAP_SYS_PTRACE) where one applies. A bare host no longer reads as “the tool does not work here”.

  • The keyboard-help overlay marks planned-vs-wired bindings (docs/internal/archive/gui/18-breach-stops.md T1): it is generated from a wired flag and greys any not-yet-mapped binding “planned” instead of advertising it as live.

  • The desktop views are now real dockable panes — the docking layout manager, presets and Reset act on visible windows (docs/internal/archive/gui/19-dockable-panes-keystone.md). Before this, the layout manager docked five named windows (Home, Recording, Scrubber, Inspector, Timeline) and the View menu offered Reset + three presets, but no view was Begin()’d under any of those names — the dockspace, every preset, tear-out and Reset acted on phantom windows and were inert, and the whole view surface was a single window nesting three exclusive tab levels (recording → views → observer), so only one view was ever visible. The shell now Begin()s each region as a real pane hosting the active recording, the inner exclusive views tab bar is deleted, and the 3-deep nesting is flattened to at most two (a pane and, for the Loom and the Observer deck, their one data-gated inner bar). The concrete payoff: the timeline, the scrubber and the Observer’s disassembly can be shown at the same time. View-menu presets and Reset now rearrange visible panes, and panes can be torn out and restored — the bottom region was split so a preset holds the timeline and the scrubber together, and Reset rebuilds the default split (the recovery path for a stale/corrupt persisted dock .ini, whose auto-fallback lands with doc 18 T2.2). Every view keeps its exact body and its fidelity placards — the chrome was restructured into the panes, never removed (D7). The non-docked path (the null test backend’s default) still draws the single-window tab layout unchanged, so nothing regresses without a dockspace. The panes are asserted headlessly (make desktop-test flips docking on and checks each region window exists and is active, that three siblings are active at once, and that a preset switch moves real panes) and the tear-out → Reset round-trip is driven by the interaction lane (make desktop-ui-test).

  • One semantic colour palette in the desktop’s theme.h (docs/internal/archive/gui/24-one-visual-language.md T1). The good / bad / maybe / changed / cone (back·fwd·both·off) / selected / statistical axis was being re-invented inline in every draw file — three barely-distinguishable yellows meant three different things and two reds both meant “refused”. They now live in one place as named accessors (dt_good_col … dt_statistical_col, each with a paired _u32), each documenting its ONE meaning, so a colour can no longer drift its meaning between panes; dt_bad_col() is now literally dt_refuse_col() (the 0.90-vs-0.95 refuse-red split is gone). Every per-pane colour literal at the drift sites — the Inspect verdicts, the scrubber / ABI-x-ray “changed” highlight, the slice cone hues, the 3D-HUD chips, and the Loom refusal placard — is deleted and routed through the accessors. A shared ui/legend component renders the palette (with a non-colour second-channel token beside each swatch) so a legend can never disagree with what a view draws; the slice explorer uses it. The fidelity chrome (D7) is unchanged — the refuse red and caution amber keep their meaning; statistical is merely named so a later graded-fidelity change can move it off the amber without touching a call site. Header-only and engine-free, so it links into the full app, the render-only viewer, and the null test backend alike.

  • AMD manual pre-release validation shrunk to the runner-uncoverable residue (self-hosted-ci-runners.md T4). docs/internal/amd-hardware-validation.md was the “one validation step that cannot run in CI”; now that the self-hosted hwtrace-privileged-zen lane runs the exact LbrExtV2 + live-IBS paths on a Zen 4/5 runner, that tier moved to CI and the doc is reframed as the residue — the four AMD paths the Zen 4/5 runner cannot reach, each with its command, hardware/privilege gate, and owning doc: Zen 2 IBS-without-LBR degradation (make docker-hwtrace-ibs on the Ryzen 9 4900HS), MSR-direct (make docker-hwtrace-msr, --privileged + host msr module — kept off the CI runner by security policy), the status live-EPERM path (unreachable root-in-container), and Zen 3 BRS (a link to amd-branchsnap-lbr-docs.md#T8, not a checklist item). The call_auto non-escalation regression signal is now enforced by the CI lane’s assert rather than a manual eyeball.

  • Linux Python wheels now build on the manylinux_2_28 floor (install on older distros). The two Linux legs of the release python job build inside quay.io/pypa/manylinux_2_28_{x86_64,aarch64} (AlmaLinux 8, glibc 2.28) instead of the ubuntu-latest glibc, via scripts/build-manylinux-wheel.sh — which source-builds the four native engines the image lacks (unicorn/keystone/capstone/libipt, pinned; libopencsd is a dead link-only dep and skipped) + fetches DynamoRIO, runs make python-package, and auditwheel repair --plat manylinux_2_28 with the load-bearing tier --exclude list. make docker-python-manylinux proves it end to end with no credentials: the manylinux_2_28 wheel installs and imports (asm + disas) on a clean AlmaLinux 8. (distribution-packaging.md T5.)

  • The out-of-process whole-window stepper block-steps where PTRACE_SINGLEBLOCK is functional (~4–10× fewer stops). The §D3 stealth whole-window helper now drives asmtest_ptrace_trace_attached_window_stop_blockstep (one #DB per taken branch) instead of one stop per instruction, degrading to the exact per-instruction stepper on a DEBUGCTL.BTF-masking hypervisor (asmtest_ptrace_blockstep_available() false). The output is byte-identical either way — a cost upgrade, not a fidelity one — which is what makes a managed whole-window affordable out of process. ASMTEST_STEALTH_NO_BLOCKSTEP=1 forces the per-instruction stepper.

  • Registry-publish pipeline moved toward keyless publishing (scaffolding; gated on registry setup + a real tag). release.yml now publishes PyPI via a dedicated OIDC Trusted Publishing job (pypa/gh-action-pypi-publish, collecting every matrix leg’s wheel — the action is Linux-only) and crates.io via rust-lang/crates-io-auth-action, both with job-scoped id-token: write and no stored token; npm publishes with --provenance. bindings/java/pom.xml gains the Maven Central metadata + dormant source/javadoc/gpg/central-publishing plugins (activated only by a real mvn deploy; make java-package still uses javac + jar). The manylinux wheel floor is recorded as manylinux_2_28. All of these are credential/registry-gated — a trusted-publisher registration (PyPI/crates.io), MAVEN_* secrets + a namespace/PGP key (Maven Central), and a CI dispatch (manylinux) must land before they go live; see releasing.md.

  • asmtest_pt_decode_window gained a trailing uint64_t *base_ip_out parameter (src/pt_backend.c) reporting the first decoded IP, so the whole-window PT drain can re-base its recorded offsets to ABSOLUTE addresses. Source-incompatible for a direct C caller of this internal decode entry (pass NULL to keep the prior offset-origin behavior); the facade (asmtest_hwtrace_begin_window/_end_window) and every language binding are unaffected. See intel-pt-whole-window-substrate.md.

  • Internal plan docs reconciled against the code; four completed plans archived. An audit of all 20 active plans against the source, Makefile lanes, CI, and git history found ~30 stale status markers whose drift was entirely one-directional — every one under-reported, marking shipped and tested work as “planned” or “forward-look”; nothing claimed landed was missing. The mechanism was visible in the artifacts: status was recorded by appending a dated block while section headers and tables kept their original marker. The markers are corrected (provenance preserved), and the four complete/closed plans move to docs/internal/archive/plans/ per the repo convention: live-attach data-flow (7/7), DynamoRIO taint tier (9/9, band-gated at ~11x bare), Zen2 IBS (8/8), and the managed-attach safepoint spike (closed NO-GO). Two corrections are load-bearing rather than cosmetic: data-flow-tracing-plan.md’s Phase-5 stub told readers the taint tier stopped at Increment 3 when all nine had landed; and F4’s blocker was retired — both it and Phase 4 stated that live GC-move canonicalization needed an out-of-process EventPipe consumer (“its own lift”), but that was an assumption and it was disproved: an in-process MovedReferences2 profiler delivers the exact {old,new,len} triples at a suspended-EE GC fence, is proven to coexist with DynamoRIO, and already ships in the taint tier. F4 is now wiring a proven feed to a landed transform, and is the recommended next milestone in that plan.

  • AMD-LBR reconstruction fills the entry-block prologue (fidelity fix). On a live AMD host a too-fast tiny routine’s frozen branch stack can carry spurious mid-routine landing edges, so amd_replay anchored at the landing offset and skipped the entry prologue [base_ip, landing) — a complete-reported reconstruction of a small routine undercounted its retired instructions (e.g. insns=4 vs the block-step baseline 5). amd_entry_fill now prepends the clean straight-line prologue all-or-nothing (a branch/ret/overshoot in that run faithfully truncates instead), with a symmetric trailing fill; anti-fabrication tests confirm no phantom instructions. New test_amd_live_smallroutine hard-asserts complete⇒full-count on a live AMD host. Verified across 3 privileged runs + a 30-iteration loop: every complete reconstruction now yields the full count, and the batch-3 case-(b) escalation invariant is preserved 30/30. (The residual case-(a) advisory from the Zen 5 findings doc is resolved.)

  • Multi-exit deterministic BPF boundary snapshot, default-on (followup Phase 5 / F13). hwtrace_begin_amd now plants one hardware breakpoint per region exit (1–4 exits, one debug register each) via asmtest_amd_all_exits + asmtest_amd_snapshot_begin_multi, so whichever ret/tail-call a multi-exit routine leaves through hits a boundary — the old single-exit gate missed earlier exits and truncated. A BPF-side drop counter drives an faithful truncated-on-drop contract (F13: a dropped ring record marks the result truncated, never silently complete — verified with a 1670-drop overflow fixture). >4 exits or any arm failure falls through to the sampled path unchanged. First live-validated on Zen 5 via the new privileged docker lane.

  • First-class privileged hardware-capture docker lane (make docker-hwtrace-privileged). Runs hwtrace-test ibs-test under --cap-add=PERFMON alone (no --privileged, no SYS_ADMIN, default seccomp) — the first live validation of the exact AMD LBR (LbrExtV2) and IBS capture paths on the Zen 5 dev box: the previously-skipping AMD/IBS live tests (LBR capture, Tier-B stitch, per-thread concurrent fds, sample_window, IBS survey_pid/survey_process) all run and pass (test_hwtrace 389/389, test_ibs 23/23).

  • make BUILD=<abs> test / usecases / valgrind now work. The suite-loop recipes ran ./$$t where $$t already holds a $(BUILD)/-prefixed path, which broke under an absolute BUILD override (.//tmp/...); dropping the ./ prefix (the path always contains a slash) completes the earlier out-of-tree-build fix.

  • jit_trace’s JIT lanes prefer the byte-identical block-step rung (review F18). The no-descent lanes select asmtest_ptrace_trace_attached_blockstep when asmtest_ptrace_blockstep_available(), falling back to the per-instruction stepper otherwise; the *-descend lanes intentionally stay per-instruction (block-step has no descent parameter). Verified byte-identical on a live V8 target: the ASLR-normalized disasm stream from block-step matches the pre-change single-step stream exactly.

  • asmtest_amd_snapshot_end drains the BPF ring without blocking (followup Phase 8 / F15). ring_buffer__poll(rb, 200) epoll-waited 200 ms on the no-hit / faithful-truncation path (the common case) before draining; ring_buffer__consume reads the producer position directly and returns at once — every record is already committed by the time the events are disabled.

  • One cached amd_lbr_v2 cpuinfo probe (followup Phase 9 / F35/F11). The duplicated /proc/cpuinfo parse in amd_backend.c and msr_lbr.c now shares a single internal cached probe.

  • CI / tooling hardening. The documentation site now builds warnings-as-errors in a dedicated docs CI job (sphinx -W), so a broken cross-reference fails the PR in-repo instead of only reddening Read the Docs after merge. The format gate is pinned to clang-format-18 (matching make docker-fmt’s ubuntu:24.04), so it no longer risks flagging the whole canonically-formatted tree as drift the day ubuntu-latest advances past 24.04. -Werror now guards the hwtrace and cli jobs (previously only the base test/check), catching warnings in the newest, highest-churn code. A new make fix-perms target reclaims root-owned build/ artifacts a docker-* lane can leave behind (which otherwise break make clean). .dockerignore now excludes every root Dockerfile* and the actual bindings/Dockerfile.lang (the stale bindings/*/Dockerfile glob matched nothing).

  • Internal working docs moved under docs/internal/ — docs/plans/, docs/analysis/, docs/reviews/, and docs/archive/ are now docs/internal/{plans,analysis,reviews,archive}, with one archive rule (done plan / fully-actioned review → archive/; see docs/internal/README.md). The four completed scoped-tracing plans and the fully-actioned 2026-07-04 review moved to archive/ accordingly, every in-repo reference (docs, comments, workflows, this changelog) was repointed — including a dozen references left stale by the earlier archiving commit — and the Sphinx exclude_patterns/docs-gate now exclude internal/** wholesale.

  • Docs/README accuracy pass from a full docs-vs-code review. The entry-point pages no longer claim the language packages are published (nothing is on a public registry yet — the release pipeline is ready but uncredentialed); the README slimmed to pitch + highlights + links (the capability list’s single source is now docs/reference/features.md) and its DynamoRIO link points at the tracing guide; --bench-format=text|json and --help joined the runner/benchmark flag tables; installation.md gained Keystone/Capstone rows and the real --emu dependency set; integration.md shows the -x assembler-with-cpp assemble step and the asmtest-emu pkg-config module; api-reference.md gained ASSERT_ABI_PRESERVED_VEC/asmtest_check_abi_vec; java.md’s JDK requirement, rust.md’s shipped Tier-2 asserts, Zig’s raw-@cImport status, and the single-step tier’s macOS support are stated consistently; and CONTRIBUTING gained “Adding an example suite” and “Building the docs” sections plus a per-language lib-setup cheatsheet in the bindings overview.

  • CI: the pinned Keystone/Capstone source builds are now cached (K1 of the 2026-07-04 review — the ~20-identical-multi-minute-LLVM-compiles-per-push item). Host builds cache via actions/cache + a new scripts/thirdparty-cache.sh (exact cmake-installed file set, keyed on OS/arch + the pinned versions); docker builds of asmtest-bindings-base cache via buildx type=gha behind a new overridable DOCKER_BASE_BUILD in mk/docker.mk. ci.yml only — release.yml deliberately stays cache-free. Non-fatal by design on any cache miss or backend outage; the warm-cache path still needs a real Actions run to confirm.

  • Docs: “asm-test vs. alternatives” (P4 of the 2026-07-04 review). A maintained comparison against the four workflows people actually use instead — a C unit framework with .s files linked in (cmocka/Criterion/Unity/gtest), raw Unicorn scripting, qemu+gdb, and asmUnit-style in-asm macros — including a truthful “when the alternative is the better choice” for each, a “what asm-test does not try to be” calibration list, and a capability matrix. Linked from the README reference funnel and the docs index “Where to start”.

  • Call-descent built-in default denylist — asmtest_descent_use_default_denylist (the one unshipped Phase-5 deliverable of the call-descent plan). Arms the L3 DESCEND_ALL safety set the plan promised: at trace start the backend populates the handle’s deny pool from the tracee — the dynamic linker’s executable mappings (the lazy-binding PLT resolver) and [vdso]/[vsyscall], managed-runtime GC/JIT modules by mapping name (CoreCLR/Mono/HotSpot/ART/V8/BoehmGC), and, on the fork path (tracee shares the tracer’s layout), dlsym-resolved entry points of the classic blocking libc/pthread calls as one-byte deny regions. Denied callees are stepped over and recorded as edges; caller-supplied deny regions/callbacks compose. Wrapped in all ten bindings (parity gate green, 99 symbols × 10); a new fork-path fixture asserts a call landing exactly on poll becomes an edge, not a frame (hwtrace suite 259 → 260).

  • Emulator snapshot/restore — emu_snapshot / emu_restore / emu_snapshot_free (E5 of the 2026-07-04 review). Captures the full register context (uc_context_save) plus the extents, permissions, and contents of every mapped region; restore reinstates the bytes and the mapping set itself (a region mapped after the snapshot is unmapped again). Mapped memory deliberately persists across emu_call_* — so fuzz/mutation sweeps previously ran each candidate against memory dirtied by its predecessors; bracketing the sweep with snapshot/restore makes killed/survived classification independent of handle history. Handle-level arming (watchpoints, register guards, preloads, the fuzz corpus) survives a restore by design. Emu suite 50 → 52.

  • ASM_MIXCALL — mixed integer + FP argument capture (A6 of the 2026-07-04 review). The canonical ptr+len+scalar shape gets a first-class macro: ASM_MIXCALL(&r, fn, (buf, n), (0.5)) marshals each parenthesized group into its register file via the existing asm_call_capture_fp — no more hand-built compound literals (the repo’s own test_structparam.c hand-roll is converted). Covered by a new mix_scale example routine (x86-64 + AArch64 + NASM bodies; shared with the bindings’ mixed-capture fixtures) and the strict-c11/C++ header-portability gate.

  • FP reference models — ASSERT_MATCHES_FREF{1,2,3} (A7). Differential testing now covers the FP surface where rounding/NaN/lane bugs live: double tuples from an asmtest_fgen_fn generator run through the FP register file (asm_call_capture_fp) and the C model, judged by ULP distance (ulps = 0 = bit-exact; NaN matches only NaN). Example property tests pin fp_add/fp_mul against C models over dyadic rationals and a specials table (±0, ±inf, NaN, DBL_MIN/MAX).

  • Failing-input shrinking in ASSERT_MATCHES_REF* (E7). On a mismatch the tuple is greedily shrunk toward 0 / ±1 / LONG_MAX / LONG_MIN (else halved toward zero) while the disagreement persists, so the report leads with the boundary value that triggers the bug — shrinks to [0, 1] — alongside the original draw. Deterministic, bounded, and self-tested (the negative suite’s mismatch now asserts the exact shrunk tuple).

  • Runner flow control — --fail-fast, --repeat=N, --shard=K/N (R5 of the 2026-07-04 review). --fail-fast stops dispatching at the first failing test (forces the serial path; the TAP plan moves to the end of the stream so it covers exactly what ran). --repeat=N block-replicates the selection N times — with --shuffle/--seed, the flake-hunting loop. --shard=K/N runs the K-th of N round-robin slices of the filtered selection, so N CI jobs can split one suite with no test lost or duplicated (self-tested: shards 1/2 + 2/2 reassemble --list exactly). Self-tests 43 → 49.

Fixed

  • “Reset layout” is a real, always-available action with auto-fallback (docs/internal/archive/gui/18-breach-stops.md T2). It was inert — behind a View menu that only appeared with docking on, acting through phantom windows. It now rebuilds the shipped default from any state, is bound to Ctrl+Shift+R (fires with or without the menu bar, in both binaries), and a corrupt or collapsed persisted build/desktop-imgui.ini that would leave zero visible panes now auto-falls-back to the default instead of stranding the user in an empty window.

  • Author output is no longer lost on close (docs/internal/gui/18-breach- stops.md T3, F24). Closing an Author tab (or a workspace recording) with an unsaved run raised no prompt and simply erased the in-memory entry; a dirty tab now raises a save/discard/cancel guard and cannot be closed with a single silent click.

  • The pinned Keystone and Capstone source builds no longer fail on a modern CMake or GCC. Keystone 0.9.2 (Feb 2019) is upstream’s newest release, so there is no version to move to; on a CMake 4 / GCC 15 host it failed three separate ways, each hiding the next. (1) Both engines declare a cmake_minimum_required() below 3.5, whose compatibility CMake 4.0 removed — configure aborted before doing anything. New tp_cmake_compat in lib-thirdparty.sh supplies CMAKE_POLICY_VERSION_MINIMUM=3.5, emitted only for cmake >= 3.31 where the variable exists, and used by both build scripts. (2) Two of Keystone’s CMakeLists set cmake_policy(SET CMP0051 OLD), which no flag can re-enable — CMake 4 removed OLD outright — so they are patched to NEW; the policy only governs whether generator expressions appear in a target’s SOURCES property, which Keystone’s build never reads. (3) Keystone’s bundled LLVM fork predates GCC 13’s stricter header transitivity and uses intptr_t without including <cstdint>, now supplied by -include cstdint in the C++ flags alone.

    The patches are applied after the git rev-parse HEAD assertion against the recorded commit, which is what keeps the B5 pin meaningful: integrity is still proven against unmodified upstream, and every subsequent change is visible in the build script rather than baked into a vendored tarball. sed writes through a temp file rather than using sed -i, whose in-place spelling differs between GNU and BSD/macOS sed.

  • .NET: the unwarmed/PT compose in-window-JIT premise check no longer permanently self-skips on PT silicon (dotnet-pt-inwindow-jit-premise.md T1). On the only host class where the Intel PT whole-window ctor arms, the >=1 method JIT'd inside the window check took its timing self-skip on every run (MethodsObserved == 0 consistently on the i7-8559U — the dotnet-managed-pt-concurrency-plan T4 residue): the fixture’s first-call JIT is compiled inside the window, but the method-load event reaches the JitMethodMap on the EventPipe dispatch thread asynchronously, and a native-speed PT window closes sub-millisecond after the call — losing the delivery race the slow single-step sibling wins for free. New AsmTrace.WaitMethodObserved(nameSubstring, timeoutMs) polls the live map’s thread-safe CountFor so the test holds the window open (2 s bound) until the delivery lands; the premise now asserts deterministically, the self-skip arm survives only for a genuine delivery stall (never-flake rule kept), and the close-time trackBytes image now contains the fresh method so the versioned decode resolves it directly. No C-core change, no new tier symbol.

  • macOS (Intel) native build correctness (fifth pass): the hwtrace suite failed to compile at the new MSR-rung commit-decision seam (amd-review-followup-2 T4, landed from the Zen box the same day). examples/test_hwtrace.c hand-declared asmtest_trace_auto_msr_commits inside the #if defined(__linux__) && defined(__x86_64__) perf_event declaration block, but test_msr_commit_decision — deliberately pure, the test that “pins the decision itself everywhere” — calls it unconditionally, so make WERROR=1 hwtrace-test on macOS died on an implicit-declaration error. The seam is defined unconditionally in src/trace_auto.c and links on Darwin; the prototype now sits above the Linux-only block. Verified on the macOS 14.8.7 / Intel host: make WERROR=1 hwtrace-test 149 passed 0 failed (suite grew 145→149 with the four new seam checks now running on macOS too). The rest of the pass was clean at the same tree — WERROR=1 test/check (57/57), asm-test 16/16 (the new statement-drop guard green under the host’s Keystone 0.9.2), the cpp 58 / ruby 57 / python 15+12-skip hwtrace binding lanes, WERROR=1 dataflow-test build + transparent self-skip, the make cli OS-gate self-skip intact after the asmspy T2/T7 wave, and mach-stepper-test 25/25.

  • The in-line assembler no longer returns machine code with a statement silently missing (assemble-silent-statement-drop.md T1-T3). Keystone drops a statement it can only partially parse — a bad, truncated or wrong-dialect operand — leaves ks_errno at KS_ERR_OK and still returns success, so asmtest_assemble handed back ok=true with an instruction the caller wrote missing from the bytes. For a library whose purpose is asserting on machine code that is the worst failure shape available: the assertion runs and reports on code the caller never wrote. The most reachable trigger needed no typo — AT&T source assembled under the ASM_SYNTAX_INTEL default that every binding passes when the caller does not name a syntax (asmtest_assemble(…, ASM_SYNTAX_INTEL, "movq $42, %rax\nret\n", …) returned ok=true with a bare {0xc3}); an ARM-style immediate on x86 (mov rax, #42) and a truncated operand (mov rax,) took the same path, and the truncated-operand shape drops silently in all five x86 dialects and on ARM64/ARM32 as well. asmtest_assemble now counts the statements its source contains and fails the whole assemble when Keystone reports fewer — “assembler skipped N of M statements (check the syntax argument)” — through the existing error path, so all ten bindings inherit it with no ABI change (the guard is in the core; asmtest_asm_bytes never exposed stat_count, so no binding could have worked around it). The header now states the all-or-nothing contract. Callers with a test that passes today while asserting on short code will start seeing this failure — that is the point. The counter is deliberately a lower bound: it tracks string and character literals, #////block comments and the two measured dialect rules (; separates statements everywhere except x86/NASM where it starts a comment; # is a comment on x86 but the immediate prefix on ARM), and never special-cases labels, directives or comment lines because Keystone counts those as statements too — every construct it walks past can only lower the count, so a false rejection of valid code is not reachable through them. Verified by a sweep over every assembler source in the tree, each in the dialect it is written for: the tagged call-site literals, the doc’s adversarial separator/comment/literal shapes, and the examples/*.s corpus files whole — the last preprocessed the way the build preprocesses them (cpp -C, comments kept), which is what actually reaches the assembler, 14 of them landing 9-74 counted statements against Keystone’s 36-131. Zero false rejections, alongside an anti-vacuity pass confirming the guard still fires on all eight defect fixtures.

  • IBS surveys can no longer report a complete, empty capture when the edge export OOMs (amd-review-followup-2 T1), plus the round-2 review’s smaller residues (T3/T4/T5). All four IBS lanes (survey_pid, survey_process, window_end, survey_fetch_pid) discarded eh_export/fh_export’s return, so an OOM’d export surfaced as OK with n==0 beside branch_samples>0 — indistinguishable from a genuinely-empty survey; they now return EUNAVAIL, mirroring the software-clock lane, with a test seam (asmtest_ibs_test_set_export_fail) proving the contract pure everywhere and live under injected OOM on an IBS host. T3: g_open_errno now resets on every capture entry (no stale ESRCH reason after a later success, test-pinned); the two live survey drains gained the seam parser’s short-SAMPLE h.size floor (F7’s last two sites); the retired-freeze-gate comment in test_hwtrace.c now describes the substrate probe it heads. T4: the trace_auto MSR rung commits a NONEMPTY truncated partial as HW_OK+truncated — the same contract a fast-tier truncated partial returns under — instead of discarding it and, with both steppers absent, returning EUNAVAIL beside a usable 16-deep partial; the genuinely-empty read still falls through (decision extracted as a host-testable seam, pinned by pure tests). T5: the Phase-4 ASMTEST_HWDBG env-gated logging now reaches the two AMD TUs it never covered — ibs_backend.c (probe outcomes, perf_event_open errno, near-full/LOST, export-OOM decision) and trace_auto.c (per-rung commit/skip and the final mechanism). T2 doc drift: parity-matrix recommendation rows no longer put AMD LBR primary on Zen 3 (BRS-only silicon — the open is -EINVAL), Matrix 1 records MSR-direct + IBS as shipped (only Zen 3 BRS stays forward-look), and the 2026-07-09 orphan review page carries a SUPERSEDED banner correcting its “zero IBS code” premise.

  • Source-built Capstone dylib unloadable on macOS (benchmarks-ci-followups T1 validation). The first dispatched run of the nightly benchmarks (macos-15-intel) leg failed at make bench-check: dyld aborted emu-bench with Library not loaded: @rpath/libcapstone.5.dylib … no LC_RPATH's found. scripts/build-capstone.sh left CMake’s Darwin default @rpath install name on the dylib, while the tree links consumers with the plain pkg-config -L/usr/local/lib -lcapstone (no rpath). The script now bakes the absolute install-name directory (-DCMAKE_INSTALL_NAME_DIR, the Homebrew convention) — a no-op on non-Apple platforms, and the K1 cache key already hashes the script so CI rebuilds instead of restoring the stale dylib.

  • 2026-07-21 review — C2/C3, S2, S4, S6, B3, B5, B7, T2 (fixed 2026-07-22). The remaining review findings not covered by the S3/S5/S7 and D1–D3/T1/K5 batches, each with an anti-vacuity-checked test and lane verification:

    • asmspy CLI (C2/C3). --log --follow now sets PTRACE_O_TRACEEXEC and drops a followed child that execves a 32-bit image (the syscall-stream engine has no i386 table, so it would render i386 write(4) as x86-64 stat(4)); and a bare app-delivered SIGTRAP (executed int3 / hardware breakpoint, by si_code) is re-injected in the syscall-stream and --procs --count=syscalls engines instead of swallowed. make docker-cli PASS.

    • Core C (S2/S4/S6). g_pt_window (the single whole-window PT slot) is mutex-guarded with a reserved-arm sentinel so two INTEL_PT arms can’t both claim it (S2, host-testable seam proves exactly-one-of-8 wins); the AArch64 wrong-depth hardware-breakpoint resume single-steps over the re-matching PC instead of relying on x86 EFLAGS.RF (S4, x86 byte-identical, aarch64 compile-checked, behavioural test gated on bare-metal NT_ARM_HW_BREAK); and ~20 growable pools route through a shared overflow-checked asmtest_grow/asmtest_grow_pow2 helper so a capacity double can no longer wrap size_t to 0 (S6, new tests/grow_overflow unit test in make check). make docker-hwtrace green (621/0 here; the PT-window mutex + pool guards join the S3/S5/S7 batch).

    • Bindings (B3/B5/B7). The .NET long-lived native-handle types (DrTrace/HwTrace NativeCode, NativeTrace, HwTrace recorder) are now IDisposable with a finalizer backstop, and the Java equivalents plus AddrChannel/CodeImage register a java.lang.ref.Cleaner, so a dropped handle no longer leaks its native mapping/trace (B3; Descent’s late-bound upcall arena staged; reclamation tests reclaim ~0 of N leaked mappings). Go (a compile-time _Static_assert abicheck package) and Zig (comptime @offsetOf/@sizeOf) now fail the build on a hand-mirrored FFI struct drift (B5). Ruby reads the uint64_t asmtest_regs_ret as unsigned and Node stops Number()-narrowing koffi’s BigInt, preserving a >2⁶³/>2⁵³ return (B7). Verified: docker-hwtrace-dotnet/-java, docker-drtrace-java, docker-go, docker-zig, docker-dataflow-zig, docker-ruby, docker-node.

    • Tests (T2). tests/expect.sh documents that the standalone-TAP tier suites are gated by exit code in their own CI lanes (wiring them into make check would only self-skip, which the dependency rule forbids).

  • 2026-07-21 review batch 3 — core-C robustness S3, S5, S7 (fixed 2026-07-22). Three memory-/integer-safety hardenings in the hardware-trace core, all verified on Linux via make docker-hwtrace (514/514, 0 failed) and macOS-Intel-portable (native WERROR=1 hwtrace-test 145/145 — the fixes and their tests are Linux/ x86-64-guarded and self-skip cleanly off it). S3: render_window recorded ABSOLUTE RIPs and disassembled them straight from live self memory, so a region unmapped after capture (a JIT free, a dlclose, a torn-down stack) SIGSEGV’d the renderer. It now copies each RIP fault-safely through a new hw_read_self_live helper — process_vm_readv(getpid()), page-clamped on an unmapped straddle, mirroring pt_backend.c’s pt_read_self_live — and renders “(undecodable)” for a freed address instead of faulting. New test_wholewindow_render_unmapped mprotects the traced page PROT_NONE post-capture; mutation-checked (the raw-deref revert SIGSEGVs at that test). S5: the jitdump readers (asmtest_jitdump_find, asmtest_jitdump_debug_find) computed name_len = (long)total - 56 - (long)code_size over an untrusted jit-<pid>.dump, where a code_size > LONG_MAX cast is implementation-defined and yields a plausible-but-wrong positive name_len the <= 0 guard misses. Both now reject a code_size overflowing the declared record (unsigned, underflow-guarded) before the cast. New test_jitdump_hostile feeds a code_size == UINT64_MAX record; mutation-checked (removing the guards flips both asserts not ok). S7: round_pages clamps the caller-controlled AUX/data-ring size to 1 GiB so (v + pg - 1) cannot wrap and the power-of-two round-up cannot shift to 0 on a hostile/garbage aux_size/data_size. See docs/internal/reviews/2026-07-21-repo-review.md; the §2 remainder S2 (PT-window race) / S4 (arm64 hw-bp resume) is gated on PT / arm64 runtime this host lacks, and S6 (the ~15-site pool-growth clamp) is left as a mechanical follow-on.

  • 2026-07-21 review batch 2 — T1, D1–D3, K5 (fixed 2026-07-22). T1: the permanent SKIP("partial-fill semantics not finalized") in the default make test set is retired — the semantics are finalized and asserted: mem.partial_fill_touches_only_first_n_bytes proves fill_bytes(buf, val, n) writes the low byte of val into buf[0..n) and leaves the tail untouched (the contract all four implementations — GAS x86-64/AArch64/riscv64 + NASM — already share); green under both syntaxes, mutation-checked. D1: the two residual Zen-3 LBR overclaims (reference/features.md, reference/diagrams.md) now state the Zen 4+ live floor per _positions.md #2. D2: the two internal-engineering pages that leaked onto the public Sphinx site moved under docs/internal/ — amd_tracing_review.md → internal/analysis/2026-07-09-amd-tracing-review-f1-f47.md (the authoritative F1–F47 edition, not a duplicate as the review first framed it) and scoped-tracing-implementation.md → internal/ with links rebased; all referrers retargeted (published guide → GitHub blob URLs); Sphinx -W clean. D3: guides/win64.md’s intro no longer calls the runner port “now underway” while the body says full parity — it is complete. K5: all six third-party GitHub Actions are SHA-pinned to their then-current release commits (pypa/gh-action-pypi-publish v1.14.1 — previously the moving release/v1 branch — plus ruby/setup-ruby, rust-lang/crates-io-auth-action, mlugg/setup-zig, msys2/setup-msys2, docker/setup-buildx-action), and ci.yml’s actions/setup-python@v5 is unified to @v6; actionlint output byte-identical to before, YAML parses. See docs/internal/reviews/2026-07-21-repo-review.md; still open there: C2/C3, S2–S7, B3/B5/B7, T2.

  • asmspy no longer orphans a planted breakpoint when the tracer is interrupted (2026-07-21 review C1 — the one hole in the “never kill the target” invariant). A PTRACE_POKETEXT 0xcc is plain memory the kernel does NOT restore on tracer death, so an unhandled Ctrl-C mid region-trace left the target to execute the orphaned int3 later and die. asmspy (TUI and headless modes) now installs a SIGINT/SIGTERM/SIGHUP handler that only sets flags: the engines see their stop flag (every headless engine call now passes one), a blocked waitpid/getch returns EINTR (no SA_RESTART — the SIGALRM quit-wake contract), and the normal unplant + two-phase-detach unwind runs, with the TUI reaching endwin(). An inherited SIG_IGN is honored (the nohup convention). --sample deliberately keeps a NULL stop — that engine’s NULL means “exactly one window”, and it plants nothing. New cli-smoke leg: --trace hotfn -1, SIGTERM mid-cycle → tracer exits 0 and the victim survives a grace period of continuous re-entry (differentially confirmed: the pre-fix binary dies 143 with no detach).

  • g_amd_snap is per-thread, like every other hwtrace lifecycle slot (2026-07-21 review S1). The AMD boundary-snapshot flag sat as a process-global int beside the __thread g_fd/g_base_map/g_active it claims to share an invariant with; a concurrent snapshot arm on one thread could flip another thread’s hwtrace_end_amd onto the wrong teardown branch (leaking that thread’s perf fd/ring). Now __thread. make docker-hwtrace 697 ok / 0 failed.

  • Java HwTrace’s availability-QUERY family self-skips instead of throwing when the native library is absent (2026-07-21 review B2). status(), resolve(), auto(), resolveTiers() and autoTier() threw RuntimeException with the library unloaded, contradicting the class’s own “callers never see a throw” self-skip contract (masked only because the lib is always bundled); each now degrades to its faithful unavailable value (EUNAVAIL status/empty cascade/empty Optional). New hwtrace-java-test leg runs HwTraceTest --not-loaded-contract in a second JVM with the resolver pointed off a cliff (bogus ASMTEST_HWTRACE_LIB, cwd outside the repo), with an anti-vacuity guard that fails if the library loaded anyway.

  • Rust binding: the in-line assembler/disassembler is reachable without ASMTEST_LIB (B1), and rv64 gets its capture struct (B4) (2026-07-21 review). asm_fns() gated on dlopen($ASMTEST_LIB) alone, so the common no-env downstream case reported “not in this build” despite the crate dylib-linking libasmtest_emu (which carries Keystone + Capstone); it now falls back to dlopen(NULL) on the already-linked image — while an explicitly SET but unloadable override still surfaces as unavailable rather than being silently substituted (new own-process tests/asm_no_env.rs proves an assemble+disas round trip with the env var removed). The missing riscv64 Regs mirror of asmtest.h’s rv64 regs_t is added (a0/a1 return pair, s0–s11, always-0 flags — no flag-mask constants, as rv64 has no flags register), and the test files’ carry-fixture externs/tests gained the same arch gate the corpus itself uses; cargo check --target riscv64gc-unknown-linux-gnu --all-targets now passes.

  • Supply-chain: every DynamoRIO image fetch is digest-verified (K1), and the manylinux wheel base is pinned (K2) (2026-07-21 review). 12 Dockerfiles (plus the drtrace CI job) fetched the DynamoRIO tarball with a raw curl, bypassing the repo’s own scripts/third-party-digests.txt gate; all now route through scripts/fetch-dynamorio.sh, which refuses a download whose SHA-256 does not match the manifest, and check-thirdparty-versions.sh gates every image’s ARG DR_VERSION (was 2 of 12). Dockerfile.manylinux-wheel / release.yml built the published PyPI wheel on a floating quay.io/pypa/manylinux_2_28_* base; both now pin the dated tag 2026.07.19-1 (one knob across arches — the pypa repos publish identical dated tags per arch), with a new checker group keeping the pair in sync.

  • Build/CI mechanics (2026-07-21 review K3/K4/B6). build/asmtest_nomain.o had two competing recipes (mk/fuzz.mk vs mk/bindings.mk) — GNU make warned “overriding recipe” on every run and the winning recipe silently dropped the fuzz object’s .build-flags prerequisite (a latent stale-rebuild bug); one canonical recipe now carries the union (.build-flags dep + -Wno-unused-function). The libFuzzer/AFL++ coverage-shim lane existed but no workflow ran it — a new fuzz CI job runs make docker-fuzz (bounded, fails unless both engines find their planted crash). The Go conformance test used the exact uintptr→unsafe.Pointer round-trip HwNativeCode.Ptr() exists to avoid; go vet ./... is clean again.

  • The shared code-image is now thread-safe — a use-after-free that crashed the .NET managed multi-threaded live-PT suite ~100% of the time on real Intel PT silicon (dotnet-managed-pt-concurrency-plan.md T1/T2/T4). With libipt present on a bare-metal Intel PT box, make hwtrace-dotnet-test under --cap-add=PERFMON SIGSEGV’d on 7 of 7 runs, always just past the stitched-trace block and always on a different thread. eu-stack on the createdump cores (gdb cannot unwind them, and gdb cannot run this suite at all — it is SIGTRAP/EFLAGS.TF) put the fault in asm-test’s own code, not CoreCLR: asmtest_pt_read_codeimage ← pt_insn_next ← asmtest_pt_decode_window ← asmtest_hwtrace_pt_hop_close. Cause: asmtest_codeimage_t had no synchronization at all, while the §Z4 ambient producer uses it from two threads at once — the JitMethodMap EventPipe callback calls asmtest_codeimage_track() (which reallocs img->regions, and r->vers via ci_region_add_version) on the runtime’s listener thread, while every per-tid PT hop close decodes against that same image on thread-pool threads, walking those very arrays in asmtest_codeimage_bytes_at(). A realloc therefore freed an array a decoder was mid-walk on. Fixed by guarding the region/version arrays with a mutex held across the lookup and the append only — never across a decode — which leaves the header’s “borrowed bytes valid until asmtest_codeimage_free” contract intact (per-version byte buffers are separately allocated and freed only at image teardown). A second, narrower lifetime race is fixed alongside it in the .NET binding: the ambient handler fires on pool threads the flow is still leaving, so a hop could be published after Complete()’s drain, or be inside CloseHop — reading the map’s code-image — while Dispose() freed it. AsmAmbientStitchedTrace now makes the _completed check atomic with the _openHops add/remove and waits in-flight closes out before freeing (bounded; it leaks rather than frees under a stalled hop). The ambient live half now runs green on PT silicon (ambient: >=2 stitched slices captured) instead of crashing. New hwtrace-dotnet-ambient-stress / docker-hwtrace-dotnet-ambient-stress lane loops the concurrent set (default 25×) as the regression guard, and the timing-dependent unwarmed/PT compose: >=1 method JIT'd inside the window check now self-skips instead of flaking to not ok when the runtime happens not to compile inside the PT window. That stress lane immediately surfaced a third, pre-existing defect in the same producer — confirmed present on unmodified main by re-running the lane against the untouched producer: Complete() snapshotted _parked while a detach-driven CloseHop on a pool thread was still enqueuing its slice, so a hop that was both opened and closed could be dropped from the stitch (ambient twin: every attached hop stitched (4 vs 5 opened)). Complete() now waits in-flight closes out before the snapshot. Invisible in a single pass, which is why it survived until a repetition lane existed.

  • Intel PT whole-window & foreign-pid decode now works on real silicon — the tier had never once run on a live Intel PT box until now (intel-pt-whole-window-substrate.md T5, intel-pt-attach-foreign-pid.md T1/T2/T4, dataflow-pt-replay-tier.md T4). First live run on a bare-metal Intel PT host (Core i7-8559U, Coffee Lake) via CAP_PERFMON in-container surfaced that the PT capture was fine but the decode produced ZERO instructions on every real capture — the synthetic-fixture tests hid it because they place trace-enable (TIP.PGE) at the region base, whereas a real unfiltered capture enables tracing in the caller. The decoder’s code image only covered the tracked target region, so it hit -pte_nomap at the first caller IP and stopped before reaching the region. Fixes: (1) new asmtest_codeimage_read_live() serves the static caller/loader bytes from the target’s live memory (process_vm_readv, page-clamped) as a fallback in read_recorder, and the region-keyed read_region gained the same self-memory fallback — the decode loop still records only in-region offsets, so the temporal-JIT guarantee and the synthetic-fixture path stay byte-identical. (2) pt_aux_open now explicitly requests the pt + branch (COFI) config bits from the PMU’s sysfs format rather than relying on a kernel default enabling RTIT_CTL.BranchEn (the kernel rejects branch without pt; timing bits stay off since the decoder carries no MTC/CYC calibration). (3) pt_capture_one_region keeps its foreign victim alive through attach_end so the decoder can read the victim’s own .text (where PGE landed). (4) Corrected the hwtrace-pt-live capture-side address-filter checks to the verified silicon behavior: the filter size must not overlap the adjacent pt_filter_sibling (perf traces the whole [start, start+size) range, and the two functions sit ~0x1f bytes apart), and a @file filter naming an anonymous region is accepted-but-unmatched on this kernel (not rejected) — either way un-filterable, which is why the decode-time fallback exists. Result: make hwtrace-pt-live 631/631 and make dataflow-pt-live 29/29 — the live PT replay matching the single-step oracle with zero single-steps of the target — both green and stable, first-ever on real PT silicon. The non-PT lanes are unchanged (docker-hwtrace 625/625 with PT self-skipping; docker-dataflow-pt synthetic 19/19 -Werror). Recorded in the internal docs/internal/intel-hardware-validation.md note.

  • asmspy CLI victims build on AArch64 — completes the cli (ubuntu-24.04-arm) build (follows the include-comment fix below). That fix revealed asmspy-aarch64 T5’s arm64 cli leg had never actually built: two victims carried unguarded x86 asm. Ported both to AArch64 — cli/int3_victim.c int3 → brk #0 with an SA_SIGINFO handler that advances uc_mcontext.pc past the 4-byte brk (AArch64 brk is a fault, not a trap: the handler must step PC or the return re-executes it forever), and cli/exec_stage2.c’s freestanding x86 syscall stub + _start → a svc #0 stub, AArch64 syscall numbers, and an AArch64 _start. Verified under qemu-user: both compile + link with WERROR on aarch64, the whole cli builds to an arm64 build/asmspy, exec_stage2 prints its freestanding banner (its svc/_start run), and int3_victim survives its own breakpoints with no SWALLOWED (the PC-advance works). The ptrace tracer interaction (asmspy re-injecting the brk) is validated by the native arm64 CI leg.

  • tools/asmfeatures links on Apple-Silicon macOS — benchmarks (macos-latest) bench-report (the second of the two macOS issues; follows the @rpath fix below). src/mach_backend.c’s whole body — including the MIG catch_mach_exception_raise* callbacks — is #if x86_64 && __APPLE__, but the generated MIG server mach_excServer.o is compiled on every macOS arch and references those callbacks unconditionally, so any executable linking the Mach objects (asmfeatures, via make bench-report) failed to link on arm64: “Undefined symbols for architecture arm64: _catch_mach_exception_raise*”. Added __APPLE__-guarded stub definitions (return KERN_FAILURE; the stepper never arms on arm64, so they are never invoked) in the non-x86_64-Darwin branch — excluded on Linux and on x86_64-macOS’s real branch. Reasoned fix (no local macOS host); the macos-latest CI leg confirms the link. The Mach OOP stepper itself stays Intel-macOS only — single-stepping on arm64-macOS is a separate port.

  • asmspy cli-smoke: deterministic unknown-arity arg-decode assertion (de-flake). The --log syscall arg-decode smoke (cli/cli_smoke.sh) asserts that at least one syscall in a 400-event window renders the faithful unknown-arity form (…) rather than a fabricated arity-of-three. The only undescribed syscall the victim’s stream produced was the incidental restart_syscall the kernel emits when a signal interrupts a blocking call — nondeterministic, and it flaked to zero matches under some kernels (measured 1-of-N on Docker-Desktop’s LinuxKit kernel), so make docker-cli failed at that step. cli/argdecode_victim.c now makes one DELIBERATELY-undescribed syscall (sysinfo, absent from asmspy’s arg_shape table), so the … rendering appears every iteration; the assertion keys off that specific call — a strengthening, not a weakening. make docker-cli cli-smoke PASS restored end-to-end.

  • Five red main CI jobs unbroken — regressions the individual lanes that introduced them did not catch. Each surfaced on the shared ci.yml matrix after an unrelated lane landed; all five were failing on every recent push.

    • cli (both ubuntu-latest and ubuntu-24.04-arm): a -Werror build break in cli/asmspy_engine.c. Two #include lines carried comments whose body contained */ (/* … TIOC*/FIO* … */, /* … S_IF*/STATX_* … */), which closes the block comment early and leaves the tail as error: extra tokens at end of #include directive [-Werror]. Only the WERROR CI leg (make WERROR=1 cli-smoke) is gated on it, so the non-WERROR docker-cli path stayed green and the break went unnoticed. Reworded both comments (TIOC*, FIO* / S_IF*, STATX_*); make docker-cli now builds and the smoke passes end-to-end.

    • test (riscv64 container): an x86-only opcode in a supposedly-portable asm stub. asmtest_sve_rdvl in src/capture.s (added with the SVE trampolines) zeroed its return via xorl %eax, %eax in a catch-all #else that also covers RV64 — src/capture.s:2786: Error: unrecognized opcode. Split the arm to match the file’s own 4-way pattern (#elif __x86_64__ → xorl, #elif __riscv → li a0, 0, #else → #error); the riscv64 container builds and runs green again.

    • dataflow (analysis lib + bindings): a legitimate new hardware self-skip not on the gate’s by-name allowlist. Chaining F5’s PT replay suite into dataflow-test added a # SKIP pt live replay: no intel_pt PMU … line, which tripped the anti-vacuity gate (it allowed only the BTF block-step skip). Extended the allowlist to permit the Intel-PT skip by name — a host gate exactly like the BTF one — with a matching ::notice.

    • dataflow (F4 GC-move canon): a stale exact-count gate. The lane grew from 37 to 43 assertions when the object-identity alias phase (1..6) landed, but the gate still asserted -ne 37. Updated to 43 with the phase noted in the message.

    • benchmarks (macos-latest): Capstone @rpath dylib not found at runtime (one of two macOS issues; partial). The pinned Capstone build installs to /usr/local and its dylib install-name is @rpath/libcapstone.5.dylib; recent macOS dropped /usr/local/lib from dyld’s default fallback search, so emu-bench aborted (Library not loaded: @rpath/libcapstone.5.dylib). Set DYLD_FALLBACK_LIBRARY_PATH=/usr/local/lib on the benchmarks job (ignored on Linux); the CI run confirmed this resolves the bench-check abort. It then surfaced a separate, pre-existing failure the abort had masked: bench-report cannot link build/asmfeatures on the Apple-Silicon macos-latest runner — Undefined symbols for architecture arm64: _catch_mach_exception_raise*. Those MIG callbacks are defined in src/mach_backend.c, whose body is x86_64-only (the Mach out-of-process stepper was ported to Intel macOS only), while the MIG server mach_excServer.o references them unconditionally. Porting the Mach tier to arm64-macOS (or excluding it from the arm64 asmfeatures link) is a follow-on for a macOS-capable agent — tracked against macos-oop-mach-stepper / benchmarks-ci-followups.

  • macOS (Intel) native build correctness (fourth pass): the asmspy CLI lane + two ungated example lanes, surfaced by building the lanes outside the nightly test-macos-x86 contract on a macOS 14.7.5 / Intel host — make cli, make cli-smoke, make WERROR=1 codeimage-test, make WERROR=1 build/jit_trace — none of which a Linux or Docker-on-Mac (Linux) lane exercises. A tree-wide make WERROR=1 sweep of every other macOS-buildable native lane (test/check/emu-test/asm-test/usecases/usecases-emu/dataflow-test/dataflow-pt-test/ hwtrace-test and the C/C++/Ruby/Python binding lanes) was already clean; these three lanes were the gap.

    • mk/cli.mk gated cli / cli-smoke on architecture only (x86_64 aarch64 arm64), not OS. asmspy is a Linux-only out-of-process tracer (ptrace / process_vm_readv / personality / /proc / <linux/futex.h> / <sys/user.h> / the glibc extension pthread_timedjoin_np, plus <sys/prctl.h> in every victim), so on macOS-x86_64 the arch gate passed and the build fell through and hard-failed — cli/asmspy.c at undeclared process_vm_readv / pthread_timedjoin_np, and the whole cli/ tree at <elf.h> / <linux/futex.h> / <sys/prctl.h>. Per-file include guards can’t fix this (cli/asmspy_engine.c alone carries ~473 Linux-only ptrace/user_regs_struct/SYS_* references); macOS’s single-step tracer is the separate Mach-exception tier (src/mach_backend.c, make mach-stepper-test). Added a Linux OS gate checked before the arch gate to both targets, mirroring the existing arch-gate self-skip idiom, so make cli / cli-smoke now print # SKIP … this is an OS gate on non-Linux instead of a compile cascade. make docker-cli (Linux in-container, UNAME_S=Linux) falls through the gate and builds asmspy + drives the cli-smoke sequence exactly as before — the gate is host-OS-keyed, not container-keyed (verified: asmspy built and the smoke ran its full sequence).

    • examples/test_codeimage.c: the BLOB_A / BLOB_B file-scope static const routines are referenced only inside the two #if defined(__linux__) bodies, so off Linux they drew -Werror,-Wunused-const-variable — and the ungated codeimage-test lane compiles this C TU under -Werror (the test itself is designed to compile everywhere and self-skip at runtime via its #else stub). Guarded the two definitions with #if defined(__linux__) to match their use.

    • examples/jit_trace.c: static int checks, failures; and the CHECK macro are used only from the #if defined(__linux__) && defined(__x86_64__) body (the #else is a self-skip stub main that reports neither), so off that target (macOS, and Linux-arm64) they drew -Werror,-Wunused-variable. Guarded both with the same condition as their callers, honouring the file’s own compile-and-skip design.

  • macOS (Intel) native build + binding self-skip correctness (third pass), surfaced by building the binding conformance corpus and the per-language dataflow-* / hwtrace-* lanes natively on a macOS 14.7.5 / Intel host — a surface no Linux or Docker-on-Mac (Linux) lane exercises (make python-test, make WERROR=1 dataflow-test, make dataflow-cpp-test / -python-test / -ruby-test, make hwtrace-python-test):

    • bindings/conformance/conformance.c included <sys/mman.h> / <unistd.h> only under #if defined(__linux__), but the CL_HAVE ptrace_descent fixture that calls mmap / mprotect / munmap is gated on ARCH (x86-64 / aarch64), not OS — so it compiles on macOS and hit implicit-function-declaration errors there, breaking make python-test. Broadened the include guard to __linux__ || __APPLE__, matching src/hwtrace.c’s asmtest_hwtrace_exec_alloc W^X path.

    • examples/test_dataflow_ptrace.c: nine static const fixtures (df_chain_v2 + the call-out / overflow set) used only inside the __linux__ && __x86_64__ block drew -Werror,-Wunused-const-variable under make WERROR=1 dataflow-test on macOS. Guarded their definitions to match their use (completing the earlier #else-stub fix to the same file).

    • bindings/dataflow_victim.c — compiled by every dataflow-<lang> lane, which run on macOS — unconditionally included <sys/prctl.h> and called prctl(PR_SET_PTRACER, …) (a Linux-only Yama trace opt-in), so the shared victim failed to build on macOS ('sys/prctl.h' file not found). Guarded both under __linux__; the live-attach lanes self-skip off Linux regardless.

    • bindings/python/tests/test_hwtrace.py’s test_window_region_free_whole_window still hard-asserted w.armed — the one binding missed when the second pass guarded the C++/Ruby/Lua/Zig/Rust window tests. Now guards the arming-dependent checks on armed and notes the transparent self-skip, matching node’s and cpp’s shape.

    • bindings/ruby/dataflow.rb used Ruby-3.0 endless-method syntax (def steps = …) that fails to parse on Ruby 2.6, violating the binding’s own required_ruby_version >= 2.6 (asmtest.gemspec) — and stock macOS ships Ruby 2.6. Rewrote as classic single-line defs (def steps; …; end), matching every sibling ruby file.

  • macOS (Intel) native build + whole-window self-skip correctness (second pass), surfaced by building the wider native tier set on a macOS 14.7.5 / Intel host (make hwtrace-test, make dataflow-test, make WERROR=1 hwtrace-test, make hwtrace-cpp-test hwtrace-ruby-test):

    • asmtest_hwtrace_pt_hop_open’s non-Linux #else returned ASMTEST_HW_ENOSYS while the tier’s classifier reports Intel PT as ASMTEST_HW_EUNAVAIL off libipt — so test_pt_hop_surface’s self-skip failed on macOS. Now returns EUNAVAIL (the same fix already applied to pt_begin_window / pt_attach_begin; the per-tid PT hop pair had reintroduced it).

    • examples/test_dataflow_ptrace.c called five test_window_* functions unconditionally in main, but their definitions live inside the __linux__ && __x86_64__ guard — no non-Linux stubs, unlike every sibling test. Added the missing #else stubs.

    • examples/test_dataflow_blockstep.c used Linux-only memfd_create unconditionally (the F2 sc_pread fixture), breaking the macOS compile even though the whole suite runtime-self-skips off Linux via asmtest_dataflow_blockstep_probe(). Added a compile-only non-Linux stub (the caller is unreachable there — the fixture’s own fd < 0 SKIP covers it).

    • examples/test_hwtrace.c’s map_exec helper drew an unused-function -Werror under make WERROR=1 hwtrace-test on macOS: every caller sits inside a Linux guard. Guarded the definition to match its callers, exactly like the adjacent frame_insns_eq.

    • The region-free §Z1 whole-window scope is Linux/x86-64-only (begin_window self-skips on macOS single-step, where the region-based tier still works). The C++, Ruby, Lua, Zig, and Rust binding tests hard-asserted w.armed, failing on macOS; they now guard the arming-dependent checks on armed and note the transparent self-skip — matching node’s already-correct shape and the C test_wholewindow_singlestep skip.

  • macOS (Intel) native build + PT self-skip correctness, surfaced by validating the out-of-process Mach single-step tier natively on a macOS 14.7.5 / Intel host (make mach-stepper-test, 25/25 live):

    • examples/test_hwtrace.c included <unistd.h> only under #if defined(__linux__), so the portable test_pt_attach_selfskip (it calls getpid() on every host) failed to compile on macOS; the include moved to the unconditional POSIX block.

    • asmtest_hwtrace_pt_begin_window / asmtest_hwtrace_pt_attach_begin returned ASMTEST_HW_ENOSYS from their non-Linux #else arms, but the tier’s single availability classifier reports Intel PT as ASMTEST_HW_EUNAVAIL on any host without libipt — so the PT begin() self-skip envelope diverged from the status/skip_reason contract on macOS. Both #else arms now return EUNAVAIL, matching the classifier and the Linux !available path.

    • tests/glob_parity.c compared asmtest_glob_match (pinned to glibc fnmatch) against the host fnmatch on undefined-behavior patterns (unterminated [, trailing \), which BSD/macOS fnmatch resolves differently — failing make check 11/44 on macOS. The divergent cases now assert the glibc-pinned contract directly, cross-checking the host fnmatch only under __GLIBC__; well-defined cases keep the live host differential everywhere.

  • parallel runner (-jN): a non-EINTR poll() failure no longer abandons the run and reports never-run tests as passed; the scheduler degrades to blocking reaps.

  • --filter on Win64: the portable glob matcher now matches POSIX fnmatch on unterminated [, backslash escapes inside classes, and trailing backslashes.

  • guard-page allocators return NULL instead of a guard-page pointer for sizes within a page of SIZE_MAX.

  • emulator fuzzing: corpus nudge no longer has signed-overflow UB at LONG_MIN/LONG_MAX range extremes.

  • Zig conformance: vec256/vec512 capture tests no longer under-fill the 8-slot vargs array.

  • Win64 --no-fork: a fault on a non-test thread no longer hijacks the test thread’s recovery stack; it takes the normal unhandled-exception path.

  • docs: the emulator guide no longer claims --emu installs only libunicorn.

  • Cross-alias register def-use edges resolved. asmtest_defuse_build (the shared, tier-neutral last-writer builder in src/dataflow.c) keyed its register axis on the raw Capstone id, so a write to one GP sub-register alias and a later read of another — mov eax, ... then a read of ax, mov r8d, ... then a read of r8 — produced no def-use edge at all, even though the value trace correctly captured both. This is the shared builder’s counterpart to dfp_alias_shape (src/dataflow_ptrace.c, added for the F6 gap barrier): a new reg_slice helper canonicalizes a Capstone GP register id to its 64-bit container plus a byte offset/length, and apply_write/emit_read now key a mappable register per CONTAINER BYTE — exactly as memory is already keyed per address byte — so a partial-overlap write/read resolves to the right last writer instead of missing the edge. AH/BH/CH/DH stay pinned to byte offset 1 of their container (not offset 0, which is AL/BL/CL/DL’s own byte): a write to ah reaches a later ah/ax/eax/rax read but never a later al read, which is the discriminator against a container-collapsing implementation that folds by container alone and ignores the byte offset (proven by temporarily mutating reg_slice that way and observing the new synthetic fixture fail, then restoring it). A 32-bit GP write (eax, r8d, …) additionally marks the FULL 8-byte container as written, not just its own 4 bytes — x86-64 defines a 32-bit write as implicitly zero-extending bits 32-63, unlike a 16/8-bit write, which leaves the untouched bytes exactly as they were — so a later full-width read resolves its upper-half producer to that same write instead of a stale one from before the zero-extension (also proven by mutation: reverting the widened write range makes a dedicated fixture fabricate exactly that phantom edge). Vector registers, segment selectors, EFLAGS, and RIP fall through to the pre-existing raw-id keying unchanged (none of them alias with anything else, so raw-id keying was already exact for them). Two new live fixtures in examples/test_dataflow_ptrace.c exercise the windowed gap barrier end-to-end through this change: a glue excursion that clobbers a sub-register alias of a register the survey recorded, and — closing F6 known-limit (4) — a glue excursion that clobbers a whole XMM register the survey recorded, the first fixture anywhere to exercise the barrier’s vector path at all.

  • The scoped ptrace dataflow producer’s call-out step-over can no longer fabricate a def-use edge across a stepped-over helper. dfp_step_loop (src/dataflow_ptrace.c) runs a call-out at native speed and records nothing over it — correct for cost, but a helper that clobbers a location the region already wrote (and a later in-region read relies on) previously left the read’s edge pointing at the stale in-region writer instead of the elided helper, silently wrong at a passing rc. Every scoped entry point (_run, attach, attach_pid*, attach_jit) now feeds a dfp_riskset (mirroring the windowed survey’s existing gap barrier), and the call-out branch snapshots it immediately before the native run and diffs it after: a synthetic GAP step is appended carrying exactly what changed (register alias-sliced, memory per byte), so a post-call read correctly resolves to the barrier instead of the stale writer. A risk-set cap hit is deferred in scoped mode — it only promotes to truncated at the first real gap (and is discarded on a gap-free exit), so a region that never calls out is never falsely flagged. Precision, not a blanket invalidation, is load-bearing here (F6 measured that a blanket shadow deletes true cross-gap edges): a helper that touches nothing at risk appends no record for that location, even though the gap step itself is still present (the call/ret round trip through any helper unavoidably moves rsp, which was already at risk from the call’s own write).

  • make install / make install-shared-hwtrace now ship asmtest_ibs.h — the hardware-tracing guide’s documented #include <asmtest_ibs.h> could not previously compile against an installed package (it was the only guide-referenced header missing from all three install lists). scripts/clean-room-test.sh gained a header-install compile check (a fresh make install + cc -fsyntax-only against every guide-referenced header) so this omission class cannot silently recur.

  • The ptrace block-step reconstructors now mark the capture truncated when a block contains a rep-prefixed string op. A rep movs/stos/… retires once per iteration under per-instruction stepping (RIP parks on it) but a static block-step reconstructor records it exactly once, so the block-step stream silently under-counted it. New asmtest_disas_is_rep_string lets bs_record_run and window_block_walk downgrade such a block to BS_AMBIGUOUS (faithful truncation), bounding the “byte-identical to per-instruction stepping” promise accordingly.

  • The ptrace block-step reconstructors no longer record never-executed instructions when the traced code contains an application int3. A JVM safepoint poll or .NET breakpoint inside a block-stepped region was misread as a BTF #DB block completion, so the region, attached, and windowed drivers fabricated the instructions after it with truncated=false. They now classify the trap via si_code (SI_KERNEL / TRAP_HWBKPT), record the executed run up to and including the trap byte, mark the capture truncated, and forward the signal — the region (owned) driver via PTRACE_CONT, the attached (foreign) driver by leaving the target in its SIGTRAP delivery-stop for the caller, and the windowed driver by handing off to the per-instruction window loop, which runs the frame to its window end at native speed (run_until_sig) and recovers *result there instead of discarding the signal.

  • No per-instruction ptrace loop in src/ptrace_backend.c swallows an application SIGTRAP any more. run_until (the call-out step-over primitive, now run_until_sig plus a 2-arg wrapper), the per-instruction region driver (asmtest_ptrace_trace_call), the foreign attached driver (asmtest_ptrace_trace_attached), the windowed per-instruction loop shared by asmtest_ptrace_trace_attached_windowed[_window_stop], the fork-owned window driver (asmtest_ptrace_trace_window_call), and call descent (asmtest_ptrace_trace_call_ex/_attached_ex) each either deliver an application int3/breakpoint via PTRACE_CONT (owned tracee) or end faithfully with the target left at its SIGTRAP delivery-stop (foreign) — never PTRACE_SINGLESTEP/PTRACE_SINGLEBLOCK with the signal attached (measured fatal: the re-armed trap fires inside a masked handler). bs_sigtrap_is_app (the si_code classifier introduced for the block-step drivers) is now a file-wide helper shared by every loop, on both x86-64 and AArch64.

  • The call-out step-over is now depth-aware (code review finding #19’s real fix). run_until (now run_until_sp, with run_until_sig/run_until kept as thin wrappers) previously resumed the trace at the FIRST arrival at a call-out’s return-address breakpoint, so a stepped-over helper that called BACK into the traced region (a callback, or a tiering/OSR stub re-invoking the method) hit its own return-address breakpoint from a deeper stack frame first and hijacked the resume into that nested invocation. classify_region_exit (shared by all four region drivers — per-instruction, block-step, attached per-instruction, attached block-step) now also passes the callee-entry stack pointer, and run_until_sp rejects a same-address hit at the wrong depth: it steps past the premature hit at native cost (a single-step over a software breakpoint, or a bare PTRACE_CONT for a hardware one — EFLAGS.RF keeps the CPU from re-trapping on it) and keeps waiting for the matching depth. A new differential fixture — a region that calls a helper which calls back into the region’s own entry exactly once before returning — proves the trace now resumes at the true, outer completion instead of the inner one.

  • asmtest_ibs.h no longer describes the shipped system-wide capture flag as a future phase — the survey_process residual-race note now names the ASMTEST_IBS_OPT_SYSTEM_WIDE flag directly.

  • The pure IBS-Op decoder now validates the record’s own caps word (BrnTrgt) before trusting the branch-target register. Two 68-byte record shapes are length-identical (BRNTRGT=0/OPDATA4=1 vs BRNTRGT=1/OPDATA4=0) and only the caps word disambiguates reg[7]; the decoder previously trusted length alone and could misread an IbsOpData4 value as a branch destination. The RipInvalid read is now gated on caps bit 7, and asmtest_ibs_available() requires CPUID IBSFFV (EAX[0]) so it cannot disagree with the caps the kernel samples with.

  • IBS ring-loss heuristic now bounds the callchain worst-case record (was 112 bytes, ~10× short — silent sample loss with lost==0 && throttled==0); ibs_fill_attr pins sample_max_stack so the bound is sound, and the internal window lane no longer opens with callchain (no in-tree consumer, and a callchain stream can overrun the single end-of-window drain).

  • ASMTEST_IBS_OPT_CALLCHAIN is documented as consumer-less: it enables kernel-side capture only; nothing in the tree decodes the stack (the drain parses past it to reach RAW), and the window lane ignores it.

  • ibs_probe and the ibs-test live skips now attempt a real perf open and report the real refusal reason instead of claiming AVAILABLE from the CPUID/sysfs substrate probe alone. On a locked-down AMD host (perf blocked by perf_event_paranoid/seccomp) the substrate is present but no sampling can open — ibs_probe prints substrate present but sampling is BLOCKED — <reason> (Op and Fetch lanes) and the five test_ibs EUNAVAIL skips print the real asmtest_ibs_unavail_reason() instead of a hardcoded guess. The AMD manual-validation checklist no longer inverts the call_auto regression signal: post-5d8e0d2 a truncated=0 where escalation must fire is a regression, not a known finding.

  • The guides and the public header no longer claim AMD LBR live capture works on Zen 3. The live-capture floor is Zen 4+ (LbrExtV2) — Zen 3 BRS exists in silicon but this tree cannot open it (the generic sample_period=1 open is rejected by the kernel’s amd_brs_hw_config; the raw-0xc4 arm is a hardware-gated follow-up). Swept every “Zen 3+” floor claim in the tracing guides, asmtest_hwtrace.h, and the AMD backend comments to “Zen 4+”, each with the one-line cannot-open explanation.

  • The dead AMD freeze-on-PMI probe (asmtest_amd_freeze_available) and its false PRESENT/ABSENT diagnostic are removed. The probe had zero live consumers after 5d8e0d2 replaced the freeze-conditional window-trust gate with an unconditional exit-presence check that runs on every part (asmtest_amd_ring_parse_decode). test_hwtrace printed a trust statement (“PRESENT (single-window Tier-A trusted)” / “ABSENT (…)”) that was false in both branches — Tier-A completeness is exit-anchored regardless of the freeze bit. The freeze test is retired; the snapshot-substrate/depth probe checks stay (renamed test_amd_snapshot_substrate_probe).

  • The AMD deterministic boundary snapshot no longer flags a provably complete 15-branch window as truncated. The depth-ceiling check in asmtest_amd_decode_reach counted the total decode-array length, but branchsnap.c prepends a synthetic boundary edge (a deterministic completion, not a captured hardware slot), so a full 15-hardware-slot window (15 + 1 synthetic = 16) tripped the 16-deep ceiling and escalated to a needless real re-execution of the routine under test (src/trace_auto.c). The new asmtest_amd_decode_reach_hw gates truncation on the hardware slot count; the asmtest_amd_decode / asmtest_amd_decode_reach wrappers pass hw_nbr == nbr so every other caller is byte-identical.

  • make cli / make cli-smoke on arm64 now self-skip like the other tiers instead of dumping raw compile errors. asmspy’s single-step engines are x86-64-hardcoded (rip/eflags-TF/orig_rax), so on aarch64 the build died mid-compile with SYS_open undeclared / no member named 'rip' — or worse, fell into the missing-dependency branch and advised installing libncurses-dev, which cannot fix an architecture. A uname -m gate (checked before CLI_MISSING, for exactly that reason) now prints a truthful # SKIP naming the open ARM64-abstraction plan row and exits 0. Measured in a real linux/arm64 container: skip + rc 0 both targets; x86-64 unchanged. a Yama/seccomp skip.** The victim called the region once and _exit(0)’d, so on a slow host the child finished and died before the parent’s PTRACE_SEIZE landed (3/3 GitHub runs today), and the resulting ESRCH surfaced as # SKIP … PTRACE_SEIZE unavailable here (yama/seccomp) — a double lie, since SEIZE worked for every other attach test in the same job — which the lane’s anti-vacuity gate rightly turned into a hard failure. The victim now LOOPS the region at the same 2 ms cadence as every other attach victim in the file, so the attach always finds a live process and a fresh entry. Proven discriminating in the docker lane: the once-and-exit victim plus a deliberate 200 ms pre-attach sleep reproduces the exact CI skip; the looping victim passes under the same handicap. The gate’s stale bookkeeping was recalibrated in the same change: the 8 suites total 389 on bare metal (the comment said 257), the VM runs ~293, and the floor moved 230 → 285 — preserving the original tightness (the smallest suite vanishing still trips it).

  • cli/asmspy.c’s new picker-sort code failed the clang-format gate. The Tab-cycle sort landed verified by make docker-cli (build + smoke) but not by make fmt-check; the format job caught 7 violations. Mechanical reflow, plus one comment hoisted above its if so the formatter keeps the condition on one line.

  • Data-flow --dataflow’s call-out step-over lied about WHY it truncated, and two test suites carried assertions that could not fail. Three small, independently diagnosed defects closed together:

    • dataflow_ptrace.c’s call-out step-over conflated a BOUND with a FAILURE (the sibling site to the --max fix: 9d55611/0129b1e). Hitting the whole-run step backstop mid-call-out and dfp_run_to actually failing shared one || and one DF_PTRACE_ETRACE, so a region that simply ran a lot of call-outs surfaced as “ptrace/attach failure (permission? ptrace_scope? … W^X JIT page)” — sending an operator to Yama/seccomp for a budget, not a bug. The backstop is now its own branch (DF_PTRACE_OK, truncated=true, same shape --max already got right); dfp_run_to failing (the callee exited, faulted, or its return byte could not be trapped) keeps DF_PTRACE_ETRACE. The 2^20-step backstop is now also overridable via ASMTEST_DF_STEP_BACKSTOP (mirroring ASMTEST_DF_ENTRY_WAIT_MS), which is what makes the bound reachable in a test at all — a real 2^20-step fixture was exactly the kind of “never exercised, needs 1M hits” gap this codebase already flags elsewhere. A new attach-based test (test_callout_step_backstop, needs the attach path’s exact pre_positioned entry so the trip point is deterministic rather than a coin flip on the fork prologue’s step parity) proves it: mutation (reverting the split) turns the check back into ETRACE.

    • test_branchsnap.c’s multi-exit test asserted covered(t, 0) as its entry evidence — vacuously. amd_replay appends block 0 unconditionally (amd_backend.c:267), so covered(t, 0) is always true by construction; the check was carried entirely by the ni > 0 conjunct beside it, same fact the Phase 9 tail-jmp tests in the same file had already found and correctly stopped relying on. snap_default_run now asserts the PATH-SPECIFIC block instead (covered(want_off) && !covered(other_off), the two exits’ own mov blocks) — real evidence that the default-on snapshot captured the exit that actually ran, not just “some” data regardless of which path executed.

    • test_dataflow_blockstep.c re-declared asmtest_blockstep_info_t with no layout guard. The tier ships no header by design (keeps the producer off the public ABI), so the suite hand-copies the struct — exactly the skew that cost F6’s sibling telemetry struct 3 green checks before a sizeof+offsetof guard caught it. asmtest_dataflow_blockstep_info_layout() (mirroring asmtest_dataflow_ptrace_win_info_layout) now lets the suite check its copy against the producer’s real layout before trusting any info.* field.

    All three were filed as open follow-ups (2026-07-17-dataflow-tier-open-followups.md) after the same day’s F1/F2/F6/F7 batch landed, deliberately deferred out of that diff to avoid scope creep. Verified: make docker-dataflow-attach (126+118 checks, 0 skips, 0 failures) and make dataflow-blockstep-test natively on the Zen 5 dev box (119/119). test_branchsnap.c’s live leg needs the BPF toolchain (clang/libbpf-dev), absent on this host and gated behind a sudo password this session could not supply — verified by compilation + the ENOSYS stub path only.

  • asmspy --dataflow on a symbol that is not running HUNG instead of erroring. The producer’s step backstop counts single-steps, and a region that never arrives burns zero steps — so the blocking wait never advanced (measured: rc=124, killed by timeout, where --trace on the identical target answered “never executed” in 4 s). The entry wait is now bounded by a CLOCK_MONOTONIC deadline (default 10 s, ASMTEST_DF_ENTRY_WAIT_MS overrides, 0 restores the old unbounded wait) and reports “ not seen entering in pid N (waited M ms)” — an outcome, not a failure: the symbol resolved and the tracer worked; the code just is not being called right now. The unwind re-establishes the all-running invariant (restore the entry byte, rewind a thread stopped at base+1, continue), which also fixes a latent hang on the never-exercised step-backstop disarm path.

  • asmspy --dataflow --max=<n> failed for every n below the region’s step count — and blamed ptrace for it. The truncation branch did everything right (partial trace appended, truncated:true) and then returned the generic ptrace-failure code, so a valid cap surfaced as “ptrace/attach failure (permission? ptrace_scope?…)”. The flag worked only when it did nothing (measured: --max=3 rc=1, --max=200 rc=0 on an 83-step region). It now returns OK with the truncated partial trace; the smoke asserts the EXACT step count per cap, so “cap ignored” cannot pass either.

  • Three asmspy --trace fidelity defects: a bound that wasn’t, a diagnosis thrown away, and a documented self-skip that never happened. (1) The entry wait’s idle window reset on EVERY waitpid event — a target that stops more often than the window (a 1 Hz timer, a chatty clone) reset the budget forever and --trace blocked indefinitely; a 30 s CLOCK_MONOTONIC wall bound now sits alongside the idle rule, checked unconditionally. (2) The entry race’s four distinct outcomes were bare integers collapsed by a bare break, so “the region never ran”, “the target exited”, and “the entry could not be armed” all rendered as “never executed”; the outcomes are now named and each maps to its own answer (“pid N exited before was seen executing” for an exit, an attach/ETRACE report for an unarmable entry). (3) That fix makes asmspy.h’s promised W^X/JIT self-skip real: an entry page refusing the breakpoint now reports “possibly a W^X JIT page refusing the entry breakpoint” instead of the confidently-wrong “never executed” — verified against an unmappable explicit range, the same failure shape a genuinely W^X page produces.

  • asmtest_trace_call_auto could report a window-overflowing AMD-LBR capture as complete, so escalation never fired (a real Zen 5 silicon finding). The Tier-A completeness check in asmtest_amd_ring_parse_decode — trust a single sampled window as complete only if it contains the region-exit branch — was gated behind !asmtest_amd_freeze_available() and thus skipped on freeze-capable parts (Zen 5). With sample_period=1 the capture picks the richest-in-region window, often an arbitrary mid-run fragment that never held the exit, so a 25-back-edge loop reconstructed a 4-edge fragment and reported truncated=0 — trace_call_auto returned it as complete instead of escalating to block-step (and test_call_auto passed vacuously). The exit-presence requirement now runs on every part; combined with the existing overflow flag it is the airtight “complete iff a non-overflowed exit-anchored window exists” invariant. Verified deterministic across 16 privileged AMD runs; test_call_auto case (b) hardened to fail hard on a fragment-reported-complete. Surfaced only because the new docker-hwtrace-privileged lane runs the exact AMD paths live.

  • The shared libasmtest_hwtrace shipped with an undefined asmtest_ibs_window_end. The Zen-2 F6 IBS survey fallback made hwtrace.c call the IBS window primitives, but the shared-lib link recipe never included ibs_backend.o — every binding’s dlopen failed on every host (the static test binaries link HWTRACE_OBJS, which carries it, masking the gap in hwtrace-test). pic/ibs_backend.o is now compiled and linked, and tracked by the knob-flip rebuild sentinel.

  • A corrupt/huge nr in an AMD branch-stack sample could drive the ring parse out of bounds (review F5/F7, followup Phase 2). The sampled-branch count from the perf ring is now clamped (nr <= 64, comfortably above the 32-deep hardware maximum) before the nr * sizeof(perf_branch_entry) size check can wrap, and a short-tail sample too small to hold the 8-byte nr itself is rejected instead of read. Exercised by the new synthetic-ring tests on every host.

  • asmtest_trace_call_auto could return ASMTEST_HW_OK with an empty trace (review F24). Each escalation rung’s reset discards the prior rung’s partial capture, but ran kept reading 1 from that earlier rung — so a rung that reset and then failed at runtime (seccomp/ptrace_scope, ENOMEM) reported a successful empty trace. ran is now cleared at every reset site and re-earned only when a rung actually commits; the legitimate truncated-but-OK partial (no downstream rung runs) is preserved.

  • asmspy’s single-step engines could kill a V8/Node target seconds after a clean detach. V8 worker threads park in blocking futex syscalls; a PTRACE_SINGLESTEP that completes across a syscall defers its #DB debug exception until the syscall returns, so a parked worker carried a queued trap through detach — when its futex later woke, the trap fired with no tracer attached and terminated the whole process (reproduced: ~1 detach in 2–6 fatal on an 11-thread V8 target). Two prior defenses missed it: the two-phase detach orders resumes, and the trap-flag clear was gated on a read-back TF bit that a kernel-forced TF masks out of GETREGS. detach_threads now clears TF unconditionally and drains the pending step — each stopped thread is single-stepped once more to consume its queued #DB while we are still the tracer (skipping threads poised on a syscall instruction so the drain cannot block); the PTRACE_SYSCALL engines skip both phases, since draining them would inject step state into a target that had none. 30 + 25 consecutive attach/trace/detach cycles on the 11-thread V8 target now survive.

  • asmspy swallowed a target’s own int3 breakpoints (and could have killed it re-injecting them). The single-step engines treated every SIGTRAP stop as their own step, so an application-executed int3 (a JIT/debugger breakpoint, e.g. V8’s IMMEDIATE_CRASH) was mis-decoded and silently dropped, breaking the app’s own breakpoint logic. Stops are now split by si_code (PTRACE_GETSIGINFO): only SI_KERNEL (an executed int3 on x86) and TRAP_HWBKPT are delivered back to the target — via PTRACE_CONT, never SINGLESTEP, because re-arming the trap flag fires a #DB inside the (SIGTRAP-masked) handler and the kernel force-kills the target. Everything else (TRAP_TRACE, TRAP_BRKPT from a step completing across a syscall, the SI_USER exec trap) is still absorbed. New int3_victim + smoke prove a self-breakpointing target survives tracing with its handler intact.

  • Fork-based tracers aborted the whole trace when an unrelated signal interrupted the post-fork handshake. trace_call, trace_call_blockstep, trace_window_call, and trace_call_descend waited for the child’s initial raise(SIGSTOP) with a bare waitpid that treated EINTR as a failed handshake (rc=ETRACE, zero frames). A host runtime’s repeating timer — or the descent stale-alarm test’s deliberate 200 µs SIGALRM storm, which failed ~70 % of runs on a fast box — could land in that window. All four handshake sites now retry across EINTR exactly as the step loop always has; a genuine child death still surfaces. The stale-alarm test passes 20/20.

  • Four defects found by a deep multi-agent audit of the whole tree (each survived double adversarial verification; the newest subsystem — asmspy — and the AMD/LBR, PT, single-step, orchestration and FFI-binding layers came back clean). (1) Reused-handle determinism leak in two emulator guests. The x86 and arm64 setups zero the GP + vector register file before every call so a routine that reads a register the caller did not set gets a deterministic 0; the RISC-V (emu_riscv_setup) and ARM32 (emu_arm_setup) setups omitted it, so a long-lived handle (how every binding holds it) leaked the previous call’s callee-saved / FP-lane state into the next call and returned a stale result with ok=true. Both now zero registers (x1..x31 + f0..f31 / r0..r12 + d0..d31 + condition flags) like the other two guests. (2) In-process stealth stepper could busy-hang forever. asmtest_hwtrace_stealth_trace’s while (!sc->ready) spin only checked for early helper death under if (use_exec); in the in-process fork fallback (common under the ptrace-restricted container/CI posture this project targets) a helper killed by seccomp/OOM/watchdog before publishing ready left the caller spinning at 100% CPU. The waitpid(WNOHANG) death check now runs unconditionally, matching the two windowed spins. (3) Block-step tracer leaked its owned tracee on overflow. asmtest_ptrace_trace_call_blockstep broke out on a blockstep_reconstruct failure (stream full / undecodable insn / no in-region terminator) with rc still OK, so the post-loop cleanup — which only reaps on rc != OK — left the forked child alive, ptrace-stopped and unreaped; repeated calls could exhaust PIDs. It now kill+waitpids on that path like the other overflow breaks. (4) DynamoRIO recording stack could pop a live region on deep nesting. The client’s on_begin pushed only while depth < MAX_DEPTH but on_end always decremented, so nesting past 16 distinctly-named regions desynced the per-thread stack and silently dropped coverage with truncated left 0. on_begin now tracks the true nesting depth unconditionally (matching the app side) and flags the trace truncated when a region falls outside the storable window — upholding the never-present-a- partial-trace-as-complete invariant.

  • Native-trace “fidelity” gaps — three places a partial capture could escape without its truncated flag. The framework’s core invariant is that an incomplete trace is never presented as complete; a review found three leaks and they are now closed. (1) The whole-window single-step loops (asmtest_ptrace_trace_attached_windowed and the fork-internal asmtest_ptrace_trace_window_call) treated any loop exit as clean — so a window whose tracee died or exit()ed before reaching the return address was reported complete; they now flag the stream truncated unless the one clean terminator (pc == win_ret, or the async *stop) was reached. (2) The AMD MSR-direct LBR path (asmtest_amd_msr_trace) returned a partial branch stack as complete when a mid-stack MSR read failed; a short read now sets truncated. Both err toward false-truncated over false-complete.

  • asmtest_hwtrace_arm_tid() reported a stale thread id after a single-step scope closed. The accessor’s documented contract is “the OS thread id that armed the active capture, or -1 when none is active”, but the single-step end() path (the default, most-portable backend, freshly wired into eight bindings) returned without clearing it — so it kept reading the last arming tid instead of -1. It now resets like the PT / AMD / whole-window paths already do; a regression assertion covers it.

  • asmspy --trace silently produced nothing for a function that runs only on a worker thread. The region engine attaches only the thread-group leader (unlike the whole-process syscall/stream engines, which SEIZE every thread), so a function executing on another thread was never single-stepped and the command exited cleanly with zero output. It now reports the region was never observed executing and points at --stream (which follows all threads).

  • asmspy ptrace-lifecycle hardening. A job-control group-stop (^Z / SIGSTOP / tty stop) is now handled with PTRACE_LISTEN instead of being resumed, so a traced target can actually be suspended while watched. On OOM while seizing threads, an already-seized thread is now detached rather than left stranded seize-stopped. The ELF section-header walk in the symbol resolver now strides by e_shentsize (not sizeof(Elf64_Shdr)), so a non-standard object with a larger entsize resolves correctly instead of reading misaligned headers.

  • Latent AMD-LBR test flake in nine bindings’ auto-select hwtrace test. Each binding mirrors the C reference test_auto_resolve_traces_live: pick auto(BEST), trace a tiny five-instruction routine, assert the result. On a privileged AMD Zen 3+ host auto picks AMD LBR, which faithfully truncates a too-fast-to-sample single-shot routine (so covered(0) is false) — the C reference and .NET already asserted covered(0) || truncated, but the fix was never ported, so rust/cpp/python/lua/ruby/zig/node/java/go still asserted only covered(0) and would fail on such a host. All nine now assert the faithful invariant.

  • The hwtrace options struct under-allocated the AMD-LBR fields in all seven FFI-mirroring bindings (8-byte OOB read in asmtest_hwtrace_init). When lbr_period/branch_filter were appended to asmtest_hwtrace_options_t (Zen 4/5 LBR work), only the .NET binding’s struct was updated; every other binding that hand-mirrors the struct still described the old 40-byte layout — Node koffi.struct, Java OPTIONS_LAYOUT, Python ctypes.Structure, Rust #[repr(C)], Ruby’s Fiddle packer, Go’s cgo typedef, and Lua’s ffi.cdef. HwTrace.init passed a 40-byte buffer, and asmtest_hwtrace_init’s g_opts = *opts copies the full 48 bytes — reading 8 bytes past the buffer on every init. Harmless for the SINGLESTEP backend (it ignores those fields), but for an AMD_LBR init the garbage read could seed a spurious sample period / reduced branch filter, silently altering capture. All seven now mirror the 48-byte C layout (verified: koffi.sizeof == 48, ctypes.sizeof == 48; the rust/ruby/go/lua docker-hwtrace-<lang> lanes green). C++ (#include "asmtest_hwtrace.h") and Zig (@cImport) use the real header and were never affected. Surfaced by the adversarial review of the whole-window attribution work.

  • Node binding: 64-bit trace-call results above 2^53 were silently rounded. Every fork/attach/stealth trace entry in the Node binding read the routine’s return (its RAX at the ret) as Number(readBigInt64LE(...)), which rounds any value past Number.MAX_SAFE_INTEGER through the double mantissa — so a routine returning a full 64-bit hash/id/pointer came back wrong, contradicting the documented “BigInt out of safe range” contract and the OOP capture forms’ exact-result guarantee. Added a _safeInt helper (Number when it fits the safe-integer range, else the exact BigInt) and applied it to all twelve result reads (callScoped, stealthTrace, windowCall, stealthWindow, traceCall/traceCallBlockstep/traceCallEx, and the traceAttached* family). Surfaced by the adversarial review of the whole-window work; regression-tested with a leaf returning 0x0102030405060708.

  • Stealth stepper seized the wrong thread on a managed runtime (getpid → SYS_gettid). asmtest_hwtrace_stealth_trace reverse-attached the helper to getpid() (the process leader), but on HotSpot the thread invoking the region is a JVM-created thread whose tid ≠ pid — so the helper single-stepped the wrong (idle primordial) thread and the run_to breakpoint fired on the untraced calling thread, killing the JVM with a fatal SIGTRAP (exit 133). Node and CoreCLR were unaffected only because their calling thread happens to be the leader (tid == pid). Fixed to seize (pid_t)syscall(SYS_gettid) — the calling thread — matching what the windowed variant asmtest_hwtrace_stealth_trace_windowed already did. Surfaced while adding the Java stealthTrace wrapper; after the fix Java captures a complete, exact stealth trace.

  • bindings-parity gate restored to green. The block-step / whole-window / snapshot commits added eight tier symbols wrapped only in the .NET binding, leaving the check-bindings-parity CI gate failing with 75 missing (binding, symbol) pairs. The BTF block-step pair (asmtest_ptrace_blockstep_available, asmtest_ptrace_trace_call_blockstep) — siblings of the universally-wrapped asmtest_ptrace_trace_call — is now genuinely wrapped in all ten bindings, each with a self-skipping parity test asserting the block-step stream is byte-identical to the single-step stream. The managed-tier / C-level symbols (the §Z1 whole-window trio, §3.1(c) attribute_window, §D3 trace_attached_windowed, and the AMD boundary snapshot) carry reasoned allow-list exemptions following the file’s existing conventions (the .NET tier keeps its real window-trio wraps).

  • test_descent_stale_alarm_flag no longer flakes on loaded CI runners. The test spams the tracer with a 200 µs SIGALRM storm to prove a stale L3 watchdog flag + EINTRs cannot abort a healthy L2 descent — but it left the descent’s real-time deadline at the 2 s default, which a loaded 2-core runner can legitimately exceed under 5000 interrupts/sec (a correct truncation, misread as the regression). The descent under test now carries an explicit 60 s deadline, so only the stale-flag bug it guards can fail it; the EINTR pressure is unchanged. (asmtest_hwtrace_call_scoped also joins the parity allow-list under the same dotnet-only posture as the window trio, restoring the gate the lazy-arm commit tripped.)

  • The scoped/windowed data-flow producers resolve r8d–r15b sub-register aliases. gp_value (the register-file value reader in both src/dataflow_ptrace.c and src/dataflow_blockstep.c) and dfp_alias_shape (the F6 gap barrier’s alias-slice classifier) had cases folding eax/ax/al/ah etc. to their 64-bit container but none for r8d/r8w/r8b .. r15d/r15w/r15b — a step that wrote one of those aliases produced a def-use record with no captured value (value_valid stayed false), and the gap barrier could not decide whether glue at risk had changed such a location (truncated). Both now fold every GP sub-register alias Capstone can emit on x86-64 to its container, exactly like the existing high-byte/32/16/8-bit cases.

  • The block-step tier no longer compares or records architecturally undefined EFLAGS bits as if silicon defined them. A new explicit mnemonic(+count)-keyed table in src/dataflow_blockstep.c (dfb_undef_flags) masks the undefined bits an instruction leaves out of both the coherence canary (regs_coherent, accumulated per replayed instruction and reset per block) and every captured EFLAGS write record (finalize_step), on both the single-step oracle and the block-step+replay paths — preserving their byte-identical property by construction, since both flow through the same shared classification. Covers and/or/xor/test (AF), mul/imul (SF/ZF/AF/PF), div/idiv (all six), bsf/bsr (CF/OF/SF/AF/PF), count-dependent shl/shr/sal/sar and rol/ror/rcl/rcr, and bt/bts/btr/btc (OF/SF/AF/PF); an instruction outside the table that touches flags at all is treated as fully flag-defining, matching every other x86 arithmetic instruction. New test hooks no_undef_mask (disables both mask sites — the negative control) and inject_flag_bit (forces a chosen bit to disagree right before the canary) land with the tier’s first opts-struct layout guard (asmtest_dataflow_blockstep_opts_layout). A dedicated xor eax,eax fixture — the AF-undefined case the tier’s primary oracle fixture deliberately avoided — proves AF reads 0 in every post-xor EFLAGS record on both paths while the trace stays byte-identical, and the canary discrimination checks prove the mask, not luck, is what tolerates it (an injected AF divergence is tolerated; the same injection with no_undef_mask set is caught).

1.1.0 — 2026-07-06

Fixed

  • Review-driven defect sweep (2026-07-02). Resolved the full backlog from the code-level review (54 findings) and the still-open 2026-07-01 / 2026-07-02 repo-review items, with a per-batch implementation note under docs/summaries/. Highlights: AArch64 callee-saved d8–d15 ABI checking + a corrected vm.s/structparam.s; SKIP() in SETUP/TEARDOWN reported as skip; JUnit XML made well-formed and no longer preceded by test stdout; hardware-trace truncation contract honored across the single-step / AMD-LBR / Intel-PT / code-image backends (block partition matches Unicorn/PT/DR); emulator SysV/AArch64 stack-and-register argument marshaling and a deterministic register reset per call on a reused handle; ptrace signal-forwarding, jitdump-truncation and tracee-reaping fixes; memory-safety and 64-bit-precision fixes across the Rust/Python/Node/Lua/C++/Java bindings; and Win64 runner teardown/DF/watchdog fixes. Build/CI: knob-aware object identity (SAN/COV/ASM_SYNTAX), header-prerequisite and PIC-object completeness, a check-version + third-party-version CI gate, publish tokens scoped to their step, the GPL corresponding-source release step, and a self-sufficient BTF-less eBPF fallback header. Added .mailmap.

Added

  • Scoped in-process tracing for .NET — the managed tier (§Z0–§Z5, §D0). The zero-config scope construct over the single-step hardware-trace tier: using (new AsmTrace()) { … } captures whatever the thread executes (no region, no HwTrace.Init — the ctor auto-inits the portable backend and self-skips with a faithful SkipReason where it cannot run). byMethod: true labels the captured window by managed method via an in-process MethodLoadVerbose listener (JitMethodMap), and withRundown: true also names warm + ReadyToRun BCL methods through a dependency-free DOTNET_IPC_V1 jitdump rundown over the runtime’s own diagnostics socket (no NuGet package, no launch knob). Results are data-first (Addresses, Methods, Disassembly, AsmMethod.Assembly/.Tier), with renderPath: true as the rendered opt-in. Labelling decodes against the code-image version live in the window (the map feeds asmtest_codeimage_track per method load), so bodies that re-tier/move after the scope still render the bytes that ran.

  • Named-method form — AsmTrace.Method(delegate) (§D0.3). Trace one managed method’s own JIT’d body: resolution via PrepareMethod + the listener (jitdump rundown fallback for warm/R2R bodies), a region + step-over capture with exact offsets, and Invoke(args…) as the library-owned non-inlinable call site. outOfProcess: true (§D3) routes Invoke through the concealed ptrace-stealth stepper — a bundled helper reverse-attaches and steps the body out of band, so the calling thread is never armed with EFLAGS.TF.

  • Faithful-degradation surface. HwTrace.DegradationNote() composes the tier ladder (Intel PT → AMD LBR → single-step → CoreSight, each with its skip reason, plus the ptrace fallback); cross-thread closes and overflows flag Truncated (native OS-tid assert + a complementary managed-thread guard); Disas.IsCall/IsBranch/IsRet/TryCallTarget classify live instructions structurally. The packable AsmTest NuGet now ships AsmTrace and the whole hwtrace wrapper alongside the bundled native payload.

  • Eleven runnable .NET examples under examples/dotnet/ (whole-window, region, methods, rundown, assemblies, annotated, tiers, hotspots, coverage, callgraph, ptrace_native — plus the out-of-process ptrace_dotnet attach demo), each split Program/Report, wired into make hwtrace-dotnet-example and make dev-dotnet. Validated on .NET 8 and .NET 9 (no diagnostics-IPC or MethodLoadVerbose drift).

  • Single-step native-trace tier: macOS-Intel front-end. The exact, unprivileged EFLAGS.TF (#DB → SIGTRAP) single-step backend now runs in-process on x86-64 macOS, not just Linux — the first Phase-5 front-end a Linux CI host (or Docker-on-Mac, whose containers are Linux) cannot exercise. XNU delivers the single-step trap as a BSD SIGTRAP, so re-asserting TF in the saved thread state re-arms stepping across sigreturn exactly as on Linux; the only platform deltas are the feature-test macro (_DARWIN_C_SOURCE) and the mcontext field access (uc_mcontext->__ss.__rip/__rflags vs. gregs[REG_RIP]/[REG_EFL]), both isolated behind shims in src/ss_backend.c. asmtest_hwtrace_available(SINGLESTEP) now returns 1 on x86-64 Darwin and the whole asmtest_hwtrace_* facade (region table, init/ register, begin/end/begin_scope/render_scope) drives it; the src/hwtrace.c gate HWTRACE_LIFECYCLE is a superset of __linux__, so the Linux path is unchanged (verified: make hwtrace-test 61 pass on macOS; make docker-hwtrace 178 pass on Linux). The binding-facing W^X executable-memory helper (asmtest_hwtrace_exec_alloc/_exec_free, src/hwtrace.c) now runs on x86-64 Darwin too — its PROT_NONE→RW→RX mmap/mprotect path is plain POSIX and identical to the Linux one — so the per-binding single-step lanes are reachable natively on macOS, not just the C suite (verified on this host: make hwtrace-{python,cpp,ruby}-test pass, with the Linux-only ptrace/codeimage backends self-skipping). Off-platform hosts self-skip with “single-step backend is x86-64 Linux/macOS only (Windows/AArch64 planned)”.

  • Scoped in-process tracing (the using/RAII/with model). A cooperative, developer-ergonomics face of the tracing machinery: bracket a region of a program’s own code with a scope construct — using (new AsmTrace()) in C#, RAII in C++/Rust, with in Python, defer in Go/Zig, a block/try-with-resources elsewhere — and get back the assembly that executed inside it, rendered on scope close. Implemented across all ten language bindings over a small shared C/decode core (error-returning asmtest_hwtrace_try_begin, arming-thread assert, asmtest_hwtrace_render, idempotent-by-name region registration, per-thread single-step state, a recorder-backed image adapter, symbolize-and-bucket, and the asmtest_hwtrace_stitch async-hop merge core). Linux-only; self-skips to a recorded no-op where no faithful backend is available. See docs/scoped-tracing-implementation.md and docs/internal/archive/plans/scoped-inprocess-tracing-plan.md.

    • §Z0/§Z1 the aspirational empty-ctor form — using (new AsmTrace()). A region-free whole-window scope with no NativeCode and no [base,len): new C entry points asmtest_hwtrace_begin_window/_end_window/_render_window over a whole-window frame mode in asmtest_ss_begin_window (the single-step handler records ABSOLUTE RIPs into the bounded ring, overflow → truncated), rendered from live self memory. The .NET reference shim gains the parameterless new AsmTrace() ctor + SkipReason (transparent self-skip). This is the single-step WEAK tier — native-leaf only, on any x86-64 Linux (test_wholewindow_singlestep, make docker-hwtrace → 201/0; .NET make docker-hwtrace-dotnet → 33/0). The STRONG whole-window PT / AMD LBR tiers, arbitrary-managed-method capture, and the other nine binding shims remain forward-look. See docs/internal/plans/scoped-tracing-zeroconfig-plan.md.

    • §D3 concealed ptrace-stealth stepper — now a bundled standalone binary. The hardware-free scope path (Zen 2 / Docker-on-Mac) reverse-attaches a helper to the caller (PR_SET_PTRACER + PTRACE_SEIZE) and single-steps the region out of band. Its stepping body + discovery moved to src/stealth_helper.c so the same code runs either as an in-process forked child (the fallback) or as the standalone asmtest-stealth-helper binary — a real separate process the managed packages can ship — which the caller discovers via a dladdr-sibling lookup (mirroring the DynamoRIO payload) or the ASMTEST_STEALTH_HELPER override, handing the shared trace over a memfd. New $(BUILD)/asmtest-stealth-helper build target + install-stealth-helper; test_ptrace_scoped_stealth asserts both paths reconstruct byte-identical offsets on any ptrace-capable Linux. The helper is bundled into every managed package payload (NuGet runtimes/<rid>/native, npm/Maven/gem/rock, the Python wheel _libs/) beside libasmtest_hwtrace, $ORIGIN-rpath’d so it resolves the co-vendored Capstone in-package, and asserted present + rpath’d (and not leaked into a darwin slot) by a fail-closed package-libs-verify gate. Only the live-JIT cross-process address channel (needs a running managed runtime) remains forward-look.

  • Call descent for the out-of-process ptrace tracer. The single-step tracer (asmtest_ptrace.h) can now optionally FOLLOW the calls a traced region makes instead of only stepping over them, at four opt-in levels (asmtest_descent_t): OFF (today’s behaviour), RECORD_EDGES (record each call-site → callee edge, still step over), DESCEND_KNOWN (single-step into resolvable callees — an allow-set of method regions or an optional resolver callback — stepping over the rest), and DESCEND_ALL (into everything, default off, denylist + instruction-budget + real-time-watchdog gated). The flat asmtest_trace_t is unchanged — it is always frame 0, byte-identical across all levels; descent records into a separate opaque handle read through scalar accessors (edges + nested per-callee frames), so asmtest_trace_t stays ABI-frozen and every binding adds accessor calls, not a struct layout. New entry points asmtest_ptrace_trace_call_ex / _trace_attached_ex / _trace_attached_versioned_ex thread the handle through the existing loops; the non-_ex symbols are unchanged (descent == NULL).

    • The descender is a return-address shadow stack with an exact pop predicate (PC == ret_addr && SP == caller-pre-call-SP && the just-stepped insn is a return) plus an SP-sweep for non-local exits (longjmp/unwind/sigreturn), same-region recursion as a distinct frame (with a recursion + max_depth cap), per-instruction byte windows via process_vm_readv, benign-signal forwarding on the live path, and a backend-owned ITIMER_REAL/SIGALRM watchdog so a blocked syscall in a descended callee self-truncates rather than hanging. AArch64 gained a NT_ARM_HW_BREAK hardware-breakpoint step-over path (the W^X JIT-heap fallback x86-64 already had). L3 is documented as best-effort / expected-to-perturb on a live managed runtime (the cross-thread lock-inversion deadlock vector is not fully mitigable) — see analysis/jit-runtime-tracing.md.

    • Surfaced in all ten language bindings (a Descent wrapper + descending trace_call_ex, with idempotent free and the per-FFI address/upcall hazards handled), pinned by a new ptrace_descent conformance-corpus tier and the header-grep parity gate; the resolver callback ships to the six upcall-safe FFIs (Python/Go/Node/Java/.NET/Lua) and Rust/Ruby/ C++/Zig expose the allow-set only. New jit_trace *-descend / *-descend-all demo lanes (make docker-hwtrace-jit-dotnet-bcl-descend, …). See docs/native-tracing.md (“Call descent levels”) and docs/internal/archive/plans/call-descent-plan.md.

  • Clean-room install test — every bundled binding, on Linux and macOS, in CI. make clean-room-test (any host) / make macos-clean-test (darwin alias) packages each binding that ships a native payload, installs it fresh into a throwaway prefix, loads it with every ASMTEST_*/DYLD_*/LD_* override scrubbed and the cwd outside the checkout, then asserts the native library it actually resolved lives under that fresh install — never a leaked dev build/ tree, a Homebrew dylib, or /usr/local. So “install fresh, no ASMTEST_LIB” is proven, not trusted: the prior per-binding smokes only checked a tier was available, which a leaked build/ or Homebrew dylib also satisfies. Bindings whose toolchain is absent self-skip; a real leak fails the run.

    • All six dlopen bindings are covered — Python, Ruby, Node, Java, Lua, and .NET. Each core loader gained a resolved-path accessor: library_path (Ruby/Lua), libraryPath() (Node/Java), Emu.LibraryPath (.NET, via Process.Modules so it reports the real loaded path however P/Invoke resolved the name), and Python’s existing python -m asmtest --where. The link bindings (C++/Rust/Go/Zig) ship source and link libasmtest themselves — no bundled payload to leak-check — so they are intentionally out of scope.

    • Verified in Docker per language: make docker-clean-<lang> builds the binding’s isolated image and runs the clean-room test in it with CLEANROOM_ONLY=<lang> — so a self-skip fails the lane (a missing toolchain can’t pass vacuously); make docker-clean-room runs the set. A new clean-room CI job (matrix over Ruby/Node/Java/.NET/Lua) gates every push, complementing the conformance bindings job (which loads the dev build/ tree). Python’s clean-room stays in the existing release.yml python job (which asserts on the repaired wheel — self-containing the wheel needs auditwheel/build the lean test image omits).

    • New reusable pieces: scripts/clean-env.sh — a sourceable env scrubber that pins DYLD_FALLBACK_LIBRARY_PATH to /usr/lib rather than unsetting it (unsetting reverts to a dyld default that includes /usr/local/lib, where a Homebrew copy could still satisfy a bare-leaf load); scripts/assert-clean-path.sh — the leak guard (rejects the checkout, /opt/homebrew, $HOMEBREW_PREFIX, /usr/local; allows a temp extraction, e.g. the jar’s); and scripts/clean-room-test.sh — the cross-platform per-binding orchestrator (the first reusable local one; the release.yml smokes can call it next, per the plan’s Track E). (macOS clean-test plan, Track A)

  • The native-trace tiers now ship inside the packages. Both optional tiers — DynamoRIO (libasmtest_drapp + libasmtest_drclient + the pinned libdynamorio) and hardware trace (libasmtest_hwtrace) — are staged into the Linux payload slots by make package-libs, so a fresh pip install / gem / npm / nupkg / jar / rock runs NativeTrace / HwTrace on a capable host with no manual make shared-* and no DYNAMORIO_HOME, exactly as the emulator/Keystone/Capstone tiers already do. drtrace is linux-x86_64 only (DynamoRIO auto-fetched via scripts/fetch-dynamorio.sh); hwtrace bundles on every Linux slot (single-step + ptrace always; the Intel PT / AMD / CoreSight decoders self-skip off the hardware they need). macOS/arm64 slots simply omit the Linux-only tier and the wrapper self-skips (available() → false) — no API or available() behavior change, bundling only removes the build step.

    • Every binding’s drtrace/hwtrace loader learned a bundled-package candidate (env override → bundled slot → dev build/ → system) and a library_path() self-report (python -m asmtest --where, and the equivalent accessor in the Go / Rust / Ruby / Node / Java / .NET / Lua / Zig wrappers) so a clean-room test can assert the tier resolved from the package, not a leaked checkout.

    • A package-bundled libdynamorio self-locates next to libasmtest_drapp (via dladdr), so the DynamoRIO tier works with zero configuration — dlopen does not consult a library’s own RUNPATH, so drapp finds its sibling explicitly.

    • Licensing unchanged in character: DynamoRIO (BSD-3-Clause core), and the “full” hwtrace’s libipt/OpenCSD/libbpf, are all permissive — collect-licenses.sh emits each only when the lib is actually staged, adding no copyleft beyond the existing Unicorn/Keystone GPL-2.0. The four source-distributed bindings (Rust/Zig/C++/Go) ship no binary payload, so their consumers build shared-drtrace/shared-hwtrace themselves (documented, not bundled). (bundle-native-trace-tiers plan)

  • Native runtime tracing (two optional tiers). A third execution tier that traces code running natively, in-process, complementing the Unicorn emulator trace. Both fill the same engine-neutral asmtest_trace_t shape (now extracted into include/asmtest_trace.h + src/trace.c, shared by all backends) and the Capstone annotation layer renders any backend’s offsets. (Native runtime tracing)

    • DynamoRIO in-process tier (asmtest_drtrace.h, libasmtest_drapp + CMake-built libasmtest_drclient.so): dr_app_* in-process attach with an enforced lifecycle state machine, begin/end region markers, basic-block and instruction coverage, and host-native W^X executable-memory allocation (asmtest_exec_alloc / asmtest_asm_exec_native). Uses DynamoRIO’s BSD core API only — no drmgr/drwrap, so no LGPL-2.1 obligation. Native-trace wrappers for every language binding — Python (asmtest.drtrace), C++, Rust, Go, Node, Java, .NET, Ruby, Lua, and Zig — each exposing the same NativeTrace/NativeCode surface and dlopen-loading libasmtest_drapp at run time, so the core binding never link-depends on DynamoRIO and each wrapper self-skips (available() → false) when the tier is absent. Targets drtrace-test, shared-drtrace, drtrace-client, drtrace-<lang>-test, drtrace-bindings-test, docker-drtrace, and docker-drtrace-bindings (container lanes with DynamoRIO installed). Gated on DYNAMORIO_HOME; self-skips when absent. All wrappers are verified against a real in-process DynamoRIO in Docker: C++/Ruby/Java/Lua/Zig/Rust/Go trace live; Node and .NET self-skip there (in-process DynamoRIO can’t take over a JIT/GC runtime’s threads — the managed-runtime limitation, where Intel PT is the recommended backend). Linux x86-64.

    • Hardware-trace tier (asmtest_hwtrace.h, libasmtest_hwtrace): four backends behind one API, one available() gating chain, and one asmtest_trace_t sink. Intel PT capture via perf_event_open + libipt decode with branch-boundary block normalization; AMD LBR (Zen 3 BRS / Zen 4 LbrExtV2, 16-deep, exact within window then truncated); ARM CoreSight (OpenCSD) scaffold; and single-step (EFLAGS.TF → #DB/SIGTRAP), the portable backend that records the same exact/complete offsets on any x86-64 Linux host (Intel, any-Zen AMD, VM, CI, plain container) with no PMU, perf_event, privilege, or decoder library. asmtest_hwtrace_available() encodes the full detect-and-skip chain; the PT/AMD/CoreSight backends self-skip off the bare-metal hardware they need (the common case). Targets hwtrace-test, shared-hwtrace, hwtrace-<lang>-test, hwtrace-bindings-test, docker-hwtrace, and docker-hwtrace-bindings (plain unprivileged container lanes); auto-detects libipt/OpenCSD via pkg-config.

    • Hardware-tier backend auto-selection. asmtest_hwtrace_resolve(policy, out, cap) returns the host’s available backends most-faithful first (Intel PT > AMD LBR > single-step > CoreSight); asmtest_hwtrace_auto(policy) returns the single best pick ready to init (or ASMTEST_HW_EUNAVAIL). policy is ASMTEST_HWTRACE_BEST or ASMTEST_HWTRACE_CEILING_FREE (drops the one fixed-window backend, AMD LBR — what a caller re-resolves under after a trace comes back truncated). On any x86-64 Linux host the cascade is non-empty (single-step is the floor), so auto() never fails there. Exposed through the C API and every language wrapper — Python, C++, Rust, Go, Node, Java, .NET, Ruby, Lua, Zig — each surfacing resolve/auto (C++ uses auto_select, since auto is a keyword) with BEST/CEILING_FREE policy constants, plus a per-binding self-test of the selection invariants and a live auto-picked trace. Scope is the hardware tier’s own backends; a cross-tier fall to DynamoRIO/the emulator stays a deliberate, fidelity-aware caller decision.

    • Cross-tier trace orchestration. asmtest_trace_resolve(policy, out, cap) / asmtest_trace_auto(policy, &choice) (asmtest_trace_auto.h, src/trace_auto.c) are the front-end over all three tiers, not just the hardware backends: they walk the full descending-fidelity cascade — Intel PT → AMD LBR → DynamoRIO → single-step → CoreSight → emulator (DynamoRIO ranks above single-step because its code cache runs at native speed while single-step pays a per-instruction kernel round-trip) — and return asmtest_trace_choice_t descriptors {tier, backend, fidelity}. It calls asmtest_hwtrace_available() directly and dlopen-probes libasmtest_drapp (via $ASMTEST_DRAPP_LIB) for the DynamoRIO tier, so it hard-links neither the DynamoRIO nor the emulator library — the three stay decoupled. The policy bitmask composes ASMTEST_TRACE_BEST, ASMTEST_TRACE_CEILING_FREE (drop AMD LBR; re-resolve under it after truncated), and ASMTEST_TRACE_NATIVE_ONLY — the flag that forbids the native→emulator fidelity crossing: under it the emulator floor is dropped, so a host with no native tier resolves to ASMTEST_HW_EUNAVAIL rather than silently downgrading real-CPU execution to an isolated guest. Shipped in libasmtest_hwtrace and exposed through every language wrapper (Python/Rust/ Go/Lua/Ruby resolve_tiers/auto_tier, camelCase resolveTiers/autoTier for C++/Node/Java/.NET/Zig, ResolveTiers/AutoTier for Go), each with a per-binding self-test of the cross-tier invariants. This is the cross-tier front-end the trace parity matrix flagged as the remaining gap.

    • Out-of-process single-step backend (W2). asmtest_ptrace_trace_call(code, len, args, nargs, &result, trace) (asmtest_ptrace.h, src/ptrace_backend.c) is the out-of-process sibling of the in-process EFLAGS.TF stepper: a tracer parent PTRACE_SINGLESTEPs a forked tracee that runs the registered routine, reads the program counter from the child’s register file at each stop, and reconstructs the same exact offsets in the parent — ordered in-region instruction offsets and the identical single-entry/ends-at-branch block partition the in-process stepper, Unicorn, DynamoRIO, and Intel PT produce — with no shared memory (the parent observes every step) and no library or privilege beyond ptrace of one’s own child. Because it touches none of the tracee’s signal disposition or code cache, it is the exact path for a JIT/GC managed runtime (JVM/.NET/Node) and the recommended managed-runtime backend on AMD (no Intel PT), and is the only single-step form possible on AArch64 (whose MDSCR_EL1.SS is kernel-only). It runs on Linux x86-64 and AArch64 off one body: the AArch64 arm reads the program counter + integer return register via PTRACE_GETREGSET/NT_PRSTATUS (AArch64 has no PTRACE_GETREGS) and decodes block lengths with ASMTEST_ARCH_ARM64 Capstone, while the fork/SIGSTOP/step/wait flow and the SysV/AAPCS64 register-arg call are shared. Built into libasmtest_hwtrace; make hwtrace-test exercises it live — including in a plain unprivileged container — asserting byte-for-byte parity with the in-process stepper plus a 62-instruction loop (no depth ceiling). This lands the Linux x86-64 front of the Zen 2 single-step plan Phase 5 (W2). asmtest_ptrace_trace_attached(pid, base, len, &result, trace) extends it to the foreign-process case — the building block for tracing a managed runtime: it traces a region in a separate, already-running process you have attached to externally (the caller owns the PTRACE_ATTACH/DETACH policy), single-stepping the target from its current stop and reading the region bytes from the target via process_vm_readv (so the tracer does not share the target’s memory). A live test attaches to a child that never called PTRACE_TRACEME and reconstructs the same [0,3,6,c,11] stream out of band. Two region resolvers turn the attach primitive into “point it at a running process”: asmtest_proc_region_by_addr(pid, addr, &base, &len) finds the executable mapping containing addr in /proc/<pid>/maps (one interior address → the whole region to trace), and asmtest_proc_perfmap_symbol(pid, name, &base, &len) parses /tmp/perf-<pid>.map — the text format V8/Node, .NET, and OpenJDK (+perf-map-agent) write so perf can symbolize generated code — to recover a JIT method’s (base, len) by name. A live test discovers a foreign process’s region from /proc/<pid>/maps using only an interior address and traces that region (no hardcoded base). asmtest_ptrace_run_to(pid, addr) closes the uncontrolled-timing gap that the rest of the flow left open: trace_attached needs the target stopped at the method entry, but a real managed runtime calls a JIT method on its own schedule, so you cannot attach at the right instant. run_to plants a software breakpoint at addr (PTRACE_POKETEXT — int3 on x86-64, brk on AArch64 — which patches an r-x text page the way a debugger does), PTRACE_CONTs until the program itself next calls in, then removes the breakpoint and rewinds the PC, leaving the target stopped exactly at addr for trace_attached (also hardened to record the entry instruction from either stop convention). A live test now drives the complete real-JIT flow with no cooperative go-flag — a child publishes a perf-map and calls its routine in a loop, the tracer resolves it by name, attaches, run_tos the entry, and traces that invocation to the exact [0,3,6,c,11] stream — completing the managed-runtime flow: resolve → attach → run_to → trace_attached → detach. Call-depth awareness removes the last “leaf only” restriction: trace_call and trace_attached previously treated the first step out of the region as the return, so a routine that called a runtime helper (GC barrier, allocation, PLT stub — the norm for real JIT methods) truncated at the first call. The stepper now decodes the region-exit instruction (asmtest_disas_is_call, Capstone CS_GRP_CALL) and runs a call-out to its return address at native speed (a breakpoint-cont over the callee, not a per-instruction step), resuming recording after it; only a genuine return ends the trace. The region’s own instructions are recorded, the helper skipped, the real return still found (test_ptrace_callout); without Capstone it falls back to the prior leaf-only behaviour. Real-runtime validation lanes point the whole pipeline at a live JIT (not a fixture) via one argv-driven harness (examples/jit_trace.c): make docker-hwtrace-jit traces Node.js (V8) — node --perf-basic-prof --no-turbo-inlining on a hot function, resolving the method from V8’s real perf-map, attaching to the live multi-threaded GC’d runtime, run_toing the entry, and single-stepping one invocation to recover the actual TurboFan code for (a+b)|0; make docker-hwtrace-jit-dotnet traces .NET (CoreCLR) — Program::Add → lea eax,[rdi+rsi]; ret (with DOTNET_TieredCompilation=0 for a stable address), tracing .NET’s W^X code heap as-shipped via the hardware- breakpoint fallback below; and make docker-hwtrace-jit-java traces OpenJDK (HotSpot) — Hot.asmtjit → lea eax,[rsi+rdx] inside the real C2 nmethod (entry barrier and stack-bang and all), JIT’d once via -XX:-TieredCompilation and kept a standalone callable body with -XX:CompileCommand=dontinline. HotSpot needs two wrinkles the others don’t, both handled in the harness: it does not stream a perf-map, so the lane drives jcmd <pid> Compiler.perfmap to materialize one for the live process; and the java launcher runs Java main() on a secondary OS thread (not the primordial one V8/CoreCLR use), so the harness picks the spinning loop thread by CPU delta and PTRACE_ATTACHes exactly it — a software-breakpoint trap on an un-traced thread is fatal. A watchdog makes a re-tiered/moved address self-skip rather than hang, so the lanes never flake (resolve + attach are asserted against the runtime’s real output; the trace is asserted-or-skipped). Hardware-breakpoint run_to: run_until (behind run_to and the call-out step-over) defaults to a software int3 but transparently falls back to an x86-64 hardware execution breakpoint (DR0/DR7 via PTRACE_POKEUSER) when PTRACE_POKETEXT is refused — i.e. on a W^X JIT code heap whose executable page is not writable (.NET’s default double-maps it; POKETEXT fails EIO). The hardware breakpoint writes no code (so it traces W^X as-shipped) and is per-thread (so it never traps a sibling runtime thread, unlike a process-wide int3); software stays the default and ASMTEST_PTRACE_HW_BP forces the hardware path (used to validate it deterministically on ordinary memory — test_ptrace_callout). AArch64 hardware breakpoints (NT_ARM_HW_BKPT) are a follow-on. A third lane, make docker-hwtrace-jit-jitdump, validates the binary jitdump byte source against real output: node --perf-prof writes a real jit-<pid>.dump, and the lane recovers a method’s recorded code bytes with asmtest_jitdump_find and checks them three ways — the address agrees with V8’s own perf-map, the bytes disassemble to real x86-64, and they match the live code at that address (jitdump’s temporal-capture guarantee) — the first validation of asmtest_jitdump_find against a real jitdump rather than a synthetic fixture. A second jitdump producer lane, make docker-hwtrace-jit-java-jitdump, validates the reader against a jitdump from a different runtime and encoder: OpenJDK HotSpot has no native jitdump, so the lane loads the perf project’s JVMTI agent (libperf-jvmti.so, from linux-tools) with -agentpath, which records every C2 method to a real jitdump. It names methods in JVM descriptor form (LHot;asmtjit(II)I, unlike V8’s symbol) and interleaves debug/unwinding records the reader must skip, so it exercises asmtest_jitdump_find on a genuinely independent encoder; the recovered bytes are checked against the live code and against HotSpot’s own jcmd Compiler.perfmap address (two independent HotSpot outputs). A third jitdump producer lane, make docker-hwtrace-jit-dotnet-jitdump, covers .NET CoreCLR — which, unlike HotSpot, writes a real /tmp/jit-<pid>.dump natively (no agent) under DOTNET_PerfMapEnabled=1, naming the method identically in the perf-map and the jitdump. So it shares the same trace_jitdump path as V8 (one routine parameterized by runtime): recover Program::Add’s recorded bytes (lea eax,[rdi+rsi]; ret) and validate them four ways — disassemble, match the live code, and agree with CoreCLR’s own perf-map address/size. The binary jitdump is now validated against all three managed runtimes (V8, HotSpot, CoreCLR). asmtest_jitdump_find(path, pid, name, &entry, bytes, cap, &len) reads the richer binary jitdump image (jit-<pid>.dump — CoreCLR, HotSpot, V8; what perf inject --jit consumes), resolving a method to its (code_addr, code_size) and its recorded native code bytes (which the text perf-map cannot give). Because each JIT_CODE_LOAD record is timestamped, a method re-emitted at a reused address (tiered/OSR recompilation) resolves to the latest body — the temporal same-address-different-bytes problem; endianness is auto-detected and non-LOAD records skipped. (The text perf-map remains the portable lowest common denominator for JITs that only emit symbols.) The full foreign-process toolkit — available/skip_reason, trace_call, trace_attached, run_to, region_by_addr, perfmap_symbol, jitdump_find — is exposed through every language wrapper (a Ptrace class / module surfacing the same methods, idiomatic per language), each with a per-binding self-test of the live-testable subset (out-of-process trace_call parity, /proc/maps and perf-map resolution, and a binary-jitdump round-trip), and asmtest_ptrace.h is now covered by the binding function-surface parity gate. The /proc + jitdump code-region readers (asmtest_proc_*, asmtest_jitdump_find) are pure file parsing, so they build and run on any Linux arch and are validated live on AArch64 (in a linux/arm64 container, where the perfmap and jitdump tests pass). The single-step trace capture on AArch64 awaits a real host: asmtest_ptrace_available() is a cached, hang-proof self-probe (bounded WNOHANG polling) that returns 0 under qemu-user — which does not emulate the ptrace tracer/tracee relationship at all — so the stepper self-skips on emulation just as the PT/CoreSight tiers self-skip off their hardware; the AArch64 fixtures are decode- and execute-validated under qemu, with only the live single-step stream pending Apple-Silicon / Linux-ARM64 / Windows-on-ARM hardware.

    • Time-aware code-image recorder (asmtest_codeimage). A userspace PERF_RECORD_TEXT_POKE: it records a timestamped timeline of a process’s code regions so asmtest_codeimage_bytes_at(img, addr, when, …) returns the bytes that were live at trace-position when — the correct answer for a JIT whose code is patched, freed, or has its address reused mid-trace, where a single late process_vm_readv snapshot returns the wrong bytes (the temporal problem the JIT-runtime-tracing analysis calls “the innovative, buildable core”, approach #2). Change detection is soft-dirty (/proc/<pid>/clear_refs to arm + the soft-dirty PTE bit to detect, read via the PAGEMAP_SCAN ioctl where available, else by parsing /proc/<pid>/pagemap), which works cross-process — the foreign-JIT case — needing only permission to read the target. The W2 stepper consumes it: asmtest_ptrace_trace_attached_versioned(pid, base, len, img, when, &result, trace) decodes a foreign region’s blocks against the time-correct bytes instead of a live snapshot (the existing trace_attached is the img == NULL case, unchanged). An optional eBPF emission detector (a CO-RE program on mprotect/mmap/memfd_create, filtered to the target PID namespace via bpf_get_ns_current_pid_tgid, events drained from a bpf_ringbuf) tells the recorder when code appears so it snapshots on the PROT_EXEC edge instead of polling; it is built only when clang+libbpf+bpftool are present (-DASMTEST_HAVE_LIBBPF) and self-skips otherwise, with the userspace soft-dirty path as the always-available fallback. Validated live: the same-address-different-bytes temporal proof and the versioned W2 trace run in make hwtrace-test / make codeimage-test on any x86-64 Linux host (no privilege); the eBPF detector is validated in make docker-hwtrace-codeimage (a --cap-add=BPF,PERFMON container — not privileged), observing a real mprotect(PROT_EXEC) emission edge. Exposed across all ten language bindings (a CodeImage wrapper) and covered by the binding function-surface parity gate. (asmtest_codeimage.h, src/codeimage.c, bpf/codeimage.bpf.c)

    • ARM CoreSight reconstruction core (host-validated). src/cs_backend.c is now split like the AMD backend: its decoder-independent reconstruction core asmtest_cs_reconstruct(arch, ranges, n, base, len, trace) turns the ordered instruction ranges an ETM/ETE decoder emits (OCSD_GEN_TRC_ELEM_INSTR_RANGE) into the same instruction-offset stream and single-entry/ends-at-branch block partition the Intel PT backend produces. It is host-validated without a CoreSight board (examples/test_hwtrace.c test_cs_reconstruction, the analogue of the AMD synthetic-branch-stack test), asserting byte-for-byte parity with the PT/AMD/single-step backends over the shared fixture. The remaining half — the live OpenCSD decode tree (ocsd_create_dcd_tree + ETMv4/ETE decoder + memory accessor feeding ranges to the core) — needs libopencsd and a real AArch64 CoreSight board to write and validate, so per the project’s no-untested- hardware-code rule it is not yet implemented; asmtest_cs_decoder_present() still returns 0 and the tier self-skips on every host, but the half the board glue will feed is now proven (CoreSight advances from bare scaffold to validated reconstruction).

    • AMD LBR Tier-B stitching (host-validated). AMD’s branch stack is 16 deep, so a single-snapshot (Tier-A) reconstruction sets truncated past 16 taken branches. Tier-B lifts that ceiling: asmtest_amd_stitch(samples, nrs, n, out, cap, &gap) splices the overlapping windows sample_period=1 emits (one per taken branch, consecutive windows overlapping by 15 edges) into one gapless taken-branch sequence — for each window it takes the smallest shift that still overlaps the accumulated tail and appends only the new edges, so a loop’s repeated identical edges stitch correctly; lost overlap (≥ a full window dropped to throttling) sets *gap. asmtest_amd_decode_stitched(...) replays the stitched sequence through the shared amd_replay loop (factored out of asmtest_amd_decode) without the 16-entry overflow flag. Host-validated without hardware (like the Tier-A reconstruction): test_amd_stitch synthesizes an 18-iteration loop’s windows, stitches them, and reconstructs the complete trace (55 instructions, two blocks, not truncated) where a single Tier-A 16-window truncates — plus gap detection. Now wired into the live capture and Zen 5-validated: hwtrace_end_amd collects every branch-stack sample in the perf data ring (time order) and, when the richest single window overflowed (best_nr >= 16), stitches them and decodes past the ceiling; the small-routine path (best_nr < 16) is unchanged. Completeness is gated on the precise loss signals — a stitch gap OR a PERF_RECORD_LOST/PERF_RECORD_THROTTLE record (the non-overwrite ring drops the newest samples on overflow and emits LOST, the signal the gaplessly-stitching survivors cannot otherwise reveal) → faithfully truncated. On a Zen 5 (Ryzen 9 9950X, make docker-hwtrace-amd) a 20000-trip loop reconstructs ~290 instructions (≈95 stitched branches, far past one 16-deep window’s ~49) and stays truncated, as the perf ring size and sample_period=1 throttling require; the live path is complete only for runs that fit the ring and survive throttling, beyond which DynamoRIO (no ceiling) remains the answer. (AMD LBR plan, Phase 5.)

  • Win64 wide-vector (AVX2 256-bit) capture. The Win64 capture trampoline topped out at 128-bit (xmm); a routine’s full ymm result under the Microsoft x64 ABI couldn’t be inspected past its low 128 bits. New asm_call_capture_vec256_win64 is the Win64 analog of the SysV asm_call_capture_vec256: it marshals four 256-bit args into ymm0..3, calls the routine, and captures the whole ymm0..15 file into a vec256_t[16], saving/restoring the callee-saved low 128 of xmm6..15 (the upper 128 is volatile per the ABI) and vzeroupper-ing on exit. A win64_vaddpd_ymm routine

    • a test_capture_win64 case assert the full 256-bit return (all four doubles, exercising the upper-128 lanes the 128-bit path can’t see), self-skipping off-AVX2 via a local CPUID/XCR0 probe. Verified on both Win64 lanes — the native ms_abi lane and the PE/Wine lane (Wine runs PE instructions on the host CPU, so it is real AVX2). Track D of the post-v1.0 expansion plan (the “Win64 wide path” follow-on); AArch64 SVE remains staged, hardware-gated on a runner that can execute it.

  • AVX-512 512-bit (zmm) capture — across the core, Win64, and all ten bindings. The wide-vector path now reaches 512 bits, validated on real AVX-512 silicon (a Zen 5 / Ryzen 9 9950X). A new vec512_t (64 bytes) and asm_call_capture_vec512 marshal eight zmm args into zmm0..7 and capture the full zmm0..31 file — AVX-512 doubles the register count as well as the width, so the capture is a vec512_t[32] (vs vec256_t[16] for AVX2) — using the EVEX-encoded vmovdqu64 (required to reach zmm16..31). It is gated on a real asmtest_cpu_has_avx512f() (CPUID + XCR0 0xe6: opmask + ZMM_Hi256 + Hi16_ZMM); the ASM_VCALL512* macros and every binding wrapper self-skip where AVX-512 is absent, so the same suite runs everywhere. Shipped in both the GAS and NASM trampolines (src/capture.{s,asm}) and the Win64 path (asm_call_capture_vec512_win64, Microsoft x64 ABI, low-128 xmm6..15 saved/restored), with ASSERT_VEC512_EQ + asmtest_assert_vec512_eq for the 64-byte lane compare. A vec_add8d corpus routine (vaddpd zmm, 8 packed doubles) and win64_vaddpd_zmm assert the full 512-bit return — the 8th double lane proves the bits neither the 128- nor 256-bit path can see. Exposed across all ten language bindings (Python, C++, Rust, Go, Node, Java, .NET, Ruby, Lua, Zig) as capture_vec512/cpu_has_avx512f analogs with per-binding parity tests. Closes the AVX-512 half of gap #4 in the post-v1.0 expansion plan (its “no AVX-512 silicon” caveat is now lifted on this host); AArch64 SVE remains the staged remainder.

  • Publishable packages (Track A): library-exposing artifacts, self-locating native libs, and a dry-run release workflow. The packaging scaffolding now produces artifacts a registry could ship. Each make <lang>-package exposes the reusable library module rather than the conformance test runner (the Ruby gem ships asmtest.rb, npm asmtest.js, the rock asmtest.lua, the JAR the Asmtest classes, and the NuGet package AsmTest.dll via a new SDK-style asmtest-lib.csproj / dotnet pack), and bundles one native slot per platform present in build/dist/native/ (the host slot locally; all four when a release has the CI native-all). The dlopen bindings self-locate their bundled native lib when ASMTEST_LIB is unset — Node/Ruby/Lua/Java fall back to native/<os>-<arch>/ next to the module (Java extracts the jar resource to a temp file; .NET resolves via the runtimes/<rid>/native/ RID layout; Python already used _libs/) — so an installed package works out of the box. A new release.yml builds the cross-platform native-all, then per binding packages → installs the artifact fresh → smoke-tests the bundled-native load (ASMTEST_LIB unset) → dry-run publishes (twine check, npm publish --dry-run, cargo publish --dry-run); the live push is gated behind per-ecosystem token secrets, so it runs end to end with no credentials. Python wheels are built per platform: a setup.py tags the wheel py3-none-<platform> (platlib, since it bundles a native lib), and the workflow repairs each into a self-contained manylinux / macOS wheel — auditwheel / delocate vendoring libunicorn — so pip install pulls no system libs. Every package + fresh-install + bundled-load smoke was verified in the per-language Docker images (the manylinux wheel checked to load with system libunicorn removed). Track A of the post-v1.0 expansion plan; see docs/packaging.md.

  • Binding parity, round 2: the new emulator/capture capabilities reach all ten bindings. The Track F mid-execution guards, Track E coverage-guided fuzzing / mutation testing, and Track D AVX2 256-bit capture were C-core-only; they now have a binding ABI and a wrapper in every language. The C side adds opaque-handle FFI in src/ffi.c (emu_watch_t / emu_reg_guard_t and the fuzz/mutation stat structs — alloc + by-field accessors; the arming/driver functions take plain pointers, so a binding calls them directly), and fuzz.o now ships in the emulator shared lib so emu_fuzz_cover1 / emu_mutation_test1 are reachable. Each binding gained the ergonomic surface in its own idiom — Python (ctypes), C++ (header structs), Ruby (Fiddle), Lua (LuaJIT ffi), Node (koffi), Go (cgo), Rust (#[repr(C)] + extern), Zig (@cImport), Java (FFM/Panama), .NET (P/Invoke) — e.g. Emulator.watch_writes / guard_reg / fuzz_cover / mutation_test and a capture_vec256 + cpu_has_avx2 gate, with the vector path self-skipping where AVX2 is absent. Done by hand (a binding-FFI codegen PoC was evaluated and reverted — it only covers the mechanical ~20%); each binding’s conformance runner gained native checks over the same byte-literal routines, verified on the host (Python/C++/Ruby) or the Docker matrix (the other seven).

  • Track C disassembly reaches all ten bindings (via one superset lib). Capstone disassembly was C-core-only — binding it naïvely would pull Capstone into every binding lib, or spawn a combinatorial lib matrix. Instead a single libasmtest_emu is the superset: make shared-emu links the emulator (-lunicorn) plus both optional native tiers (Keystone assembler and Capstone disassembler) into that one lib, so any binding that points ASMTEST_LIB at it gets disassembly and the assembler with no extra flag. Every binding gained disas / disas_available in its own idiom — Python (ctypes), C++ (header, gated ASMTEST_ENABLE_DISAS), Ruby (Fiddle), Lua (LuaJIT ffi), Node (koffi), Go (cgo dlsym), Rust (dlsym + extern), Zig (@cImport, free), Java (FFM/Panama), .NET (P/Invoke) — wrapping emu_disas (decode one instruction at an offset to "mnemonic operands") with a probe that self-skips against an older lib that lacks the tier. The per-binding *-asm-test checks now drive libasmtest_emu, so a single run exercises CallAsm and disas; the bindings-asm base image and install-deps --asm gained Capstone. Each binding’s conformance decodes known x86-64 bytes (xor rax, rax / ret / nop), verified on the host (Python/C++/Ruby) and the Docker matrix (the other seven). Track C of the post-v1.0 expansion plan.

  • Win64 runner parity — per-test isolation, -jN, and benchmarks. The Win64 tier ran every test in one process (--no-fork); the POSIX runner’s fork-based per-test isolation, -jN pool, and benchmark mode were gated off. All three now work on Win64. With no fork(), the runner re-execs itself per test — a hidden --asmtest-child=<index> runs exactly one test and writes its result to a temp file the parent reads back — driven through the existing asmtest_win32_run / _run_pool primitives (CreateProcess + WaitForSingleObject / WaitForMultipleObjects). Isolation is now the default (matching POSIX); --no-fork selects the in-process facility. A crash is contained in the child (caught there, or backstopped by the child’s death); a hang is killed by the parent’s deadline. --bench (rdtsc cycles per call) runs on Win64 too (a BENCH body is trusted, so it runs unguarded). tests/win64/suite_win64.c gained a BENCH, and make win64-runner-test exercises all four modes (isolation, -jN, --no-fork, --bench) under Wine; an optional windows-latest CI job signs the same suite off on a genuine Windows host with no Wine. Track B of the post-v1.0 expansion plan. See docs/win64.md.

  • Wide-vector capture — AVX2 256-bit (ymm). Vector capture was strictly 128-bit (vec128_t / ASM_VCALLn); a routine’s ymm result couldn’t be inspected past its low 128 bits. New vec256_t (the 256-bit analog union) and asm_call_capture_vec256 marshal ymm0..7 args and capture the whole ymm file into a vec256_t[16] (out[0] = return), with ASM_VCALL256n / ASSERT_VEC256_EQ and the existing ASSERT_DEQ/FEQ over the doubled lane counts. A runtime CPUID probe (asmtest_cpu_has_avx2 / asmtest_cpu_has_avx512f, checking both the feature bit and OS XCR0 enablement) makes the path self-skip (SKIP) on a host without the feature instead of executing an unsupported instruction. The trampoline is in both backends (src/capture.s GAS + src/capture.asm NASM, vzeroupper on exit), vec256_t is pinned in the manifest and by a _Static_assert, and an vec_add4d AVX2 example + test_simd case assert the full 256-bit result (including the upper-128 lane) on both backends. Track D of the post-v1.0 expansion plan; AVX-512 (zmm), AArch64 SVE, a Win64 wide path, and binding parity are staged follow-ons — and the emulator wide path self-skips because its bundled Unicorn exposes YMM/ZMM but does not execute AVX (UC_ERR_INSN_INVALID). See docs/floating-point-simd.md.

  • Coverage-guided fuzzing & mutation testing in the emulator. The emulator already recorded basic-block coverage but only ever fed a report; it now feeds input generation, and a mutation tester proves an input set actually catches a perturbed routine. Both run a one-int-arg routine inside the emulator (the instruction cap + fault hooks contain a pathological input or a broken mutant), are seedable for reproducibility, and live in a new dependency-free src/fuzz.c:

    • emu_fuzz_cover1 — coverage-guided generation: keeps inputs that grow the block-coverage union, drawing candidates fresh or by mutating a corpus member (the feedback), so it reaches blocks fixed vectors miss (for the classify example, 5 blocks vs a positive vector’s 3).

    • emu_mutation_test1 — mutation testing: flips bits of the routine, runs each mutant and the original on an input set, and counts mutants the set fails to distinguish (survivors = test-gap). A weak suite over classify leaves 100 of 192 mutants alive; a path-covering suite leaves only the ~16 equivalent mutants — a stronger input set demonstrably kills more. Reuses the framework’s seedable splitmix64 RNG (asmtest.h) and the emulator’s coverage trace; no new dependency. Track E of the post-v1.0 expansion plan. See docs/emulator.md.

  • Mid-execution guards in the emulator (watchpoints + register invariants). Assert properties while a routine runs, not just on its result — introspection no ABI-boundary tool can do. Armed on the emu handle and persisting across emu_call_* until cleared, recording the first violation as data (no host crash), x86-64 guest. Memory-write watchpoints (emu_watch_writes with EMU_WATCH_ONLY / EMU_WATCH_NEVER, ASSERT_NO_WRITE_VIOLATION / ASSERT_WRITE_VIOLATION) catch a logical scribble into mapped memory that does not fault — where a guard page sees nothing — and name the offending store (emu_watch_describe reuses the Track C disassembler: write to 0x400800 (8 bytes): mov qword ptr [rdi + 0x800], rax  (@0x3)). Register invariants (emu_guard_reg, ASSERT_REG_INVARIANT) assert a register holds a value at every basic-block entry — a callee-saved / stack-pointer guard that catches mid-routine corruption even when the value is restored by return (which ABI capture cannot see). Step-bounded assertions need no new API: run with max_insns=N and inspect out->regs. New hooks (UC_HOOK_MEM_WRITE, a second UC_HOOK_BLOCK) in src/emu.c; types, arming functions, and assertions in include/asmtest_emu.h. Track F of the post-v1.0 expansion plan. See docs/emulator.md.

  • Disassembly in emulator diagnostics (Capstone). The emulator records faults, traces, and coverage as raw byte offsets — @0x2f, never the instruction. With Capstone linked (the disassembler counterpart to the Keystone in-line assembler) those offsets now carry the instruction at them, across all four guests (x86-64, AArch64, RISC-V, ARM32). New helpers in include/asmtest_emu.h / src/disasm.c: emu_disas (one instruction at an offset, with PC-relative targets resolved to absolute), emu_fault_describe (a fault line that names the offending instruction — read fault accessing 0xdead0000: mov rax, qword ptr [rdi]  (@0x0)), and disassembling counterparts to the reporters — emu_trace_disasm, emu_trace_report_disasm, emu_coverage_uncovered_disasm (turning uncovered: 0x2f into uncovered: 0x2f  cmp rax, 0). Optional and auto-detected (pkg-config --exists capstone): every helper degrades to bare offsets when Capstone is absent (emu_disas_available() reports which), so the same call works either way and the core library / shared libs / binding images stay Capstone-free — only build/test_emu links it, exactly as the assembler tier keeps Keystone in its own object. make deps DEPS_ARGS=--emu now installs libcapstone-dev, so the CI emu job exercises the annotated diagnostics on every matrix OS. RISC-V disassembly needs Capstone ≥ 5 and self-skips on older builds. This is Track C of the post-v1.0 expansion plan. See docs/emulator.md.

  • CI builds the cross-platform native payloads for the bindings. make package-libs only ever staged the build host’s shared libs, so a release shipped a single-platform payload. A new payloads CI matrix runs the native staging on each {x86-64, AArch64} × {Linux, macOS} runner (the Intel-macOS corner nightly, as test-macos-x86 does) and uploads each build/dist/native/<os>-<arch>/ as an artifact; a payloads (collect + verify) job merges them into one tree, runs the new make package-libs-verify to assert every platform slot carries both the core and the libasmtest_emu lib, and re-uploads the combined set as a single native-all artifact a publish step would consume. This is the “multi-platform native payloads” step the packaging scaffolding stopped short of — no registry credentials or extra hardware needed. See docs/packaging.md.

  • Static Mach-O verification of the macOS payloads, on Linux. make package-libs-verify-macho (scripts/verify-macho.sh, folded into package-libs-verify) catches the most common macOS packaging regressions — a wrong/missing arch slice or a leaked absolute install-name/dependency — at build time on the Linux release collector, with no Mac needed, via llvm-otool / llvm-lipo. For every build/dist/native/darwin-*/ .dylib it asserts: the slot’s arch is present (llvm-lipo -archs), the install-name (LC_ID_DYLIB) is @rpath/@loader_path-relative and neither it nor any dependency bakes in /Users, /opt/homebrew, or /usr/local (a dev-build or Homebrew leak; system /usr/lib and /System are fine), and a min-OS load command is present (and <= MACOS_MIN_FLOOR when that var is set). It self-skips where the llvm tools are absent (a dev host), so package-libs-verify stays green everywhere; the package-libs-collect CI job installs llvm so it runs there for real. This is Track B of the macOS clean-test plan — the independent cross-check that scripts/package-native.sh’s macOS-side install-name rewrites actually produced correct Mach-O.

  • The full emulator surface reaches every binding. A review found four core emulator capabilities that no binding could reach (the FFI lacked an opaque-handle wrapper), plus an assembler tier the corpus did not anchor. All are now exposed across all ten bindings, driven by a widened binding ABI in src/ffi.c / src/emu.c:

    • Cross-arch emulator guests — run raw AArch64 / RISC-V / ARM32 machine-code bytes on any host (Guest/GuestEmulator + per-arch register reads through asmtest_emu_{arm64,riscv,arm}_reg), not just the x86-64 guest.

    • Emulator FP / vector args and >2 integer args — call_fp / call_vec / call_bytes over raw bytes, beyond the old two-integer call2.

    • Execution trace / basic-block coverage — an opaque Trace handle (asmtest_emu_trace_*) recorded by call_traced, with covered(off).

    • Win64 calling convention — call_win64, to test a Win64 routine on a System V host. Anchored in the shared conformance corpus: new emu_bytes / emu_trace cases (cross-arch int, x86 wide/FP/vector, Win64, two-block coverage) run on every host via checked-in pre-assembled byte literals, and the assembler tier is now emitted into corpus.json and executed by a new make conformance-asm build. The C++ binding also gains the previously missing sum_via_rbx / clear_carry cases. See the binding-parity plan.

  • In-line assembler tier (Keystone). Pass a routine as an assembly string and run it, instead of only as pre-assembled object code. asmtest_assemble() (in the new include/asmtest_assemble.h) turns text into machine code for the emulator’s guest set — x86-64 (Intel or AT&T syntax), AArch64, ARM32, and RISC-V where the linked Keystone supports it — with errors reported as data and output the caller frees via asmtest_asm_free(). Bridge wrappers emu_call_asm / emu_arm64_call_asm / emu_riscv_call_asm / emu_arm_call_asm assemble at the emulator’s load base (so PC-relative and branch targets resolve) and run through the matching emu_*_call in one call. Optional and pkg-config gated like the emulator tier, and folded into the superset libasmtest_emu (built by make shared-emu, which links libkeystone + libcapstone + libunicorn): make asm-test builds the standalone in-line-assembler suite, with make docker-asm and a CI asm job on both x86-64 and arm64. Keystone has no Linux distro package, so make deps DEPS_ARGS=--asm points at scripts/build-keystone.sh (a pinned source build the CI job and Docker image use). RISC-V in-line assembly self-skips until a Keystone release ships a RISC-V backend (none does yet). See the implementation plan.

  • In-line assembler reaches every binding, with a widened shim. All ten bindings now expose the assembler — the original five (.NET, Ruby, Lua, Node, Java) plus Python, Go, Rust, C++, and Zig — bound optionally so they self-skip against an older lib that lacks the assembler and pay no cost in the normal binding images. The dlopen bindings probe the symbol; Go and Rust resolve it through the libc dynamic loader (they statically link the plain lib); C++ and Zig link the assembler-carrying libasmtest_emu directly. The opaque-handle shim is widened from the original Intel-only, two-integer-arg, error-blind asmtest_emu_call_asm: the new asmtest_emu_call_asm6 takes Intel or AT&T syntax, up to six integer args, and an instruction cap (max_insns); asmtest_asm_last_error() surfaces the Keystone diagnostic so a failed assemble reports why instead of a bare false; and asmtest_asm_bytes() exposes multi-arch text→bytes (x86-64/AArch64/RISC-V/ARM32) so a binding can assemble guests its x86-only emulator handle can’t run. Each binding presents this as a callAsm/assemble pair with a uniform failure contract — an assemble error raises/returns the diagnostic (it is never a silent miss). The bindings-asm CI matrix grows from five to all ten (make <lang>-asm-test), each case now also covering the failure path and a multi-arch assemble; the C make asm-test suite adds asmtest_emu_call_asm6 / asmtest_asm_last_error / asmtest_asm_bytes coverage. The original asmtest_emu_call_asm stays as a thin compatibility wrapper.

  • Native Win64 tier (capture). A Microsoft x64 (“Win64”) capture trampoline (src/capture_win64.asm) mirrors all eight System V asm_call_capture* variants on real x86-64 silicon — integer/FP/vector args, the 32-byte shadow space, struct return and by-reference struct args, and ABI-preservation over the larger Win64 callee-saved set (rdi/rsi plus the callee-saved xmm6–15). The captured state has a first-class regs_t layout in include/asmtest.h (selected by -DASMTEST_ABI_WIN64, LLP64-correct, with _Static_assert offset pins) and a machine-readable manifest (make manifest-win64 → asmtest_abi_win64.json). It runs with no Windows host, two ways: the native lane via GCC/Clang __attribute__((ms_abi)) (make win64-msabi-test), and a real Windows PE built with nasm -f win64 + MinGW-w64 and run under Wine in an isolated image (Dockerfile.win64, make docker-win64). A new CI win64 job runs both on every push; the capture suite doubles as the native Win64 conformance check. This is the capture tier (suite runs --no-fork); the Win32 runner port is now underway (see below). See docs/win64.md and the implementation plan.

  • Native Win64 tier — runner port. The framework’s process-level guarantees now have Win32 equivalents for the Win64 tier, each in src/platform_win32.c (plus the platform-neutral src/glob_match.c), compiled only for the Win64 target and verified under Wine: per-test isolation + timeout via CreateProcess / WaitForSingleObject / TerminateProcess (asmtest_win32_run, classifying OK / CRASH-as-NTSTATUS / TIMEOUT), the -jN parallel pool via WaitForMultipleObjects (asmtest_win32_run_pool), the guard-page allocator via VirtualAlloc + VirtualProtect(PAGE_NOACCESS), in-process crash-to-failure via a vectored exception handler + __builtin_longjmp (asmtest_win32_guard, no SEH unwinding), and a portable --filter glob matcher (*, ?, [...] classes, \ escaping) replacing MinGW’s missing fnmatch. New make win64-{guard,isolate,pool,filter,seh}-test targets exercise each under Wine and join make win64-check / the CI win64 job. A thin platform seam (src/platform.h, ASMTEST_FNMATCH) wires the --filter and guard-page paths into src/asmtest.c with no POSIX regression. The runner itself is then built for Win64: a Win32 run_one (the per-test facility’s vectored handler + watchdog, mapping the recovery reason to fail/skip/crash/timeout), main() running --no-fork with the fork/pipe/poll isolation, parallel pool, signal handlers, and SysV-trampoline helpers gated to POSIX. make win64-runner-test builds src/asmtest.c with MinGW and runs a real TEST() suite (tests/win64/suite_win64.c) under Wine: the runner discovers and runs the suite, asserts real Win64 captures, and contains a crashing and a hanging test as reported failures while surviving. Still POSIX-only: a forked/-jN mode on Win64 and benchmarks. See docs/win64.md.

  • Packaging scaffolding for all ten bindings. Each binding now has a publish-ready registry manifest and a make <lang>-package target that assembles a distributable bundling the host’s prebuilt native libs: asmtest.gemspec (RubyGems), asmtest-1.0.0-1.rockspec (LuaRocks), pom.xml (Maven), asmtest.nuspec (NuGet), CMakeLists.txt (a find_package-able C++ INTERFACE target), build.zig.zon (Zig package), plus upgraded pyproject.toml (wheel package-data over a bundled asmtest/_libs/), Cargo.toml (crates.io metadata), and package.json (npm files). make package-libs stages the shared libs into build/dist/native/<plat>/; the dlopen bindings (Python/Ruby/Lua/Node/Java/.NET) bundle libasmtest_emu, while the link bindings (Rust/Zig/C++/Go) ship as source. A new docs/packaging.md is the release guide (native-lib split, version pinning, per-language commands, the multi-platform caveat). Scaffolding only — no registry credentials or cross-OS build matrices.

  • Go binding (Track G). A cgo wrapper in bindings/go/ over the opaque-handle FFI layer — no struct layout mirrored: it declares the binding-ABI entry points (asmtest_corpus_routine, asmtest_capture6/_fp2 + asmtest_regs_*, asmtest_check_abi, asmtest_emu_call2 + accessors) and links the prebuilt shared libs. Exposes Regs (capture / ABI / flags / FP), Emu + EmuResult (faults as data), and Tier-2 Assert* helpers over a small TB interface that *testing.T satisfies (so the helpers are themselves testable — the suite proves each one bites). make go-test runs go test; conformance_test.go replays the corpus, built + run in its own asmtest-go image (make docker-go) and the bindings CI matrix. This closes the last language track — all ten bindings (Python, Rust, C++, Zig, Node, Java, .NET, Ruby, Lua, Go) now ship Tier 1 + Tier 2.

  • Tier-2 idiomatic assertions (all ten bindings). Optional assertion layers over the Tier-1 result objects, with legible failure messages, idiomatic to each language: Python (asmtest.assertions, raising AssertionError), Rust (methods on Regs/EmuResult, panicking), C++ (asmtest::assert_* throwing assertion_error, for GoogleTest/Catch2), Zig (error-union helpers over std.testing), Node/Ruby/Lua/Java/.NET (throwing/raising assert_* helpers in the conformance runner), and Go (Assert* helpers failing a *testing.T). Each covers both the pass paths and the failure paths (the assertion fails when it should — pytest raises, Rust should_panic, Zig expectError, a recording TB stub in Go, try/catch elsewhere). assert_ret, assert_abi_preserved, assert_flag, assert_fp, assert_no_fault, assert_reg, ….

  • Node, Java, .NET, Ruby & Lua bindings (Tracks N/J/D/C). Five more language wrappers, all over a new opaque-handle FFI layer (src/ffi.c + emu helpers in emu.c): asmtest_regs_new + asmtest_capture6 / _fp2 + asmtest_regs_* accessors for the capture tier, asmtest_emu_call2 + asmtest_emu_* accessors for the emulator, and asmtest_corpus_routine(name) for routine addresses — so a dynamic binding needs no C struct layout. Bindings: Node (koffi), Java (FFM/Panama), .NET (P/Invoke), Ruby (stdlib Fiddle), Lua (LuaJIT ffi); each replays the conformance corpus (make node-test / java-test / dotnet-test / ruby-test / lua-test).

  • Isolated per-language Docker images. Each wrapper is built and tested in its own image (bindings/<lang>/Dockerfile on a shared Dockerfile.bindings-base), so toolchains never mix. make docker-<lang> builds + runs one language; make docker-bindings does all ten. The CI bindings job is now a per-language matrix running make docker-<lang>.

  • Zig binding (Track Z). The lowest-ceremony wrapper: bindings/zig/ consumes the C headers directly via @cImport — no separate binding layer — and replays the conformance corpus (make zig-test → zig build test, build.zig targets Zig 0.13.x). Added to the Docker bindings image and the bindings CI job.

  • Rust binding (Track R). A no-crates-io crate in bindings/rust/: #[repr(C)] mirrors of regs_t and the emulator structs (arch-selected via cfg) plus extern "C" declarations of the binding-ABI entry points, linked against the prebuilt shared libs by build.rs. Exposes capture / capture_fp / capture_vec → Regs, abi_preserved (native verdict shim), and an Emulator whose EmuResult carries faults as data. make rust-test runs cargo test; tests/conformance.rs replays the conformance corpus.

  • C++ binding (Track X). The C headers now carry extern "C" guards (and a portable ASMTEST_STATIC_ASSERT), so a C++ TU both compiles and links against the framework. bindings/cpp/asmtest.hpp adds an RAII Emu, initializer-list capture*, vector-lane helpers, and abi_preserved / flag_set predicates; make cpp-test runs an example suite that drives the framework from C++. New ASMTEST_NO_MAIN knob omits the runtime’s main() for embedding.

  • Docker per-language wrapper testing. Dockerfile.bindings bundles the Python, C++, and Rust toolchains plus libunicorn; make docker-bindings (and docker-python / docker-cpp / docker-rust) build and test every wrapper in one reproducible image — verifying a binding on any host, including a language not installed locally. A bindings CI job runs the same tests natively on x86-64 and arm64 Linux.

  • Python binding (Track P). A pure-ctypes package in bindings/python/ (no cffi/compile step) loads the shared library and the asmtest_abi.json manifest and exposes capture() / capture_fp() / capture_vec() (returning a Regs snapshot with ret, flags, fret, vector lanes, abi_preserved, and flag_set) plus an Emulator context manager whose EmuResult surfaces faults as data. Struct layout is read from the manifest, so the binding is correct for whatever architecture the library was built for. make python-test builds the shared libs, manifest, corpus, and a routine fixture lib, then runs pytest; the suite replays the same corpus.json the C reference emits and reproduces every case. A new bindings-python CI job (x86-64 + arm64 Linux) runs it — the reusable per-language CI template (bindings plan 0.5), which completes Track 0.

  • Shared libraries + ABI manifest (Track 0). The first slice of the multi-language bindings substrate. make shared builds libasmtest.{so,dylib} (framework runtime + capture trampoline, from -fPIC objects in a separate build/pic/ tree) and make shared-emu builds libasmtest_emu.{so,dylib} (adds emu.o, links -lunicorn), both with platform-correct versioned filenames, soname/install-name, and dev symlinks; make install-shared / install-shared-emu install them plus a new asmtest-emu.pc. make manifest emits asmtest_abi.json — a machine-readable struct layout (sizes, field offsets, host arch, sentinels, flag masks) compiled from the real headers via scripts/gen-manifest.c — so FFI bindings consume offsets instead of hand-transcribing them. _Static_asserts in asmtest.h / asmtest_emu.h pin regs_t and the emulator register structs to offsetof, preventing the headers, the trampoline’s stores, and the manifest from drifting apart. make install (static + headers) is unchanged. See docs/internal/archive/plans/multi-language-bindings-plan.md.

  • Binding ABI + conformance corpus (Track 0). Non-jumping verdict shims asmtest_check_abi / asmtest_check_flag return a verdict + reason instead of longjmp-ing into the runner, so an FFI binding can validate a capture with no C runner present (the existing ASSERT_ABI_PRESERVED / ASSERT_FLAG_* now delegate to them). ASMTEST_NO_MAIN builds the runtime without its main() for embedding. make conformance runs bindings/conformance/conformance.c — the C reference for a fixed corpus of canonical routines (int / FP / SIMD / flags / ABI capture + an x86-64 emulator case), checked against expected literals — and emits corpus.json, the portable expected-results table every language binding must reproduce. The binding-ABI contract symbols are designated in the API reference.

  • Parallel execution (Track E). -jN / --jobs=N runs up to N tests concurrently as forked children (a pool over the existing per-test fork model), while output stays in registration order regardless of finish order. Per-test timeout and crash containment are unchanged; --no-fork forces serial. New expect.sh self-tests pin the ordering, failure reporting, and crash containment under -j4.

  • libc-callback example (Track E). examples/callback.s / .asm with examples/test_callback.c: sum_map(arr, n, fn) and count_if(arr, n, pred) call a C function pointer per element, demonstrating an assembly routine calling back into C with correct callee-saved/stack-alignment discipline.

  • Valgrind story (Track E). make valgrind runs the example suites under memcheck (--no-fork) to catch bugs in the routine under test, complementing the always-on guard-page allocator; make docker-valgrind and the --valgrind flag of scripts/install-deps.sh round it out. Documented alongside the guard-page approach in the README.

  • Emulator FP/SIMD (Track C). The x86-64 emulator guest marshals double args (emu_call_fp) and 128-bit vector args (emu_call_vec) into xmm0..7 and captures the whole XMM file (emu_x86_regs_t.xmm[]). The AArch64 guest gains the same (emu_arm64_call_fp / emu_arm64_call_vec, NEON v[]); the RISC-V (emu_riscv_call_fp, f[]) and ARM32 (emu_arm_call_fp, q[]) guests gain scalar FP, with their FP units enabled at open (RISC-V mstatus.FS, ARM32 CPACR + FPEXC). Generic ASSERT_EMU_VEC128_EQ works across guests.

  • Emulator assertions (Track C). ASSERT_NO_FAULT, ASSERT_FAULT, ASSERT_FAULT_AT, ASSERT_EMU_REG_EQ, ASSERT_EMU_FP_EQ, ASSERT_EMU_VEC_EQ, and coverage ASSERT_BLOCK_COVERED / ASSERT_BLOCKS_AT_LEAST.

  • Coverage reporting (Track C). emu_trace_report, emu_coverage_uncovered (lists the blocks a run missed against a universe trace), emu_trace_lcov (offset-level lcov export), and the emu_trace_covered predicate.

  • Emulator vector parity & source-line coverage (Track C, leftovers). emu_arm_call_vec marshals 128-bit NEON vectors into ARM32 q0..q3 and captures the whole q0..q15 file, matching the x86-64/AArch64 vector path. Source-line coverage: a caller-supplied emu_line_map_t (ascending (offset, line) rows, produced out-of-band) drives emu_line_lookup, emu_trace_source_report, and emu_trace_lcov_source, which report block coverage against source lines (hit and missed) — no DWARF parsing, no new dependency. The RISC-V “V” extension has no counterpart: Unicorn’s RISC-V guest exposes no vector registers, so it stays scalar-FP (documented in src/emu.c and docs/emulator.md). Closes Track C’s open C items.

Changed

  • Consolidated the AMD native-tracing docs (design-doc curation). The AMD tracing story was split across four overlapping files; the three genuinely-AMD ones — the AMD LBR snapshot backend plan, the improvement analysis, and the improvement implementation plan — are merged into a single docs/internal/plans/amd-tracing-plan.md with three parts (shipped LBR backend / improvement analysis / improvement roadmap), collapsing two duplicated “governing constraint” + “implementation status” preambles into one. All ~9 references (the src/amd_backend.c header comment and the sibling plans / parity matrix) repoint to the merged file; the vendor-neutral single-step (“Zen 2”) plan stays separate. The docs/internal/analysis/trace-parity-matrix.md overlap the review also flagged was assessed and left intact: its matrices are trace-specialized (not restatements of docs/features.md) and its “Matrix N” numbering is cited from src/trace_auto.c / include/asmtest_trace_auto.h, so trimming would lose detail and dangle those code references. Addresses review finding #15.

Fixed

  • DRAPP_KEYSTONE is now part of drtrace_app.o’s build identity. The app-side object was compiled with $(DRAPP_KS_DEF) (-DASMTEST_HAVE_KEYSTONE, gated by the DRAPP_KEYSTONE knob) but that flag was not among the rule’s prerequisites, so flipping the knob between sub-makes in the same build/ tree reused the stale object. This is not hypothetical: the per-binding native-trace lanes invoke $(MAKE) shared-drtrace … DRAPP_KEYSTONE=0 precisely because a Keystone-enabled drapp .so has unresolved emu_* symbols and won’t dlopen — so a reused Keystone-on object silently breaks exactly those lanes (CI only escaped it by running each in a clean tree). Both drtrace_app.o rules (static + PIC) now depend on a $(BUILD)/.drapp-flags sentinel that records the full compile-flag string and is rewritten only when it changes (cmp guard, so no spurious rebuilds), folding Keystone — and by extension SAN/COV via CFLAGS — into the object’s identity. Verified: no rebuild when nothing changes, a rebuild the moment the Keystone define or any CFLAGS knob flips. (The related make -j race on the aggregate drtrace-bindings-test targets — concurrent sub-makes sharing one build/ — is a separate concern and not addressed here.)

  • Version sync now covers the C header, not just the binding manifests. include/asmtest.h pins the version in four macros (ASMTEST_VERSION_MAJOR/MINOR/ PATCH + the ASMTEST_VERSION string), and scripts/amalgamate.sh derives the single-header version from that header — but scripts/sync-version.sh / make check-version only touched the nine binding manifests. So a version bump plus make sync-version && make check-version passed green while ASMTEST_VERSION, ASMTEST_VERSION_NUM, and the amalgamated asmtest_single.h all stayed at the old version — an unchecked second source of truth. sync-version.sh now writes and checks the four header macros too (splitting VERSION into numeric MAJOR/MINOR/PATCH, and requiring an exact 3-part numeric semver). Verified by round-tripping a bump: the header, ASMTEST_VERSION_NUM, and the amalgamation all follow, and check-version fails on a stale header.

  • RNG asmtest_rng_range divided by zero on ranges wider than LONG_MAX. asmtest_rng_range(rng, LONG_MIN, LONG_MAX) — a natural draw in differential / property testing — computed the span as (uint64_t)(hi - lo) + 1, where hi - lo signed-overflows long (UB) and the wrap makes (uint64_t)(hi-lo)+1 evaluate to 0, so the following % span was an integer division by zero → SIGFPE (a hard crash, or a spurious “crashed by signal 8” under --fork). The span is now computed entirely in uint64_t (wraps to 0 only for the full 2^64 width, which is special- cased to “every value is in range”), and the offset is added in uint64_t too so a large offset can’t overflow the lo + … back through signed. New self-tests posit.rng_range_full_width_no_sigfpe / _wide_in_bounds / _narrow_in_bounds.

  • Signed-overflow UB in the floating-point ULP-distance helpers. fp_ulp_distance / fp_ulp_distance_f computed the gap between the sign-magnitude-mapped keys with a signed ia - ib, which overflows int64_t/int32_t for far-apart operands (e.g. ASSERT_DNEAR(-DBL_MAX, DBL_MAX, …), ~1.84e19 ULPs). It yielded the right magnitude on two’s-complement hardware but is UB — and halted this repo’s own make sanitize (UBSan, halt_on_error=1) lane, so it was self-inconsistent. The larger-minus- smaller gap is now taken in uint64_t/uint32_t (bit-identical result, no UB; the signed comparison that picks the larger key is unchanged, so -0.0/+0.0 still collapse to a zero-ULP distance). New self-test posit.near_full_range_no_overflow.

  • Out-of-process ptrace stepper leaked the tracee when the traced routine faulted. Tracing a routine that takes a real signal (SIGILL/SIGSEGV) — exactly what the out-of-process single-step tier exists to trace — hit a break in the step loop (src/ptrace_backend.c) that returned ASMTEST_PTRACE_OK without reaping the forked tracee. Because PTRACE_O_EXITKILL fires only when the tracer exits (not when the trace function returns), the child was left stopped in signal-delivery and unreaped, so a suite of faulting routines slowly exhausted PIDs. The non-SIGTRAP break now kill(SIGKILL) + waitpids the tracee like every other exit path, still recording the partial trace as truncated. New live regression test_ptrace_faulting_no_leak traces a ud2 fixture eight times and asserts each leaves no child to reap (waitpid(-1) → ECHILD); it fails on the old code and passes on the fix (make docker-hwtrace, 95 tests green).

  • AMD LBR live capture (first verified on real hardware, Zen 5). Running the branch-record capture on an actual AMD LbrExtV2 host (Ryzen 9 9950X, Zen 5 — the project’s real dev box, long mis-documented as Zen 2) surfaced a capture bug: hwtrace_end_amd kept the last PERF_SAMPLE_BRANCH_STACK sample, which for a small routine is all post-routine glue branches — it decoded to an empty in-region trace yet reported it complete (truncated=0). It now keeps the sample richest in in-region branches (the one taken at/just after the routine, whose 16-deep window still holds its branches) and sets truncated when none is found — the faithful dynamic-fallback signal. A branch-heavy loop now reconstructs exactly from the live LbrExtV2 stack; a tiny single-shot routine (too fast for an in-region PMU sample) faithfully truncates. New live regression test_amd_live + make docker-hwtrace-amd lane (hwtrace image run with --security-opt seccomp=unconfined --cap-add=PERFMON); the standard docker-hwtrace lane is unchanged (AMD self-skips without perf). Docs corrected across the trace-parity matrix, native-tracing, and the AMD-LBR / single-step plans (dev host is Zen 5 with amd_lbr_v2; AMD LBR live-verified, no longer “unverified on dev”).

  • Emulator handle reuse. Unicorn’s translation-block cache is now flushed when new code is loaded, so reusing an emu_t/guest handle for a different routine no longer re-runs the previous routine’s stale translation.

1.0.0 — 2026-06-24

First tagged release. Captures the complete framework plus the Track A self-test suite and Track B packaging.

Added

  • Core framework. Auto-discovered TEST(...) cases, a provided main(), per-suite SETUP/TEARDOWN, SKIP(reason), and colored TAP reporting with a nonzero exit on failure.

  • Assertions. ASSERT_TRUE/FALSE, signed ASSERT_EQ/NE/LT/LE/GT/GE, unsigned ASSERT_UEQ/UNE/ULT/ULE/UGT/UGE, ASSERT_STREQ, ASSERT_MEM_EQ (hexdump diff), ASSERT_REG_EQ, FP ASSERT_FP_EQ/NEAR + lane ASSERT_DEQ/DNEAR/FEQ/FNEAR, and ASSERT_VEC_EQ.

  • ABI capture. Register/flags capture via ASM_CALLn, ASSERT_ABI_PRESERVED, ASSERT_FLAG_SET/CLEAR; full call model (ASM_CALLN, ASM_SRET, ASM_FCALLn, ASM_VCALLn, struct-by-value) across the System V integer/FP/vector paths.

  • Differential / property testing. ASSERT_MATCHES_REF{1,2,3} with a seedable splitmix64 RNG (ASMTEST_SEED).

  • Robustness & CLI. Per-test fork() isolation with an alarm() timeout, crash/hang containment, and a runner CLI (--filter, --list, --shuffle/ --seed, --timeout, --no-fork, --format=tap|junit).

  • Benchmark mode. BENCH(...) cases timed in cycles/call via rdtsc / cntvct_el0, run under --bench.

  • Portability. x86-64 and AArch64, Linux and macOS; GAS (default) and NASM (ASM_SYNTAX=nasm, x86-64) backends. CI covers all four OS/arch combinations.

  • Emulator tier (optional). Unicorn-backed x86-64, AArch64, RISC-V (RV64), and ARM32 guests; Windows x64 ABI on the x86-64 engine; instruction trace and basic-block coverage.

  • Framework self-tests (Track A). tests/positive.c, tests/negative.c, and the tests/expect.sh black-box harness, run by make check and wired into CI.

  • Packaging (Track B). ASMTEST_VERSION macros; make lib builds libasmtest.a; make install/uninstall honoring PREFIX/DESTDIR; an asmtest.pc pkg-config file; and make amalgamate producing the single-header asmtest_single.h.