Changelog¶
Included verbatim from
CHANGELOG.md in the
repository root. The format follows Keep a Changelog and the project aims to
follow Semantic Versioning.
Changelog¶
All notable changes to asm-test are documented here. The format follows Keep a Changelog, and the project aims to follow Semantic Versioning.
Unreleased¶
Added¶
The Session strip desktop view — the whole session as per-thread lanes with run seams, a kernel rail, address bands, and run-length density, simplified to the top lanes by default (the rest counted, never vanished) — and its 3D companion, the session flow scene: the strip’s channels as smooth, depth-stacked ribbon rates over stream order, with seam walls at the strip’s run boundaries.
Desktop live sessions: live union weave (the 3D pane accumulates every capture the session makes, so a fresh Start adds to the scene) and stable plane layout (regions keep their first-seen slot), both ON by default; a GPU 3D rendering settings toggle (OFF takes the honest degraded 2D branch); and a two-step clear previous sessions affordance.
asmspy --serveemits the target’s address-space map as avmmapevent, re-sent only when the map changes; the desktop’s 3D plane names its regions from it (pick detail, region roster, provenance chip).asmspy --dataflow --auto --sampler=ptrace— a third, perf-free sampler rung (ptrace residency probes → call-target expansion →int3arrival confirmation), so--autokeeps working whereperf_event_openis locked down entirely; the chain advances to it on a perf refusal.The 3D plane’s simplified posture: the top-8 worldlines with the five per-event spike layers withheld and counted on a placard; the HUD’s detail toggle restores the full population. Plus occupancy and permissions plane layers (default off).
The live dataflow producer reads ymm16–31 (the EVEX registers glibc’s EVEX code paths keep vectors in), gated on XCR0 rather than CPUID alone.
asmspy --info <pid>— an attach-free process snapshot (identity, runtime, per-threadwchan+ current syscall, symbol/JIT surface, and which tracing modes will work on the target).--jsonemits it as a one-event.asmtracerecording carrying the newprocinfokind.Four standalone 3D scenes, on a substrate that is not the address plane (docs/internal/gui/59-standalone-scenes.md). The 3D pane gains a scene selector; every kind states what each of its axes is and what it is not, because the axis label is part of the host contract rather than each scene’s choice:
Divergence worldline — two recordings, one fused tube up their shared prefix, then A and B ribbons with a rib per step of architectural-state disagreement. A failed identity/basis/arch check renders a refusal card, not an empty scene; a truncated side caps the tube TORN (“we stopped looking” ≠ “they agreed to the end”); an uncomputed step is a hollow rib, never a zero-width one.
Invocation stack — one slab per invocation of a routine, stacked on a discrete invocation # axis that is never scrubbed as time. A block absent from an invocation is a hole in that slab, not a zero column; an unterminated invocation is a labelled prefix; a slab is tinted by the codeimage version in force when it started. Drill-in uses the dedicated
dt_link::invocationfield and leavesdt_link::stepuntouched.Module excursion ribbon — one depth-vs-call-order sheet per thread, banded by module, with the boundary crossings brightened. A depth-capped capture marks its lane floor CAPPED (and the cap survives the drill-in); an unresolved module gets its own hue rather than being blanked or merged; a single-threaded recording renders the honest 2D icicle form instead.
SIMD lane prism — what happens inside one wide vector register across its writes, built from the already-decoded
ValRec::wide/bytes. A stable value→hue hash makes a shuffle’s permutation visible. Element width is not recorded, so the default 16 byte-lanes is labelled as a default and only an unambiguous mnemonic subdivides; uncaptured bytes render a[wide]wireframe rather than zeros.
Inspect before you leap: hover readout and pickable overlays in the 3D scene (docs/internal/gui/47-scene-inspect-and-pickable-overlays.md). The 3D pane now answers “what is this, and where would a click send me?” before any navigation happens:
A throttled hover pick (
SceneHost::pick, at most once per actual mouse-move pixel, zero readbacks during a drag/orbit/pan) resolves to aPickHintvia the newresolve_pick_hint, which shares one classification helper withresolve_pickso the preview can never disagree with what a click does.Convergence arcs and access spurs — previously undrawn in the pick pass — are now pickable via new id bands past the vertex band, decoded through an explicit
PickBandsso band sizes are never inferred; a convergence arc resolves to whichever tid is nearer the current playhead and states which side was chosen.A hover tooltip shows identity/location/quantity/fidelity and the click destination (or “click → nothing here”); the HUD legend gains swatches for both overlay classes, a
convergencelayer toggle, and a persistent “hover to inspect, click to open the flat reader” hint.
Go there: camera pan, recentre, address/region goto, and discoverable controls in the 3D scene (docs/internal/gui/48-scene-navigation-and-goto.md). The plane is a reproducible address layout; now it is navigable:
Cameragainspan/frameplus four keyboardCamKeypan values, bound to middle-drag or shift+left-drag in the viewport (plain left-drag stays orbit).Double-click recentres on whatever is under the cursor — a cell or a placed vertex’s address — without reorienting; a background/padding double-click is a stated no-op, not silent nothing.
A “go to” row in the HUD accepts a hex address or a region from
Projection.regions; an address the recording does not map refuses with a stated reason rather than snapping to the nearest cell, and a region target frames its real cell footprint, never a bounding box. The global find bar offers the same resolver as “show in 3D” for any hit carrying an absolute address.SceneViewcomputes a stable “home” landmark once per weave — the code region’s centroid, not the plane’s fixed centre — so “reset view” returns to a place that does not drift as a live capture grows; “default view” keeps the literal fixed-centre preset as a separate button. A “you are here” readout names the region and address under the camera target.The HUD’s collapsible “controls” block is generated from the
CamKeyenum so an unadvertised key fails a test, not a silent drift.
A GL-free flat terrain surface for the 3D pane (docs/internal/gui/ 52-flat-terrain-surface.md). The 3D overview’s spatial channel was unavailable wherever GL was absent (headless tests, the null backend, a driver whose shader will not build) and had no perspective-free reading mode even with GL, despite the pane’s own rule that precise reading is a 2D job:
views/scene2d.{h,cpp}’scell_paint()mirrorskTerrainFrag’s branch chain exactly (kind hue viaspace::region_style()-> height mix -> churn -> stat -> unknown -> torn), cited back fromembedded.has the keep-in-sync C++ mirror so the two renderers cannot silently disagree about a cell’s fidelity state.build_scene2d_plan()folds a terrain slice into down-sampled blocks (max height, OR’d flags — a torn or statistical cell is never dropped to binning, and the factor is stated on screen), breaks trajectory strips at unplaced vertices instead of interpolating across them, and dims — never drops — path points past the terrain-time playhead, the same rule 49’s worldline clipping uses.The three previously GL-only degraded branches (no
scene_host, notready(), no frame texture) now draw the flat surface beside their existing placard; a new “flat surface” toggle swaps it in for the viewport on a working GL driver. Hover and click route throughscene2d_pick_cellinto the SAMEscene3d::resolve_pick/resolve_pick_hintthe 3D pick path uses, so a flat pick and a 3D pick of the same cell can never disagree — proven directly in the newtest_scene2d.cpp.Region-transition boundaries and labels, a compacted-domain-edge mark (padding is drawn as nothing, never a dark cell), and a scale strip (
scene3d::height_scale_note) round out the surface as a real reading instrument, not just a fallback.
Launch a process and trace it from birth, and target a running one by window-pick (docs/internal/archive/gui/45-launch-and-window-target.md). Two new ways onto a live target, alongside attach-by-PID:
asmspy --serve’s wire protocol gains alaunchcommand ({"cmd":"launch","mode":...,"argv":[...],"cwd":...}) that forks the target itself,PTRACE_TRACEMEs it, and hands it straight to the requested mode’s engine flagged as already-traced — the recorded session starts at the target’s true first instruction, with no detach/reattach gap.asmspy --launch <mode> -- <cmd> [args...]is the headless CLI equivalent. Scoped to the whole-process modes plustrace/watchfor v1;dataflow/auto(a different, deeper seize subsystem) are refused with a stated reason rather than silently degraded.The desktop Home rail gains a fifth entry, “Launch a new process”: a form (command/arguments/working directory/mode) landing on the same live-capture workflow every attach uses (
LiveSession::send_launch,inspect_launch_full_detail, theLaunchpane).A crosshair affordance next to “Capture a live process”: drag it onto any window on screen (even outside the app’s own) to attach to that window’s owning process by PID, Spy++-style — X11-only (
desktop/src/platform/window_picker.{h,cpp}, new), engine-free (D4), honestly unavailable under pure Wayland or over an ssh-remote capture target.libx11-dev+xvfbare pinned inDockerfile.desktop, andmake desktop-test-xvfbproves window resolution against a live (virtual) X11 display and a real second window, not a mock.
The faithful city, Phase A — the MVP terrain reskin (docs/internal/archive/gui/44-faithful-city-phase-a-mvp-terrain-reskin.md). The 3D spacetime overview now reads as a place, not an abstract density field: the terrain is zoned by memory-region kind (
Scene::set_zoning, a newuKindR8UI texture blended intokTerrainFrag), an in-domain but never-touched cell reads as a sunken fog-of-war pit (TerrainFlag::TF_UNKNOWN), the recording’s dominant fidelity tier drives a damped weather sky (scene3d/atmosphere.h, byte-identical in colour source to the 2D fidelity banner), IBS survey residency renders as a physically separate, stippled “ghost district” terrain that never scrubs with the playhead (Scene::set_stat_terrain), and a followed “citizen” vehicle glyph with a comet tail rides the placed PC vertex at a new, independently-tickingfollow_stepplayhead (SceneView::follow_play) that deliberately never fuses with the terrain-time or execution-step clocks. All five newSceneLayerstoggles (zoning/weather/ghost_fog/vehicle, plus the pre-existing five) default on, so the “city” preset is the out-of-the-box view; no.asmtraceschema change.Desktop GUI packaging: a Debian package and an AppImage for
asmtest-desktop(packaging/debian-desktop/,packaging/appimage/). The full app links the GPL-2.0 Unicorn emulator and the Keystone/Capstone engines directly, none of which is a distro package on the pinned base images, so both artifacts vendor the three privately (rpath$ORIGIN) via a newscripts/package-native.sh --appmode; GLFW/GL/FreeType stay real host dependencies (declaredDepends:for the.deb, checked at launch bypackaging/appimage/AppRunfor the AppImage — GL in particular must come from the host to match its driver). Both build straight from a checkout (no released tarball needed) and are verified end to end in CI via two new Docker lanes,make docker-syspkg-deb-desktopandmake docker-syspkg-appimage(folded intodocker-syspkgand a newsyspkg-desktopCI job). The AppImage pack step pins its own tooling (scripts/fetch-appimagetool.sh,scripts/fetch-appimage-runtime.sh) rather than lettingappimagetoolfetch its runtime stub from a rolling tag on every pack. See installation.md.Live
blame+statedifffrom theasmspyserve/dataflow leg (docs/internal/archive/gui/41-live-blame-statediff-serve-leg.md). The backward-attributionblamecone and the step-to-stepstatediffregister delta were producible only by the emulator recorder; a live attach could not emit them. Both are now emitted by the live single-step leg as pure projections over data it already captures —blameis a backward slice over the def-use graph the sink already builds (seeded at the penultimate step),statediffis a delta of the per-step register ring — spelled with the same wire builders the recorder uses, so a live artifact and a golden one are byte-identical. Opt in with serveblame:true/statediff:trueor CLI--blame/--statediff(the latter self-arms the register ring); off by default. No capture-engine, wire, or schema change (both kinds were already defined). The blame cone and state-diff views already worked live client-side; this adds the reproducible, deep-linkable precomputed artifact.Per-invocation dataflow views over a continuous capture (docs/internal/archive/gui/40-segment-dataflow-by-invocation.md). A continuous
dataflow/autocapture appends many invocation passes into one recording, each restartingdf_stepat 0 (delimited by adf_invocationmarker). The desktop’sdecode_streamsindexed those steps in one flat space, so the passes aliased — offsets were last-write-wins and operands/edges merged across passes that merely shared a step number — the last dataflow consumer that still conflated passes after the register Scrubber learned to segment its ring.build_segmented_dataflownow splits them (one decoded stream per pass, bucketed exactly as the Scrubber’sbuild_segmented_step_indexbuckets regstate),Streams::dfresolves to the latest pass (the live default), and the Slice / Timeline / Loom panes carry a per-pass invocation pager — following the latest by default, pinnable to an earlier pass. A one-shot recording stays a single pass, byte-identical to before. Pure desktop decode + a selector: no producer, wire, or schema change.df_stepstates its region base on the wire (rbase) (docs/internal/archive/gui/37-region-tag-on-df-step.md). 36 anchors a routine-relativedf_stepoffset by deriving the base from a recording’s singlecodeimagespan, and must refuse whenever a recording carries zero or ≥2 spans — which a liveautocandidate walk produces routinely. The producer already knows the base as it writes the offset, so it now states it:df_stepgains an optionalrbase(emitted only when nonzero, so arbase == 0step is byte-identical to pre-37). All four producers state it (live ptrace, both corpus recorders, the Author VM). The desktop reader places a tagged step fromrbase + off— a stated fact — with per-event precedence over 36’s recording-wide anchor, and grades HOW it was placed (wire/single-span/mixed) in the HUD, so a multi-span capture that 36 alone must refuse now places every vertex; 36’s single-span anchor remains the permanent fallback for pre-37 recordings and for reltrace. The terrain churn walk is redeemed as a sound “region as-of this step” resolver — it countsdf_stepoffsets as steps (a dataflow recording’s churn no longer pins at step 0) and keys the churn join on each step’s ownrbase. Additive optional field on a known kind — no envelope bump, no break.
Fixed¶
asmspyattach to a multi-process tree no longer kills a followed child: the teardown drained a queued single-step trap against the wrong address space, which was fatal to the child moments after detach.asmspyhardens its JIT perf-map/debug file opens against planted/tmpand/procpaths (symlink/ownership checks before trusting a path).autoreliably captures — one shot is not a policy (docs/internal/archive/gui/39-auto-capture-reliability.md). The complaint “start and arm a process, it starts then stops, and the pane saysrefused: no session is running” was two independent bugs. The region picker: the candidate walk (arm a ranked pick, and on “not seen entering” try the next) lived inline in two identical, untestable copies, and that duplication is exactly how the strong AMD-IBS/entry sampler ended up with no walk at all — it returned a single pick, soncandwas 0 and the guard never fired (on an AMD box the better sampler gave the less resilient capture). The walk is now one pure, unit-testedasmspy_autoregion_walk, andauto_pickhands back the ranked list like its sw-clock sibling, so BOTH samplers walk; an empty sample window is a retry, not a verdict (an idle target yields nothing in one 400 ms window but may in the next), and the window is settable end to end (--window=<ms>, a wiremsparam, a capture-pane input); acontinuouscapture now survives a quiet region instead of ending at the first lull, surfacing the quiet window as a 0-stepdf_invocationmarker. The session lifecycle: a capture that ends on its own is now announced from the tracer tail (an idle client learns without sending a command), the desktop frees the ptrace jack it used to hold forever (reconcilingactiveagainst the terminal-event count), astopfor a self-ended session is acked rather than refused for a precondition the host’s own reap just removed, and a stalerefused:banner no longer haunts the next healthy capture. Every policy change is proven by pure tests (no CI lane has AMD silicon), per CLAUDE.md.The 3D overview places a live routine-relative path, or says why not (docs/internal/archive/gui/36-anchor-the-3d-plane.md). A live
dataflow/autocapture (df_step+codeimage, notrace) isbasis:"rel", so its 3D-overview tab used to open onto an empty, unlabelled plane — three defects stacked: the promised “rel: routine-relative” HUD chip never fired (it keyed on the empty canvas basis), every routine-relative vertex silently failed to project (so a total placement failure was indistinguishable from success), and the value producer emits no block coverage (a flat plane even when anchored). Now a routine-relative offset is anchored to the recording’s singlecodeimagecode span (base + offis the true address — a derivation from a stated fact, not a guess): the PC path is placed on the plane (TRAJ_ANCHOREDalongsideTRAJ_RELATIVE_BASIS; the wire basis staysrel), the terrain grows a labelled single-step-residency height rung (df_step, never block coverage), placement is counted (an off-plane vertex is named, not dropped — the 4096-bytecodeimageclamp), and a recording with zero or ≥2 code spans refuses louder — no geometry and a stated reason (HEIGHTS NOT PLACED/PATH NOT PLACED), never the old silent empty plane. Convergence detection admits an anchored rel path (atracerecording withtid) while still refusing an unanchored one. Purespace/+ HUD — no engine, GL, or wire change.
Added¶
scene-df-loopgolden (docs/internal/archive/gui/36-anchor-the-3d-plane.md T4). A 3D-overview golden in the live dataflow shape no existing golden carried — absolutecodeimage+ region-relativedf_step, notrace— so the anchor fallback and the terrain’sdf_stepresidency rung are exercised end to end by a committed artifact (record_scene_df, regenerated byte-stable underdocker-cli).Author-mode arm64 run/trace (docs/internal/archive/gui/32-per-guest-value-producer.md R5 T3). The Author door’s Run button now dispatches assembled AArch64 code through the per-guest emulator L0 value-fabric producer (
asmtest_dataflow_emu_run_arch,src/dataflow_emu.c) instead of refusing it — never through the x86-64-onlyemu_call_traced/emu_result_tpath, and never touchingsrc/emu.c’s register ring. The result is a genuinely distinct shape (def-use edges + per-step operand values, no register file, no fault kind/address — the door says so explicitly rather than showing zeros) that materialises, on Save, intotrace/df_step/df_edgeevents through the existing recording writer, so an arm64 Author run opens in the Loom / Slice / Timeline like any other recording. The arch-gating table’s AArch64 row is now runnable, the shared refusal label now names only the arches still unsupported (ARM32 / RISC-V), and the capability panel states which arches Author mode runs straight from that same table.arm64 register-ring time-travel for the Scrubber (docs/internal/gui/ 32-per-guest-value-producer.md R5 T2, the regstate half). The emulator’s per-step register ring (
src/emu.c) is arch-parameterized:emu_arm64_tgains its own opt-in drop-oldest ring (emu_arm64_step_capture/_clear/_count/_dropped/_at), mirroring the x86-64 ring’s shape overemu_arm64_regs_tinstead of a union grafted onto the x86-64 handle — zero changes toemu_t,emu_x86_regs_t, oremu_snapshot/emu_restore. A newemu_arm64_regs_t@aarch64/aapcs64regstatedescriptor (docs/internal/gui/asmtrace-schema.md) names the AArch64 GP file (x0..x30,sp,pc,nzcv); thearm64-df-chaingolden now also carries aregstatering (steps_cap = 8, the arm64 analogue ofadd_signed’s worked example), and a newarm64-regstate-truncatedgolden proves the ring’s faithful truncation (D7). No desktop/reader change: the Scrubber’s existing generic “unnamed integer field” fallback already renders the new register names (in a plain sorted rather than hand-curated order, a cosmetic follow-on).A command palette (
Ctrl+Shift+P/Ctrl+P) over thedt_nav_gorouter (docs/internal/archive/gui/21-spine-navigation.md T1). A modal fuzzy finder that makes the whole spine reachable by typing: view-switch, go-to step/offset/link, open-recent (the open workspace), attach-a-process, run-walkthrough, reset-layout, and jump-to-a-recorded-offset — every command dispatching only throughdt_nav_go(or the exactwant_view/show_*intent the keymap uses), plus a discoverable hint for every advertised accelerator, enumerated fromdt_nav_bindings()so the palette can never advertise a key the app does not honour. Filtered with the “showing N of M” idiom (app-only ImSearch relevance, degrading to an unranked list under the null backend). No new dependency.A persistent wayfinding breadcrumb (recording ▸ session ▸ view + step/filter/process scope) with same-basename disambiguation (docs/internal/archive/gui/21-spine-navigation.md T2). Sourced from
nav.currentand drawn in the outer shell (the docked menu bar / the windowed top strip) so it is visible from every pane; two same-basename.asmtracefiles are now distinguishable in both the band and the recording tab titles (basename + the shortest distinguishing parent-dir segment, else a short path hash). Builds on the doc-18 back/forward affordance; the empty state prompts rather than blanks.An always-visible overview/minimap on the timeline and the Loom (whole trace, current viewport marked, click-to-jump) (docs/internal/gui/21-spine- navigation.md T3). A compressed projection of each recording’s own rows — never a fabricated or padded layout (docs 04/08 ban; a sparse trace yields a sparse strip, tested) — with the click routed through
dt_nav_goso a minimap click and a typed go-to land identically. Completes the doc-14 T5 overview-strip / timeline-windowing follow-on (theImZoomSlidernow drives a real timeline window). No new dependency (ImPlot, ImZoomSlider already vendored).Keyboard camera for the 3D overview, Tab-focusable panes and 3D viewport, and keyboard-operable slice cones (docs/internal/archive/gui/22-selection-and-search.md T2, F18). The 3D spacetime overview was a mouse-only island; because Dear ImGui exposes no OS screen-reader tree, keyboard operability is the only accessibility substitute. Arrows now orbit,
+/-(and=) dolly,Rresets andTis the faithful top-down 2D-ish fallback — routed through the SAMECameramethods the mouse drag uses (a purecamera_key), so keyboard and mouse are one code path. The 3D viewport gains a Tab-reachable focus target that exists even under the null backend (no GL), which is what makes the keyboard camera headlessly testable; the arrow keys defer to the camera only when a 3D pane holds focus. The slice cones (b/f/Enter) were already keyboard-operable.Global find (
Ctrl+F): highlight-all, match count, aggregate cost, and Enter/Shift+Enter cycling; type-to-narrow extended to disasm and hot-edges (docs/internal/archive/gui/22-selection-and-search.md T3, F17). Find is a measurement, not just a jump: it highlights EVERY hit in the timeline (with a clean seam for the doc-21 minimap), reports the match count AND the aggregate cost (summed hot-edge samples across sites — “this symbol retires N samples across M sites”), and cycles matches through the ONE navigation spine. It searches the timeline, disasm, order-preserving syscall stream and hot-edges in stream order (never relevance-ranked), and never hides a row. The doc-16 client-side “showing N of M” narrowing now also covers the disasm and hot-edges lists. The call tree is deliberately untouched — it stays engine-filtered, so surviving depths never lie (D7).App-level undo/redo (
Ctrl+Z/Ctrl+Y) over filter / cone / selection / take-set state — distinct from the Author editor’s text undo; the Loom takes gutter gains per-take remove and clear-forks (docs/internal/gui/22-selection- and-search.md T4, F12). Applying an aggressive filter, lighting the wrong cone or forking exploratory Loom takes is now reversible: an app-level command stack reverses and replays those view-model changes, deliberately DISJOINT from the Author editor’s own text undo (both guard on text-input focus and own separate state) and from the router’s back/forward history. The Loom takes gutter — which held no take set before — now accumulates forks with a working per-take remove and a clear forks action, both reversible; clearing removes a whole take node with its loud refusal intact, never silently dropping a failure (D7).Derivable
severityfidelity-chrome tier in the.asmtraceschema (docs/internal/archive/gui/23-graded-truth-layer.md T1). An optionalprovenance.severitystring (neutral | caution | integrity) names the tier a reader renders a recording’s fidelity chrome at. It is derivable from the fidelity fields the schema already carries (torn / truncated / trust / redacted / skip / basis / drops), so a producer need not emit it and every old recording still grades; additive, no new envelope major, no field on any existing kind. Landed under doc 01’s append-only rule with 01-owner sign-off recorded in the schema doc (the Phase-3-freeze checkpoint, D5). It grades loudness; it gates no truth off (D7).A uniform elapsed-time + Cancel busy signal on long operations, and a degrade-to-coarse 3D scrub (docs/internal/archive/gui/23-graded-truth-layer.md T4).
progress.hgrows a pureLongOp(elapsed clock + faithful determinate-vs- indeterminate mode + observable cancel flag) and adraw_progresshalf, so an op that can exceed a frame shows a spinner + elapsed + Cancel instead of a silent UI-thread stall an expert would mistake for a hang. The 3D scrub degrades to the labelled coarse terrain plane (a pureshould_degradedecision +coarse_slice) while a full re-slice would exceed the frame budget, then swaps to the full slice — the coarse plane is the same labelled rung the terrain shows normally, so it hides nothing. The no-fabricated-total fidelity rule is unchanged.A Queue path on the live patch bay (docs/internal/archive/gui/23-graded-truth-layer.md T3). A blocked capture can be queued as a visible, cancellable chip; it starts automatically the moment the jack frees — and only then (never an auto-swap, so the one-ptrace-jack invariant is never bypassed).
The desktop remembers your workspace, and a Settings pane (docs/internal/archive/gui/20-workspace-and-settings.md T3/T4/T5). Open recordings, the active view and each pane’s selection now restore across launches (persisted as
asmtrace-links inbuild/desktop-workspace.jsonbeside the dock.ini), with an MRU recents list on the home rail and a File ▸ Open Recent menu — each entry reopening to its exact prior position — plus drag-drop to open a file. A recording whose file has vanished is kept in recents with its load error, never silently dropped. Named dock perspectives and named saved filter presets persist in the same store. A new Settings pane adds a user text-scale (0.8×–2.0× viaFontGlobalScalewith a DPI-aware atlas re-bake on content-scale change), a remembered window size (retiring the hardcoded 1280×720), and a light theme that keeps the fidelity-chrome (warn/refuse) contrast — and states transparently that text-scale is the only in-app accessibility lever, because Dear ImGui exposes no OS screen-reader tree.In-app term registry, per-view “?” caveats, domain-term-first headings, and a searchable Terms pane (docs/internal/archive/gui/24-one-visual-language.md T3). The coined GUI lexicon (Loom, fabric, patch-bay, hollow span, born-untraced, patient-zero, hot-edges, knot, jack, worldline, Reweave, dim/hot take, terrane) is now defined in the ONE Sphinx glossary (
docs/project/glossary.md) with each term’s plain-language definition and its expert synonym;scripts/gen-terms.pygenerates the app’s registry from that one file (no hand-copied second list), so a tooltip, a legend and the docs cannot drift. Every coined surface now leads with its canonical domain term and the metaphor as a subtitle (“Data-flow lineage (Loom)”, “First divergence (patient zero)”, “Edge execution counts (hot-edges)”), carries a per-view “?” with the verbatim metric caveat (“hot-edges are edge counts, not a call stack”), and offers a searchable Terms pane (the doc-16 ImSearch idiom, guarded so the null backend degrades to a plain list). The registry/lookup/heading metadata are asserted headlessly against the same glossary the build parses (one source).First-open primer + legend on the Loom and the 3D overview (docs/internal/archive/gui/24-one-visual-language.md T5). The two heaviest surfaces now open a dismissible, in-canvas primer (“what this is / how to read it” + the shared legend) instead of dropping a learner into a raw fabric/terrain; it shows once per session, never nags again, and is re-openable from the per-view “?”. Until it is acknowledged the view holds its lean default (nothing heavier is front-loaded). State transitions are asserted headlessly.
Convention-alignment keyboard shortcuts, accurately advertised (docs/internal/archive/gui/18-breach-stops.md T1).
Shift+Ffits / frames the current selection,,/.step to the previous / next sibling,F10/F11are step / step-back aliases ofj/k,W/S/A/Ddrive camera zoom/pan when a spatial pane (timeline / 3D) holds focus — a labelled context that also keepsd’s diff meaning outside it — andCtrl+Ccopies a deep link (an alias ofy). The help overlay is now generated from awiredflag, so it greys any not-yet-mapped binding “planned” and can never again advertise a dead key as live. Every advertised binding has a passingdesktop-ui-test.Back/forward navigation history over the deep-link router (docs/internal/archive/gui/18-breach-stops.md T6):
Alt+Left/Alt+Rightwalk a bounded, serialisable stack ofasmtrace-links and land identically to a fresh navigation, with a minimal breadcrumb affordance by the status line. A new jump clears the forward branch (browser-history discipline).A reset-layout keybinding (
Ctrl+Shift+R) and a pre-commit confirm before arming a perturbing single-step capture (docs/internal/gui/18-breach- stops.md T2/T5): the confirm states the page-dirty / timing cost and, on arm64, the blocking-syscall termination that detach cannot undo; the capture picker now defaults to the least-perturbing substrate the host supports (AMD IBS where available, else the lightest ptrace mode), and arm64 single-step modes are greyed and annotated with the stated hazard.Author output can be saved, and an unsaved-work marker in the tab title (docs/internal/archive/gui/18-breach-stops.md T3): an Author run materialises into a Recording and writes to a
.asmtracefile through the same confirm-overwrite dialog live captures use; a dirty (authored + unsaved) tab shows a trailing*.Pan/zoom, fit-graph and selection for the graph views (docs/internal/gui/15- plotting-and-graph-nav.md T3, via imgui-node-editor). The process topology, the live/replay call tree and the frozen hot-edge snapshot now draw on a real pan/zoom canvas: drag to pan, wheel to zoom, “Fit graph” frames the whole graph (
NavigateToContent), and double-clicking a node navigates through the same deep-link router the buttons use (a process’s syscalls, a call’s or edge’s address in the region view). The node positions are the app’s own deterministic layout, fed every frame — node-editor’s settings file is disabled so it can neither persist a dragged position nor invent one on load: force-directed layout stays banned (docs 04/08), the fidelity guardrail is that the layout lives in a pure, tested builder, never in the drawing library. Large graphs (a 10k-routine trace) cull off-viewport nodes before emitting them. imgui-node-editor is vendored (MIT, master021aa0ea— the last release fails to compile on the pinned imgui) with its bundledcrude_json; its four TUs compile into both desktop binaries and it is in the imgui-repin compile-probe. The layout is a pure model, so the null-backend tests assert node positions, the disabled settings file, click-routing and culling without pixels.The keyboard shortcuts now all work, and are tested. The desktop help overlay advertised twelve keybindings but only
[/]were wired; the other ten are now implemented (docs/internal/archive/gui/17-interaction-testing-and-editor.md T1):1/2/3/4switch view (canvas / timeline / slice / diff),j/k(and the arrows) andPgDn/PgUpstep the selection,Enteropens the slice explorer at the selection,b/flight the backward / forward dependence cone andcclears it,dattaches or detaches a second recording for the diff,xswaps A and B,n/pwalk the divergences of an attached pair,ycopies a deep link to the current position, andCtrl+Gopens a go-to box that accepts a step number or a fullasmtrace-link. They are decided in one place, so the help overlay and the behaviour cannot drift.A headless interaction-test lane for the desktop app (
make desktop-ui-test), built on the Dear ImGui Test Engine. It drives the real UI through simulated clicks and keypresses on the null backend — the layer the golden-text tests cannot reach — and writes JUnit XML for CI. Every one of the twelve keybindings has a passing test (which is what proves they all work). The engine is fetched at build and compiled into the test binary only, never into the shippedasmtest-desktop/asmtest-viewer, so those stay MIT.Non-modal toasts for live-session events (docs/internal/gui/16-live- feedback-and-filtering.md T1, via ImGuiNotify). Events that were previously silent — a capture refusal, a session ending, a fatal, a skip (success-with-nothing-to-report), a save completing or failing — now raise a transient toast in the corner. An exact saved capture’s toast carries an “Open in Loom” button that opens it through the same door the panes use; a statistical capture’s does not (there is no Loom for it). Toasts supplement, never replace, the in-pane refusal banners (the banner is still the record). The decision of which toast to raise is a pure, unit-tested function over the live status and save outcome, so it is driven as a model by the null-backend tests. ImGuiNotify is vendored (MIT) with a minimal Font Awesome range merged into the UI font for its type icons. Bundled in both the full app and the render-only viewer.
Client-side filtering for the long lists (docs/internal/gui/16-live- feedback-and-filtering.md T2). The Learn-door walkthrough catalog narrows as you type (via ImSearch). The syscall stream gains a name filter that hides non-matching rows while keeping each row’s original number and execution order and showing “showing N of M”, so a filtered view never reads as if the trace only made those calls; it matches syscall names only, never the redacted payload. (Filtering is applied only where it can stay faithful: the time-ordered syscall stream uses an order-preserving filter rather than a relevance re-rank, and the address-ordered disassembly keeps its existing time/range controls.)
The register scrubber and the ABI x-ray are now hosted in visible shell panes (docs/internal/archive/gui/09-teaching-producers.md T3/T4 — the integration surfacing pass). The register time-travel scrubber is a per-recording Scrubber tab: its
regstateseek index (analysis/stepindex.h) is built once when the recording opens, parallel to the workspace like the decoded streams and Observer decks, and its playhead is persisted per recording so switching tabs holds each recording’s place. The ABI x-ray is a tab that locks the active recording (the SysV leg) against the attached B (the Win64 leg, thedbinding) — reusing the Diff tab’s A/B mechanism — with the walkthrough rail rebuilt only when the pair changes; no B attached shows an “attach the Win64 leg” placard, the same shape Diff shows with no second recording. Both degrade to their own fidelity placards (absent producer, unaligned pair, torn ring) exactly as their standalone draws do. Both views are backend-free, so they ship in the render-only viewer too, anddesktop/test/test_shell.cpppins the wiring end to end: the seek index is built andpresent()for aregstaterecording, amake_pairpair feeds a present and aligned x-ray through the very calls the tab makes, and the per-recording parallel vectors survive a close.The 3D spacetime overview is now hosted in a visible shell pane (docs/internal/archive/gui/10-spacetime-3d-overview.md — the integration surfacing pass). It is a per-recording 3D overview tab. The GL boundary is a new abstract
SceneHost(desktop/src/ui/scene_host.h): the shell weaves the pure, engine-freespace/models once per recording (aSceneView), draws the HUD (draw_scene_hud) and resolves picks (resolve_pick) itself — all GL-free, soui/shell.ostays out of the viewer’s engine closure and the null backend drives the whole model + HUD + placard path — and reaches the GL scene only throughShellState::scene_host.main.cppinjects the one concrete host (ui/gl_scene_host.cpp: ascene3d::Sceneover an offscreen RGBA8+depth FBO whose colour texture the shell blits withImGui::Image, re-uploading the terrain/trajectory only when the recording or playhead changes); the headless tests inject none and the pane shows a placard where the viewport would be. A left-drag orbits, the wheel dollies, and a click that did not drag picks →decode_pick→resolve_pick→ 04’s deep-link router (3D to find, 2D to read). Because the scene links no engine it ships in the render-only viewer too (D4).desktop/test/test_shell.cppdrivesdraw_scene_overviewunder the null backend (models woven, code regions placed, terrain heat, the exact trajectory, the terrain slice tracking the HUD playhead, and the faithful no-regions placard for acodeimage-less recording); the GL scene stays pinned by thetest_scene_fbosmoke and the pick router bytest_drillin. This completes the deliberately-outstanding integration item — all ten GUI docs’ views are now surfaced in the shell.Golden scenes and a CI-runnable GL lane for the 3D spacetime overview (docs/internal/archive/gui/10-spacetime-3d-overview.md T7 — the view is now complete). Three committed scenes pin the whole stack end to end. Two are generated by
make asmtrace-goldenfrom one byte literal whose listing sits beside the bytes, so the terrain’s heat is countable on paper:tests/golden-asmtrace/scene-abs-loop.asmtrace(the coarse rung —codeimage+ absolute-basistracefrom a 3-iteration loop, giving one hot cell over cold ones: coarse terrain plus one exact trajectory) andscene-abs-loop-truncated.asmtrace(the same bytes with the trace buffer holding 6 of the 13 executed steps, so the producer flipstruncated,drops.lostcounts the rest, and every populated cell isTF_TORN) — both byte-stable undermake asmtrace-golden-check. The third, the richmemscene, cannot be generated at all (memis a reserved schema kind with no producer) and is hand-authored undertests/golden-asmtrace/scenes/, marked schema-unfrozen and deliberately inert but present: its accesses fall outside every code region, so the coarse-rung projection gains no data cell and the scene renders pixel-identically to its coarse twin.desktop/test/test_scene_fbo.cppnow builds and renders all three under surfaceless EGL + software Mesa in themake docker-desktoplane, asserting terrain heat and the exact tube for the abs scene, the red gash for the truncated one, and — at the pixel level — that a survey-only scene draws its statistical layer and nothing on the exact one; the statistical scene reuses the committedobs-survey-ibsfixture rather than adding a second survey saying the same thing. Each scene’s model facts are asserted in the test’s pure half too, so they hold on a host with no GL.desktop/README.mdgains a 3D spacetime overview section (controls, the coarse-vs-rich staging, and the “3D to find, 2D to read” rule).Drill-in router + fidelity invariants for the 3D spacetime scene — every pick reaches the right 2D view, and truncation/statistical survive it (docs/internal/archive/gui/10-spacetime-3d-overview.md T6). The scene is an overview, not a reading surface:
scene3d/pick.h’sresolve_pick(pure, GL-free) now routes every pickable kind through 04’s deep-link router to the flat 2D view that actually reads it — an exact code cell → the trace canvas at that offset (or the codeimage-versioned disasm pane, 08-T7, when the region churned); a data cell (richmemrung) → the slice explorer at the step whose access last hit it; an exact PC vertex → the operand timeline at that step. Two fidelity invariants are pinned as tested behaviour: truncation survives the drill-in — opening aTF_TORNcell lands on a 2D view that still carries 04/08’s truncation banner, so the 3D tear is never the only signal; and statistical is never exact — a survey-only recording draws no exact trajectory tube, its statistical provenance is surfaced, and a statistical pick (aTF_STATcell or aTRAJ_STATISTICALvertex) opens the hot-edge view (08-T4), never the exact slice explorer. Newdesktop/test/test_drillin.cpp(headless, no GL) asserts each pickable-kind mapping and both invariants against thetruncatedandobs-survey-{ibs,sw}fixtures; builds into both binaries and stays engine-free (D4).Live-observer overlay for the 3D spacetime scene — N per-thread trajectories and cross-thread convergence hints (docs/internal/archive/gui/10-spacetime-3d-overview.md T5). A live 07
LiveSessionnow feeds the 3D overview: each thread is one coloured trajectory growing over the shared terrain in real time, fed incrementally by re-running the trajectory builder on the session’s growing recording — no second capture path, the overlay consumes the session the app already owns and opens no ptrace of its own (D6/D9). A new engine- and GL-free detectordesktop/src/space/converge.{h,cpp}surfaces where two threads meet: when two different tids place a PC vertex in the same projection cell within a sliding step window, it emits one convergence mark per thread-pair-and-cell (the closest crossing), which the scene draws as a bright magenta arc bowing above the plane on its own toggleable layer. It is explicitly a hint, never a proof — per-thread step indices are not a global clock, so it shows co-location, not a proven race or ordering — and a statistical or region-relative (single-step) path never converges, because a shared cell over a sampled or unplaced path would be a false address-space claim. Threads are placed by colour, not by physical core/NUMA (the repo has no scheduling feed — the affinity layer stays gated, recorded not faked). Builds into both binaries and stays engine-free (D4); covered bytest_converge(the detector + the incremental feed against a two-thread fake-serve fixture) and an added arc case intest_scene_fbo.ABI x-ray in the desktop GUI — SysV vs Win64 argument marshalling, side by side (docs/internal/archive/gui/09-teaching-producers.md T4). The classroom flagship: two register scrubbers (System V | Microsoft x64) LOCKED to one playhead, driven by an authored walkthrough (06’s
notestops).asmtrace_recordgrew a Win64 recording path — it now records the same corpus routine twice, once throughemu_call_tracedand once throughemu_call_win64_traced, into paired goldens (abixray-make_pair-{sysv,win64},abixray-sum3-{sysv,win64}) that carrytrace+ per-stepregstate+ the walkthrough stops. Advancing a stop seeks both panes together; the per-step register deltas do the animation, and a cross-pane diff marks the registers whose SysV and Win64 values disagree — so a0 → rdi (SysV) vs rcx (Win64), the callee-saved rsi/rdi role reversal and struct-return eightbyte classification are the register diff, made mechanical. Faithful (D7): a pane with noregstateproducer refuses and names--steps, two runs of different lengths are flagged not aligned (they cannot share a playhead), a dropped step renders UNKNOWN per pane, and a stop past the recorded window is refused not clamped. Builds into both binaries and stays engine-free (D4); covered bytest_abixray(builder) andtest_abixray_draw(the ImGui half).Save a live capture to a
.asmtracefile from the Inspect door — and open it in the Loom (docs/internal/archive/gui/07-serve-live-host.md). The live host kept its recordings in memory only; the Inspect door now has a save capture section that writes the growing (or last completed) recording to disk in the same NDJSON a--recordrun produces, via a new model-native serializerrecording_to_asmtrace/save_recording_file(desktop/src/doc/recording.h) that round-trips through the loader: header, every event in stream order, and theendfooter only when the recording had one — so a torn capture reloads torn and a truncated one stays truncated. A saved exact recording offers an Open in Loom button that loads the file into the Workspace and jumps the tab strip to its Loom (a statistical capture says why it cannot be woven instead of sending the user to the Loom’s refusal). Covered by newtest_recordinground-trip cases (string and file, clean and truncated).Register time-travel scrubber in the desktop GUI (docs/internal/archive/gui/09-teaching-producers.md T3). A playhead over a recording’s steps showing the full register file at each step, seeked O(1) through a new shared
regstateindex (desktop/src/analysis/stepindex.h, which the Loom’s now-column reads too). Registers that changed vs the previous held step are highlighted. Truthful about the ring’s limits (D7): when the producer dropped a prefix of steps the scrubber renders that region as a torn edge — seeking into it shows UNKNOWN, never zeros — and a recording with noregstateevents states the producer is absent and deep-links the docs (the re-run-with-larger-max_insnsfallback is deliberately not offered). Builds into both binaries and stays engine-free (D4); covered bytest_scrubber(builder: seek, diff-highlight, tear, absent message) andtest_scrubber_draw(the ImGui half under the null backend).The 3D spacetime overview’s GL scene — camera, terrain mesh, trajectory tubes, colour-ID picking (docs/internal/archive/gui/10-spacetime-3d-overview.md T4). The desktop now draws the address-space projection, a terrain slice and the execution trajectories as a 3D scene under the ImGui HUD, in the same GL context the shell already stands up. A
2^order × 2^ordergrid VBO is displaced by a per-cell height texture and coloured by a flags texture (aTORNtear is a red gash, statistical residency is dimmed, JIT churn is tinted); exact paths draw as opaque lines while statistical residency is translucent and stippled and is never joined into an exact tube. An orbit camera (drag to orbit, wheel to dolly, a “reset view” and a faithful “top-down 2D-ish” preset) frames it, and every pick — resolved through an offscreenR32UIcolour-ID framebuffer — leaves the 3D overview for the flat 2D view that reads it: a terrain cell opens the trace canvas at that code offset, a trajectory vertex the slice explorer at that step (3D to find, 2D to read). The scene links no engine — only OpenGL and the purespace/models — so it ships in the render-onlyasmtest-vieweras in the full app; the ImGui HUD and the pick/router logic are separate TUs that link no GL. Camera math is pinned bytest_camera(no display), and the terrain + picking bytest_scene_fbo, a gated GL smoke that renders offscreen via surfaceless EGL on software Mesa. The math comes from a newly pinned single header,linmath.h(WTFPL), fetched and digest-verified like imgui/json.Opt-in per-step register-capture ring in the emulator (docs/internal/archive/gui/09-teaching-producers.md T1).
emu_step_capture(e, cap)arms one moreUC_HOOK_CODEon the x86-64 guest that snapshots the fullemu_x86_regs_tbefore every executed instruction into a bounded, caller-sized ring — the per-step producer a register time-travel scrubber and the ABI x-ray replay from. It is drop-accounted (when a run exceedscapthe earliest entries are evicted andemu_step_dropped()counts them, so a torn timeline is genuine data, not a silent gap), read back withemu_step_count()/emu_step_at(), and never armed by default: the arming is handle-level, persists acrossemu_call_*untilemu_step_capture_clear(), and survivesemu_snapshot/emu_restore(the Track F arming discipline).The live Observer views — seven of them, plus a
codeimagekind and a PT-replay slice (docs/internal/archive/gui/08-observer-views.md T1–T8). The desktop now renders what a liveasmspy --servesession produces: the syscall stream (payloads redacted by default, revealed per row, session-wide reveal behind a second confirmation), the watchpoint timeline, the process topology, the statistical hot-edge table, the call tree with the engine-side filter panel the TUI never exposed, the region trace as discrete invocation snapshots, and a disassembly pane that resolves bytes as of trace time.The load-bearing property is that these are not “live views” at all: each is a pure function of the recording document model, so the Inspect door and a replayed
.asmtracetab draw the same deck from the same code — which is how an AMD IBS survey, an arm64 watchpoint refusal and a JIT code image are all asserted in CI on hardware that has none of them.Each view carries the fidelity rule that view can most easily lose. Redaction distinguishes hidden here from withheld at record time (and revealing cannot conjure bytes that are not in the file). A watchpoint’s direction has three values, the third being “the trap fired and the instruction did not decode”, and a value that was never read back is never rendered as
0. A refused watchpoint arm is a successful session with nothing to report, carryingasmspy_hwdebug_reason()’s measured string verbatim. The topology view states that theprocsengine holds the ptrace jack for the whole descendant tree, and never showsinvwithout saying whether it counts syscalls or calls. Hot edges are edges, not stacks — there is no flame graph, because nothing in an IBS sample observed a call stack — and IBS entry evidence is labelled differently from software-clock residency. The region view pages between invocations and never scrubs, because between two invocations the target ran unobserved for an unknown time.codeimage— captured code bytes at a version (schema; produced byasmspy --serve). A JIT patches, frees and reuses code addresses, so bytes read after the fact are not the bytes that ran. A region-scoped serve session now tracks its region throughasmtest_codeimageand streams versioned snapshots, and the desktop resolves an address at trace timetto the version with the greatestwhen≤t— never the newest, and unknown rather than the next one along, since that would be the succeeding method’s code. Where the recorder is unavailable (soft-dirty /PAGEMAP_SCAN, Linux ≥ 6.7) the session emits the measured reason as anoteand captures without it; the pane then falls back to the recordeddisasmstrings and labels them as the weaker source.make cli-smokeasserts the events are well-formed, that the tracer’s own entryint3never appears in them, and that a static region does not accumulate byte-identical versions.The PT-replay def-use slice (08 T8). On a PT host a value slice can be produced with zero single-steps of the target: the hardware records the path, and the F5 producer replays it through the emulator against the recorded code image to reconstruct the values. The result is an ordinary def-use stream, so the existing slice explorer and Loom draw it unchanged. Capture is hardware-gated with the library’s own reason; replay is not — a path decoded at capture time (
stitch) and the bytes it was decoded against (codeimage) replay anywhere Unicorn and Capstone are present, which is whatdesktop-testexercises on hosts with no Intel PT.asmspy --serve[=<socket>]— the live-session control loop (docs/internal/archive/gui/07-serve-live-host.md T1/T2/T6). asmspy can now be driven as a capture host rather than a one-shot command: it reads NDJSON commands (start/pause/stop/quit) on stdin or aunix(7)socket and streams back the events of whichever engine is running. This is how the desktop GUI captures — it spawnsasmspy --serveas a subprocess, locally or viassh <host> asmspy --serve, so remote capture is the same code path as local and the viewer links no tracer at all.It is a thin wrapper over
libasmspy, and deliberately so: every mode drives one engine through the public header with the same sinks and the same writer TU--recorduses. There is no serve-specific event body anywhere, so a session’s events are exactly a recording’s events — slice[header … end]out of the stream, drop the control lines, and any reader in this tree parses it. The protocol is specified normatively (Serve protocol in docs/internal/gui/asmtrace-schema.md) rather than defined by the C code, including the three lifecycle kinds (session/cmd/err) that were reserved in the kind registry and are now defined.Two properties carry over rather than being re-argued. One session at a time, because a target has one ptrace jack: a second
startis refused with anerrnaming the rule. And a session only ever ends through the engines’ two-phase detach — stop flag,SIGALRMto unblock a pendingwaitpid, join — so the target survives, whichmake cli-smokenow asserts end to end (start log → stop → start stream → quit, victim alive after). The flag matrix is the argument parser’s, verbatim, so the two front ends cannot disagree about what is legal.Fidelity is preserved where it would have been easiest to lose:
pausesuspends emission but not tracing, so the events it swallows are counted, the recording is markedtruncated, and the terminal event carriespaused_dropped. And a skip reports the measuring source’s reason — which surfaced a real defect:asmspy_strerrorhad no case forASMSPY_SAMPLE_UNAVAIL, so that positive skip code rendered as the default"attach failed"— wrong twice over, since the IBS sampler is out of band and attaches nothing. Fixed.libasmspy— the asmspy tracer engine is now a linkable library (docs/internal/archive/gui/07-serve-live-host.md T0). The ptrace engines (cli/asmspy_engine.c) and the/proc/ELF/JIT resolvers (cli/asmspy_proc.c) were loose objects compiled straight into theasmspybinary with no public header, so anything else that wanted them — the forthcoming--servecontrol loop, a future language binding — would have had to re-declare or re-implement them. They now ship asbuild/libasmspy.aplusmake shared-asmspy(build/libasmspy.so), behind one public headercli/libasmspy.h, exactly as every other tier already ships (libasmtest_emu/_dataflow/_hwtrace).This is packaging, not a rewrite: the engines, sinks, and the one-tracer-thread / two-phase-detach contracts are byte-for-byte the code that was already there, and
cli/asmspy.hnow includes the public header and re-exports it, so every existing includer compiles unchanged.cli/asmspy.hkeeps only the front end’s own entry point.What the split buys is a boundary that can be tested: the new
cli/test_libasmspy.cincludes onlylibasmspy.hand links onlylibasmspy.aplus the framework tier objects — notasmspy.oand not ncurses — then attaches to a live victim, streams syscalls through a real sink, detaches and asserts the target survived. A hidden dependency on the CLI front end, or a TUI dependency leaking into the engine, now fails there and nowhere else. It runs inmake cli-smoke.cli/asmspy_autoregion.hbecame self-contained in the same change (it usedasmspy_sample_edge_twithout including anything that declared it, forcing consumers to pin an include order by hand —desktop/src/vm_compat.cpphad that pinned against clang-format). The desktop still links nothing of this (D9): it reaches the engines only through theasmspy --servesubprocess.make desktop-setup— one command from a bare host to a runnable GUI. Nothing bootstrapped the desktop app:make depscovered the engines but knew nothing about GLFW or GL, so the app backends were reachable only by copying the apt line out of the dependency-gate’s guidance text. The new target runsinstall-deps.sh --desktop, then the pinned Capstone and Keystone source builds — not optional extras, since no Linux package manager ships either engine, so a package-manager-only setup would leavemake desktopstill gated — and finally builds both binaries.make desktop-setup-renderdoes the viewer half: app backends only, no engines, no source builds (D4).The build step is a recursive
$(MAKE), which is load-bearing rather than stylistic:DESKTOP_MISSING/DESKTOP_ENGINE_MISSINGare$(shell pkg-config)probes expanded when make readsmk/desktop.mk, so a setup target that installed the dependencies and then merely depended ondesktopwould be judged against the pre-install answers and print the guidance text it had just made obsolete. Every step is idempotent, so re-running on a set-up host is a plain incremental build.install-deps.shlearned the dependencies it was missing:--desktop(glfw + GL + unicorn + capstone + keystone + pkg-config + buildtools),--desktop-render(the app backends alone), and--buildtools(git + cmake — what the pinned source builds need in order to run, which nothing previously installed even though--asm/--emuboth point at them). Package names cover apt/dnf/yum/pacman/zypper/apk/brew;gl_pkgis empty on brew because macOS OpenGL is an Xcode framework, not a package.--alland the no-flag default pick up the new dependencies;--emu/--asm/--nasm/--tidyare unchanged.The three doors: Learn, Author, and a capability panel — plus runner record mode (desktop GUI plan, Phase 2; docs/internal/archive/gui/06-doors-and-learning.md). The GUI’s first-run promise is “no blank IDE”. The Learn door plays bundled walkthroughs, and a walkthrough is not a document beside a recording — it IS a recording, with ordered
stop:truenotes in it, so a story cannot drift away from the run it narrates. Four ship (square,demo-fail,ct_eq, and a truncated low-fidelity fixture), regenerated byte-identically bymake asmtrace-walkthroughs; the truncated one’s last stop points past the recorded window and the player must say so rather than clamp. The Author door (full app only) assembles what you type, runs it, and renders faults as data — the assembler’s own diagnostic survives verbatim, because it is the one sentence that names the AT&T-under-Intel trap. The capability panel shows what this host can do and why not, straight fromasmtest_trace_resolve/asmtest_hwtrace_status/ both IBS reasons: a greyed row always carries its measured reason, and ticking “native only” into an empty cascade says the library returns EUNAVAIL rather than silently downgrading to the emulator.ct_eqships as a real suite (examples/ct_eq.s+test_ct_eq.c,make ct-eq-test), replacing the illustrative docs snippet: block-coverage union across secret-differing inputs, withleaky_eqas a negative control that asserts the union does grow — without it, “no new blocks” would also pass for a routine that never ran.Runner record mode:
--record-dir=DIR(alsoASMTEST_RECORD_DIR). Every test gets a recording path, and a failing test’s TAP block and JUnit<failure>text carryrecording:andstep:— additive keys, so existing report parsers keep working. The runner links no engine and records nothing itself: it arms a directory and carries a producer’s note (newasmtest_record_path/asmtest_note_recording, withasmtest_rec_emu()as the emulator-tier glue), so a suite with no producer accepts the flag, writes nothing, and emits norecording:key. A directory that cannot be created is a hard exit 2 — a run asked to record that silently recorded nothing is the exact outcome the flag exists to prevent. A recording noted by a test that then crashes still reaches the report; a child that died before reporting names none.The Loom — a recorded run as a spacetime fabric (desktop GUI plan, Phase 2 flagship; docs/internal/archive/gui/05-loom-day-one.md). New
desktop/src/loom/turns a value trace and its def-use graph into lanes, worldline spans, hops and knots: a register deck in the producer’s own fixed order, memory coalesced into bands on first touch, and a pure zoom-aware draw plan (spans under 3px collapse into a live-count density ribbon; a band at >=12px per byte explodes into per-byte rows). Clicking a worldline lights its whole thread-tree,[/]walk it a generation at a time, a biography narrates how the value was born, every hop it takes and where it escapes to memory, and a zeroization audit answers “who still holds a descendant at time T” — with its own title stating that “clear” means not overwritten within the traced window. The lane annex joins what companion recordings saw about the same place under a deliberately closed two-verdict enum:corroboratesorunconfirmed, never “contradicts”, because a statistical feed’s silence proves nothing. Forks change exactly one fact — an entry argument or the routine’s source — and re-run from entry, rendering an interventionaldim/hot/neutralverdict per step, whereneutralis what you get whenever either side never captured the value. Five fidelity rules are structural and tested verbatim: a statistical producer is REFUSED with a reason and no partial fabric, an uncaptured value stays hollow, a value alive at the last recorded step gets a fade-out and never a death cap, a worldline whose first record is a read is marked born-of-untraced-state, and an unassemblable fork patch fails loudly rather than weaving code the user did not write. Four committed golden looms (including a generated truncation fixture) and ten newdesktop-testbinaries cover it;test_loom_paritypins the generation walk’s closure tosrc/dataflow.c’s slicer over 200 pseudo-random graphs.Desktop replay views — trace canvas, operand timeline, slice explorer, recording diff, deep links (desktop GUI plan, Phase 2; docs/internal/archive/gui/04-replay-views.md). The desktop app now renders what a recording contains rather than just its summary: per-offset heat with a block-granular coverage gutter, the per-step operand-value timeline, the def-use slice explorer (click a step; see everything that produced the value and everything it affects), and a two-recording diff. Slicing is computed client-side from recorded
df_edgeevents, so the engine-freeasmtest-viewerslices with no engine linked — andtest_slice_diffpins the viewer’s closure tosrc/dataflow.c’s slicer over 200 pseudo-random graphs, so a divergence between the GUI’s answer and the TUI’s fails the build. The slice layout is layered by step index and fully deterministic, never force-directed. Value annotations call the TUI’s owncli/asmspy_dataview.hhelpers, so both frontends speak one dialect. Every position is addressable asasmtrace-link:v=slice&rec=...&step=4, round-trip byte-stable, and every view and keyboard binding routes through one router. Fidelity is enforced structurally: a recording mixing region-relative and absolute events draws no rows (a placard instead), a truncated recording carries a non-collapsible banner naming how much is missing, cones over a truncated stream are labelled lower bounds, a dropped step renders as unknown rather than as offset 0, statistical hot-edge data is never merged into exact heat, a refused diff produces a reason instead of plausible numbers, and a diff bounded by truncation says “no divergence observed within the recorded window” — never “identical”.make desktop-testgrew ten binaries covering all of it.Backend-completeness panel and its data readers (docs/internal/archive/gui/02-exporters-and-readers.md T5/T6). New
desktop/src/data/readers for the three shapes the existing producers emit (a liveasmfeaturessweep, a committedbenchmarks/boxes/<box>/record, a fullasmtest-bench-report/v1) plus each box’s append-onlyperf-history.jsonl, whose torn final line is counted rather than fatal. “Not measured” survives asstd::nulloptthrough both spellings the producers use (JSONnulland an omitted key) and renders as an em dash, never as0. The panel shows tier x backend x arch withskip_reasonrendered verbatim and truncation (trace_insns < insns_truth, orcomplete:false) made loud — both pinned by byte-compared golden renders.asmtrace_recordnow emitstraceevents, so the golden corpus feeds the trace canvas with real recorded data. It deliberately emits nocoverageevent: the L0 value producer measures executed steps, not basic blocks, and block starts cannot be recovered from an offset stream without instruction lengths..asmtraceexporters — recordings open in speedscope, Perfetto, genhtml and Graphviz (desktop GUI plan, Phase 1; docs/internal/archive/gui/02-exporters-and-readers.md). Newtools/asmtrace_export.c(make asmtrace-export) reads a recording and writes a speedscope evented profile (--speedscope), Chrome Trace Event JSON for Perfetto (--chrome), a block-offset lcov record (--lcov) or the--treecall graph as Graphviz DOT (--dot-tree) — so a capture made in CI or a container can be rendered later, anywhere, without re-running the traced program. One TU, libc only: no engine objects, no Capstone, no JSON library. Fidelity is enforced rather than documented: statisticalsurveyevents are never exported as stacks (exit 2, naming the reason), truncation and dropped samples surface in every mode, the time axis is labelled as the event ordinal it is — no producer records timestamps — and a mixed address basis, a newer format major or a compressed container is refused by name instead of best-efforted.make asmtrace-export-testbyte-compares every mode against committed expected files and pins each refusal by exit code and by the reason it prints.Desktop GUI skeleton — a Dear ImGui shell over the
.asmtracedocument model (desktop GUI plan, Phase 2; docs/internal/archive/gui/03-desktop-shell.md). Newdesktop/tree building two binaries:asmtest-desktop, the full app, which links the Author-tier engines and so is GPL-2.0 as a whole; andasmtest-viewer, a render-only viewer with zero engine dependencies that stays permissively distributable. Dear ImGui 1.91.9 and nlohmann/json 3.11.3 are fetched pinned + digest-verified the way the native engines are (scripts/fetch-imgui.sh,scripts/fetch-json.sh), andmk/desktop.mkaddsmake desktop/desktop-render/desktop-testplus thedocker-desktoplane. The.asmtraceloader groups events by kind, enforces the schema’s forward-compat and fidelity rules (a stream with no provenance is refused, a newer major is refused by name, and truncation / drops / redaction / a torn file each survive into the model), and the headlessdesktop-testdrives ImGui through its null backend — no display, no GL, no engines — so it runs on any host with a C++17 compiler and opens every committed golden recording..asmtracerecordings — one NDJSON format for every headless asmspy mode (desktop GUI plan, Phase 1).--record=<file>on--log --trace --dataflow --stream --graph --tree --procs --sample --watchwrites a.asmtracerecording beside the existing text/JSON output, and--jsonon--log/--streamstreams that same format to stdout — so `asmspy –log–json x.asmtrace
*is* a recording. Every stream carries mandatory provenance (which backend, exact vs statistical, the measured skip reason), and truncation, drops, throttling and redaction are fields rather than renderer discipline: a run that SKIPS still produces a closed recording naming the gate, and a producer killed mid-record leaves a visibly torn file. Syscall recordings split the line from its payload — the recorded line keeps the syscall name, fds, flag words, counts and return value while every decoded buffer, path, sockaddr and fd-backing path becomes a placeholder — so a reader can default-redact content without losing the call. The draft schema isdocs/internal/gui/asmtrace-schema.md;make asmtrace-goldenregenerates the committedtests/golden-asmtrace/corpus (deterministic emulator L0 recordings, plus hand-authored truncated/dropped/redacted/torn fixtures) andmake asmtrace-golden-checkgates it byte-for-byte.vec512_tjoins theasmtest_abi.json` manifest, closing the named-descriptor precondition for AVX-512 register state.Standing self-hosted Intel PT runner — unattended nightly coverage (self-hosted-ci-runners.md T5). New
scripts/runner-jit-loop.sh <owner/repo> <lane> <runner-dir>: the runbook’s production JIT/ephemeral loop as a script — mint a fresh JIT config, run exactly one job, re-register (60 s backoff on mint failure) — deployed on the bare-metal i7-8559U PT box as a systemd user unit with linger,HW_RUNNER_INTEL_PTleft at1so the nightlyhw.ymlschedule runs unattended. Re-registration proven by two consecutive green dispatches on freshly minted ephemeral runner identities (run 29999081537, run 29999251602). The runbook (docs/internal/ci/runners.md) gains the unit template, the deployment record, and the standing form’s recorded posture tradeoffs; the remainingHW_RUNNER_CORESIGHT/_MACOS_TART/_KVMvariables now exist at0per the settings checklist.Self-hosted Intel PT CI lane first green run (self-hosted-ci-runners.md T5).
hw.yml’shwtrace-pt-baremetalexecuted live for the first time (run 29997961188, 2026-07-23) on an ephemeral pinnedv2.335.1runner registered--labels intel-pton the bare-metal i7-8559U PT box: the broad hardware-capture tier underCAP_PERFMON(# 649 passed, 0 failed, PT tier live) plus the fail-not-skipmake docker-hwtrace-pt-live(ASMTEST_REQUIRE_PT=1,1..644,# 644 passed, 0 failed). One-shot ephemeral per the power-down rule (runner de-registered,HW_RUNNER_INTEL_PTback to0); unattended nightly coverage still needs a standing runner. Seedocs/internal/ci/runners.md.macOS clean-room Track D lane validated green — first shakedown complete (macos-cleanroom-lanes.md T6).
make docker-osx-bindingsnow runs a vanilla x86-64 macOS 13.7.8 (Ventura) guest under KVM on a bare-metal Linux host, SSHes in, and runs the Track-A clean-room install test: on 2026-07-23 it exited rc=0 (stable ×2) withclean-room-test: OKondarwin-x86_64— ruby PASS (the freshly-installed gem resolved its bundlednative/darwin-x86_64/libasmtest_emu.dylib, proving no dev-build//Homebrew//usr/localleak) and python/node/lua/java/dotnet/hdr SKIP (toolchain-free guest, the expected shape). The one-time Ventura install was driven headless over QEMU VNC (no X11), producing a reusable prebuilt disk (build/osx/mac_hdd_ng.img,DOCKER_OSX_DISK). Getting there hardened the lane script: QEMU-display none(gtk default is fatal in an X-less container),-iondocker run(a closed stdin EOFs the-monitor stdiomonitor and quits QEMU),DOCKER_OSX_CPU/DOCKER_OSX_SMP/DOCKER_OSX_VNC/DOCKER_OSX_CPUSETpassthroughs (the image’sPenryndefault spins newer macOS userlands — useHaswell-noTSX-IBRS), and a tree-copy exclude so the lane never tars the guest’s own disk into the guest. Hard host gate: the box’scurrent_clocksourcemust readtsc— an earlier 2026-07-22/23 attempt (Ryzen 9 4900HS) froze four times on a warped-TSC boot (clocksource demoted tohpet; macOS is TSC-only and livelocks under load there, at any vCPU count); the green run was on a healthy-TSC Ryzen 9 9950X/Zen 5 box. Two corrections vs the earlier runbook: enable Remote Login via System Settings → Sharing → Remote Login (Ventura refusessystemsetup -setremotelogin onwithout Full Disk Access), and each fresh-container run re-downloads the ~850 MB recovery before QEMU starts. Full evidence + runbook:docs/internal/docker-osx-linux-host.md.Bare-metal Intel PT self-hosted lane + a fail-not-skip docker target, and the dark CoreSight placeholder (self-hosted-ci-runners.md T5).
hw.ymlgainshwtrace-pt-baremetal(runs-on: [self-hosted, linux, x64, intel-pt]) andhwtrace-coresight-board, both carrying the workflow’s required guard pair — anHW_RUNNER_*variable that is0/absent by default and thegithub.actor == github.repository_owneractor guard — so they land and stay green with zero runners. The PT job’s non-vacuity check is not a skip-string grep (the PT skip text has a known cosmetic misreport, so a grep would assert on a string that lies): newmake docker-hwtrace-pt-liveruns the require-mode targethwtrace-pt-live(ASMTEST_REQUIRE_PT=1) in the hwtrace image under--cap-add=PERFMONwith default seccomp, turning the PT tier’s availability self-skip into a hard failure — verified on a non-PT host to fail for the right reason (661 passed, 1 failed, non-zero exit,not a GenuineIntel x86-64 host). It also prints/sys/bus/event_source/devices/intel_pt/typefrom inside the container each run: that shakedown answer was already measured on the bare-metal i7-8559U box (PMU visible in-container;CAP_PERFMONbypassesperf_event_paranoid=4, so no--privilegedand no host-native fallback) and is now recorded in the runner runbook. The CoreSight job is deliberately dark — flippingHW_RUNNER_CORESIGHTis the acceptance step of the OpenCSD decode work, not something to do early.asmspy names why a hardware watchpoint/breakpoint could not arm, and the AArch64
cliCI leg is now gating (asmspy-aarch64-support.md T7). “Hardware watchpoint unavailable” was three different host facts behind one guessed message (“qemu / seccomp / permission”). Measured on the hostedubuntu-24.04-armrunner:NT_ARM_HW_BREAKreports 6 slots andNT_ARM_HW_WATCH4 (debug_arch=8), andPTRACE_SETREGSETon either returnsENOSPC— slots exposed, reservation refused, so nothing can arm and nothing can fire.asmspy_hwdebug_reason()now records what actually happened at the arm site, so--watchskips with “host reports 4 watchpoint slots but refused to reserve one: No space left on device” and--trace --tidadds the same note. The smoke’s--trace --tid=block takes the host-capability split the--watchblock already had (named skip on AArch64, strict on x86-64) and still asserts the property that survives it: an unarmable entry is reported as an attach failure, never as a false “never executed”. A--watchsuccess line that printed unconditionally — including on the skip path — now says which happened. With those the arm64clileg is green end to end andcontinue-on-erroris removed.Fixed:
asmspy --stream/--graph/--treekilled every multi-threaded AArch64 target they traced (asmspy-aarch64-support.md T2). On AArch64 a single step armed on a thread parked in a blocking syscall survivesPTRACE_DETACH— arm64’sptrace_disable()setsSPSR.SSand clears onlyTIF_SINGLESTEP— and fires as a fatalSIGTRAPwhen that syscall returns, 200-400 ms after asmspy has exited. Whole-process tracing therefore left the target dying: measured on Neoverse-N2, a second trace of the same process found every thread already dead withtermsig=5. The whole-process engines now resume a thread poised on ansvcwithPTRACE_SYSCALLinstead of a step (step_resume), so nothing is armed across a call that may block and the thread is still stepped again from the syscall-exit stop;PTRACE_O_TRACESYSGOODtags those stops on AArch64 only, leaving the x86 stop stream unchanged. The teardown drain gained the matching guard (/proc/<tid>/syscall, since a thread at a syscall-entry stop has its pc past thesvc— stepping it hung asmspy unkillably) and now keeps stepping until it consumes a real trap rather than a queuedPTRACE_EVENT_STOP. The detach-survival smoke assertion, which checkedkill -0immediately and so was blind to a delayed kill, now polls for ~2 s.The macOS drgate gate now runs in CI — CPython signal chaining under attach is regression-protected on an independent host (macos-dynamorio-signal-chaining.md ✅ closure). The nightly
drtrace-macosjob (macos-15-intel) had run only the C harnesses (drtrace-test-macos+test_drtrace), so the very case whose wedge motivated the fork’s macOS sigreturn fix —test_drgate.py::test_signal_chaining, a signal delivered to a live CPython interpreter while DynamoRIO is attached — was validated only on the dev box that landed the fix. The job now installs pytest and runsmake drtrace-python-testagainst the freshly built pinned fork (tests/test_drtrace.py3 cases +tests/test_drgate.py4 cases, including takeover scope, signal chaining, tracing under the managed host, and start/stop bracketing), with the lane’s standing self-skip-is-failure posture: a DR-not-found skip, a pytest skip, or a short collection fails the job rather than silently retiring the managed-host gate. gate (aarch64-sve-capture.md T8).** The SVE execution sign-off had been recorded as hardware-gated (“no SVE host in this environment”); the hostedubuntu-24.04-armrunner is Azure Cobalt 100 / Neoverse-N2, which carriessveandsve2in HWCAP atsve_default_vector_length= 16 bytes onLinux 6.17.0-1020-azure— soasm_call_capture_sveand itsptrue/faddcorpus body had been executing on real silicon in every CI run, not only under qemu-user TCG (ok 5 - simd.sve_adds_doubles_at_any_vl,0 skipped;make check57 passed / 0 failed). Thetestjob’s arm64 leg gains an SVE silicon sign-off step that prints the silicon facts (kernel,CPU part, VL) and fails if that test ever self-skips there, so a runner-fleet change or a regressedHWCAP_SVEprobe cannot silently retire the only real-silicon validation the SVE path has. Still gated, and reported as such: a native VL other than 16 B (Graviton3 at 32 B, A64FX at 64 B) — the sweep lane’s 48/128/256 B legs remain qemu-TCG emulation.Self-hosted AMD Zen CI lane ran green live on real silicon (self-hosted-ci-runners.md T3). The
hwtrace-privileged-zenjob in.github/workflows/hw.ymlrunsmake docker-hwtrace-privilegedon a registered AMD Zen 4/5 runner and asserts the exact LbrExtV2 branch-stack + live-IBS paths RAN (a self-skip there is a hard failure). It was exercised end to end on the Ryzen 9 9950X (Zen 5) box on 2026-07-22: a pinnedv2.335.1ephemeral runner (tarball SHA-256 verified against GitHub’s published digest) picked up aworkflow_dispatchand went green —# 667 passed, 0 failed, zero AMD-LBR/IBS self-skips,call_autoescalating off the LBR window (insns=77 truncated=0), assert step passing.hw.ymlwas additionally hardened with an explicitgithub.actor == github.repository_owneractor guard on the self-hosted job (on top of the existing no-push/no-pull_requestposture), so the lane runs only for a trigger the repo owner initiated. The hostedhwtrace-privilegedbitrot-gate comment inci.ymlwas corrected: thecall_autonon-escalation finding is FIXED (5d8e0d2) rather than open, and the self-hosted counterpart is no longer “future” — it now exists inhw.yml. See the runbookdocs/internal/ci/runners.md(standing unattended-nightly coverage needs a persistent runner via the JIT/ephemeral loop).macOS x86-64 DynamoRIO native-trace tier — M0 is GO (macos-dynamorio-fork-build.md FB1–FB3 driving macos-dynamorio-port.md T3–T5). DynamoRIO publishes no macOS release asset (0 across all 455 releases), so the macOS tier now builds the runtime from a git-commit-pinned source fork:
scripts/build-dynamorio-macos.sh+make dynamorio-macosproducelib64/release/libdynamorio.dylib(plus thedrmgr/drreg/drxextension set the drclient sub-build links), pinned inscripts/third-party-digests.txtwith the license vendored, Darwin-x86-64-gated (clean skip everywhere else). Three root-caused fixes in the fork (wilvk/dynamorio, branchasmtest/macos-fixes) make thedr_app_*embedding functional on modern macOS: thedr_app_setupstartup fault (exec path read fromKERN_PROCARGS2instead of the above-envp walk; baseline was 10/10 SIGSEGV), the invisible environment in dylib embeddings (dyld passes no envp to initializers, soDYNAMORIO_OPTIONSwas never read and no client ever loaded; now captured live via_NSGetEnviron, the STATIC_LIBRARY approach), anddr_get_proc_addressreturning NULL on every modern binary (LC_DYLD_EXPORTS_TRIEunparsed, plus a trie rebase that was only correct for preferred-base-0 dylibs — never the main executable). On the asm-test side,libdynamorioresolution is dylib-aware (DR_LIBNAME), the DR make tier resolves.dylibclient names on Darwin, and the newmake drtrace-test-macosM0 harness traces a normally-compiled__TEXTfunction end to end — attach, Mach-O marker resolution, coverage accumulation, symbol mode, clean detach — 13/13 twice in a row on the macOS-14.7.5/Intel host, against the pinned script-built home. M1a (macos-dynamorio-port.md T6): the same target now also runs the generated-bytestest_drtraceharness on Darwin — thePROT_NONE → RW → RXexec_allocW^X path, coverage accumulation, the exact instruction-mode offset stream, and the truncation bit all pass on native Intel silicon (18/18 on the same host; Rosetta stays must-verify, gated on Apple Silicon hardware). M1b groundwork (T7/T8): executable-memory allocation now sits behind a platform seam whose arm64-macOS arm usesMAP_JIT+ per-threadpthread_jit_write_protect_np(byte-identical elsewhere;asmtest_asm_exec_nativereturns ENOSYS on non-x86-64 hosts), and on Darwin thetest_drtraceharness is ad-hoc signed with the newdrtrace.entitlements(com.apple.security.cs.allow-jit; never--options runtime, whose one-MAP_JIT-region limit the harness’s three allocations would trip). arm64 acceptance itself stays gated on Apple Silicon hardware plus an arm64 DR runtime (upstream i#5383); the guide’s new “macOS arm64” section records the hardened-interpreter limitation as a property of the OS model. M2 bindings (T9): the binding lanes resolve platform-correct library names on Darwin (drtrace_env→.dylib+DYLD_LIBRARY_PATH), withdrtrace-cpp-test,drtrace-ruby-test, anddrtrace-python-testgreen on macOS x86-64 — compiled-function/symbol mode is the documented macOS binding path. Signal chaining under an attached trace now works on macOS (a fourth fork fix): a signal raised while attached is delivered to the host runtime’s handler and the process resumes cleanly, so the Pythontest_drgate.py::test_signal_chaininggate runs on Darwin (previously Darwin-skipped). Root cause (confirmed single-threaded, not the multi-thread i#58 first suspected): the app handler is delivered fine, but its return through libsystem_sigtramp→ macOS 3-argsigreturn(uctx, infostyle, token)needs a per-delivery kernel token DR cannot forge for the frames it synthesizes, so the realsigreturnwas rejected and the thread resumed into DR gencode →ud2→ SIGILL → terminate. The fork (pinb8785a5d8) now restores the app context inhandle_sigreturnand skips the realsigreturnon macOS x86-64, exactly as the VMX86 path does. Forkapi.startstop/api.detachrun 10/10 crash-free (their multi-thread takeover assertions are upstream-NYI on macOS, i#58, and upstream macOS CI never runs them — they are outside theOSXctest label set). arm64 stays gated on the upstream arm64 port (i#5383). A nightly/dispatch-onlydrtrace-macosCI job onmacos-15-intelbuilds the pinned fork from source (commit-stamp cached) and runs the M0 harness, outside thetestmatrix so a DR failure never blocks the emulator tier — with a self-skip-is-failure guard, since the runner that just built the runtime must actually test it.AArch64 out-of-process single-step stream validated live on real silicon (aarch64-ptrace-single-step-validation.md T1–T6). The out-of-process
ptracetracer’s AArch64 arm — written and decode/execute-validated under qemu, but whose live capture had never run — now runs LIVE on GitHub’subuntu-24.04-armhosted runners (Azure Cobalt 100 / Neoverse-N2 VMs) via a gatinghwtrace-arm64CI job: 205okassertions, the whole ptrace tier (trace_call/trace_attached/run_tosoftware-brkplant / call-out step-over) exercised on real hardware, with an anti-vacuity step that fails the build if the tier self-skips. A companionhwtrace-bindings-arm64job runs the Go and Java wrappers’ AArch64 ptrace fixtures live (the other wrappers gate their whole hwtrace suite on the x86-only in-process single-step backend and self-skip), plus a live Python host step. The first arm64benchmarks/boxes/record (arm-linux-arm64-gha) is committed with a livenative-oopcapability row. Measured answer to the previously-undetermined hardware-breakpoint question:NT_ARM_HW_BREAKis armed-but-silent on this hypervisor (slots reported, arm accepted, the debug exception withheld), so the forced-hardwarerun_toself-skips there with that named reason and bare-metal arm64 hw-breakpoint firing moves to self-hosted-runner territory. qemu-user keeps self-skipping transparently. Also brought the asmspy CLI toward AArch64: fixed theSYS_stat/dup2legacy-syscall gaps (arm64 usesnewfstatat/dup3— the latter now decoded), the host-arch disassembly of traced code, andR_AARCH64_JUMP_SLOT/ 2-slot-PLT0 stub resolution (asmspy-aarch64-support.md T2/T7).F5 live foreign-pid PT replay wired to the landed capture (dataflow-pt-replay-tier.md T4). The out-of-band data-flow value tier’s live case now CONSUMES the
asmtest_hwtrace_pt_attach_*foreign-pid capture that landed with intel-pt-attach-foreign-pid:examples/test_dataflow_pt.cforks a deterministic victim, captures ONE in-region invocation over Intel PT with zero single-steps of the target, replays the decoded offset path through F5, and asserts the value trace matches both the emulator L0 and theforce_singlestepblock-step oracle’s executed path + result. It was previously a dead-code stub behind a never-defined macro; it is now a runtime-probed body (asmtest_hwtrace_available(ASMTEST_HWTRACE_INTEL_PT)) that self-skips off Intel PT and, underASMTEST_REQUIRE_PT=1(make dataflow-pt-live), fails rather than skips. Gated on bare-metal Intel PT silicon for the live oracle match;make docker-dataflow-ptruns the synthetic decode→replay bridge + both block-step-layout guards green (19/19) with the live half self-skipping.Java binding publishable to Maven Central (distribution-packaging.md T6).
make java-packagenow runs a realmvn packageagainstbindings/java/pom.xml(the same POM Central publishes) instead of rawjavac+jar cf, emitting the binding jar plus matching-sources/-javadocjars. A new tag-gated, secret-guardedmavenjob inrelease.ymlmvn deploys them — GPG-signed — to the Central Portal staging, no-opping unless bothMAVEN_CENTRAL_TOKENandMAVEN_GPG_KEYare set (mirroring the crates/npm publishes). Maven is a version-pinned Apache tarball in the java docker image;make docker-java-packageproves the whole build locally with no credentials, anddocs/reference/releasing.mdcarries the Central + LuaRocks runbooks.Cross-system benchmark legs on real Windows and Intel macOS. The deterministic golden gate now runs on every OS the framework targets: a per-push
benchmarks-windowsleg (windows-latest, mingw/MSYS2) builds the PE benchmark producers and runsmake win64-bench-check+win64-bench-reporton a genuine Windows kernel, and a nightlybenchmarks-macos-x86leg (macos-15-intel) produces the Intel-macOS report.benchmarks-comparemerges all five OS × arch reports.Nightly auto-commit of per-box benchmark records. On the nightly schedule (and manual dispatch) each benchmark leg records its per-box history and a new
benchmarks-recordjob commitsbenchmarks/boxes/gh-**back tomainasgithub-actions[bot](aGITHUB_TOKENpush, so it never re-triggers CI). Golden emu counts stay human-reviewed — never auto-committed — so a real count drift fails a leg’sbench-checkinstead of being laundered into history. Newmake win64-bench-recordpersists the Windows box record.Ambient stitched operations (.NET, opt-in, Intel PT).
AsmAmbientStitchedTracefollows one logical operation acrossawait/ thread hops with zero calls in the body — anAsyncLocalvalue-changed handler opens a per-threadintel_ptslice when the flow lands on a thread and decodes-at-disable when it leaves, then stitches the slices at close (op.Hops/op.Path). PT-only by construction; off Intel PT it self-skips and runs the body uninstrumented. Built on the new per-tid PT hop capture primitive (asmtest_hwtrace_pt_hop_open/_hop_close) on the one sharedintel_ptperf-AUX arm. The default whole-window capture is unchanged (it stays the faithful per-thread window that flagstruncatedon a cross-thread hop — never auto-stitch).Runtime-enabled jitdump byte recovery on an already-running process — turn jitdump emission on in a live foreign runtime and recover a JIT method’s recorded bytes with
asmtest_jitdump_find, with no launch flag on the target. For CoreCLR (make docker-hwtrace-jit-dotnet-attach-jitdump) the harness sendsEnablePerfMap(All)over the runtime’s diagnostics IPC socket (the documentedDOTNET_IPC_V1wire, hand-rolled in C — no NuGet), so an already-JITtedProgram::Addis rundown-emitted into/tmp/jit-<pid>.dumpeven though the victim was launched withoutDOTNET_PerfMapEnabled. For HotSpot (make docker-hwtrace-jit-java-attach-jitdump) it loads a new in-tree attach-capable JVMTI jitdump agent (examples/jvmti_jitdump_agent.c, test-support only, never shipped) into the running JVM withjcmd JVMTI.agent_load; on attach the agent replays every already-compiled method viaGenerateEvents(COMPILED_METHOD_LOAD)into the dump. Correction: the linux-toolslibperf-jvmti.soexports onlyAgent_OnLoad(noAgent_OnAttach), so HotSpot refuses to load it viajcmd— it is-agentpath-only and cannot serve the attach case; hence the bespoke agent. V8/Node has no runtime-enable path (--perf-profis wired once at isolate init), so there is deliberately no Node attach lane. Both lanes run on any host with the runtime — no Intel PT / hardware gate.Opt-in safe-managed whole-window policy.
ASMTEST_WHOLEWINDOW_SAFE_MANAGED=1makesasmtest_hwtrace_begin_windowrefuse an in-processEFLAGS.TFwhole-window arm when a managed runtime (CoreCLR / JVM / Mono) lives in the process, returning the new distinctASMTEST_HW_EMANAGEDstatus instead of single-stepping code whoseSIGTRAPdisposition the runtime’s signal layer owns. The .NET empty-ctornew AsmTrace()builds on it to route a managed window to Intel PT where the silicon exists, else the §D3 out-of-process stepper, else a transparent self-skip — never in-process TF (ww.Routereports the chosen route). Default (env unset / nosafeManaged) is byte-identical to today;asmtest_hwtrace_managed_runtime_present()exposes the probe.Native trace-point → IL / bytecode / source-line attribution. A captured native offset of a managed method now resolves to a source line, a .NET IL offset, or a JVM bytecode index — not just a method name — from feeds the runtimes already emit.
asmtest_jitdump_debug_find(include/asmtest_ptrace.h) recovers the per-address(line, column, file)table from a jitdump’sJIT_CODE_DEBUG_INFOrecords the byte reader skips, andasmtest_jitdump_debug_line_mapbridges it into the shippedemu_line_map_t(works today for V8node --perf-prof). A widened, backend-neutral schema (asmtest_srcmap_*,asmtest_srcreg_*,include/asmtest_trace.h) carriesoffset → {kind, value, file, column}with enclosing-point lookup and a version-keyed registry stamped on the code-image capture sequence, so a method re-JIT’d at a reused address resolves against the body live when the trace ran. For CoreCLR, a newIlToNativeMapEventListener (bindings/dotnet/hwtrace/HwTrace.cs) subscribes theMethodILToNativeMapJIT keyword (0x20000) in-process — no launch knob, no Intel PT — and resolves an address to(method, nativeOffset, ilOffset). For HotSpot, a newmake docker-hwtrace-jit-java-bcilane loads an in-tree JVMTI agent (examples/jvmti_bci_agent.c, test-support only, never shipped) that capturesCompiledMethodLoad’s address→bytecode-index map into anasmtest_srcregand proves a native address resolves to a real bytecode index. The V8/HotSpot/CoreCLR jitdump debug lanes (make docker-hwtrace-jit-jitdump,-jit-java-jitdump,-jit-dotnet-jitdump) print the per-method attribution table and CHECK the reader against each real encoder. Attribution covers JIT-compiled code only: interpreted code keeps its bytecode index in VM state the native PC stream cannot see — see il-bytecode-attribution.md.libdft64 differential taint oracle (
make docker-taint-oracle). The shipped DynamoRIO in-band taint client had only an offline emulator/Capstone forward-slice as an independent check. This lane cross-validates it against a second, live, independently-implemented byte-level taint engine — libdft64 on Intel Pin — by running the same seed/sink fixtures (examples/taint_fixtures.h) through both and asserting byte-for-byte sink agreement on the general-purpose / integer-memory subset both cover (branch-condition, call-argument, and mem-copy-length sinks), with passing negative controls. A pinned, digest-gated Pin 3.20 kit (libdft64’s only tested pin) + git-commit-pinned libdft64 are fetched at build time and are test/oracle-only — never linked intolibasmtestor any shipped binding. libdft64’s documented blind spots (SIMD isbasic SSE/AVX, rules unverified; no eflags; no ZMM; no implicit flow; no x87/ternary) are enumerated as named skips, never a blanket pass — see data-flow-capture.md (even libdft punts on SIMD). The lane is CI-gated (thetaint-oraclejob) and self-skips only on non-x86 hosts (Pin is x86-64 gcc-linux).
Intel SDE future/absent-ISA test lane (
make docker-sde/make sde-test SDE_HOME=$(scripts/fetch-sde.sh)). Assembly that uses an ISA extension the host CPU lacks — APX’sr16-r31, AVX10.2, AMX, or AVX-512 on an AVX2-only box — was untestable by this framework (the DynamoRIO tier runs on real silicon; the Unicorn tier’s vendored QEMU 5.0.1 predates AVX TCG). A pinned, digest-gated Intel SDE 10.8.0 (scripts/fetch-sde.sh+Dockerfile.sde, with APX-capable GAS from binutils 2.46.1 and NASM 3.02, all SHA-256-pinned) emulates those extensions for the whole process, so an unmodified suite binary runs undersde64 -futureand gets the full register/flag/memory/ABI assertion battery on any x86-64 host, including CI runners. The lane proves SDE is byte-for-byte transparent to correct baseline code (native vs SDE TAP identical), adds an APX fixture suite (examples/apx_basic.s+test_apx_basic, gated on a newasmtest_cpu_has_apx()CPUID probe) that skips on real pre-APX silicon and runs green under emulation, asserts the AVX-512-on-AVX2 un-skip (an existingtest_simdcapability skip becomes a real execution under-future), cross-checks SDE against the Unicorn tier on overlapping baseline ISA, and offers an optional-mixinstruction-mix report (make sde-mix) folded into the canonicalasmtest_trace_tshape. SDE is proprietary freeware, fetched + digest-verified at build/test time and never bundled into a shipped artifact — test-lane only. Documented in the SDE testing guide and the implementation doc.Live whole-window compose lane (.NET). The
hwtrace-dotnetself-suite gains checks proving the zero-config compose seam (MethodLoadVerbose→codeimage_track→ close-time versioned decode): over the in-process WEAK single-step tier (new AsmTrace()) against a managed method whose first JIT happens inside the window (genuinely-compiled-in-window, not a pre-warmed body); over the crash-proof §D3 out-of-process inlineusing-scope (new AsmTrace(outOfProcess: true)) against a resident method — the first suite coverage of that bare inline OOP whole-window ctor; and (self-skipping off Intel-PT silicon) the STRONG PT tier — plus a mid-window re-tier decode-at-version check. Runs as the namedmake docker-hwtrace-dotnet-unwarmedlane. Tasks T1–T5 of the managed-wholewindow-compose implementation doc. Two re-verification findings refine the doc’s original premise: (1) forcing a stop-the-worldGC.Collect(0)on the single-stepped thread inside the window is intermittently fatal — the in-process EFLAGS.TF window dies (SIGTRAP) if the runtime spawns a thread in-window, per the repo’s own degradation note — so it is omitted (the unwarmed JIT itself supplies the live-window instruction noise); (2) the inline OOP ctor’s region-freewindow_stopstepper single-steps everything, so a first-call JIT inside that window steps the whole compiler and aborts CoreCLR (exit 134) — the unwarmed mid-window-JIT compose is therefore proven on the range-basedAsmTrace.Windowfactory (whose stepper runs the JIT at native speed), while the inline OOP ctor is for resident (warm) code, matching thecrashproof-showdownexample.Data-flow F5: an out-of-band PT + code-image + Unicorn-replay value producer (
make dataflow-pt-test/make docker-dataflow-pt). The least-perturbing L0 value tier: it reconstructs an Intel PT trace’s executed instruction stream (captured with zero single-steps), supplies the bytes live at trace time from the code-image recorder, and replays that exact path through Unicorn to derive per-instruction values into the sameasmtest_valtrace_tthe shared def-use (L1) and slice (L2) analysis consume — byte-identical to the emulator L0 oracle on a deterministic region.src/dataflow_pt.copens no perf event (it consumes a captured AUX blob + code-image); it reuses the block-step tier’s purity/replayability verdicts and truncates faithfully on an impure, VEX/EVEX, or nondeterministic region (no single-step fallback), a per-step path cross-check catching a divergence. The synthetic-AUX decode→rebase→materialize→replay bridge is validated in CI with no PT hardware (libipt’s own encoder,libipt-devadded toDockerfile.dataflow-attach); live foreign-pid capture is silicon-gated (bare-metal Intel PT + theintel-pt-attach-foreign-pidcapture arm) — wiring-complete, hardware-unvalidated, with amake dataflow-pt-livefail-not-skip target for a runner that claims PT. Documented in the native-tracing guide and the F5 implementation doc.Whole-window in-process guards for the zero-config scope (deny regions / instruction budget / wall-clock watchdog). The
using (new AsmTrace())region-free whole-window scope, powered by the in-processEFLAGS.TFsingle-step tier, took the descent tier’s “step into everything” semantics without any of its safety guards: a window over code that reached a blocking libc call (areadon an empty pipe, apoll) stepped the runtime forever, bounded only by the capture ring’s memory, never by time. The three out-of-process descent guards are now ported onto that in-process path — a per-frame instruction budget (default 4x the ring cap), a process-global deny-region table with an opt-in blocking-libc default set (a stepped RIP inside a denied region ends the capture and the denied call then runs at native speed), and a repeatingITIMER_REAL/SIGALRMwatchdog (default 10 s; breaks a blocked syscall viaEINTR) — all malloc/lock-free inside the SIGTRAP handler. Plainasmtest_hwtrace_begin_windowgets safe defaults (budget + watchdog on, denylist off), so a hung zero-config window is always bounded; the new additiveasmtest_hwtrace_begin_window_ex(with the F27/F36struct_sizeidiom and anasmtest_hwtrace_window_guards_tconfig) configures them per window, andasmtest_hwtrace_window_guardreports which guard fired (ASMTEST_HW_GUARD_*) render-on-close style. C-level knobs (parity-exempted); the zero-config defaults flow through thebegin_windowevery binding already wraps.Zero-config whole-window scope docs. The hardware-tracing guide gains a “zero-config whole-window scope (region-free)” section — the C
begin_window/end_window/render_windowsurface, the .NETusing (new AsmTrace())form, the WEAK / STRONG / CEILING tier ladder (single-step / Intel PT / AMD LBR Zen 4+), and how the STRONG-tier PT decode is validated on a synthetic PT-packet fixture with no silicon. The troubleshooting reference gains aSkipReasontable and a one-time-provisioning table (perf_event_paranoid/setcap cap_perfmon/--cap-add, ptraceCAP_SYS_PTRACE, eBPFCAP_BPF), and both it and portability record the whole-window facility’s Linux-only floor.asmspy runs on AArch64 Linux. The out-of-process tracer’s register / single-step / detach reads are lifted behind an architecture shim (
cli/asmspy_arch.h: PC / return / SP / LR / syscall-number accessors overPTRACE_GETREGSon x86-64 andPTRACE_GETREGSET(NT_PRSTATUS)on AArch64), so every engine —--stream,--graph,--tree,--region,--log,--procs— runs on both arches. The single-step teardown honours AArch64’s kernel-owned step model (no user trap flag;svc #0syscall-instruction guard), the call-graph / call-tree frame logic uses AArch64bl-writes-LR semantics (frame identity(entry_lr, sp)), and--logdecodes the AArch64 syscall ABI (number inx8, args inx0-x5,*at-only name table).--watchgains an AArch64 arm over theNT_ARM_HW_WATCHregset (DBGWCR/DBGWVR/BASencoding, pinned by a purecli/test_arch.cunit test on every host); it self-skips where the host exposes no watchpoint slots (qemu-user, some hypervisors). Validated on the nativeubuntu-24.04-armCI runner (a real VM, not qemu) alongside the existing x86-64 leg; see the asmspy guide.libFuzzer / AFL++ external-engine fuzzing shim (
make docker-fuzz). Drive an x86-64 guest routine under the emulator with an industrial fuzzer, feeding the emulator’s basic-block coverage into the engine’s feedback channel without compiler-instrumenting the guest bytes (they run under Unicorn). A new tested seamemu_cover_hits(src/fuzz.c) reports one input’s distinct executed block offsets;examples/fuzz_libfuzzer.cregisters them as SanitizerCoverage 8-bit counters (+ the PC table clang-18’s libFuzzer requires), andexamples/fuzz_afl.c(native persistent-mode forkserver) plus an aflpp_driver reuse of the libFuzzer harness write them into AFL++’s shared-memory bitmap via a plain-compiled helper (examples/fuzz_afl_map.c).Dockerfile.fuzz(clang 18 + afl++ 4.09c on the bindings base) runsmake fuzz-shim-test, which fails unless both engines steer to a planted crash — a real test, never a self-skip. Node (per-block) coverage; documented in the fuzzing-shim guide.src/capture.s;retstubs elsewhere, incl. the NASM twin) marshals scalable-vectorz0..z7and predicatep0..p3arguments per AAPCS64 and captures the wholez0..z31/p0..p15file into two new max-size containers —svec_t(256-byte VLmax) andspred_t(32-byte PLmax), of which only the lowasmtest_sve_vl()(resp./8) live bytes are written.ASM_SVCALL_1/_2call and self-skip without SVE, andASSERT_SVEC_EQ/ASSERT_SPRED_EQcompare exactly the live VL. AHWCAP_SVEruntime probe (asmtest_cpu_has_sve/asmtest_sve_vl) returns 0 everywhere SVE is absent (x86-64, and macOS arm64 — Apple silicon has no non-streaming SVE); the ABI manifest pins both containers, and the corpus routinesve_adddexercises the path. The newmake docker-sve-sweeplane runs the SIMD suite under qemu-user at several vector lengths (VQ 1/3/8/16 → VL 16/48/128/256 bytes, including a non-power-of-two) to flush out VL-assumption bugs before any SVE silicon is available; execution sign-off on real SVE hardware (Graviton3/Grace/A64FX-class) remains pending — see aarch64-sve-capture.md.XED-decoded Intel Pin trace lane (
make docker-pintool/make pintool-test). A pinned, digest-gated Intel Pin 4.2 kit (scripts/fetch-pin.sh, SHA-256 pinned inscripts/third-party-digests.txt) drives a Pintool (pintool/asmtest_pintool.cpp) that fills the sharedasmtest_trace_toffset model over POSIX shared memory. The lane asserts byte-for-byte instruction/block offset parity with both the in-process single-step backend and the DynamoRIO backend (Pin ≡ DynamoRIO ≡ single-step), and carries an Intel APX (EGPR/REX2) fixture whose bytes Pin’s XED decodes on any x86-64 host while the pinned DynamoRIO decoder rejects them — the decoder-currency gap the tier exists to close (DR #6226 is open; the APX execution halves are gated on APX silicon). Test-lane only: Pin is digest-verified at build/test time and never bundled into a shipped package.Intel Pin probe-mode argument/return capture lane (
make pin-probe-test, inmake docker-pintool). A Pin probe-mode tool (pintool/probe_capture.cpp) splices a jump at a named routine’s entry/exit and records the SysV integer/FP argument registers, the return register(s) + flags, and up to a 4 KiB cap of a pointed-to buffer — at native speed (no code cache) — intoat_val_rec_trecords over a POSIX shm channel (include/asmtest_valtrace_shm.h). A pointer is validated against the target’s mapped ranges and the read clamped to the cap and the mapping end, so an invalid pointer is refused, never faulted; a routine too short or non-relocatable to probe is reported as an explicit per-target skip with a reason. The capture is proven by diffing it against the independent out-of-process ptrace stepper on the same routine (two producers agree on the arg + return registers). Captured buffers may contain secrets — a sensitive artifact. x86-64 Linux, test/oracle only (Pin is proprietary freeware, digest-verified at test time and never bundled), the same handling DynamoRIO gets; documented in the data-flow tracing guide.Real object identity for managed memory def-use on the live-attach tier. A heap snapshot of
{Address, Size, TypeID}nodes from the runtime’s GCBulkNode / GCBulkEdge / GCBulkType events, joined with theMovedReferences2move feed, keys each captured memory record on (object, offset) where the snapshot has evidence and degrades to the landed address identity where it does not (asmtest_objid_canonicalize,src/dataflow_objid.c; unit suitetest_dataflow_objid). On a live attach (make docker-gccanon-attach, newaliasphase) the false def-use edge that address identity forges when a GC slides a live object onto a dead object’s vacated slot is reproduced under address identity and then eliminated by object identity.Weighted cross-ISA cost proxy
BM_MODEL_COSTin the cross-system benchmark.emu-benchnow emits amodel_costrow beside each deterministicinsnsrow: each executed instruction is classified with Capstone (newasmtest_disas_class→ OTHER/MEM/BRANCH/MULDIV) and summed against a fixed weight table (1/3/2/8), a faithful cross-architecture cost model — comparable by construction, not silicon cycles.bench-comparerenders it as its own Model cost matrix (never mixed with raw counts or real cycles). The metric needs Capstone: without it the bench emits theinsnsrows alone and every gate still passes. Model values depend on the Capstone version, so they are kept out of the golden file —bench-golden-checkfiltersmodel_costrows (and fails loudly if anyinsnsrow is missing its model sibling). See cross-system benchmarking.Native RISC-V (rv64) host tier — the capture framework now runs on a RISC-V machine, not just as an emulator guest. A
regs_tbranch and trampolines (src/capture.s) for the RV64GC / LP64D psABI:a0/a1return pair, integer callee-saveds0–s11, and FP callee-savedfs0–fs11(checked byASSERT_ABI_PRESERVED/ASSERT_ABI_PRESERVED_VECafter an_fp/_fp_ncapture, since rv64gc has no vector file), plus rv64 bodies for every example suite and the framework self-tests. Two ISA facts are surfaced faithfully rather than faked: RISC-V has no condition-flags register, soASMTEST_NO_FLAGSis set andASSERT_FLAG_*is a compile error on rv64 (flag-only suites self-skip with a printed reason); and there is no 128-bit vector capture (ASM_VCALL*self-skips viaasmtest_cpu_has_vec128()— RVV is a possible future arm). Amake docker-riscv64lane builds alinux/riscv64image and runs the core suites + self-tests under QEMU binfmt (make binfmt-riscv64, pinnedtonistiigi/binfmt), wired into CI as thetest-riscv64job. The tracing tiers stay x86-64/AArch64. See riscv-native-tier.md.Live Intel PT whole-window smoke + a
make hwtrace-pt-livelane that FAILS rather than skips where PT is claimed.test_pt_live_selfjit(examples/test_hwtrace.c) self-JITs the canonical routine, arms the region-free PT capture, decodes the REAL AUX stream throughasmtest_pt_decode_window, and exercisesPERF_EVENT_IOC_SET_FILTER(via the newasmtest_hwtrace_pt_set_filterknob), the anonymous-JIT decode-time fallback, and AUX-ring truncation on a 4 KiB ring. Off theintel_ptPMU (AMD/VMs/ containers) it self-skips with the specific reason — one of the two legitimate hardware gates — whilemake hwtrace-pt-livesetsASMTEST_REQUIRE_PT=1to convert that skip into a build failure on a runner that is supposed to expose bare-metal Intel PT. The live capture is silicon-gated (nointel_pton the reachable dev boxes); inmake docker-hwtracethe test prints a clean# SKIP pt live: …everywhere. See intel-pt-whole-window-substrate.md.System-package specs for the C core — Homebrew, Debian, AUR, vcpkg and Conan — each built, installed and consumed in a Docker CI lane.
packaging/now holds a Homebrew formula, a Debianlibasmtest-devsource package, an AURPKGBUILD+.SRCINFO, a vcpkg overlay port and a Conan 2 recipe for the MIT static core (libheaders +
asmtest.pc; the GPL engines stay in the dlopen binding packages only, never here).make docker-syspkgruns all five lanes and an additivesyspkgCI job runs them as a matrix — each builds the package, runs its native linter (brew audit/style,lintian,namcap, vcpkg post-build validation), installs it, and compiles a pkg-config/CMake consumer against it. The lanes are hermetic on the reproduciblemake package-sourcetarball; per-manager index publication is a maintainer step (runbook in releasing.md). See distribution-packaging.md.
.NETinlineusing (new AsmTrace(HwBackend.IntelPt))now arms the STRONG whole-window Intel PT capture (bindings/dotnet/hwtrace/HwTrace.cs), replacing the reserved"forward-look (not wired)"self-skip. The backend-keyed ctor gates onHwTrace.Available(IntelPt)(self-skip names the PT gate off bare-metal Intel PT), sets up theJitMethodMap+ perf-map rundown before arming, and drives the nativeasmtest_hwtrace_pt_begin_window/_end_windowpair via a finalizablePtWindowCtx(a leaked scope’s fd + AUX mappings are reclaimed drain-less by the finalizer) and a newKind.PtWindowDispose that decodes on close against the map’s code-image, fillingAddresseswith ABSOLUTE addresses andIsStatistical == false.hwtrace-dotnet-testcarries the invariant-envelope case on any host; live capture is silicon-gated (make hwtrace-pt-live+make hwtrace-dotnet-teston a bare-metal Intel PT box). OnlyHwBackend.CoreSightkeeps the forward-look self-skip. See intel-pt-whole-window-substrate.md.Whole-window Intel PT STRONG tier wired behind the empty-ctor scope, with a runtime WEAK/STRONG decode-trust ladder.
asmtest_hwtrace_begin_window/_end_window(src/hwtrace.c) now arm and drain a real region-free Intel PT capture on an initedINTEL_PTtier — the ONE shared perf-AUX arm (pt_aux_open) that also serves the region path, exposed as the nativeasmtest_hwtrace_pt_begin_window/_end_windowpair the .NET inline ctor uses.asmtest_hwtrace_window_autoauto-selects STRONG only when theintel_ptPMU is present andasmtest_hwtrace_pt_window_trusted()proves the whole-window decode on the §Z2 synthetic fixture at runtime; else the WEAK single-step tier. The CEILING AMD LBR tier is deliberately never auto-selected for the exact whole-window contract (a sampled branch survey cannot meet it; live floor Zen 4+) — the quiet sampled complement stays explicit (new AsmTrace(HwBackend.AmdLbr)). The .NET empty-ctorusing (new AsmTrace())consults the ladder (AutoInitWindowBackend), andDegradationNote()names the PT probe outcome (present-but-untrusted vs no PMU). Live PT capture is silicon-gated; the ladder, the runtime trust probe, and the synthetic-fixture decode all run inmake docker-hwtrace(green on this AMD host — nointel_ptPMU, so the ladder resolves to WEAK and the native PT pair self-skips). See intel-pt-whole-window-substrate.md.Reproducible source tarball (
make package-source) attached to every release.git archiveof HEAD piped throughgzip -nemitsbuild/dist/asm-test-<version>.tar.gz+SHA256SUMS, byte-identical for a given commit across machines, so its digest is known ahead of the tag. Therelease.ymlcorresponding-source job builds it, uploads it as a dry-run artifact, and attaches both files to the tagged GitHub release — the digest-pinned source the system-package specs (Homebrew/Debian/AUR/vcpkg/conan) and Debian’s orig-tarball flow consume. See distribution-packaging.md.Socket-syscall
sockaddrcontents are decoded (asmspy--log).connect/bind/sendtorender theirstruct sockaddr *as{AF_INET, 127.0.0.1:8080}/{AF_INET6, [::1]:80}/{AF_UNIX, "/path"}(abstract sockets as"@name"), andaccept/accept4/recvfromdecode their OUT pointer on success (raw pointer on failure);socket()’s domain renders asAF_INET/AF_UNIX/… An unknown family prints{family=N, len=M}, never a guessed name.make docker-clicli-smoke PASS. See asmspy-cli-enhancements.md.ioctlrequests andfcntlcommands are named (asmspy--log).ioctlrendersTIOCGWINSZ-style names, or a faithful_IOC(dir, type, nr, size)decomposition for an unknown request (never a guessed name);fcntlrendersF_GETFL/F_SETFD/… with correct conditional arity (an arg-less command such asF_GETFLshows no third slot).make docker-clicli-smoke PASS. See asmspy-cli-enhancements.md.futexoperations are named (asmspy--log). The op renders asFUTEX_WAIT/FUTEX_WAKE_PRIVATE/… withFUTEX_PRIVATE_FLAGandFUTEX_CLOCK_REALTIMEmasked off before naming and re-rendered as suffixes, never silently dropped; an unknown op keeps its number plus those suffixes.make docker-clicli-smoke PASS. See asmspy-cli-enhancements.md.stat/statxresult buffers are decoded (asmspy--log).fstat/stat/lstat/newfstatatrender{st_mode=S_IFREG|0644, st_size=18}on success (a raw pointer on failure), andstatxrenders its mask-honoring{stx_mode=…, stx_size=…}(a field the kernel did not fill is omitted, not invented). The path decode these calls already had is preserved.make docker-clicli-smoke PASS. See asmspy-cli-enhancements.md.Hot-edges → data-flow drill-in (asmspy TUI, mode 7). In the frozen hot-edges view, arrows select an edge and
Enteropens a data-flow capture (mode 9) of the function containing the edge’sto_addr(falling back tofrom_addr), reusing the call-graph drill-in idiom. The pure decision (asmspy_edge_drill) requires a sized function and accepts a mid-function landing (drill ≠ rank); it is unit-tested intest_autoregion(6 checks) so it is covered on every host, not just an AMD IBS box. The ncurses wiring is pty-driven (manual-only); the decision logic runs in CI. See asmspy-cli-enhancements.md.Block-step replay record-and-inject for rdtsc/rdtscp/rdrand/rdseed/cpuid, gated per block rather than per region.
src/dataflow_blockstep.c’sstep_blocknow injects each site’s recorded post-state (read from the T5 DR exec-breakpoint boundary) into the Unicorn replay and terminates the block there — the same record-and-inject shape assyscall/int 0x80, minus the producer-local write-set synthesis (Capstone already reports the complete architectural write set for all five mnemonics).region_scan’sinjectableverdict now admits regions whose only impurities are syscall/int80/HWREC (subject to the existing 4-slot DR0-3 cap — a 5th+ distinct site still falls back to single-step, reasonhwrec-overflow), and a region no longer forfeits the replay’s perturbation win for a hwrec site its real run never reaches (hw_hits/injectedstay 0 for an unexecuted site while the region still replays). New opts test hookno_hw_recordskips arming the DR breakpoints, reproducing the pre-injection fail-closed truncation on demand.make dataflow-blockstep-test191/191 (was 186/186, +5), stable across 5 consecutive runs on this Zen 2 host. See dataflow-producer-correctness.md.Extents-driven block-step region scan: a caller-vouched list of real instruction extents lets
region_scanskip an embedded constant-pool island instead of desyncing on it (BSVS-2). Newasmtest_blockstep_extent_t({off, len}, blob-absolute) plus opts fieldsextents/nextents(src/dataflow_blockstep.c) — NULL/0 keeps today’s whole-region sweep.region_scanis split into a per-extent inner sweep (region_scan_extent, reused for both the implicit whole-region case and each real extent) whose verdicts aggregate across extents; bytes outside every extent are never decoded, so a data island sitting between two extents costs nothing.run()validates extents are sorted, non-overlapping, and fully inside[region_off, code_len)before any tracee is spawned (DF_BLOCKSTEP_EINVALotherwise); the publicis_pure/is_replayable/is_injectableclassifiers stay whole-blob (extents are arun()-only capability). New fixtureisland_sse(the same constant-pool-island shape as the existingislandfixture, with a legacy-SSEpaddqin place ofisland’s VEX-128vpaddqso extents can actually recover it into the replay path) proves the positive case byte-identical to the single-step oracle with stops cut, while desyncing exactly likeislandwithout extents — the negative control.make dataflow-blockstep-test199/199 (was 191/191, +8), stable across 5 consecutive runs;make docker-dataflow-attach520/520 across all 8 suites, 0 skips;make docker-docsclean. See dataflow-producer-correctness.md.Def-use graph and forward/backward slice surface in the Ruby, Lua, Zig, Rust, Go, Java and .NET data-flow bindings (previously producer-only), via a by-pointer slice-seed entry point (
asmtest_slice_forward_seed/_backward_seed,include/asmtest_valtrace.h— a 72-byteat_val_rec_tis SysV MEMORY-class and several of these FFIs, Ruby Fiddle chief among them, cannot pass it by value). Each binding’sValueTracenow exposesdefuse()/forward_slice(step)/backward_slice(step)alongside the existing live-attach producer, so all ten language bindings (with the prior Python/C++/Node) share one surface. Round-trip-tested with a hand-built r10→r11→r12 register-move chain (forward_slice(0)/backward_slice(2)both{0,1,2}) and, over the shared live-attachdf_chainfixture, the memory def-use edge these seven could never reach before —backward_slice(4)andforward_slice(0)both equal{0,1,2,3,4}(the store at step 1 reached through the load at step 2), excluding the trailingret. All sevendocker-dataflow-<lang>lanes green at their new 40/40 (was 36/36), 0 skips. See dataflow-bindings-slice-codeimage.md.Code-image recorder wrapper and versioned (time-correct) operand decode in all ten data-flow language bindings (
asmtest_codeimage.h’snew/track/now/bytes_at/free, plus the newasmtest_dataflow_ptrace_attach_pid_versionedentry point). Each binding gains aCodeImagewrapper (Python/Ruby/Lua’savailable()/skip_reason()/track()/now()/bytes_at(), C++’s RAIICodeImage, Node’s mirroring the existing hwtrace-binding class, Zig/Rust/Go/ Java/.NET’s thin function wrappers) and aValueTrace.attach_pid_versionedmethod;attach_jitno longer unconditionally passesNULL/null/nilfor the versioned-decodeimgargument — a caller with a recorder now gets time-correct operand decode across a mid-capture JIT patch/free/reuse instead of NULL/live-snapshot bytes. Verified live: each binding tracks a recorder over a real victim’s published region and decodes anattach_pid_versionedcapture through it (result and step-count assertions tied to that run’s own arguments, so a stubbed capture cannot pass). All tendocker-dataflow-<lang>lanes green with the new assertions, 0 skips. See dataflow-bindings-slice-codeimage.md.Block-step pre-cover: an IBS covered-block table that memoizes the ptrace block-step reconstructors’ decode.
asmtest_bs_precover_build/_free(include/asmtest_blockstep_internal.h, internal — no new public ABI symbol) pre-walks eachasmtest_ibs_normalize_blocks-covered leader’s straight-line run ONCE and caches the per-instruction factsclassify_branchwould otherwise recompute on every#DBstop;blockstep_reconstruct(shared by the region and attached block-step drivers) resolves a cache hit by replayingasmtest_bs_scan_terminator’s exact decision procedure over the cached run — same bytes, same primitives, zero Capstone calls — so a hit is provably identical to a fresh scan, and a miss (including a leader that is not a real instruction boundary — the hostile case) falls back to the shipped path unconditionally. IBS stays statistical: pre-cover only memoizes tracer-side decode, never lets coverage skip recording anything (the exact parity contract inasmtest_ibs.h’s INVARIANT stands). A differential over theLOOP_X86fixture (precover NULL vs. covering the loop head) proves byte-identicalinsns[]/blocks[]/truncated/result while cutting branch- probe decode calls;asmtest_bs_stats/_resetare test hooks that count cumulative probe calls and cache hits. See ptrace-blockstep-tracer-correctness.md.ASMTEST_TRACE_IBS_PRECOVER: an opt-in policy bit that wires the block-step pre-cover table above into the cross-tier auto cascade (asmtest_trace_call_auto,include/asmtest_trace_auto.h). When set and IBS-Op is available, the block-step rung forks an isolated warm-up child that re-runs the routine for a bounded ~30ms while the parent surveys it out of band withasmtest_ibs_survey_process(no ptrace, no perturbation), builds a pre-cover table from the resulting live histogram, and installs it around the oneasmtest_ptrace_trace_call_blockstepcall the rung already made — so a trace produced with the bit set is byte-for-byte identical to one produced without it; any survey/build failure degrades silently to the plain rung. On a live Zen 2 host a 25-iteration loop’s block-step branch-probe decode calls dropped from 101 to 0 (a real live warm-up survey covering both of the routine’s basic blocks). Off AMD, or wherever the auto-cascade’s fast in-process single-step backend already completes the capture before reaching block-step, the bit is a proven no-op — never a behavior change. See ptrace-blockstep-tracer-correctness.md.In-process, branch-granular single-step (W3):
asmtest_ss_btf_available/asmtest_ss_btf_trace(src/ss_btf.c), the missing third single-step form. asm-test already had branch-granular stepping out of process (PTRACE_SINGLEBLOCK) and per-instruction stepping in-process (EFLAGS.TF,ss_backend.c); this armsDEBUGCTL.BTFalongsideEFLAGS.TFover the same thread-pinned/dev/cpu/N/msrrouteasmtest_amd_msr_traceuses, so a taken-branch retiring — not every instruction — is what traps, with no 16-entry ceiling on the reconstructed stream (unlike AMD LBR). Gated by a hang-proof functional probe (some hypervisors silently maskDEBUGCTL.BTFand degrade to per-instruction stepping — the probe catches this, a build check cannot); deliberately scoped to a pinned leaf-routine envelope with per-trap re-arm (BTF is a hardware one-shot the CPU clears on every#DB) and faithful truncation on any observed context switch (Linux does not preserve a user-written BTF across one) — the general, context-switch-proof case stays owned by the shippedPTRACE_SINGLEBLOCKtrio. Rides the existingdocker-hwtrace-msr--privilegedlane, no new capability. Live- verified on a Zen 2 host: the sharedROUTINEfixture reproduces the single-step baseline’s exact[0,3,6,c,11]stream and{0,0x11}block partition byte-for-byte, and a 20-trip loop (19 taken back-edges, past any 16-entry LBR window) reconstructs all 62 instructions, complete, 10/10 stable runs. See inproc-btf-block-step.md.macOS out-of-process single-step tracer (
asmtest_mach_*), completing the W2 foreign-process story on macOS.asmtest_mach_trace_call/_trace_attached/_run_to(asmtest_mach.h,src/mach_backend.c) mirror the Linuxptraceout-of-process tracer’s exact shape and offsets, but throughtask_for_pid+ a MachEXC_MASK_BREAKPOINTexception port +thread_set_stateinstead — macOSptracecannot editRIP/RFLAGSat all.run_to’s breakpoint arm falls back from a softwareint3to aDR0/DR7hardware execution breakpoint on W^X code, same as the Linux tracer. x86-64 only for now.make mach-stepper-test(needsscripts/codesign-debugger.sh’s ad-hoc self-sign, or root) runs the lane live; self-skips (ASMTEST_MACH_EPERM) without either.Pure tests for the IBS sample-period rounding/clamp and the additive-ABI
struct_sizeguard (first coverage). The/16rounding, the<16 → defaultclamp, and theperiod_jittertail-guard had never executed against non-default values anywhere; internalasmtest_ibs_effective_period/asmtest_ibs_effective_jitterseams (shared with the live attr fill, so tested == shipped) now pin them.Intel Pin vs. DynamoRIO analysis + a four-track umbrella plan — what Pin makes possible that the shipped DynamoRIO tier cannot. 2026-07-17-intel-pin-vs-dynamorio.md and intel-pin-capabilities-plan.md. Most “Pin advantages” are maturity, not impossibility (DR has
drrun -attach, supports Windows, and the taint ground is already held in-tree), and the note separates those out. Four items survive as separable plan tracks: PIN-1 an Intel SDE lane that runs the existingTEST()suites undersde64 -futureso APX / AVX10.2 / AMX / AVX-512 assembly gets full register/flag/memory/ABI assertions on any x86-64 host including CI — a true impossibility for both DR (executes on real silicon) and the Unicorn tier (QEMU 5.0.1 predates AVX TCG;vaddps ymm→UC_ERR_INSN_INVALIDand VEX-128 is silently mis-run as SSE), converting CLAUDE.md’s “specific CPU generation” hardware self-skip into an installable, pinnable dependency; PIN-2 an XED-decoded Pin trace tier for the newest extensions DR’s own decoder rejects (APX is open — DR #6226; the once-broken AVX-512 VNNI is fixed — DR #5440 closed 2022-04-25, and the pinned DR post-dates it — so the case rests on APX alone); PIN-3 Pin probe-mode arg/return capture (original code runs native, no code cache — thecapture-args-returns.mdmiddle tier DR has no equivalent for); and PIN-4 libdft64 as an independent taint oracle diffed byte-for-byte against the shipped DR taint client — theASSERT_MATCHES_REFcross-validation idiom. The reverse gaps are recorded so Pin is not mistaken for a superset (x86-only — no AArch64; no in-process no-IPC model; proprietary freeware). Every track is fetched-and-pinned, test/oracle-only, never shipped — DR’s exact handling — so none adds a bindings-parity obligation.AMD hardware review + follow-up plan — an adversarial pass over the AMD tiers in which four of nine candidate findings were REFUTED and recorded as such. 2026-07-17-amd-hardware-review.md and amd-review-followup-plan.md. Every claim resting on kernel or silicon behaviour was checked against fetched primary source at pinned tags (v6.10/v6.12/v6.14), not recall — one verifier caught
mastershifting line numbers mid-review and re-pinned. Confirmed: Zen 3 BRS cannot open (the probe/capture usePERF_COUNT_HW_BRANCH_INSTRUCTIONS→0x00c2, butamd_brs_hw_configdemands raw0xc4andsample_period > lbr_nr(16);EINVALis then reported asAMD_NOHW, so a real Zen 3 owner is told “no AMD branch records”) — code fix hardware-blocked per the house rule, but three docs assert the working path, including a Phase 0 marked landed that specifies the exact missing arm;branchsnap’s synthetic boundary edge inflates the depth check so a completeuse == 15window is spuriouslytruncated→ a real re-execution (n_dec = use + 1vs a check counting hardware slots only);IBS_MAX_RECORDis pre-callchain (112B vs ~1032-1184B — it landed in68b53850, callchain ina266b91two days later and never touched it), leaving the loss heuristic ~10× short wherePERF_RECORD_LOSTprovably cannot cover the gap →lost==0 && throttled==0silent loss in a fidelity-first lane; andasmtest_amd_freeze_available()is dead (nm: one definition, zero undefined refs) with a flatly false string printed to a human. The organising finding is process, not code: no CI lane exercises AMD silicon — the one AMD-targeted job is namedhwtrace-privileged (PERFMON; AMD-exact self-skips off Zen)and runs onubuntu-latest— so the only gate is a manual checklist with exactly one commit, timestamped identically to the fix for the bug it still calls open, which now instructs treating atruncated=0regression as a known issue and gates on two signals this project measured false five days later. Refuted and recorded so they are not re-raised (the Matrix 3 convention): the MSR TOS-rotation claim (transplants Intel architecture — LbrExtV2 pinsFrom[0]/To[0]by register renaming and hard-codeshw_idx = 0because rotation is impossible; the linear read is correct), “IBS opts are unreachable” (NULL is the designed contract;SYSTEM_WIDEis tested live under--cap-add=PERFMON), theRipInvalidChkimpact (the affected silicon population is empty — Family 10h hardwiresBrnTrgt = 0, so the existing gate rejects it first), and the unread-IBS-regs “gap” (a dated non-goal indata-flow-tracing-plan.md:94; the original grep was a false negative on “address sampling”).asmspy --dataflow --autoworks off AMD now: a portable software-clock sampler (--sampler=ibs|sw), with the residency hazard owned out loud. Auto-targeting was AMD-IBS-only;asmtest_swclock_survey_process(new in the IBS lane’s backend:PERF_COUNT_SW_TASK_CLOCK+PERF_SAMPLE_IP, no PMU at all) samples any-vendor hosts and VMs under the same out-of-band, unprivileged envelope, with availability probed BY DOING and an errno-carrying reason (asmtest_swclock_unavail_reason) from birth — the--samplelesson, applied rather than re-learned. The pure ranking half (asmspy_autoregion_rank_ip) is transparent about being the WEAKER rule: an IP histogram measures residency, and residency’s winner is often the entered-once-never-returns shape whose entry breakpoint can never fire — test_autoregion #15 pins the two rules DISAGREEING on the same behaviour. So the sw path ranks up to 3 candidates andcmd_dataflowWALKS them: a winner never seen entering is refused at the bounded entry wait and the next-ranked candidate is tried, each refusal reported. Proven live indocker-cli-ibson the built-to-disagreeauto_victim: sw picksgrind_forever, the wait refuses it, the walk capturesentered_often. 9 new pure checks (31 total); both cli lanes PASS.cli/asmspy_ghash.h+test_ghash.c— the graph’s hash index gets the collision test its faithful-gap note demanded. The--graphengine’s open-addressed index shipped with a measured blind spot: a probe loop that trusts the hash and skips the key compare emits byte-identical smoke output, because ≤7-node graphs in a 128-slot table never collide (on a larger graph that mutant silently over-merged an edge). The mechanism now lives in a pure header (asmspy_gh_find’s eq callback owns the key compare; the engine’s node/edge lookups supply it) and a unit test brute-forces three keys into one slot, so exactly that mutant fails. Mutation-proven 3/3 (measured): accept-first-slot → 5 FAILs, idx-not-idx+1 slot encoding → 3 FAILs, grow-without-rehash → 1 FAIL. Engine behavior unchanged (make docker-cliPASS,--graphuniqueness e2e included).asmspy TUI symbol picker (modes 2 and 9) gains a
Tab-cycle sort: address -> hot edges -> name. The picker used to be a flat, address-ordered list with no way to find “what’s actually running” short of guessing a name to filter by. Hot edges reuses the exact entry-arrival rule--dataflow --autopicks with (asmspy_autoregion_rank, one AMD IBS-Op window) rather than minting a second definition of “hot” — so the picker’s ranking and--auto’s pick always agree. An IBS-less host (or one where perf is locked down) falls back to address order with an on-screen reason rather than silently doing nothing. Name sorts case-insensitively. The row permutation is kept separate from the symtab’s own address-sorted storage (asmspy_symtab_at’s binary search depends on it), mirroring the process picker’s existingorder[]indirection.asmspy --dataflow <pid> --auto [--module=<m>]— trace what a process is DOING, no symbol needed. Auto-targeting samples the target OUT OF BAND (AMD IBS-Op, no ptrace, no perturbation) for 400 ms, ranks the hottest entry edge — an edge whose target is a function’s start is a direct observation of the exact event the data-flow producer blocks on, where the intuitive rules fail: the hottest raw edge is usually a mid-function loop back-edge no entry breakpoint can catch, and a PC histogram picks the functions entered once and never again (main, every event loop), which HANGS the producer — then hands the winner’s(base, len)to the existing--dataflowengine. Entry-edge counts are quantitative (measured: they reproduce a victim’s true 8× loop trip count and 2× call-site ratio). The ranking is a pure header (cli/asmspy_autoregion.h, 21-check unit test) so its correctness is covered on every host while the AMD-hardware live leg runs in thedocker-cli-ibslane. The resolver layers the JIT map over the ELF symtab, so a hot JIT’d/managed method (perf-map or jitdump, real size) can win the pick and names the capture — proven live against a perf-map victim whose only entry arrival is an anonymous-mapping function the ELF symtab cannot see. An idle target gets a truthful refusal (“no function was observed being ENTERED”), not a guess; zero-size symbols cannot win on their exact-start-only resolution technicality;--auto --tidis a usage error (the sampler carries no tid, so pinning could only arm a breakpoint on a thread that never arrives);--module=scopes the pick with the same substring rule as--tree --module=.make docker-cli-ibs— the lane that actually tests--sample, and accurate skip reasons everywhere. The plaindocker-clilane runs under Docker’s default seccomp, which blocksperf_event_open— so every--sampleassertion had self-skipped since the view landed (a green gate over an untested view, on hosts where IBS works fine). The new lane reruns the same image and smoke with--cap-add=PERFMONunder the default profile (CAP_PERFMON bypassesperf_event_paranoid; no sysctl change needed), so the--sampleand--autoblocks execute their else-branches for real on an AMD IBS host. Alongside it, a new public APIasmtest_ibs_unavail_reason()carries the realperf_event_openerrno out of the backend:# SKIP --sample:used to print an empty string precisely when the substrate probe passed but perf was blocked — the one case an operator on their own AMD box needed the reason most. EACCES now says paranoid/CAP_PERFMON, EPERM names seccomp, instead of one indistinguishable silence.asmspy region view samples WORKER threads (
--trace, TUI mode 2) + a new--tid=<t>filter.asmspy_engine_regionused toPTRACE_ATTACHthe thread-group LEADER and run it to the region, so a function that runs on a worker thread was never observed and the view reportedASMSPY_REGION_NEVER_RAN— structurally blind to exactly the code asmspy exists to show, since a managed method almost never runs on the leader. It now SEIZEs every thread (PTRACE_O_TRACECLONE, so a thread spawned mid-run can win a later round) and races them all to the entry, sampling whichever arrives first.--tracegains--tid=<t>to pin one thread, matching--stream/--graph/--tree/--dataflow. The design reuses the data-flow tier’s oracle-validated race (dfp_seize_all/dfp_run_to_multi) rather than inventing a second one, reimplemented over asmspy’s own thread table per the standing precedent that an engine stays incli/and leavessrc/ptrace_backend.cuntouched.--tidpins via a per-thread hardware execution breakpoint, not the shared int3: a shared int3 traps every thread, and stepping hot non-target threads back over it was measured not to converge.cli_smoke.shasserts the worker sample, both--tiddirections, and target survival past a settle;make docker-cli→ cli-smoke PASS. Known gaps, unchanged: the any-thread entry breakpoint isPOKETEXT-only (a W^X JIT page self-skips; no DR0 fallback), and the TUI has no thread picker.CI gate for the DynamoRIO attach tier (
taint-attach). Increments 1 and 3-5 — cooperative attach, the marker-less interactive nudge, externaldrrun -attachinto a running native process with taint capture, and K-round attach/capture/detach cycling — had docker lanes but no CI job, so five increments of capability had no regression gate while the launch-only taint tier next door had four. The external-attach image needs--cap-add=SYS_PTRACEand its comment still described it as the Increment-2 research probe (“a manual diagnostic, not in the main gate”) — true when written, but Increments 4-5 were later added to the same image, so landed capability inherited a probe’s CI posture. Both lanes now run on every push. The two MANAGED probes stay out by design: they record a reproducible NO-GO, not capability.Data-flow tracing Phase 6 Increment 1 —
libasmtest_dataflowshared lib + Python binding.make shared-dataflowbuilds the pure analysis pipeline (L0 value sink + L1 def-use + L2 slice + method identity + GC-move canonicalization + runtime-helper summaries; the emu/ptrace/DR producers stay separate tiers) into a dlopen-able shared library — the packaging target the language bindings consume. First bindings: Python (asmtest.dataflow, ctypes), C++ (bindings/cpp/asmtest_dataflow.hpp, header-only), and Node (koffi,bindings/node/dataflow.js) all wrap the pure GC-move canonicalizerasmtest_gcmove_canonand the tiered-re-JIT-aware method resolverasmtest_method_resolve_pc;make dataflow-{python,cpp,node}-testbuild and run their TAP suites (8 / 16 / 16 checks, mirroring the Ctest_dataflow_gcmove/test_dataflow_methodsemantics). Each self-skips cleanly when the lib is not built, so none reddens a general binding job. A newdataflowCI job builds the lib and runs the Python + C++ bindings on every push (the Node binding runs in the docker bindings lane). All three host bindings (Python, C++, Node) additionally wrap the full L0->L1->L2 pipeline (ValueTrace: value-trace build -> def-use -> forward/backward slice), round-trip-validated against the C semantics (register move chains, load-after-store through memory, no spurious cross-links — which also validate the 13-field at_val_rec_t marshalling). The remaining seven language bindings are later increments.Live GC-move detection feed for the data-flow tier (
GcMoveMap, .NET). An in-procEventListeneron the CoreCLR runtime provider that enables the GCHeapSurvivalAndMovement keyword and capturesGCBulkMovedObjectRangesfrom a compacting GC — the live source the pure GC-move canonicalizer (asmtest_gcmove_canonicalize) was built to consume. Validated viamake docker-hwtrace-dotnet: an induced compacting gen2 collection is captured as 11 events / 20474 moved ranges (suite 169 → 177). Truthful scope: in-procEventListenersurfaces the reliable scalar range count but not the manifest struct-arrayValuespayload, so the concrete{old_base, new_base, len}triples that drive the canonicalizer end-to-end are deferred to a raw EventPipe/nettrace path — the keyword, compacting-GC inducement, and listener wiring (the uncertain parts) are proven.Data-flow tracing Phase 5 Increment 1 — DynamoRIO in-band L0 value producer (
src/dataflow_dr.c,src/dataflow_dr_client.c). The in-band, whole-process analog of the scoped ptrace L0 producer: a DynamoRIO client instruments a target under real DR and captures per-instruction operand values into the sameasmtest_valtrace_tsink, so the L1 def-use builder and L2 slicer work unchanged on an in-band capture. Cross-validated against the emulator L0 oracle on a shared fixture (in-band def-use edges and forward/backward slices equal the oracle’s), validated live viamake docker-drtrace(dr-valtrace-test14/14; DR is a software DBI engine, so the lane needs no privilege or special hardware). Self-skips (exit 0) withoutDYNAMORIO_HOME. Store values and RIP-relative/segmented/VSIB memory EAs are deferred to a later increment (the current fixture avoids them); DR-side taint shadowing and whole-process breadth are also later increments.Data-flow tracing Phase 4 Increment 3 — runtime-helper summary edges (
src/dataflow_helpers.c).asmtest_defuse_build_summarizedrecognizes a .NET runtime-helper call in an L0 value trace (via the Increment-1 method resolver → a helper table matched by name/prefix) and collapses the helper run into a summary node — emitting only its declared input reads and output writes and dropping the body — so caller dataflow connects across the helper (arg def → summary → return use) without instrumenting CoreCLR internals. Supports reg→reg helpers (allocation, generic-dict lookup) and aMEM_AT_REGwrite-barrier output. Conservative by construction: an unrecognized call is descended normally, never given a fabricated edge. Pure, host-independent suitetest_dataflow_helpers(36 checks).Data-flow tracing Phase 4 Increment 2 — GC-move canonicalization (
src/dataflow_gcmove.c). A pure, host-independent transform (asmtest_gcmove_t, the shape of EventPipeGCBulkMovedObjectRanges:old_base/new_base/len) that remaps memory addresses across a heap compaction to a stable canonical identity, so a managed value’s def-use survives the move without pre/post-move false aliasing — the plan’s Phase-4 exit criterion. Synthetic suitetest_dataflow_gcmove(26 checks) proves the exact trap: a def at the old address and a use at the new address unify into one object, while an unrelated object that later reuses the freed old address does not alias it. The live EventPipe feed is a later increment.IBS-Fetch front-end coverage lane (AMD IBS statistical lane Phase 7). A second AMD IBS producer beside the retired-op edge sampler:
ibs_fetch(PMU type 10) samples fetch addresses (front-end / i-cache / ITLB view). A pure, host-independent decoder turns onePERF_SAMPLE_RAWfetch record into{fetch_addr, valid, complete, icache_miss, itlb_miss, latency}(unit-tested with synthetic records on every CI host), plus an availability probe and a headless fetch-coverage survey with faithful throttled/lost provenance, self-skipping off IBS/permission. Kept fully internal (src/ibs_backend.h) — no publicasmtest_ibs_*surface added, so no binding flag day. Verified live on Zen 5.Data-flow tracing Phase 4 Increment 1 — PC→method-identity+version resolver (
src/dataflow_method.c). A pure, host-independent resolver that labels each step of an L0 value trace with its owning method + version from a jitdump/perf-map-shaped method-map, correctly handling tiered re-JIT (newestcode_indexwins for an address; a re-JIT to a new address is a new version) — the managed-taint prerequisite. Synthetic suitetest_dataflow_method(29 checks, incl. the moved-re-JIT version distinction); the hard GC-move canonicalization is deferred to a later increment.hwtrace-privilegedCI job + AMD hardware-validation doc. A CI job exercisesmake docker-hwtrace-privilegedso the--cap-add=PERFMONlane can’t bitrot (the AMD-exact tests self-skip on GitHub’s non-AMD runners — transparent by design; it lights up on a future AMD runner).docs/internal/amd-hardware-validation.mddocuments the manual pre-release validation on real Zen 3+/Zen 5 silicon — closing the gap that let thecall_autoLBR truncation bug hide (the exact AMD paths never ran in CI).Size-negotiated hwtrace options ABI + machine-readable status surface + escalation mechanism, across all ten bindings (the AMD-followup API flag day — Phases 1, 3, and F22/F26/F37).
asmtest_hwtrace_options_tnow leads with asize_t struct_sizethe caller sets (theINIT_OPTSidiom, or explicitly after a zero-fill);asmtest_hwtrace_initcopiesmin(struct_size, sizeof)and zero-fills the tail, so an older/newer caller is never read out of bounds — and a caller that fails to self-describe (struct_size == 0or too small to reachbackend) is rejected withEINVALrather than having a set field silently dropped. Newasmtest_hwtrace_status()(available / code / stage / probe errno /perf_event_paranoid/ reason) andasmtest_hwtrace_perf_event_paranoid()distinguishASMTEST_HW_EPERM(substrate present, permission denied — e.g. AMD LBR on an unprivilegedparanoid > 2host) fromEUNAVAIL(missing silicon), backed by one shared classifier sostatus()andskip_reason()cannot drift.asmtest_trace_choice_tgrew amechanismfield (HW_BRANCH/TF_STEP/MSR_LBR/BLOCKSTEP/PER_INSN/DBI/EMULATOR/STATISTICAL) plusASMTEST_FIDELITY_STATISTICAL, sotrace_call_autoreports which rung actually won and a statistical result is structurally unmistakable for an exact one. All ten wrappers mirror the new layouts and wrap the new calls; the parity gate passes with zero allow-list changes. Suite 358 → 383 (ABI guard, status incl. the live-EPERM assertion, mechanism).Data-flow tracing gains a live scoped ptrace L0 producer (Phase 3 — real values, out of band).
src/dataflow_ptrace.csingle-steps a routine (fork+PTRACE_TRACEME, orPTRACE_SEIZEattach to a live victim that survives detach) and emits the sameasmtest_valtrace_tstream the Phase-0/1 analyzers consume, so def-use + slicing work unchanged on live captures — reading each step’s registers (GETREGS,GETFPREGSfor XMM,NT_X86_XSTATEfor 256-bit YMM) and the memory its operands touch. Cross-validated edge-for-edge and value-for-value against the emulator L0 oracle; RIP-relative effective addresses resolve against the next instruction (a bug an adversarial verify caught before merge), gs-based and wide-vector operands captured.dataflow-test26 → 36.Data-flow tracing tier, Phases 0–2 (
include/asmtest_valtrace.h,make dataflow-test) — the CI-runnable milestone of the data-flow plan. Phase 0: the shared L0 value-trace sink (asmtest_valtrace_t: caller-owned buffers, append/stash-wide/truncate discipline mirroringasmtest_trace_t) plus the Capstone operand read/write-set enumerator (explicit register + memory operands with base/index/scale/disp/segment, and the implicit ones —eflagswrites,rspread+write on push/pop — viacs_regs_access). Phase 1: the L1 def-use graph over a recorded value trace (register moves and load-after-store memory edges) and the L2 forward/backward slicer on top of it. Phase 2: the emulator (Unicorn) L0 producer — replay a routine underuc_hookinstrumentation and emit the value trace the pure phases analyze; validated live (per-step values, def-use edges, and both slice directions asserted against a hand-traced fixture). Pure phases run on every host; the emulator cases self-skip without Unicorn. Newmk/dataflow.mk; suitestest_dataflow/test_operands/test_dataflow_emu(53 checks). The known raw-address aliasing false positive (pre-GC-canonicalization) is asserted AS a false positive, per the plan’s Phase 4 note.asmspyreads binary jitdump files — the bytes-accurate, tiered-recompile-aware JIT symbol source (asmspy plan Theme A). Ajit-<pid>.dumpreader (LE header +JIT_CODE_LOAD/JIT_CODE_MOVErecords; unknown record types skipped viatotal_size; a truncated in-flight tail ends the parse keeping what’s whole) is now tier 1 of the JIT resolve chain — discovered the way perf does (a mapped marker in/proc/<pid>/maps, then/tmpand the target’s cwd), parsed ahead of the text perf-map (tier 2, the LCD), re-JIT and code-motion aware (newestcode_indexwins,CODE_MOVErelocates). Same rate-limited refresh-on-miss discipline as the perf-map path. Hardened against hostile files (bounded name reads, zero/shorttotal_sizerejection — fuzzed under ASan/UBSan). Also new:--tree --json/--dotexports mirroring the--graphexporters, the extracted call-graph sort comparator (cli/asmspy_graphsort.h) with its ordering/tiebreak unit test, and ajitdump_victimend-to-end smoke..NET
AsmTrace.Windowcaptures methods JIT’d mid-window — the sibling-thread live JIT publish (extensions plan E3), closing the deep-BCL gap.JitMethodMap.SetPublishChannelnow starts a dedicated, never-stepped publisher thread: theMethodLoadVerbosecallback only enqueues(base,len)onto a lock-free queue (publishing inline could fire on the single-stepped thread and re-enter the runtime under step — the observed SIGABRT that kept this OFF), and the sibling drains it and P/Invokes each record into the shared address channel while the window runs. Stop joins the publisher before the channel is freed (no use-after-free window); the §E1 hybrid keeps live publish off by design (it must capture only the surveyed hot slice). NewWholeWindowScope.LiveJitPublishedcounter; suite grows 161 → 169 checks including a ptrace-free mechanism test (native ring-head readback) and a mid-window-JIT integration case (52 records live-published in the docker lane).Env-gated debug logging for the hwtrace/AMD tier (
src/debug.{c,h}, followup Phase 4 / F32).ASMTEST_HWTRACE_DEBUG=1(orASMTEST_AMD_DEBUG=1) turns on stderr tier diagnostics; unset costs one cachedgetenvper process. Covered by two suite checks (silent when off, emits when on).Host-independent synthetic-ring tests for the AMD branch-stack parse (review F43/F44). The
hwtrace_end_amdring-parse now has an internal, linkable entry (asmtest_amd_ring_parse_decode) driven by craftedPERF_RECORD_SAMPLEbuffers — the nr-clamp / LOST / Tier-A-vs-Tier-B logic andamd_span_decodable’s dropped-jmp follow finally run on every CI host, AMD or not (+368 lines inexamples/test_hwtrace.c, suite 341 → 358).AMD LBR window-reach tuning guide (
docs/guides/tracing/amd-lbr-tuning.md, review F47 / followup Phase 10) — what bounds a 16-deep LbrExtV2 window, the sizing/splitting levers, whattruncatedmeans and how it is reported, when the statistical IBS lane is the better tool, and the privileged-vs-unprivileged lanes includingperf_event_paranoid.asmspy --sample+ TUI mode 7 — a live statistical hot-edge view, out of band (IBS lane Phases 2–3, the flagship deliverable). Built on the newasmtest_ibs_survey_process(pid, ms, opts, out): whole-process IBS-Op coverage that opens one perf event + ring per thread of the target (enumerating/proc/<pid>/task, with one mid-window rescan for threads spawned after start; the residual born-and-died-in-window race and the privileged system-wide remedy are documented, not hidden) and merges everything into one hot-edge histogram.asmspy_engine_sampleresolves both endpoints of each edge through the existing ELF-symtab → JIT-perf-map chain, so managed Node/.NET/Java frames are named. Headlessasmspy --sample <pid> [ms] [--json]prints the histogram —count from -> towith[misp N%]/[ret]tags and faithfulbranch/total samples/throttledprovenance — or machine-readable JSON; TUI menu item “7) Hot edges (sample)” shows the same table live, pausable + scrollable + Tab-sortable (count / mispredicts). Unlike the stream/graph/tree views this never attaches ptrace and never single-steps — the target runs at full speed — making it the only rich view that is safe on a live JIT, exactly the targets single-stepping can crash. Self-skips (# SKIP, exit 0) off IBS; new busy victimcli/sample_victim.c+ a--samplesmoke incli/cli_smoke.sh; the TUI view is driven end-to-end through a pty harness. Verified live on Zen 2: both surfaces name the victim’s hot back-edge without perturbing it.IBS-Op fallback for the AMD whole-window statistical survey (IBS lane Phase 4; fixes the AMD review’s F6). On Zen 2 the branch stack (BRS / LbrExtV2) does not exist, so
asmtest_hwtrace_sample_window_amd(and its begin/end split) returnedEUNAVAILon the one AMD host class that most needs a crash-proof survey. New internal window primitives (asmtest_ibs_window_begin/_end,src/ibs_backend.h) arm IBS-Op on the calling thread around the caller’s window body, reusing the channel/drain/edge-hash machinery; on branch-stackperf_openfailure — or whenASMTEST_FORCE_IBS_SURVEYis set, for cross-validation on Zen 3+/CI — the survey delegates to IBS and flattens each sampled edge’s target into theips[]endpoint histogram weighted by count, so the caller’s bucket-by-method hotness view is unchanged in shape. Purely STATISTICAL: a separate producer that never feeds the exactinsns[]/blocks[]parity cascade; the branch-stack path is byte-identical when the stack is present and the env unset. Covered bytest_amd_sample_window_ibs(self-skips off IBS); verified live on Zen 2 (~468/468 endpoints in the hot loop, fullhwtrace-test341/341).asmspygained three whole-process structure views and a per-thread lens since the entry below. A call graph (TUI mode 4, headless--graph <pid> [n] [--sort=invocations|fanout]): everycallattributed caller→callee across all threads, aggregated per function, with--jsonexport (nodes and{caller,callee,count}edges, addresses as0xstrings) and--dotemitting a Graphviz digraph (kind-coloured nodes, count-labelled edges) ready fordot -Tsvg. A call tree (TUI mode 5, headless--tree <pid> [n]): the same feed with nesting/order preserved, indented by depth, in a two-pane TUI. A process/thread topology view (headless--procs, plus anF2flat-list ↔ tree toggle in the process picker): the process forest drawn with├─/└─/│box glyphs, threads then child processes nested under each process. A--tid=<t>filter for--stream/--tree/--graphseizes and steps only that thread, leaving the rest of the process at full speed. The call-graph and region (assembly & funcs) TUI views now pause + scroll like the log views (spacefreezes a stable snapshot, arrows/PgUp/PgDn/Home/End move,Tabswitches pane focus in the region view). And all single-step engines resolve JIT frames through the runtime’s perf map (/tmp/perf-<pid>.map— Node/V8--perf-basic-prof, .NETDOTNET_PerfMapEnabled=1, OpenJDK perf-map-agent), refreshed rate-limited on miss so a compiling JIT keeps getting named: managed frames rendername [jit]in the stream/tree and[JIT]-tagged internal nodes in the graph.Statistical AMD IBS-Op tracing lane (
asmtest_ibs.h,src/ibs_backend.c) — Phases 0–1. A new, self-contained statistical trace producer for AMD hosts where every branch-stack facility is absent (Zen 2 has no BRS / LbrExtV2, so every exact hwtrace backend self-skips and the machine falls back to ~1000×-slower single-stepping). IBS-Op (Instruction-Based Sampling) is the one branch-tracing facility this silicon has: it tags a retired op per NMI window and, for taken branches, delivers both the source (IbsOpRip) and target (IbsBrTarget) — a statisticalfrom → tocontrol-flow edge, sampled out of band, against a running thread, unprivileged (the kernelswfiltbit makes user-only sampling open atperf_event_paranoid=2) and without perturbing the target — exactly the case the single-step views are dangerous on (a live JIT / managed runtime). It needs no external library (rawperf_event_open+ a pure decoder). Public surface:asmtest_ibs_available()/asmtest_ibs_skip_reason()(the full AMD/IBS/BrnTrgt/swfiltdetect-and-skip chain),asmtest_ibs_decode_op()(a pure, host-independent decode of one IBS-OpPERF_SAMPLE_RAWrecord into an edge — unit-tested with synthetic records on every CI host, AMD or not), andasmtest_ibs_survey_pid()(attach IBS-Op to one thread, drain for N ms, return an aggregated hot-edge histogram sorted by count with faithful provenance —samples/branch_samples/lost/throttled). INVARIANT: statistical only — it can prove a block was seen, never that one was not, so it never feeds the exactinsns[]/blocks[]parity contract; it is a separate diagnostic producer, not a member of the exact-trace cascade. Built intolibasmtest_hwtrace; validated bymake ibs-test(also folded intomake hwtrace-test) and the containerizedmake docker-hwtrace-ibs; probe binaryexamples/ibs_probe.c. The live path is validated on an AMD Ryzen 9 4900HS (Zen 2, kernel 6.14): the test captures a spin loop’s back-edge out of band from a separate thread. Plan:docs/internal/plans/zen2-ibs-tracing-plan.md. Phases 2–4 landed subsequently — the whole-process survey, theasmspy --sampleview (headless + TUI mode 7), and the statistical survey fallback; see their own entries above.asmspy— an interactive process tracer (newcli/subsystem, Linux x86-64) — a small ncurses front-end over the out-of-process (ptrace) tracer: attach to any running process and watch it live and out of band. Three live views: syscalls with data (a ministrace; every syscall named from a table generated against the host’s own<sys/syscall.h>,read/writebuffers and path arguments decoded,read/writefile descriptors resolved to their path/socket/pipe via/proc/<pid>/fdlikestrace -y, decoded strings split into their own pane; the syscall stream and the whole-process instruction stream follow every thread of the target —PTRACE_SEIZEof all tasks plusPTRACE_O_TRACECLONEfor threads spawned later, each line tagged[tid]when more than one is followed, and (for syscalls) entry/exit read fromPTRACE_GET_SYSCALL_INFOso seizing a thread mid-syscall never desyncs), a chosen function’s assembly with per-instruction execution heat counts plus its callees ranked by call count (resampled each time the target calls it), and a whole-process live instruction stream (every instruction as it executes, resolved to its function). The two log feeds (syscall log, live stream) pause + scroll back through their history (spaceto freeze,↑/↓/PgUp/PgDn/Home/Endto move,End/spaceto resume the tail), and scrollback survives target exit. The process picker filters as you type and sorts by pid, recent CPU activity, or string-scan density (Tabcycles,rrescans,bnavigates back). Every view is also a headless subcommand for scripts and CI:--list [active|scan],--syms <pid> [filter],--log <pid> [n],--trace <pid> <sym|0xADDR[:LEN]> [n](an explicit0xADDR:LENrange reaches stripped code or a JIT region no symbol covers),--stream <pid> [n]; a negativenruns until the target exits, and malformed arguments are rejected up front. Built bymake cli(needs libncurses + Capstone; self-skips with guidance) or containerized viamake docker-cli(Dockerfile.cli); carries its own/proclister and ELF.symtab/.dynsymfunction resolver. End-to-end headless smoke (cli/cli_smoke.sh,make cli-smoke) drives all five subcommands against the example victims and is gated in CI (clijob). Guide:docs/guides/tracing/asmspy.md.asmtest_trace_call_auto— the auto-escalating, call-owning cross-tier trace, now in all ten bindings. A single entry point that owns the invocation and traces a native routine under the fastest exact tier, then automatically escalates to a ceiling-free tier when the trace comes backtruncated— walking the ladder fast HWTRACE backend → MSR-direct AMD-LBR rung → BTF block-step → per-instruction single-step until the capture is complete (or the tiers are exhausted, transparently flagged). This closes the “arm → detect truncation → re-resolve → re-run” loop that was previously only a documented idiom. Landed C-first (src/trace_auto.c, the MSR rung folded in later), then wrapped in every binding (python/cpp/rust/zig/node/java/dotnet/ruby/lua/go), removing its formerALLparity exemption.*usedreports the tier that produced the final trace, so a caller can see whether escalation fired. Covered bytest_call_auto*in the hwtrace suite.asmtest_hwtrace_arm_tidwrapped in the remaining seven bindings (go, java, lua, node, ruby, rust, zig) — the §0.2 thread-scope assert accessor (the OS thread id that armed the active hardware-trace capture,-1if none) was python/cpp/dotnet-only; it is now surfaced in all ten bindings with an idiomatic accessor (HwTrace.armTid()/arm_tid/HwTraceArmTid()per language), closing its seven per-binding parity exemptions. Each wrapper was built and its hwtrace test suite run green in the per-language docker lanes.Whole-window attribution, version-aware render, and async-hop merge in the Node and Java bindings — dotnet-parity Phase 2, the remaining CI-runnable clusters. Wraps six .NET-lead C symbols across both bindings:
CodeImage.renderVersioned(when, trace)(asmtest_hwtrace_render_versioned) — disassemble a trace’s absolute addresses against a code-image timeline AS OF a capture sequence, not live memory. Version-aware (unlikerender_window): tracking a region asaddthen rewriting it tosubrendersaddat the old sequence andsubat the new. PlusNativeTrace.appendInsn(wrapstrace_append_insn, a non-tier symbol) to build such an absolute-address trace.HwTrace.regionName/symbolizeBuckets/attributeWindow(asmtest_hwtrace_region_name/_symbolize_bucket/_attribute_window) — whole-window noise attribution: reverse-resolve an address to its mapped-region name, bucket a list of IPs by JIT symbol (perf-map) or region, and attribute a live whole-window capture’s absolute addresses to caller-named regions first (so two identical-byte leaves in distinct mappings split into separate buckets — what symbol/disasm attribution cannot do).AddrChannel-free; range classification, no Capstone.HwTrace.stitchHandles(hops, …)(asmtest_hwtrace_stitch_handles) — the §D0.4 async-hop merge: order N already-captured hop traces byseqand concatenate into one logical trace with per-hop slice bounds. Host-independent (pure merge — runs on every lane, arm64 included); the hops must outlive the call (shallow-copy, not duplicated).asmtest_hwtrace_stitch(the C core) stays binding-internal.Struct marshalling is pinned to the exact SysV layouts (
bucket_t136 B,slice_bound_t32 B,named_region_t80 B) and cross-checked against the dotnet[StructLayout]s. Validated in thedocker-hwtrace-node/-javalanes against the C oracles (test_render_versioned,test_symbolize_bucket,test_wholewindow_buckets,test_stitch_slices). All sixALLallow-list lines stay (seven-eight bindings still don’t wrap them);trace_append_insnis a non-tier symbol (ungated).
Crash-proof WHOLE-WINDOW out-of-process capture in the Node and Java bindings — dotnet-parity Phase 2, increment 3, the out-of-process analog of the in-process
window()form (which single-steps the calling thread and is fatal for arbitrary managed code). Wrapsasmtest_ptrace_trace_window_call(Ptrace.windowCall/HwTrace.ptraceTraceWindowCall— fork-internal: a forked child runs the window frame and is stepped, so it asserts unconditionally on any ptrace lane) andasmtest_hwtrace_stealth_trace_windowed(HwTrace.stealthWindow— a helper child reverse-attaches and steps the calling thread’s window body out of band, mirroring dotnet’sAsmTrace.Window; self-skips on a refused reverse-attach), plus the fiveasmtest_addr_channel_*FFI shims behind a newAddrChannelclass. Pre-publish the code regions the window frame calls into (its leaves/methods) on the channel; the capture records the frame plus every published region as ABSOLUTE addresses (classify by range — no Capstone), stepping over everything else. Validated in thedocker-hwtrace-node/-javalanes against the C oracle’s driver-blob ceremony (a 35-byte frame calling two 7-byte leaves): resultm2(7,3)==4, driver + both leaves recorded in call order, complete. TheALLallow-list lines for both windowed symbols stay (seven bindings still don’t wrap them); the addr_channel shims live in a non-tier header (ungated).Crash-proof out-of-process stealth capture (
stealthTrace) in the Node and Java bindings — dotnet-parity Phase 2, increment 2.HwTrace.stealthTrace(code, a, b)(Node) /HwTrace.stealthTrace(NativeCode, long...)(Java) wrapasmtest_hwtrace_stealth_trace: a helper child reverse-attaches (PR_SET_PTRACER+PTRACE_SEIZE) and single-steps the native leaf out of band, so noEFLAGS.TFis ever armed on the runtime’s own (V8 / JVM) thread — the crash-proof counterpart to the in-processcallScoped/windowforms, mirroring dotnet’sAsmTrace.Method(..., outOfProcess: true). Theresultis exact (the helper reads the caller’s RAX at theret); the instruction stream is best-effort over a live runtime (its async signals can truncate the per-instruction walk — faithfully reported viatruncated), so the tests assert the exact[0,3,6,c,11]stream only when not truncated. Validated in thedocker-hwtrace-node/-javalanes; self-skips cleanly where a Yamaptrace_scoperefuses the reverse-attach.AMD LBR Zen 4/5 coverage: slot-efficient branch filtering (#2B), period-spaced stitch validation (#2A), and single-exit snapshot-by-default (#3). Three improvements that stretch how much of a routine each 16-deep AMD branch-record window reconstructs, all respecting the silicon ceiling and the “never emit corrupt as complete” rule.
#2B slot-efficient branch filtering (opt-in, SCOPE-SAFE). New
asmtest_hwtrace_options_t.branch_filter(default 0 =PERF_SAMPLE_BRANCH_ANY, unchanged). Nonzero requests a reduced HW filter (COND | IND_JUMP | ANY_CALL | ANY_RETURN) that drops only the direct unconditionaljmp— its target is statically decodable, so it need not consume a scarce LBR slot — and the reconstructor follows it from the region bytes for a byte-identical trace over a longer window. Dropping direct call too was deliberately rejected (an out-of-region-callee return strands the pre-call in-region code — a silent-corruption risk). The decoder is unified/no-flag:amd_replayfollows a dropped jmp only when one appears mid-straight-line-walk, which under the default full filter can never happen (a taken jmp is the recordedfrom), so the follow path is provably dead code on the tested default. New primitiveasmtest_disas_is_uncond_jump; the capture retries the full filter onEOPNOTSUPP/EINVALso the tier stays available. Applies to both the sampled and the deterministic-snapshot paths (the statistical WindowHot survey keepsBRANCH_ANY). Host-independently validated (test_amd_reduced_filterF1–F5: dropped-jmp equivalence, back-edge-cycle termination, region-exit truncation, chained follow); two independent adversarial reviews confirmed the classify/follow logic exhaustive over every x86-64 CTI. Live-validated + reach-measured on Zen 5 (Ryzen 9 9950X,test_branchsnap): the deterministic snapshot withbranch_filter=1follows a dropped jmp to its target block on real LbrExtV2, and reconstructs 1.86× more executed instructions per 16-deep window (65 vs 35) on a loop whose body has a direct jmp plus a conditional back-edge.#2A period-spaced Tier-B stitching — host-independent validation + documented caveat.
test_amd_stitch_period_spacedproves the landedlbr_periodpath stitches period-spaced (P=4) windows of a distinct-edge path back to the exact sequence, and asserts the flip side: a self-similar loop silently undercounts underperiod>1(the smallest-overlap heuristic can’t tell 1 iteration from P) — which is why the default stayslbr_period=0(period=1, universally exact). Live-measured on Zen 5 (test_amd_reach_period): confirms the finding on real hardware —period=4reconstructs fewer instructions thanperiod=1(231 vs 297) on a loop, since every loop is edge-self-similar, so period-spacing’s reach benefit is confined to (inherently short) distinct-edge paths, not loops.#3 deterministic snapshot by default for single-exit regions.
hwtrace_begin_amdnow selects the Phase-3 boundary snapshot by default on the supporting substrate (amd_lbr_v2+perfmon_v2+ Linux ≥ 6.10), but only when the region has a lone ret (amd_last_ret_offnow counts rets) — the one exit breakpoint is then guaranteed hit, so the common small routine gets deterministic capture with no richest-window guessing. Multi-exit routines (which an earlier ret could make the breakpoint miss) keep the sampled path; explicitopts.snapshotis honored for any region and every arm failure falls through to sampling. Validated:docker-hwtrace-amd(328 decoder checks green) anddocker-hwtrace-codeimage(branchsnap marker path green on the Ryzen 9 9950X). Onlybindings/dotnetmirrors the newbranch_filterfield (matching the shippedlbr_periodposture); the field is an ABI-safe tail append (struct stays 48 bytes).
Whole-window scope (
begin_window/end_window/render_window) in the Node and Java bindings — the region-free, empty-ctorusing (new AsmTrace())§Z1 substrate (Phase 2, increment 1 of the dotnet-parity roadmap).HwTrace.window(fn)(Node) /HwTrace.window(Runnable)(Java) arm a single-step capture on the calling thread with NO registered region, run the body, disarm, and render the executed absolute addresses from live memory — returning{path, truncated, insns[]}. It is FAITHFUL-BUT-NOISY by design: single-stepping the managed runtime records everything between begin and end (the FFI dispatch + runtime), so the traced routine’s own addresses appear as a subset. A single V8-dispatched call runs ~100k instructions (captured cleanly, subset verified); a HotSpot + FFM call exceeds the single-step whole-window’s internalSS_WINDOW_CAP(1<<20), so the Java capture faithfully reportstruncated(best-effort). Validated indocker-hwtrace-node/-java. Thebegin_window/end_window/render_windowALLexemptions stay inscripts/bindings-parity-allow.txt, now consumed by the seven bindings that don’t wrap them.call_scoped— a registry-free traced native call — now in ALL TEN bindings. The Python/Ruby/Node/Java bindings shipped it first; the remaining five (C++, Rust, Zig, Lua, Go) now wrap it too. Each wrapsasmtest_hwtrace_call_scoped_ex+asmtest_hwtrace_render_scope: arm, call the native leaf, and disarm entirely in native code — a tighter window than thescopeform (whose FFI dispatch ofcode.callis stepped, though region-filtered) — returning the call’s result, the executed body’s disassembly, and the truncation bit in one step. Registry-free, so it is safe in a tight loop (noMAX_REGIONSexhaustion).HwTrace.call_scoped(code, *args)(Python/Ruby),HwTrace.callScoped(code, …args)(Node),HwTrace.callScoped(code, long…)(Java),HwTrace::callScoped(code, args…)(C++),HwTrace::call_scoped(&code, &[args])(Rust),HwTrace.callScoped(&code, args)(Zig),HwTrace.call_scoped(code, ...)(Lua),CallScoped(code, args…)(Go); each returns{result, path, truncated, rc}and each is validated in its Docker lane (result 42, body renders toretin 5 insn lines, a 40-call loop with no exhaustion). Struct-by-value for the 8-byteasmtest_hwtrace_scope_thandle is native in the five new bindings (C++ POD, Rust#[repr(C)], Zigcallconv(.C), LuaJIT FFI, cgo) — no packing, unlike the Ruby/Java bridges. With all ten now wrapping the pair, theALLexemptions forcall_scoped_ex/render_scopeleavescripts/bindings-parity-allow.txt.§D0.4 async-hop stitching now has a LIVE producer —
AsmStitchedTrace(.NET). The shippedasmtest_hwtrace_stitchmerge core previously had no live producer (only synthetic-slice host tests). Newasmtest_hwtrace_stitch_handles(traces[], scope_ids, seqs, tids, versions, n, out, bounds, nbounds)(src/hwtrace.c) is the binding-facing bridge — it merges N already-captured trace handles (the slice struct embeds heap pointers a binding can’t marshal by value). On top of it, the .NETAsmStitchedTracecarries anAsyncLocal<scopeId>acrossawait/thread hops and feeds each hop’s managed-safe lazy-arm capture to the core, so one logical operation traced across a realTask.Runthread hop stitches its per-thread slices in seq order. Each hop uses the new registry-freeasmtest_hwtrace_call_scoped_ex([base,len)direct, noMAX_REGIONSslot) so a long-running operation with many hops cannot exhaust the fixed 32-slot region table process-wide. Validated on the single-step tier (no Intel PT needed): hosttest_stitch_handles/test_call_scoped_ex(incl. a 64-call no-exhaustion check) and the .NET lane (scope id flows across the hop; two different-thread hops merge with correct bounds; 40 operations all capture).FP shim family for the lazy-arm scope —
(double…)->doublemethods trace in-process.asmtest_hwtrace_call_scoped_fp(src/ss_backend.c/src/hwtrace.c) dispatches a homogeneous double signature through the SysV FP ABI (xmm0..7 args, 0-8 arity). The .NETAsmTrace.Method(...).Invokenow tries the integer(long…)->longshim, then the FP family, before falling back out-of-process — so adouble-signature method is captured in-process instead of degrading. Host-tested (test_call_scoped_fp) and on the .NET lane. See managed-singlestep-lazy-arm-plan.md.Managed single-step is now safe by construction —
AsmTrace.Method()lazy-arms only the method body. Newasmtest_hwtrace_call_scoped(name, fn, args, nargs, result, out)(src/ss_backend.c/src/hwtrace.c) arms the single-step window, calls the target through the SysV integer ABI, and disarms — all in native code — so the region filter keeps only the body’s offsets and NONE of the caller’s or a managed runtime’s machinery is ever underEFLAGS.TF. The .NETInvokeno longer stepsDynamicInvokein-process (the crash surface where an in-windowpthread_createthat blocksSIGTRAPforce-killed the process on slow hosts): it marshals through a(long…)->longshim table and, for signatures the shims can’t express, auto-falls back to the out-of-process stepper with a loudSkipReason— never a silent miss.HwTrace.DegradationNote()gains the faithful managed-window warning. Host-tested (test_call_scoped, byte-for-byte parity with begin/end) and validated on the .NET lane; see managed-singlestep-lazy-arm-plan.md.Slow-host crash-avoidance stress lane (
make hwtrace-dotnet-stress, CI:docker-hwtrace-dotnet-stressin thehwtrace-bindingsjob) — the lazy-arm plan’s “Sharpening 1”. The ONE lane that runs with CoreCLR’s tiering worker unpinned (noDOTNET_TC_BackgroundWorkerTimeoutMs): it parks past the worker’s idle-exit, churns tier-up enqueues on the invoking thread (freshDynamicMethods driven past the call-count threshold), and interleaves lazy-armInvokes — recreating on the loaded CI runner the exact environment where the old stepped-DynamicInvokepath died with exit 133. Surviving with every capture intact is the pass signal.The zig toolchain tarball is now integrity-pinned — the one third-party fetch P2’s supply-chain pass left unverified.
DOCKER_SETUP_zigverifies a per-arch sha256 (ZIG_SHA256_x86_64/_aarch64inmk/docker.mk) before extracting, the anchors are recorded inscripts/third-party-digests.txt, andcheck-thirdparty-versions.shnow asserts both anchors exist for the declaredZIG_VERSION— so a version bump that forgets the digests fails loudly.Example suites are now auto-discovered. Every
examples/test_foo.c+examples/foo.spair (foo.asmunderASM_SYNTAX=nasm) links through atest_%pattern rule — drop the two files in andmake testpicks the suite up, with no Makefile edit. Legacy pairs whose routine object doesn’t match the test name (test_arith→add.o,test_capture→flags.o,test_struct→structs.o) keep explicit link rules, andSUITE_EXCLUDESlists thetest_*.cfiles owned by other targets (bench, usecases, demos, the emulator/trace tiers). This makes the long-standing docs claim in writing-tests.md true instead of correcting it downward.asm_call_capture_vec256_win64andasm_call_capture_vec512_win64are now declared inasmtest.h(under-DASMTEST_ABI_WIN64), completing the Win64 mirror of the System V capture surface. Both existed insrc/capture_win64.asmand were exercised by the Win64 suite, but a consumer following the win64 guide had to hand-declare the prototypes; the guide’s entry-point table now lists_vec512_win64too.Wide-arity, mixed-FP, and struct-return capture reachable from all ten bindings (N4 of the 2026-07-04 review — previously the array-form C entry points existed but no binding referenced them). Three FFI-friendly shims join
asmtest_capture6/_fp2/_vec_f32in the opaque-handle layer:asmtest_capture_args(stack-spilling wide arity),asmtest_capture_mix(integer + FP register files together), andasmtest_capture_sret(hidden-pointer struct return). The struct-layout bindings (C++/Rust/Zig/Python) call the array forms directly per their existing idiom; the opaque-handle bindings (Node/Java/.NET/Ruby/Lua/Go) wrap the shims. Every binding gained wide-arity (sum8), mixed (mix_scale), and struct-return (make_big) conformance tests — all ten docker lanes green; fixtures registered in the corpus name table (no repeat of N7); NASM counterpart included.Docs: Teaching with asm-test (the in-repo scope of P5 from the 2026-07-04 review). The instructor recipe the primitives always supported but nothing documented: a three-file assignment layout (student
.s, rubricgrade.c, grade-timeMakefile), a complete GitHub Classroom autograding config (one scored step per rubric item via--filter+--fail-if-no-tests, so a deleted rubric test is a scored zero, not a free pass), and instructor notes on hidden tests, timeouts, and grading non-x86 courses through the emulator tier. The separate “Use this template” assignment repository remains a maintainer action..NET examples roadmap — the full remaining tail (11 items from dotnet-examples-roadmap.md, all instruction-count-faithful, all green in the docker lane). Five new reports:
flatprofile(perf-report parity: self / Overhead % / cumulative %),amplification(user vs BCL vs native-runtime split + the WEAK-tier factor),runtimegaps(largestRuntimeBeforebursts by the method they precede),footprint(code working-set pages + jump-distance locality), andruntimebuckets(the ~1M-insn runtime lump named by module — resolved per 4 KB page, not per address, so ~hundreds of/proclookups instead of ~1M). Six new example projects:instructionmix,perfannotate,loops(backedge trip counts),descent(native call-descent tree with self/inclusive counts),descent_dotnet(out-of-process call descent into a live CoreCLR — descendsProgram::Leaftwice as nested frames;jit_dotnetgained an additivechainmode), andcodeimage(one address, two code bodies over logical time). The binding gainedHwTrace.SymbolizeBucketsover the already-exportedasmtest_hwtrace_symbolize_bucket(.NET suite 123 → 126).Consumer-facing CI integration (P2 of the 2026-07-04 review). A composite GitHub Action at the repo root (
action.yml, “Setup asm-test”: POSIX-sh steps; inputsversion/prefix/optional-tiers/test-command; exportsPKG_CONFIG_PATHand library paths), an includable GitLab CI template (ci/asmtest.gitlab-ci.yml,.asmtest-install+ a documented consumer job with JUnit wiring), and a CI integration guide covering both plus the rawmake installfallback. The wrapped install recipe is proven end-to-end locally; the Action’suses:path needs a real Actions run.macOS clean-room plan — Track E finished, Tracks C/D written.
release.yml’s seven smoke blocks nowsource scripts/clean-env.shinstead of ad-hoccd /tmp && env -uscrubbing (behavior-preserving; interpreters resolved to absolute paths before the PATH scrub), and the methodology is documented in docs/clean-room-testing.md. Tracks C (scripts/osx-vm.sh+make osx-vm-test, tart VM) and D (scripts/docker-osx-bindings.sh+make docker-osx-bindings, Docker-OSX/KVM) are written per the plan’s spec and clearly banner-marked UNVALIDATED — they need Apple-Silicon-tart / bare-metal-KVM hosts this environment lacks.AMD tracing plan Phase 2 & 3 follow-ups — attached block-step + snapshot marker routing. Completes the two sub-items the earlier block-step / snapshot commits left open:
asmtest_ptrace_trace_attached_blockstep— the third public block-step symbol. Block-steps a SEPARATE, externally-attached process (one debug exception per taken branch, intra-block instructions reconstructed with Capstone), reading foreign bytes viaprocess_vm_readvand leaving the target stopped past the region for the caller — the rootless managed-runtime completeness fallback. Wrapped in all ten bindings; a newtest_ptrace_attach_blockstepasserts the stream is byte-identical to the per-instruction attached tracer over a true external attach.opts.snapshotbegin/end routing on AMD — the deterministic boundary LBR snapshot (bpf_get_branch_snapshotat a region-exit hardware breakpoint) is now reachable through the ordinarybegin/endmarkers, not just the standaloneasmtest_amd_snapshot_trace. The capture split intoasmtest_amd_snapshot_begin/_end(armed single-slot); the AMD marker path derives the exit from the region’s lastretand falls back to thesample_period=1sampled path when the BPF toolchain/caps/LbrExtV2 substrate is absent.
AMD hardware-trace improvements — Phases 0, 4, 5 of the AMD tracing plan. Completes the P0/P1 near-term work on the AMD LBR backend, all validated live on the Zen 5 dev box (Ryzen 9 9950X,
amd_lbr_v2) viamake docker-hwtrace-amd:Phase 4 — LbrExtV2 speculation-bit filtering.
amd_replaynow drops aperf_branch_entrywhosespec == PERF_BR_SPEC_WRONG_PATH(a speculative, never-retired phantom edge) before reconstruction; dropping it is expected, so it does not settruncated. Thespecfield (Linux ≥ 6.1) is gated behind a-fsyntax-onlystruct-member build probe (ASMTEST_HAVE_PERF_BR_SPEC), so the filter compiles out cleanly on older headers / Zen 3 BRS.amd_edge_eq(the stitcher’s from+to overlap key) is untouched.Phase 5 — Tier-B stitch hardening.
asmtest_amd_stitchgained a decodable-distance guard: a smallest-overlap match is accepted only if the adjacency it splices is real straight-line code, so a dropped/throttled-sample mis-stitch becomes a faithful gap instead of a silently-wrong trace. The AMD data ring default grew 64 KB → 256 KB to extend gapless stitch reach before the kernel drops samples; thedata_sizeheader comment now documents both backend defaults.Phase 0 — runtime branch-stack depth.
asmtest_amd_lbr_depth()reads the true depth from CPUID0x80000022EBX[9:4] (lbr_v2_stack_sz), replacing the hardcoded 16 in the Tier-A/Tier-B split, stitch bound, and LOST heuristic. A no-op today (every shipping Zen reports 16) that removes the assumption.Phases 6 (Zen 3 BRS period-adjust) and 7 (IBS-Op coverage) remain forward-look — they require Zen 3 / Zen 2 silicon the dev box lacks, and the project does not ship hardware code it cannot self-validate.
Docker-OSX clean-room lane (Track D): containerized
sshpass, repointed at surviving upstream tags. NewDockerfile.sshpass+make docker-sshpassbuild a smallasmtest-sshpassimage;scripts/docker-osx-bindings.shnow runs every ssh/scp-equivalent call through it instead of requiring a hostsshpassinstall (and thesudothat would need), per CLAUDE.md’s “add it where the work runs” rule. Separately,sickcodes/docker-osxdeleted every tag but:latest/:masterfrom Docker Hub in 2024 (:venturaand friends now 404) —DOCKER_OSX_IMAGEdefaults to:latest, and the script gainedDOCKER_OSX_DISKsupport (-v <disk>:/image -e IMAGE_PATH=/image) plus a one-time-install recipe in its header, since a virgin:latestboots the macOS installer rather than a headless system. See macos-cleanroom-lanes.md.
Changed¶
Selection is now one shared brushing-and-linking model — a pick in any pane cross-highlights the same entity in detail/disasm/Loom/3D at once (docs/internal/archive/gui/22-selection-and-search.md T1, F7). Every view used to hold its own selection, so an analyst re-found the same address by hand in each pane. Selection is now ONE Workspace/shell-level entity (
{rec, step, offset, lane}plus a bumpedepoch), held distinctly from navigation (nav.currentpoints a view; the selection brushes an entity): a pick in any pane — the timeline, the slice explorer, the Loom, a 3D drill — cross-highlights that same entity in every pane it appears in, and only there. A pane that cannot show the entity shows nothing selected rather than a fabricated row (D7), and cross-highlighting brushes in place without yanking every pane’s viewport.Fidelity chrome is now a graded 3-tier system over a derivable
severityfield (docs/internal/archive/gui/23-graded-truth-layer.md T1, F5). The proliferating fidelity forms — a redaction placard, a statistical chip, a coarse chip, a bounded-window note and a torn banner, all equally loud and non-collapsible — collapse into ONE vocabulary: one banner, one inline chip, one glyph set, with mandated placement (a banner is pane-top, a chip is on a header row, enforced by the component API). Loudness follows the schema’s own severity gradient: neutral (skip / statistical / redacted / coarse / bounded) is a quiet chip; caution (truncated-but-usable / paused gap) is an amber banner that collapses to its chip after first read; integrity (torn / mixed-basis refusal / a drop on an exact capture) is a loud red banner that never collapses. This restructures the fidelity layer and removes no truth — every field still renders, graded against the committed low-fidelity fixtures (D7).The live session-end state is now a persistent, cause-distinguished in-pane placard with an inline fix (docs/internal/archive/gui/23-graded-truth-layer.md T2, F20). The single collapsed “ended” is fanned into its cause — stopped-clean, torn (host crashed), torn (EOF), or a PROTOCOL-MISMATCH (a stale
build/asmspythat printed a usage banner and exited 0) — each with the trust of the on-screen data, and the protocol-mismatch placard carries the verbatim one-line fix (make cli; Disconnect + reconnect). It persists across frames until the next Connect/Start; toasts (doc 16) supplement it, never replace it.“Paused” is split into an operator pause and a budget block (docs/internal/archive/gui/23-graded-truth-layer.md T3, F23). The bare word named two states with disjoint recoveries; now an operator pause reads “PAUSED (you) — Resume” and a budget preemption reads “BLOCKED — jack held by <session> on <target>” with explicit Swap (a named two-step confirm), Queue (a cancellable chip that starts when the jack frees), and Cancel. No path auto-swaps without confirmation.
Desktop view tabs are now data-driven, and first run is a task rail (docs/internal/archive/gui/20-workspace-and-settings.md T1/T2). Only the views a recording can actually fill are shown — a bare-log recording no longer presents empty Loom / ABI-x-ray / 3D / Scrubber tabs; the views it cannot fill collapse into ONE faithful “unavailable views (N)” affordance that still names each absent view and its verbatim machine reason (D7 — restructured, never removed), and the set is scoped by the active mode. First run replaces the “choose a door” chooser with a persistent task rail (Learn how assembly runs / Open a trace / Capture a live process / Author a routine); an empty workspace auto-lands in Learn (the dependency-free path), and the chosen mode drives its dock perspective so the label and the layout agree. Resolves the plan-says-3-doors vs build-renders-4 drift (F13): there are four task modes, named as tasks, and the “door” vocabulary is retired.
CVD-safe categorical palette + a second channel on every colour-coded distinction (docs/internal/archive/gui/24-one-visual-language.md T2). Every categorical distinction now also carries a NON-colour channel so ~5% of users who cannot read the hue still read the axis: cone direction gains a shape glyph beside each node (◄ inflow / ► outflow / ● selection / · off-cone), the Loom take axis reads by pattern (solid = hot / hollow = dim / dashed = unaligned) with a named inline legend, and the src×dst hot-edge heatmap uses a CVD-safe, perceptually-uniform ImPlot colormap (Viridis) whose labelled
ColormapScaleis the magnitude channel. A pureui/cvd.{h,cpp}simulates protan/deuter/tritan and computes WCAG contrast; the palette is verified in test (text ≥4.5:1, fills/borders ≥3:1 at the smallest font), and the caution amber (dt_warn) is marked large-text-only. The shared legend renders from ONE encoding table, so the legend is itself the proof no distinction rides on colour alone.One filter affordance + one time-position widget across the desktop views (docs/internal/archive/gui/24-one-visual-language.md T4). A single type-to-narrow filter with a “showing N of M” count (
ui/filter.h) replaces the ad-hoc client-side idioms, free ImGui column-sort landed on the hot-edge table (reordering the view, never the recorded model order), and ONE time-position widget (ui/timepos.h) now carries two faithful variants — a continuous scrub where a real total exists (the Loom/3D playheads) and a discrete step where it does not (Invocations, the disassembly logical-time control), the discrete case VISIBLY MARKED as an intentional fidelity choice with its verbatim reason. The counts, sort order and discrete-reason registry are asserted headlessly.The capability panel leads with what the host can do (docs/internal/archive/gui/18-breach-stops.md T4). It opened with a wall of red errno rows; it now leads with a one-line positive summary (the available backends) and the standing “Learn and Author work here — no root, hardware, or attach needed” floor, and demotes the unavailable backends under an expandable “why can’t I capture X?” that keeps every verbatim machine reason and adds the shared
attach_verdictremedy (paranoid / Yama / i386 / CAP_SYS_PTRACE) where one applies. A bare host no longer reads as “the tool does not work here”.The keyboard-help overlay marks planned-vs-wired bindings (docs/internal/archive/gui/18-breach-stops.md T1): it is generated from a
wiredflag and greys any not-yet-mapped binding “planned” instead of advertising it as live.The desktop views are now real dockable panes — the docking layout manager, presets and Reset act on visible windows (docs/internal/archive/gui/19-dockable-panes-keystone.md). Before this, the layout manager docked five named windows (
Home,Recording,Scrubber,Inspector,Timeline) and the View menu offered Reset + three presets, but no view wasBegin()’d under any of those names — the dockspace, every preset, tear-out and Reset acted on phantom windows and were inert, and the whole view surface was a single window nesting three exclusive tab levels (recording →views→ observer), so only one view was ever visible. The shell nowBegin()s each region as a real pane hosting the active recording, the inner exclusiveviewstab bar is deleted, and the 3-deep nesting is flattened to at most two (a pane and, for the Loom and the Observer deck, their one data-gated inner bar). The concrete payoff: the timeline, the scrubber and the Observer’s disassembly can be shown at the same time. View-menu presets and Reset now rearrange visible panes, and panes can be torn out and restored — the bottom region was split so a preset holds the timeline and the scrubber together, and Reset rebuilds the default split (the recovery path for a stale/corrupt persisted dock.ini, whose auto-fallback lands with doc 18 T2.2). Every view keeps its exact body and its fidelity placards — the chrome was restructured into the panes, never removed (D7). The non-docked path (the null test backend’s default) still draws the single-window tab layout unchanged, so nothing regresses without a dockspace. The panes are asserted headlessly (make desktop-testflips docking on and checks each region window exists and is active, that three siblings are active at once, and that a preset switch moves real panes) and the tear-out → Reset round-trip is driven by the interaction lane (make desktop-ui-test).One semantic colour palette in the desktop’s
theme.h(docs/internal/archive/gui/24-one-visual-language.md T1). The good / bad / maybe / changed / cone (back·fwd·both·off) / selected / statistical axis was being re-invented inline in every draw file — three barely-distinguishable yellows meant three different things and two reds both meant “refused”. They now live in one place as named accessors (dt_good_col…dt_statistical_col, each with a paired_u32), each documenting its ONE meaning, so a colour can no longer drift its meaning between panes;dt_bad_col()is now literallydt_refuse_col()(the 0.90-vs-0.95 refuse-red split is gone). Every per-pane colour literal at the drift sites — the Inspect verdicts, the scrubber / ABI-x-ray “changed” highlight, the slice cone hues, the 3D-HUD chips, and the Loom refusal placard — is deleted and routed through the accessors. A sharedui/legendcomponent renders the palette (with a non-colour second-channel token beside each swatch) so a legend can never disagree with what a view draws; the slice explorer uses it. The fidelity chrome (D7) is unchanged — the refuse red and caution amber keep their meaning;statisticalis merely named so a later graded-fidelity change can move it off the amber without touching a call site. Header-only and engine-free, so it links into the full app, the render-only viewer, and the null test backend alike.AMD manual pre-release validation shrunk to the runner-uncoverable residue (self-hosted-ci-runners.md T4).
docs/internal/amd-hardware-validation.mdwas the “one validation step that cannot run in CI”; now that the self-hostedhwtrace-privileged-zenlane runs the exact LbrExtV2 + live-IBS paths on a Zen 4/5 runner, that tier moved to CI and the doc is reframed as the residue — the four AMD paths the Zen 4/5 runner cannot reach, each with its command, hardware/privilege gate, and owning doc: Zen 2 IBS-without-LBR degradation (make docker-hwtrace-ibson the Ryzen 9 4900HS), MSR-direct (make docker-hwtrace-msr,--privileged+ hostmsrmodule — kept off the CI runner by security policy), thestatuslive-EPERM path (unreachable root-in-container), and Zen 3 BRS (a link to amd-branchsnap-lbr-docs.md#T8, not a checklist item). Thecall_autonon-escalation regression signal is now enforced by the CI lane’s assert rather than a manual eyeball.Linux Python wheels now build on the
manylinux_2_28floor (install on older distros). The two Linux legs of the releasepythonjob build insidequay.io/pypa/manylinux_2_28_{x86_64,aarch64}(AlmaLinux 8, glibc 2.28) instead of the ubuntu-latest glibc, viascripts/build-manylinux-wheel.sh— which source-builds the four native engines the image lacks (unicorn/keystone/capstone/libipt, pinned; libopencsd is a dead link-only dep and skipped) + fetches DynamoRIO, runsmake python-package, andauditwheel repair --plat manylinux_2_28with the load-bearing tier--excludelist.make docker-python-manylinuxproves it end to end with no credentials: the manylinux_2_28 wheel installs and imports (asm + disas) on a clean AlmaLinux 8. (distribution-packaging.md T5.)The out-of-process whole-window stepper block-steps where
PTRACE_SINGLEBLOCKis functional (~4–10× fewer stops). The §D3 stealth whole-window helper now drivesasmtest_ptrace_trace_attached_window_stop_blockstep(one#DBper taken branch) instead of one stop per instruction, degrading to the exact per-instruction stepper on aDEBUGCTL.BTF-masking hypervisor (asmtest_ptrace_blockstep_available()false). The output is byte-identical either way — a cost upgrade, not a fidelity one — which is what makes a managed whole-window affordable out of process.ASMTEST_STEALTH_NO_BLOCKSTEP=1forces the per-instruction stepper.Registry-publish pipeline moved toward keyless publishing (scaffolding; gated on registry setup + a real tag).
release.ymlnow publishes PyPI via a dedicated OIDC Trusted Publishing job (pypa/gh-action-pypi-publish, collecting every matrix leg’s wheel — the action is Linux-only) and crates.io viarust-lang/crates-io-auth-action, both with job-scopedid-token: writeand no stored token; npm publishes with--provenance.bindings/java/pom.xmlgains the Maven Central metadata + dormant source/javadoc/gpg/central-publishing plugins (activated only by a realmvn deploy;make java-packagestill uses javac + jar). The manylinux wheel floor is recorded asmanylinux_2_28. All of these are credential/registry-gated — a trusted-publisher registration (PyPI/crates.io),MAVEN_*secrets + a namespace/PGP key (Maven Central), and a CI dispatch (manylinux) must land before they go live; see releasing.md.asmtest_pt_decode_windowgained a trailinguint64_t *base_ip_outparameter (src/pt_backend.c) reporting the first decoded IP, so the whole-window PT drain can re-base its recorded offsets to ABSOLUTE addresses. Source-incompatible for a direct C caller of this internal decode entry (passNULLto keep the prior offset-origin behavior); the facade (asmtest_hwtrace_begin_window/_end_window) and every language binding are unaffected. See intel-pt-whole-window-substrate.md.Internal plan docs reconciled against the code; four completed plans archived. An audit of all 20 active plans against the source, Makefile lanes, CI, and git history found ~30 stale status markers whose drift was entirely one-directional — every one under-reported, marking shipped and tested work as “planned” or “forward-look”; nothing claimed landed was missing. The mechanism was visible in the artifacts: status was recorded by appending a dated block while section headers and tables kept their original marker. The markers are corrected (provenance preserved), and the four complete/closed plans move to
docs/internal/archive/plans/per the repo convention: live-attach data-flow (7/7), DynamoRIO taint tier (9/9, band-gated at ~11x bare), Zen2 IBS (8/8), and the managed-attach safepoint spike (closed NO-GO). Two corrections are load-bearing rather than cosmetic:data-flow-tracing-plan.md’s Phase-5 stub told readers the taint tier stopped at Increment 3 when all nine had landed; and F4’s blocker was retired — both it and Phase 4 stated that live GC-move canonicalization needed an out-of-process EventPipe consumer (“its own lift”), but that was an assumption and it was disproved: an in-processMovedReferences2profiler delivers the exact{old,new,len}triples at a suspended-EE GC fence, is proven to coexist with DynamoRIO, and already ships in the taint tier. F4 is now wiring a proven feed to a landed transform, and is the recommended next milestone in that plan.AMD-LBR reconstruction fills the entry-block prologue (fidelity fix). On a live AMD host a too-fast tiny routine’s frozen branch stack can carry spurious mid-routine landing edges, so
amd_replayanchored at the landing offset and skipped the entry prologue[base_ip, landing)— a complete-reported reconstruction of a small routine undercounted its retired instructions (e.g.insns=4vs the block-step baseline5).amd_entry_fillnow prepends the clean straight-line prologue all-or-nothing (a branch/ret/overshoot in that run faithfully truncates instead), with a symmetric trailing fill; anti-fabrication tests confirm no phantom instructions. Newtest_amd_live_smallroutinehard-asserts complete⇒full-count on a live AMD host. Verified across 3 privileged runs + a 30-iteration loop: every complete reconstruction now yields the full count, and the batch-3 case-(b) escalation invariant is preserved 30/30. (The residual case-(a) advisory from the Zen 5 findings doc is resolved.)Multi-exit deterministic BPF boundary snapshot, default-on (followup Phase 5 / F13).
hwtrace_begin_amdnow plants one hardware breakpoint per region exit (1–4 exits, one debug register each) viaasmtest_amd_all_exits+asmtest_amd_snapshot_begin_multi, so whicheverret/tail-call a multi-exit routine leaves through hits a boundary — the old single-exit gate missed earlier exits and truncated. A BPF-side drop counter drives an faithful truncated-on-drop contract (F13: a dropped ring record marks the result truncated, never silently complete — verified with a 1670-drop overflow fixture). >4 exits or any arm failure falls through to the sampled path unchanged. First live-validated on Zen 5 via the new privileged docker lane.First-class privileged hardware-capture docker lane (
make docker-hwtrace-privileged). Runshwtrace-test ibs-testunder--cap-add=PERFMONalone (no--privileged, noSYS_ADMIN, default seccomp) — the first live validation of the exact AMD LBR (LbrExtV2) and IBS capture paths on the Zen 5 dev box: the previously-skipping AMD/IBS live tests (LBR capture, Tier-B stitch, per-thread concurrent fds,sample_window, IBSsurvey_pid/survey_process) all run and pass (test_hwtrace 389/389, test_ibs 23/23).make BUILD=<abs> test/usecases/valgrindnow work. The suite-loop recipes ran./$$twhere$$talready holds a$(BUILD)/-prefixed path, which broke under an absoluteBUILDoverride (.//tmp/...); dropping the./prefix (the path always contains a slash) completes the earlier out-of-tree-build fix.jit_trace’s JIT lanes prefer the byte-identical block-step rung (review F18). The no-descent lanes selectasmtest_ptrace_trace_attached_blockstepwhenasmtest_ptrace_blockstep_available(), falling back to the per-instruction stepper otherwise; the*-descendlanes intentionally stay per-instruction (block-step has no descent parameter). Verified byte-identical on a live V8 target: the ASLR-normalized disasm stream from block-step matches the pre-change single-step stream exactly.asmtest_amd_snapshot_enddrains the BPF ring without blocking (followup Phase 8 / F15).ring_buffer__poll(rb, 200)epoll-waited 200 ms on the no-hit / faithful-truncation path (the common case) before draining;ring_buffer__consumereads the producer position directly and returns at once — every record is already committed by the time the events are disabled.One cached
amd_lbr_v2cpuinfo probe (followup Phase 9 / F35/F11). The duplicated/proc/cpuinfoparse inamd_backend.candmsr_lbr.cnow shares a single internal cached probe.CI / tooling hardening. The documentation site now builds warnings-as-errors in a dedicated
docsCI job (sphinx -W), so a broken cross-reference fails the PR in-repo instead of only reddening Read the Docs after merge. Theformatgate is pinned toclang-format-18(matchingmake docker-fmt’subuntu:24.04), so it no longer risks flagging the whole canonically-formatted tree as drift the dayubuntu-latestadvances past 24.04.-Werrornow guards thehwtraceandclijobs (previously only the basetest/check), catching warnings in the newest, highest-churn code. A newmake fix-permstarget reclaims root-ownedbuild/artifacts adocker-*lane can leave behind (which otherwise breakmake clean)..dockerignorenow excludes every rootDockerfile*and the actualbindings/Dockerfile.lang(the stalebindings/*/Dockerfileglob matched nothing).Internal working docs moved under
docs/internal/—docs/plans/,docs/analysis/,docs/reviews/, anddocs/archive/are nowdocs/internal/{plans,analysis,reviews,archive}, with one archive rule (done plan / fully-actioned review →archive/; seedocs/internal/README.md). The four completed scoped-tracing plans and the fully-actioned 2026-07-04 review moved toarchive/accordingly, every in-repo reference (docs, comments, workflows, this changelog) was repointed — including a dozen references left stale by the earlier archiving commit — and the Sphinxexclude_patterns/docs-gate now excludeinternal/**wholesale.Docs/README accuracy pass from a full docs-vs-code review. The entry-point pages no longer claim the language packages are published (nothing is on a public registry yet — the release pipeline is ready but uncredentialed); the README slimmed to pitch + highlights + links (the capability list’s single source is now
docs/reference/features.md) and its DynamoRIO link points at the tracing guide;--bench-format=text|jsonand--helpjoined the runner/benchmark flag tables;installation.mdgained Keystone/Capstone rows and the real--emudependency set;integration.mdshows the-x assembler-with-cppassemble step and theasmtest-emupkg-config module;api-reference.mdgainedASSERT_ABI_PRESERVED_VEC/asmtest_check_abi_vec; java.md’s JDK requirement, rust.md’s shipped Tier-2 asserts, Zig’s raw-@cImportstatus, and the single-step tier’s macOS support are stated consistently; and CONTRIBUTING gained “Adding an example suite” and “Building the docs” sections plus a per-language lib-setup cheatsheet in the bindings overview.CI: the pinned Keystone/Capstone source builds are now cached (K1 of the 2026-07-04 review — the ~20-identical-multi-minute-LLVM-compiles-per-push item). Host builds cache via
actions/cache+ a newscripts/thirdparty-cache.sh(exact cmake-installed file set, keyed on OS/arch + the pinned versions); docker builds ofasmtest-bindings-basecache via buildxtype=ghabehind a new overridableDOCKER_BASE_BUILDinmk/docker.mk. ci.yml only —release.ymldeliberately stays cache-free. Non-fatal by design on any cache miss or backend outage; the warm-cache path still needs a real Actions run to confirm.Docs: “asm-test vs. alternatives” (P4 of the 2026-07-04 review). A maintained comparison against the four workflows people actually use instead — a C unit framework with
.sfiles linked in (cmocka/Criterion/Unity/gtest), raw Unicorn scripting, qemu+gdb, and asmUnit-style in-asm macros — including a truthful “when the alternative is the better choice” for each, a “what asm-test does not try to be” calibration list, and a capability matrix. Linked from the README reference funnel and the docs index “Where to start”.Call-descent built-in default denylist —
asmtest_descent_use_default_denylist(the one unshipped Phase-5 deliverable of the call-descent plan). Arms the L3DESCEND_ALLsafety set the plan promised: at trace start the backend populates the handle’s deny pool from the tracee — the dynamic linker’s executable mappings (the lazy-binding PLT resolver) and[vdso]/[vsyscall], managed-runtime GC/JIT modules by mapping name (CoreCLR/Mono/HotSpot/ART/V8/BoehmGC), and, on the fork path (tracee shares the tracer’s layout), dlsym-resolved entry points of the classic blocking libc/pthread calls as one-byte deny regions. Denied callees are stepped over and recorded as edges; caller-supplied deny regions/callbacks compose. Wrapped in all ten bindings (parity gate green, 99 symbols × 10); a new fork-path fixture asserts a call landing exactly onpollbecomes an edge, not a frame (hwtrace suite 259 → 260).Emulator snapshot/restore —
emu_snapshot/emu_restore/emu_snapshot_free(E5 of the 2026-07-04 review). Captures the full register context (uc_context_save) plus the extents, permissions, and contents of every mapped region; restore reinstates the bytes and the mapping set itself (a region mapped after the snapshot is unmapped again). Mapped memory deliberately persists acrossemu_call_*— so fuzz/mutation sweeps previously ran each candidate against memory dirtied by its predecessors; bracketing the sweep with snapshot/restore makes killed/survived classification independent of handle history. Handle-level arming (watchpoints, register guards, preloads, the fuzz corpus) survives a restore by design. Emu suite 50 → 52.ASM_MIXCALL— mixed integer + FP argument capture (A6 of the 2026-07-04 review). The canonical ptr+len+scalar shape gets a first-class macro:ASM_MIXCALL(&r, fn, (buf, n), (0.5))marshals each parenthesized group into its register file via the existingasm_call_capture_fp— no more hand-built compound literals (the repo’s owntest_structparam.chand-roll is converted). Covered by a newmix_scaleexample routine (x86-64 + AArch64 + NASM bodies; shared with the bindings’ mixed-capture fixtures) and the strict-c11/C++ header-portability gate.FP reference models —
ASSERT_MATCHES_FREF{1,2,3}(A7). Differential testing now covers the FP surface where rounding/NaN/lane bugs live: double tuples from anasmtest_fgen_fngenerator run through the FP register file (asm_call_capture_fp) and the C model, judged by ULP distance (ulps = 0= bit-exact; NaN matches only NaN). Example property tests pinfp_add/fp_mulagainst C models over dyadic rationals and a specials table (±0, ±inf, NaN, DBL_MIN/MAX).Failing-input shrinking in
ASSERT_MATCHES_REF*(E7). On a mismatch the tuple is greedily shrunk toward0/±1/LONG_MAX/LONG_MIN(else halved toward zero) while the disagreement persists, so the report leads with the boundary value that triggers the bug —shrinks to [0, 1]— alongside the original draw. Deterministic, bounded, and self-tested (the negative suite’s mismatch now asserts the exact shrunk tuple).Runner flow control —
--fail-fast,--repeat=N,--shard=K/N(R5 of the 2026-07-04 review).--fail-faststops dispatching at the first failing test (forces the serial path; the TAP plan moves to the end of the stream so it covers exactly what ran).--repeat=Nblock-replicates the selection N times — with--shuffle/--seed, the flake-hunting loop.--shard=K/Nruns the K-th of N round-robin slices of the filtered selection, so N CI jobs can split one suite with no test lost or duplicated (self-tested: shards 1/2 + 2/2 reassemble--listexactly). Self-tests 43 → 49.
Fixed¶
“Reset layout” is a real, always-available action with auto-fallback (docs/internal/archive/gui/18-breach-stops.md T2). It was inert — behind a View menu that only appeared with docking on, acting through phantom windows. It now rebuilds the shipped default from any state, is bound to
Ctrl+Shift+R(fires with or without the menu bar, in both binaries), and a corrupt or collapsed persistedbuild/desktop-imgui.inithat would leave zero visible panes now auto-falls-back to the default instead of stranding the user in an empty window.Author output is no longer lost on close (docs/internal/gui/18-breach- stops.md T3, F24). Closing an Author tab (or a workspace recording) with an unsaved run raised no prompt and simply erased the in-memory entry; a dirty tab now raises a save/discard/cancel guard and cannot be closed with a single silent click.
The pinned Keystone and Capstone source builds no longer fail on a modern CMake or GCC. Keystone 0.9.2 (Feb 2019) is upstream’s newest release, so there is no version to move to; on a CMake 4 / GCC 15 host it failed three separate ways, each hiding the next. (1) Both engines declare a
cmake_minimum_required()below 3.5, whose compatibility CMake 4.0 removed — configure aborted before doing anything. Newtp_cmake_compatinlib-thirdparty.shsuppliesCMAKE_POLICY_VERSION_MINIMUM=3.5, emitted only for cmake >= 3.31 where the variable exists, and used by both build scripts. (2) Two of Keystone’s CMakeLists setcmake_policy(SET CMP0051 OLD), which no flag can re-enable — CMake 4 removed OLD outright — so they are patched to NEW; the policy only governs whether generator expressions appear in a target’sSOURCESproperty, which Keystone’s build never reads. (3) Keystone’s bundled LLVM fork predates GCC 13’s stricter header transitivity and usesintptr_twithout including<cstdint>, now supplied by-include cstdintin the C++ flags alone.The patches are applied after the
git rev-parse HEADassertion against the recorded commit, which is what keeps the B5 pin meaningful: integrity is still proven against unmodified upstream, and every subsequent change is visible in the build script rather than baked into a vendored tarball.sedwrites through a temp file rather than usingsed -i, whose in-place spelling differs between GNU and BSD/macOS sed..NET: the
unwarmed/PT composein-window-JIT premise check no longer permanently self-skips on PT silicon (dotnet-pt-inwindow-jit-premise.md T1). On the only host class where the Intel PT whole-window ctor arms, the>=1 method JIT'd inside the windowcheck took its timing self-skip on every run (MethodsObserved == 0consistently on the i7-8559U — the dotnet-managed-pt-concurrency-plan T4 residue): the fixture’s first-call JIT is compiled inside the window, but the method-load event reaches theJitMethodMapon the EventPipe dispatch thread asynchronously, and a native-speed PT window closes sub-millisecond after the call — losing the delivery race the slow single-step sibling wins for free. NewAsmTrace.WaitMethodObserved(nameSubstring, timeoutMs)polls the live map’s thread-safeCountForso the test holds the window open (2 s bound) until the delivery lands; the premise now asserts deterministically, the self-skip arm survives only for a genuine delivery stall (never-flake rule kept), and the close-timetrackBytesimage now contains the fresh method so the versioned decode resolves it directly. No C-core change, no new tier symbol.macOS (Intel) native build correctness (fifth pass): the hwtrace suite failed to compile at the new MSR-rung commit-decision seam (
amd-review-followup-2T4, landed from the Zen box the same day).examples/test_hwtrace.chand-declaredasmtest_trace_auto_msr_commitsinside the#if defined(__linux__) && defined(__x86_64__)perf_event declaration block, buttest_msr_commit_decision— deliberately pure, the test that “pins the decision itself everywhere” — calls it unconditionally, somake WERROR=1 hwtrace-teston macOS died on an implicit-declaration error. The seam is defined unconditionally insrc/trace_auto.cand links on Darwin; the prototype now sits above the Linux-only block. Verified on the macOS 14.8.7 / Intel host:make WERROR=1 hwtrace-test149 passed 0 failed (suite grew 145→149 with the four new seam checks now running on macOS too). The rest of the pass was clean at the same tree —WERROR=1 test/check(57/57),asm-test16/16 (the new statement-drop guard green under the host’s Keystone 0.9.2), the cpp 58 / ruby 57 / python 15+12-skip hwtrace binding lanes,WERROR=1 dataflow-testbuild + transparent self-skip, themake cliOS-gate self-skip intact after the asmspy T2/T7 wave, andmach-stepper-test25/25.The in-line assembler no longer returns machine code with a statement silently missing (assemble-silent-statement-drop.md T1-T3). Keystone drops a statement it can only partially parse — a bad, truncated or wrong-dialect operand — leaves
ks_errnoatKS_ERR_OKand still returns success, soasmtest_assemblehanded backok=truewith an instruction the caller wrote missing from the bytes. For a library whose purpose is asserting on machine code that is the worst failure shape available: the assertion runs and reports on code the caller never wrote. The most reachable trigger needed no typo — AT&T source assembled under theASM_SYNTAX_INTELdefault that every binding passes when the caller does not name a syntax (asmtest_assemble(…, ASM_SYNTAX_INTEL, "movq $42, %rax\nret\n", …)returnedok=truewith a bare{0xc3}); an ARM-style immediate on x86 (mov rax, #42) and a truncated operand (mov rax,) took the same path, and the truncated-operand shape drops silently in all five x86 dialects and on ARM64/ARM32 as well.asmtest_assemblenow counts the statements its source contains and fails the whole assemble when Keystone reports fewer — “assembler skipped N of M statements (check the syntax argument)” — through the existing error path, so all ten bindings inherit it with no ABI change (the guard is in the core;asmtest_asm_bytesnever exposedstat_count, so no binding could have worked around it). The header now states the all-or-nothing contract. Callers with a test that passes today while asserting on short code will start seeing this failure — that is the point. The counter is deliberately a lower bound: it tracks string and character literals,#////block comments and the two measured dialect rules (;separates statements everywhere except x86/NASM where it starts a comment;#is a comment on x86 but the immediate prefix on ARM), and never special-cases labels, directives or comment lines because Keystone counts those as statements too — every construct it walks past can only lower the count, so a false rejection of valid code is not reachable through them. Verified by a sweep over every assembler source in the tree, each in the dialect it is written for: the tagged call-site literals, the doc’s adversarial separator/comment/literal shapes, and theexamples/*.scorpus files whole — the last preprocessed the way the build preprocesses them (cpp -C, comments kept), which is what actually reaches the assembler, 14 of them landing 9-74 counted statements against Keystone’s 36-131. Zero false rejections, alongside an anti-vacuity pass confirming the guard still fires on all eight defect fixtures.IBS surveys can no longer report a complete, empty capture when the edge export OOMs (amd-review-followup-2 T1), plus the round-2 review’s smaller residues (T3/T4/T5). All four IBS lanes (
survey_pid,survey_process,window_end,survey_fetch_pid) discardedeh_export/fh_export’s return, so an OOM’d export surfaced asOKwithn==0besidebranch_samples>0— indistinguishable from a genuinely-empty survey; they now returnEUNAVAIL, mirroring the software-clock lane, with a test seam (asmtest_ibs_test_set_export_fail) proving the contract pure everywhere and live under injected OOM on an IBS host. T3:g_open_errnonow resets on every capture entry (no staleESRCHreason after a later success, test-pinned); the two live survey drains gained the seam parser’s short-SAMPLEh.sizefloor (F7’s last two sites); the retired-freeze-gate comment intest_hwtrace.cnow describes the substrate probe it heads. T4: thetrace_autoMSR rung commits a NONEMPTY truncated partial asHW_OK+truncated— the same contract a fast-tier truncated partial returns under — instead of discarding it and, with both steppers absent, returningEUNAVAILbeside a usable 16-deep partial; the genuinely-empty read still falls through (decision extracted as a host-testable seam, pinned by pure tests). T5: the Phase-4ASMTEST_HWDBGenv-gated logging now reaches the two AMD TUs it never covered —ibs_backend.c(probe outcomes,perf_event_openerrno, near-full/LOST, export-OOM decision) andtrace_auto.c(per-rung commit/skip and the final mechanism). T2 doc drift: parity-matrix recommendation rows no longer put AMD LBR primary on Zen 3 (BRS-only silicon — the open is-EINVAL), Matrix 1 records MSR-direct + IBS as shipped (only Zen 3 BRS stays forward-look), and the 2026-07-09 orphan review page carries a SUPERSEDED banner correcting its “zero IBS code” premise.Source-built Capstone dylib unloadable on macOS (
benchmarks-ci-followupsT1 validation). The first dispatched run of the nightlybenchmarks (macos-15-intel)leg failed atmake bench-check: dyld abortedemu-benchwithLibrary not loaded: @rpath/libcapstone.5.dylib … no LC_RPATH's found.scripts/build-capstone.shleft CMake’s Darwin default@rpathinstall name on the dylib, while the tree links consumers with the plain pkg-config-L/usr/local/lib -lcapstone(no rpath). The script now bakes the absolute install-name directory (-DCMAKE_INSTALL_NAME_DIR, the Homebrew convention) — a no-op on non-Apple platforms, and the K1 cache key already hashes the script so CI rebuilds instead of restoring the stale dylib.2026-07-21 review — C2/C3, S2, S4, S6, B3, B5, B7, T2 (fixed 2026-07-22). The remaining review findings not covered by the S3/S5/S7 and D1–D3/T1/K5 batches, each with an anti-vacuity-checked test and lane verification:
asmspy CLI (C2/C3).
--log --follownow setsPTRACE_O_TRACEEXECand drops a followed child thatexecves a 32-bit image (the syscall-stream engine has no i386 table, so it would render i386write(4)as x86-64stat(4)); and a bare app-deliveredSIGTRAP(executed int3 / hardware breakpoint, bysi_code) is re-injected in the syscall-stream and--procs --count=syscallsengines instead of swallowed.make docker-cliPASS.Core C (S2/S4/S6).
g_pt_window(the single whole-window PT slot) is mutex-guarded with a reserved-arm sentinel so two INTEL_PT arms can’t both claim it (S2, host-testable seam proves exactly-one-of-8 wins); the AArch64 wrong-depth hardware-breakpoint resume single-steps over the re-matching PC instead of relying on x86EFLAGS.RF(S4, x86 byte-identical, aarch64 compile-checked, behavioural test gated on bare-metalNT_ARM_HW_BREAK); and ~20 growable pools route through a shared overflow-checkedasmtest_grow/asmtest_grow_pow2helper so a capacity double can no longer wrapsize_tto 0 (S6, newtests/grow_overflowunit test inmake check).make docker-hwtracegreen (621/0 here; the PT-window mutex + pool guards join the S3/S5/S7 batch).Bindings (B3/B5/B7). The .NET long-lived native-handle types (
DrTrace/HwTraceNativeCode,NativeTrace,HwTracerecorder) are nowIDisposablewith a finalizer backstop, and the Java equivalents plusAddrChannel/CodeImageregister ajava.lang.ref.Cleaner, so a dropped handle no longer leaks its native mapping/trace (B3;Descent’s late-bound upcall arena staged; reclamation tests reclaim ~0 of N leaked mappings). Go (a compile-time_Static_assertabicheckpackage) and Zig (comptime@offsetOf/@sizeOf) now fail the build on a hand-mirrored FFI struct drift (B5). Ruby reads theuint64_tasmtest_regs_retas unsigned and Node stopsNumber()-narrowing koffi’sBigInt, preserving a >2⁶³/>2⁵³ return (B7). Verified:docker-hwtrace-dotnet/-java,docker-drtrace-java,docker-go,docker-zig,docker-dataflow-zig,docker-ruby,docker-node.Tests (T2).
tests/expect.shdocuments that the standalone-TAP tier suites are gated by exit code in their own CI lanes (wiring them intomake checkwould only self-skip, which the dependency rule forbids).
2026-07-21 review batch 3 — core-C robustness S3, S5, S7 (fixed 2026-07-22). Three memory-/integer-safety hardenings in the hardware-trace core, all verified on Linux via
make docker-hwtrace(514/514, 0 failed) and macOS-Intel-portable (nativeWERROR=1 hwtrace-test145/145 — the fixes and their tests are Linux/ x86-64-guarded and self-skip cleanly off it). S3:render_windowrecorded ABSOLUTE RIPs and disassembled them straight from live self memory, so a region unmapped after capture (a JIT free, a dlclose, a torn-down stack) SIGSEGV’d the renderer. It now copies each RIP fault-safely through a newhw_read_self_livehelper —process_vm_readv(getpid()), page-clamped on an unmapped straddle, mirroringpt_backend.c’spt_read_self_live— and renders “(undecodable)” for a freed address instead of faulting. Newtest_wholewindow_render_unmappedmprotects the traced pagePROT_NONEpost-capture; mutation-checked (the raw-deref revert SIGSEGVs at that test). S5: the jitdump readers (asmtest_jitdump_find,asmtest_jitdump_debug_find) computedname_len = (long)total - 56 - (long)code_sizeover an untrustedjit-<pid>.dump, where acode_size > LONG_MAXcast is implementation-defined and yields a plausible-but-wrong positivename_lenthe<= 0guard misses. Both now reject acode_sizeoverflowing the declared record (unsigned, underflow-guarded) before the cast. Newtest_jitdump_hostilefeeds acode_size == UINT64_MAXrecord; mutation-checked (removing the guards flips both assertsnot ok). S7:round_pagesclamps the caller-controlled AUX/data-ring size to 1 GiB so(v + pg - 1)cannot wrap and the power-of-two round-up cannot shift to 0 on a hostile/garbageaux_size/data_size. Seedocs/internal/reviews/2026-07-21-repo-review.md; the §2 remainder S2 (PT-window race) / S4 (arm64 hw-bp resume) is gated on PT / arm64 runtime this host lacks, and S6 (the ~15-site pool-growth clamp) is left as a mechanical follow-on.2026-07-21 review batch 2 — T1, D1–D3, K5 (fixed 2026-07-22). T1: the permanent
SKIP("partial-fill semantics not finalized")in the defaultmake testset is retired — the semantics are finalized and asserted:mem.partial_fill_touches_only_first_n_bytesprovesfill_bytes(buf, val, n)writes the low byte ofvalintobuf[0..n)and leaves the tail untouched (the contract all four implementations — GAS x86-64/AArch64/riscv64 + NASM — already share); green under both syntaxes, mutation-checked. D1: the two residual Zen-3 LBR overclaims (reference/features.md,reference/diagrams.md) now state the Zen 4+ live floor per_positions.md#2. D2: the two internal-engineering pages that leaked onto the public Sphinx site moved underdocs/internal/—amd_tracing_review.md→internal/analysis/2026-07-09-amd-tracing-review-f1-f47.md(the authoritative F1–F47 edition, not a duplicate as the review first framed it) andscoped-tracing-implementation.md→internal/with links rebased; all referrers retargeted (published guide → GitHub blob URLs); Sphinx-Wclean. D3:guides/win64.md’s intro no longer calls the runner port “now underway” while the body says full parity — it is complete. K5: all six third-party GitHub Actions are SHA-pinned to their then-current release commits (pypa/gh-action-pypi-publishv1.14.1 — previously the movingrelease/v1branch — plusruby/setup-ruby,rust-lang/crates-io-auth-action,mlugg/setup-zig,msys2/setup-msys2,docker/setup-buildx-action), and ci.yml’sactions/setup-python@v5is unified to@v6; actionlint output byte-identical to before, YAML parses. Seedocs/internal/reviews/2026-07-21-repo-review.md; still open there: C2/C3, S2–S7, B3/B5/B7, T2.asmspy no longer orphans a planted breakpoint when the tracer is interrupted (2026-07-21 review C1 — the one hole in the “never kill the target” invariant). A
PTRACE_POKETEXT0xcc is plain memory the kernel does NOT restore on tracer death, so an unhandled Ctrl-C mid region-trace left the target to execute the orphaned int3 later and die. asmspy (TUI and headless modes) now installs a SIGINT/SIGTERM/SIGHUP handler that only sets flags: the engines see theirstopflag (every headless engine call now passes one), a blockedwaitpid/getchreturns EINTR (noSA_RESTART— the SIGALRM quit-wake contract), and the normal unplant + two-phase-detach unwind runs, with the TUI reachingendwin(). An inheritedSIG_IGNis honored (the nohup convention).--sampledeliberately keeps a NULL stop — that engine’s NULL means “exactly one window”, and it plants nothing. New cli-smoke leg:--trace hotfn -1, SIGTERM mid-cycle → tracer exits 0 and the victim survives a grace period of continuous re-entry (differentially confirmed: the pre-fix binary dies 143 with no detach).g_amd_snapis per-thread, like every other hwtrace lifecycle slot (2026-07-21 review S1). The AMD boundary-snapshot flag sat as a process-globalintbeside the__threadg_fd/g_base_map/g_activeit claims to share an invariant with; a concurrent snapshot arm on one thread could flip another thread’shwtrace_end_amdonto the wrong teardown branch (leaking that thread’s perf fd/ring). Now__thread.make docker-hwtrace697 ok / 0 failed.Java
HwTrace’s availability-QUERY family self-skips instead of throwing when the native library is absent (2026-07-21 review B2).status(),resolve(),auto(),resolveTiers()andautoTier()threwRuntimeExceptionwith the library unloaded, contradicting the class’s own “callers never see a throw” self-skip contract (masked only because the lib is always bundled); each now degrades to its faithful unavailable value (EUNAVAIL status/empty cascade/empty Optional). Newhwtrace-java-testleg runsHwTraceTest --not-loaded-contractin a second JVM with the resolver pointed off a cliff (bogusASMTEST_HWTRACE_LIB, cwd outside the repo), with an anti-vacuity guard that fails if the library loaded anyway.Rust binding: the in-line assembler/disassembler is reachable without
ASMTEST_LIB(B1), and rv64 gets its capture struct (B4) (2026-07-21 review).asm_fns()gated ondlopen($ASMTEST_LIB)alone, so the common no-env downstream case reported “not in this build” despite the crate dylib-linkinglibasmtest_emu(which carries Keystone + Capstone); it now falls back todlopen(NULL)on the already-linked image — while an explicitly SET but unloadable override still surfaces as unavailable rather than being silently substituted (new own-processtests/asm_no_env.rsproves an assemble+disas round trip with the env var removed). The missingriscv64Regsmirror of asmtest.h’s rv64regs_tis added (a0/a1 return pair, s0–s11, always-0flags— no flag-mask constants, as rv64 has no flags register), and the test files’ carry-fixture externs/tests gained the same arch gate the corpus itself uses;cargo check --target riscv64gc-unknown-linux-gnu --all-targetsnow passes.Supply-chain: every DynamoRIO image fetch is digest-verified (K1), and the manylinux wheel base is pinned (K2) (2026-07-21 review). 12 Dockerfiles (plus the drtrace CI job) fetched the DynamoRIO tarball with a raw
curl, bypassing the repo’s ownscripts/third-party-digests.txtgate; all now route throughscripts/fetch-dynamorio.sh, which refuses a download whose SHA-256 does not match the manifest, andcheck-thirdparty-versions.shgates every image’sARG DR_VERSION(was 2 of 12).Dockerfile.manylinux-wheel/release.ymlbuilt the published PyPI wheel on a floatingquay.io/pypa/manylinux_2_28_*base; both now pin the dated tag2026.07.19-1(one knob across arches — the pypa repos publish identical dated tags per arch), with a new checker group keeping the pair in sync.Build/CI mechanics (2026-07-21 review K3/K4/B6).
build/asmtest_nomain.ohad two competing recipes (mk/fuzz.mk vs mk/bindings.mk) — GNU make warned “overriding recipe” on every run and the winning recipe silently dropped the fuzz object’s.build-flagsprerequisite (a latent stale-rebuild bug); one canonical recipe now carries the union (.build-flagsdep +-Wno-unused-function). The libFuzzer/AFL++ coverage-shim lane existed but no workflow ran it — a newfuzzCI job runsmake docker-fuzz(bounded, fails unless both engines find their planted crash). The Go conformance test used the exactuintptr→unsafe.Pointerround-tripHwNativeCode.Ptr()exists to avoid;go vet ./...is clean again.The shared code-image is now thread-safe — a use-after-free that crashed the .NET managed multi-threaded live-PT suite ~100% of the time on real Intel PT silicon (dotnet-managed-pt-concurrency-plan.md T1/T2/T4). With
libiptpresent on a bare-metal Intel PT box,make hwtrace-dotnet-testunder--cap-add=PERFMONSIGSEGV’d on 7 of 7 runs, always just past the stitched-trace block and always on a different thread.eu-stackon thecreatedumpcores (gdb cannot unwind them, and gdb cannot run this suite at all — it isSIGTRAP/EFLAGS.TF) put the fault in asm-test’s own code, not CoreCLR:asmtest_pt_read_codeimage←pt_insn_next←asmtest_pt_decode_window←asmtest_hwtrace_pt_hop_close. Cause:asmtest_codeimage_thad no synchronization at all, while the §Z4 ambient producer uses it from two threads at once — theJitMethodMapEventPipe callback callsasmtest_codeimage_track()(whichreallocsimg->regions, andr->versviaci_region_add_version) on the runtime’s listener thread, while every per-tid PT hop close decodes against that same image on thread-pool threads, walking those very arrays inasmtest_codeimage_bytes_at(). Arealloctherefore freed an array a decoder was mid-walk on. Fixed by guarding the region/version arrays with a mutex held across the lookup and the append only — never across a decode — which leaves the header’s “borrowed bytes valid untilasmtest_codeimage_free” contract intact (per-version byte buffers are separately allocated and freed only at image teardown). A second, narrower lifetime race is fixed alongside it in the .NET binding: the ambient handler fires on pool threads the flow is still leaving, so a hop could be published afterComplete()’s drain, or be insideCloseHop— reading the map’s code-image — whileDispose()freed it.AsmAmbientStitchedTracenow makes the_completedcheck atomic with the_openHopsadd/remove and waits in-flight closes out before freeing (bounded; it leaks rather than frees under a stalled hop). The ambient live half now runs green on PT silicon (ambient: >=2 stitched slices captured) instead of crashing. Newhwtrace-dotnet-ambient-stress/docker-hwtrace-dotnet-ambient-stresslane loops the concurrent set (default 25×) as the regression guard, and the timing-dependentunwarmed/PT compose: >=1 method JIT'd inside the windowcheck now self-skips instead of flaking tonot okwhen the runtime happens not to compile inside the PT window. That stress lane immediately surfaced a third, pre-existing defect in the same producer — confirmed present on unmodifiedmainby re-running the lane against the untouched producer:Complete()snapshotted_parkedwhile a detach-drivenCloseHopon a pool thread was still enqueuing its slice, so a hop that was both opened and closed could be dropped from the stitch (ambient twin: every attached hop stitched (4 vs 5 opened)).Complete()now waits in-flight closes out before the snapshot. Invisible in a single pass, which is why it survived until a repetition lane existed.Intel PT whole-window & foreign-pid decode now works on real silicon — the tier had never once run on a live Intel PT box until now (intel-pt-whole-window-substrate.md T5, intel-pt-attach-foreign-pid.md T1/T2/T4, dataflow-pt-replay-tier.md T4). First live run on a bare-metal Intel PT host (Core i7-8559U, Coffee Lake) via CAP_PERFMON in-container surfaced that the PT capture was fine but the decode produced ZERO instructions on every real capture — the synthetic-fixture tests hid it because they place trace-enable (
TIP.PGE) at the region base, whereas a real unfiltered capture enables tracing in the caller. The decoder’s code image only covered the tracked target region, so it hit-pte_nomapat the first caller IP and stopped before reaching the region. Fixes: (1) newasmtest_codeimage_read_live()serves the static caller/loader bytes from the target’s live memory (process_vm_readv, page-clamped) as a fallback inread_recorder, and the region-keyedread_regiongained the same self-memory fallback — the decode loop still records only in-region offsets, so the temporal-JIT guarantee and the synthetic-fixture path stay byte-identical. (2)pt_aux_opennow explicitly requests thept+branch(COFI) config bits from the PMU’s sysfs format rather than relying on a kernel default enablingRTIT_CTL.BranchEn(the kernel rejectsbranchwithoutpt; timing bits stay off since the decoder carries no MTC/CYC calibration). (3)pt_capture_one_regionkeeps its foreign victim alive throughattach_endso the decoder can read the victim’s own.text(where PGE landed). (4) Corrected thehwtrace-pt-livecapture-side address-filter checks to the verified silicon behavior: the filter size must not overlap the adjacentpt_filter_sibling(perftraces the whole[start, start+size)range, and the two functions sit ~0x1f bytes apart), and a@filefilter naming an anonymous region is accepted-but-unmatched on this kernel (not rejected) — either way un-filterable, which is why the decode-time fallback exists. Result:make hwtrace-pt-live631/631 andmake dataflow-pt-live29/29 — the live PT replay matching the single-step oracle with zero single-steps of the target — both green and stable, first-ever on real PT silicon. The non-PT lanes are unchanged (docker-hwtrace625/625 with PT self-skipping;docker-dataflow-ptsynthetic 19/19-Werror). Recorded in the internaldocs/internal/intel-hardware-validation.mdnote.asmspy CLI victims build on AArch64 — completes the
cli (ubuntu-24.04-arm)build (follows the include-comment fix below). That fix revealed asmspy-aarch64 T5’s arm64clileg had never actually built: two victims carried unguarded x86 asm. Ported both to AArch64 —cli/int3_victim.cint3→brk #0with anSA_SIGINFOhandler that advancesuc_mcontext.pcpast the 4-byte brk (AArch64brkis a fault, not a trap: the handler must step PC or the return re-executes it forever), andcli/exec_stage2.c’s freestanding x86syscallstub +_start→ asvc #0stub, AArch64 syscall numbers, and an AArch64_start. Verified under qemu-user: both compile + link withWERRORon aarch64, the whole cli builds to an arm64build/asmspy,exec_stage2prints its freestanding banner (itssvc/_startrun), andint3_victimsurvives its own breakpoints with noSWALLOWED(the PC-advance works). The ptrace tracer interaction (asmspy re-injecting the brk) is validated by the native arm64 CI leg.tools/asmfeatureslinks on Apple-Silicon macOS —benchmarks (macos-latest)bench-report(the second of the two macOS issues; follows the@rpathfix below).src/mach_backend.c’s whole body — including the MIGcatch_mach_exception_raise*callbacks — is#if x86_64 && __APPLE__, but the generated MIG servermach_excServer.ois compiled on every macOS arch and references those callbacks unconditionally, so any executable linking the Mach objects (asmfeatures, viamake bench-report) failed to link on arm64: “Undefined symbols for architecture arm64: _catch_mach_exception_raise*”. Added__APPLE__-guarded stub definitions (returnKERN_FAILURE; the stepper never arms on arm64, so they are never invoked) in the non-x86_64-Darwin branch — excluded on Linux and on x86_64-macOS’s real branch. Reasoned fix (no local macOS host); the macos-latest CI leg confirms the link. The Mach OOP stepper itself stays Intel-macOS only — single-stepping on arm64-macOS is a separate port.asmspy
cli-smoke: deterministic unknown-arity arg-decode assertion (de-flake). The--logsyscall arg-decode smoke (cli/cli_smoke.sh) asserts that at least one syscall in a 400-event window renders the faithful unknown-arity form (…) rather than a fabricated arity-of-three. The only undescribed syscall the victim’s stream produced was the incidentalrestart_syscallthe kernel emits when a signal interrupts a blocking call — nondeterministic, and it flaked to zero matches under some kernels (measured 1-of-N on Docker-Desktop’s LinuxKit kernel), somake docker-clifailed at that step.cli/argdecode_victim.cnow makes one DELIBERATELY-undescribed syscall (sysinfo, absent from asmspy’sarg_shapetable), so the…rendering appears every iteration; the assertion keys off that specific call — a strengthening, not a weakening.make docker-clicli-smoke PASS restored end-to-end.Five red
mainCI jobs unbroken — regressions the individual lanes that introduced them did not catch. Each surfaced on the sharedci.ymlmatrix after an unrelated lane landed; all five were failing on every recent push.cli(bothubuntu-latestandubuntu-24.04-arm): a-Werrorbuild break incli/asmspy_engine.c. Two#includelines carried comments whose body contained*/(/* … TIOC*/FIO* … */,/* … S_IF*/STATX_* … */), which closes the block comment early and leaves the tail aserror: extra tokens at end of #include directive [-Werror]. Only the WERROR CI leg (make WERROR=1 cli-smoke) is gated on it, so the non-WERRORdocker-clipath stayed green and the break went unnoticed. Reworded both comments (TIOC*, FIO*/S_IF*, STATX_*);make docker-clinow builds and the smoke passes end-to-end.test (riscv64 container): an x86-only opcode in a supposedly-portable asm stub.asmtest_sve_rdvlinsrc/capture.s(added with the SVE trampolines) zeroed its return viaxorl %eax, %eaxin a catch-all#elsethat also covers RV64 —src/capture.s:2786: Error: unrecognized opcode. Split the arm to match the file’s own 4-way pattern (#elif __x86_64__→xorl,#elif __riscv→li a0, 0,#else→#error); the riscv64 container builds and runs green again.dataflow (analysis lib + bindings): a legitimate new hardware self-skip not on the gate’s by-name allowlist. Chaining F5’s PT replay suite intodataflow-testadded a# SKIP pt live replay: no intel_pt PMU …line, which tripped the anti-vacuity gate (it allowed only the BTF block-step skip). Extended the allowlist to permit the Intel-PT skip by name — a host gate exactly like the BTF one — with a matching::notice.dataflow (F4 GC-move canon): a stale exact-count gate. The lane grew from 37 to 43 assertions when the object-identity alias phase (1..6) landed, but the gate still asserted-ne 37. Updated to 43 with the phase noted in the message.benchmarks (macos-latest): Capstone@rpathdylib not found at runtime (one of two macOS issues; partial). The pinned Capstone build installs to/usr/localand its dylib install-name is@rpath/libcapstone.5.dylib; recent macOS dropped/usr/local/libfrom dyld’s default fallback search, soemu-benchaborted (Library not loaded: @rpath/libcapstone.5.dylib). SetDYLD_FALLBACK_LIBRARY_PATH=/usr/local/libon the benchmarks job (ignored on Linux); the CI run confirmed this resolves thebench-checkabort. It then surfaced a separate, pre-existing failure the abort had masked:bench-reportcannot linkbuild/asmfeatureson the Apple-Siliconmacos-latestrunner —Undefined symbols for architecture arm64: _catch_mach_exception_raise*. Those MIG callbacks are defined insrc/mach_backend.c, whose body is x86_64-only (the Mach out-of-process stepper was ported to Intel macOS only), while the MIG servermach_excServer.oreferences them unconditionally. Porting the Mach tier to arm64-macOS (or excluding it from the arm64asmfeatureslink) is a follow-on for a macOS-capable agent — tracked against macos-oop-mach-stepper / benchmarks-ci-followups.
macOS (Intel) native build correctness (fourth pass): the asmspy CLI lane + two ungated example lanes, surfaced by building the lanes outside the nightly
test-macos-x86contract on a macOS 14.7.5 / Intel host —make cli,make cli-smoke,make WERROR=1 codeimage-test,make WERROR=1 build/jit_trace— none of which a Linux or Docker-on-Mac (Linux) lane exercises. A tree-widemake WERROR=1sweep of every other macOS-buildable native lane (test/check/emu-test/asm-test/usecases/usecases-emu/dataflow-test/dataflow-pt-test/ hwtrace-test and the C/C++/Ruby/Python binding lanes) was already clean; these three lanes were the gap.mk/cli.mkgatedcli/cli-smokeon architecture only (x86_64 aarch64 arm64), not OS. asmspy is a Linux-only out-of-process tracer (ptrace /process_vm_readv/personality//proc/<linux/futex.h>/<sys/user.h>/ the glibc extensionpthread_timedjoin_np, plus<sys/prctl.h>in every victim), so on macOS-x86_64 the arch gate passed and the build fell through and hard-failed —cli/asmspy.cat undeclaredprocess_vm_readv/pthread_timedjoin_np, and the wholecli/tree at<elf.h>/<linux/futex.h>/<sys/prctl.h>. Per-file include guards can’t fix this (cli/asmspy_engine.calone carries ~473 Linux-only ptrace/user_regs_struct/SYS_*references); macOS’s single-step tracer is the separate Mach-exception tier (src/mach_backend.c,make mach-stepper-test). Added a Linux OS gate checked before the arch gate to both targets, mirroring the existing arch-gate self-skip idiom, somake cli/cli-smokenow print# SKIP … this is an OS gateon non-Linux instead of a compile cascade.make docker-cli(Linux in-container,UNAME_S=Linux) falls through the gate and builds asmspy + drives the cli-smoke sequence exactly as before — the gate is host-OS-keyed, not container-keyed (verified: asmspy built and the smoke ran its full sequence).examples/test_codeimage.c: theBLOB_A/BLOB_Bfile-scopestatic constroutines are referenced only inside the two#if defined(__linux__)bodies, so off Linux they drew-Werror,-Wunused-const-variable— and the ungatedcodeimage-testlane compiles this C TU under-Werror(the test itself is designed to compile everywhere and self-skip at runtime via its#elsestub). Guarded the two definitions with#if defined(__linux__)to match their use.examples/jit_trace.c:static int checks, failures;and theCHECKmacro are used only from the#if defined(__linux__) && defined(__x86_64__)body (the#elseis a self-skip stubmainthat reports neither), so off that target (macOS, and Linux-arm64) they drew-Werror,-Wunused-variable. Guarded both with the same condition as their callers, honouring the file’s own compile-and-skip design.
macOS (Intel) native build + binding self-skip correctness (third pass), surfaced by building the binding conformance corpus and the per-language
dataflow-*/hwtrace-*lanes natively on a macOS 14.7.5 / Intel host — a surface no Linux or Docker-on-Mac (Linux) lane exercises (make python-test,make WERROR=1 dataflow-test,make dataflow-cpp-test/-python-test/-ruby-test,make hwtrace-python-test):bindings/conformance/conformance.cincluded<sys/mman.h>/<unistd.h>only under#if defined(__linux__), but theCL_HAVEptrace_descentfixture that callsmmap/mprotect/munmapis gated on ARCH (x86-64 / aarch64), not OS — so it compiles on macOS and hit implicit-function-declaration errors there, breakingmake python-test. Broadened the include guard to__linux__ || __APPLE__, matchingsrc/hwtrace.c’sasmtest_hwtrace_exec_allocW^X path.examples/test_dataflow_ptrace.c: ninestatic constfixtures (df_chain_v2+ the call-out / overflow set) used only inside the__linux__ && __x86_64__block drew-Werror,-Wunused-const-variableundermake WERROR=1 dataflow-teston macOS. Guarded their definitions to match their use (completing the earlier#else-stub fix to the same file).bindings/dataflow_victim.c— compiled by everydataflow-<lang>lane, which run on macOS — unconditionally included<sys/prctl.h>and calledprctl(PR_SET_PTRACER, …)(a Linux-only Yama trace opt-in), so the shared victim failed to build on macOS ('sys/prctl.h' file not found). Guarded both under__linux__; the live-attach lanes self-skip off Linux regardless.bindings/python/tests/test_hwtrace.py’stest_window_region_free_whole_windowstill hard-assertedw.armed— the one binding missed when the second pass guarded the C++/Ruby/Lua/Zig/Rust window tests. Now guards the arming-dependent checks onarmedand notes the transparent self-skip, matching node’s and cpp’s shape.bindings/ruby/dataflow.rbused Ruby-3.0 endless-method syntax (def steps = …) that fails to parse on Ruby 2.6, violating the binding’s ownrequired_ruby_version >= 2.6(asmtest.gemspec) — and stock macOS ships Ruby 2.6. Rewrote as classic single-line defs (def steps; …; end), matching every sibling ruby file.
macOS (Intel) native build + whole-window self-skip correctness (second pass), surfaced by building the wider native tier set on a macOS 14.7.5 / Intel host (
make hwtrace-test,make dataflow-test,make WERROR=1 hwtrace-test,make hwtrace-cpp-test hwtrace-ruby-test):asmtest_hwtrace_pt_hop_open’s non-Linux#elsereturnedASMTEST_HW_ENOSYSwhile the tier’s classifier reports Intel PT asASMTEST_HW_EUNAVAILoff libipt — sotest_pt_hop_surface’s self-skip failed on macOS. Now returnsEUNAVAIL(the same fix already applied topt_begin_window/pt_attach_begin; the per-tid PT hop pair had reintroduced it).examples/test_dataflow_ptrace.ccalled fivetest_window_*functions unconditionally inmain, but their definitions live inside the__linux__ && __x86_64__guard — no non-Linux stubs, unlike every sibling test. Added the missing#elsestubs.examples/test_dataflow_blockstep.cused Linux-onlymemfd_createunconditionally (the F2sc_preadfixture), breaking the macOS compile even though the whole suite runtime-self-skips off Linux viaasmtest_dataflow_blockstep_probe(). Added a compile-only non-Linux stub (the caller is unreachable there — the fixture’s ownfd < 0SKIP covers it).examples/test_hwtrace.c’smap_exechelper drew an unused-function-Werrorundermake WERROR=1 hwtrace-teston macOS: every caller sits inside a Linux guard. Guarded the definition to match its callers, exactly like the adjacentframe_insns_eq.The region-free
§Z1whole-window scope is Linux/x86-64-only (begin_windowself-skips on macOS single-step, where the region-based tier still works). The C++, Ruby, Lua, Zig, and Rust binding tests hard-assertedw.armed, failing on macOS; they now guard the arming-dependent checks onarmedand note the transparent self-skip — matching node’s already-correct shape and the Ctest_wholewindow_singlestepskip.
macOS (Intel) native build + PT self-skip correctness, surfaced by validating the out-of-process Mach single-step tier natively on a macOS 14.7.5 / Intel host (
make mach-stepper-test, 25/25 live):examples/test_hwtrace.cincluded<unistd.h>only under#if defined(__linux__), so the portabletest_pt_attach_selfskip(it callsgetpid()on every host) failed to compile on macOS; the include moved to the unconditional POSIX block.asmtest_hwtrace_pt_begin_window/asmtest_hwtrace_pt_attach_beginreturnedASMTEST_HW_ENOSYSfrom their non-Linux#elsearms, but the tier’s single availability classifier reports Intel PT asASMTEST_HW_EUNAVAILon any host without libipt — so the PTbegin()self-skip envelope diverged from thestatus/skip_reasoncontract on macOS. Both#elsearms now returnEUNAVAIL, matching the classifier and the Linux!availablepath.tests/glob_parity.ccomparedasmtest_glob_match(pinned to glibcfnmatch) against the hostfnmatchon undefined-behavior patterns (unterminated[, trailing\), which BSD/macOSfnmatchresolves differently — failingmake check11/44 on macOS. The divergent cases now assert the glibc-pinned contract directly, cross-checking the hostfnmatchonly under__GLIBC__; well-defined cases keep the live host differential everywhere.
parallel runner (
-jN): a non-EINTRpoll()failure no longer abandons the run and reports never-run tests as passed; the scheduler degrades to blocking reaps.--filteron Win64: the portable glob matcher now matches POSIXfnmatchon unterminated[, backslash escapes inside classes, and trailing backslashes.guard-page allocators return NULL instead of a guard-page pointer for sizes within a page of
SIZE_MAX.emulator fuzzing: corpus nudge no longer has signed-overflow UB at
LONG_MIN/LONG_MAXrange extremes.Zig conformance: vec256/vec512 capture tests no longer under-fill the 8-slot vargs array.
Win64
--no-fork: a fault on a non-test thread no longer hijacks the test thread’s recovery stack; it takes the normal unhandled-exception path.docs: the emulator guide no longer claims
--emuinstalls only libunicorn.Cross-alias register def-use edges resolved.
asmtest_defuse_build(the shared, tier-neutral last-writer builder insrc/dataflow.c) keyed its register axis on the raw Capstone id, so a write to one GP sub-register alias and a later read of another —mov eax, ...then a read ofax,mov r8d, ...then a read ofr8— produced no def-use edge at all, even though the value trace correctly captured both. This is the shared builder’s counterpart todfp_alias_shape(src/dataflow_ptrace.c, added for the F6 gap barrier): a newreg_slicehelper canonicalizes a Capstone GP register id to its 64-bit container plus a byte offset/length, andapply_write/emit_readnow key a mappable register per CONTAINER BYTE — exactly as memory is already keyed per address byte — so a partial-overlap write/read resolves to the right last writer instead of missing the edge.AH/BH/CH/DHstay pinned to byte offset 1 of their container (not offset 0, which isAL/BL/CL/DL’s own byte): a write toahreaches a laterah/ax/eax/raxread but never a lateralread, which is the discriminator against a container-collapsing implementation that folds by container alone and ignores the byte offset (proven by temporarily mutatingreg_slicethat way and observing the new synthetic fixture fail, then restoring it). A 32-bit GP write (eax,r8d, …) additionally marks the FULL 8-byte container as written, not just its own 4 bytes — x86-64 defines a 32-bit write as implicitly zero-extending bits 32-63, unlike a 16/8-bit write, which leaves the untouched bytes exactly as they were — so a later full-width read resolves its upper-half producer to that same write instead of a stale one from before the zero-extension (also proven by mutation: reverting the widened write range makes a dedicated fixture fabricate exactly that phantom edge). Vector registers, segment selectors,EFLAGS, andRIPfall through to the pre-existing raw-id keying unchanged (none of them alias with anything else, so raw-id keying was already exact for them). Two new live fixtures inexamples/test_dataflow_ptrace.cexercise the windowed gap barrier end-to-end through this change: a glue excursion that clobbers a sub-register alias of a register the survey recorded, and — closing F6 known-limit (4) — a glue excursion that clobbers a whole XMM register the survey recorded, the first fixture anywhere to exercise the barrier’s vector path at all.The scoped
ptracedataflow producer’s call-out step-over can no longer fabricate a def-use edge across a stepped-over helper.dfp_step_loop(src/dataflow_ptrace.c) runs a call-out at native speed and records nothing over it — correct for cost, but a helper that clobbers a location the region already wrote (and a later in-region read relies on) previously left the read’s edge pointing at the stale in-region writer instead of the elided helper, silently wrong at a passingrc. Every scoped entry point (_run,attach,attach_pid*,attach_jit) now feeds adfp_riskset(mirroring the windowed survey’s existing gap barrier), and the call-out branch snapshots it immediately before the native run and diffs it after: a synthetic GAP step is appended carrying exactly what changed (register alias-sliced, memory per byte), so a post-call read correctly resolves to the barrier instead of the stale writer. A risk-set cap hit is deferred in scoped mode — it only promotes totruncatedat the first real gap (and is discarded on a gap-free exit), so a region that never calls out is never falsely flagged. Precision, not a blanket invalidation, is load-bearing here (F6 measured that a blanket shadow deletes true cross-gap edges): a helper that touches nothing at risk appends no record for that location, even though the gap step itself is still present (the call/ret round trip through any helper unavoidably movesrsp, which was already at risk from the call’s own write).make install/make install-shared-hwtracenow shipasmtest_ibs.h— the hardware-tracing guide’s documented#include <asmtest_ibs.h>could not previously compile against an installed package (it was the only guide-referenced header missing from all three install lists).scripts/clean-room-test.shgained a header-install compile check (a freshmake install+cc -fsyntax-onlyagainst every guide-referenced header) so this omission class cannot silently recur.The ptrace block-step reconstructors now mark the capture truncated when a block contains a
rep-prefixed string op. Arep movs/stos/…retires once per iteration under per-instruction stepping (RIP parks on it) but a static block-step reconstructor records it exactly once, so the block-step stream silently under-counted it. Newasmtest_disas_is_rep_stringletsbs_record_runandwindow_block_walkdowngrade such a block toBS_AMBIGUOUS(faithful truncation), bounding the “byte-identical to per-instruction stepping” promise accordingly.The ptrace block-step reconstructors no longer record never-executed instructions when the traced code contains an application
int3. A JVM safepoint poll or .NET breakpoint inside a block-stepped region was misread as a BTF#DBblock completion, so the region, attached, and windowed drivers fabricated the instructions after it withtruncated=false. They now classify the trap viasi_code(SI_KERNEL / TRAP_HWBKPT), record the executed run up to and including the trap byte, mark the capture truncated, and forward the signal — the region (owned) driver viaPTRACE_CONT, the attached (foreign) driver by leaving the target in its SIGTRAP delivery-stop for the caller, and the windowed driver by handing off to the per-instruction window loop, which runs the frame to its window end at native speed (run_until_sig) and recovers*resultthere instead of discarding the signal.No per-instruction ptrace loop in
src/ptrace_backend.cswallows an application SIGTRAP any more.run_until(the call-out step-over primitive, nowrun_until_sigplus a 2-arg wrapper), the per-instruction region driver (asmtest_ptrace_trace_call), the foreign attached driver (asmtest_ptrace_trace_attached), the windowed per-instruction loop shared byasmtest_ptrace_trace_attached_windowed[_window_stop], the fork-owned window driver (asmtest_ptrace_trace_window_call), and call descent (asmtest_ptrace_trace_call_ex/_attached_ex) each either deliver an applicationint3/breakpoint viaPTRACE_CONT(owned tracee) or end faithfully with the target left at its SIGTRAP delivery-stop (foreign) — neverPTRACE_SINGLESTEP/PTRACE_SINGLEBLOCKwith the signal attached (measured fatal: the re-armed trap fires inside a masked handler).bs_sigtrap_is_app(thesi_codeclassifier introduced for the block-step drivers) is now a file-wide helper shared by every loop, on both x86-64 and AArch64.The call-out step-over is now depth-aware (code review finding #19’s real fix).
run_until(nowrun_until_sp, withrun_until_sig/run_untilkept as thin wrappers) previously resumed the trace at the FIRST arrival at a call-out’s return-address breakpoint, so a stepped-over helper that called BACK into the traced region (a callback, or a tiering/OSR stub re-invoking the method) hit its own return-address breakpoint from a deeper stack frame first and hijacked the resume into that nested invocation.classify_region_exit(shared by all four region drivers — per-instruction, block-step, attached per-instruction, attached block-step) now also passes the callee-entry stack pointer, andrun_until_sprejects a same-address hit at the wrong depth: it steps past the premature hit at native cost (a single-step over a software breakpoint, or a barePTRACE_CONTfor a hardware one —EFLAGS.RFkeeps the CPU from re-trapping on it) and keeps waiting for the matching depth. A new differential fixture — a region that calls a helper which calls back into the region’s own entry exactly once before returning — proves the trace now resumes at the true, outer completion instead of the inner one.asmtest_ibs.hno longer describes the shipped system-wide capture flag as a future phase — thesurvey_processresidual-race note now names theASMTEST_IBS_OPT_SYSTEM_WIDEflag directly.The pure IBS-Op decoder now validates the record’s own caps word (BrnTrgt) before trusting the branch-target register. Two 68-byte record shapes are length-identical (
BRNTRGT=0/OPDATA4=1vsBRNTRGT=1/OPDATA4=0) and only the caps word disambiguates reg[7]; the decoder previously trusted length alone and could misread anIbsOpData4value as a branch destination. The RipInvalid read is now gated on caps bit 7, andasmtest_ibs_available()requires CPUID IBSFFV (EAX[0]) so it cannot disagree with the caps the kernel samples with.IBS ring-loss heuristic now bounds the callchain worst-case record (was 112 bytes, ~10× short — silent sample loss with
lost==0 && throttled==0);ibs_fill_attrpinssample_max_stackso the bound is sound, and the internal window lane no longer opens with callchain (no in-tree consumer, and a callchain stream can overrun the single end-of-window drain).ASMTEST_IBS_OPT_CALLCHAINis documented as consumer-less: it enables kernel-side capture only; nothing in the tree decodes the stack (the drain parses past it to reach RAW), and the window lane ignores it.ibs_probeand theibs-testlive skips now attempt a real perf open and report the real refusal reason instead of claiming AVAILABLE from the CPUID/sysfs substrate probe alone. On a locked-down AMD host (perf blocked byperf_event_paranoid/seccomp) the substrate is present but no sampling can open —ibs_probeprintssubstrate present but sampling is BLOCKED — <reason>(Op and Fetch lanes) and the fivetest_ibsEUNAVAIL skips print the realasmtest_ibs_unavail_reason()instead of a hardcoded guess. The AMD manual-validation checklist no longer inverts thecall_autoregression signal: post-5d8e0d2atruncated=0where escalation must fire is a regression, not a known finding.The guides and the public header no longer claim AMD LBR live capture works on Zen 3. The live-capture floor is Zen 4+ (LbrExtV2) — Zen 3 BRS exists in silicon but this tree cannot open it (the generic
sample_period=1open is rejected by the kernel’samd_brs_hw_config; the raw-0xc4arm is a hardware-gated follow-up). Swept every “Zen 3+” floor claim in the tracing guides,asmtest_hwtrace.h, and the AMD backend comments to “Zen 4+”, each with the one-line cannot-open explanation.The dead AMD freeze-on-PMI probe (
asmtest_amd_freeze_available) and its false PRESENT/ABSENT diagnostic are removed. The probe had zero live consumers after5d8e0d2replaced the freeze-conditional window-trust gate with an unconditional exit-presence check that runs on every part (asmtest_amd_ring_parse_decode).test_hwtraceprinted a trust statement (“PRESENT (single-window Tier-A trusted)” / “ABSENT (…)”) that was false in both branches — Tier-A completeness is exit-anchored regardless of the freeze bit. The freeze test is retired; the snapshot-substrate/depth probe checks stay (renamedtest_amd_snapshot_substrate_probe).The AMD deterministic boundary snapshot no longer flags a provably complete 15-branch window as
truncated. The depth-ceiling check inasmtest_amd_decode_reachcounted the total decode-array length, butbranchsnap.cprepends a synthetic boundary edge (a deterministic completion, not a captured hardware slot), so a full 15-hardware-slot window (15 + 1 synthetic = 16) tripped the 16-deep ceiling and escalated to a needless real re-execution of the routine under test (src/trace_auto.c). The newasmtest_amd_decode_reach_hwgates truncation on the hardware slot count; theasmtest_amd_decode/asmtest_amd_decode_reachwrappers passhw_nbr == nbrso every other caller is byte-identical.make cli/make cli-smokeon arm64 now self-skip like the other tiers instead of dumping raw compile errors. asmspy’s single-step engines are x86-64-hardcoded (rip/eflags-TF/orig_rax), so on aarch64 the build died mid-compile withSYS_open undeclared/no member named 'rip'— or worse, fell into the missing-dependency branch and advised installing libncurses-dev, which cannot fix an architecture. Auname -mgate (checked beforeCLI_MISSING, for exactly that reason) now prints a truthful# SKIPnaming the open ARM64-abstraction plan row and exits 0. Measured in a real linux/arm64 container: skip + rc 0 both targets; x86-64 unchanged. a Yama/seccomp skip.** The victim called the region once and_exit(0)’d, so on a slow host the child finished and died before the parent’sPTRACE_SEIZElanded (3/3 GitHub runs today), and the resultingESRCHsurfaced as# SKIP … PTRACE_SEIZE unavailable here (yama/seccomp)— a double lie, since SEIZE worked for every other attach test in the same job — which the lane’s anti-vacuity gate rightly turned into a hard failure. The victim now LOOPS the region at the same 2 ms cadence as every other attach victim in the file, so the attach always finds a live process and a fresh entry. Proven discriminating in the docker lane: the once-and-exit victim plus a deliberate 200 ms pre-attach sleep reproduces the exact CI skip; the looping victim passes under the same handicap. The gate’s stale bookkeeping was recalibrated in the same change: the 8 suites total 389 on bare metal (the comment said 257), the VM runs ~293, and the floor moved 230 → 285 — preserving the original tightness (the smallest suite vanishing still trips it).cli/asmspy.c’s new picker-sort code failed the clang-format gate. The Tab-cycle sort landed verified bymake docker-cli(build + smoke) but not bymake fmt-check; the format job caught 7 violations. Mechanical reflow, plus one comment hoisted above itsifso the formatter keeps the condition on one line.Data-flow
--dataflow’s call-out step-over lied about WHY it truncated, and two test suites carried assertions that could not fail. Three small, independently diagnosed defects closed together:dataflow_ptrace.c’s call-out step-over conflated a BOUND with a FAILURE (the sibling site to the--maxfix:9d55611/0129b1e). Hitting the whole-run step backstop mid-call-out anddfp_run_toactually failing shared one||and oneDF_PTRACE_ETRACE, so a region that simply ran a lot of call-outs surfaced as “ptrace/attach failure (permission? ptrace_scope? … W^X JIT page)” — sending an operator to Yama/seccomp for a budget, not a bug. The backstop is now its own branch (DF_PTRACE_OK,truncated=true, same shape--maxalready got right);dfp_run_tofailing (the callee exited, faulted, or its return byte could not be trapped) keepsDF_PTRACE_ETRACE. The 2^20-step backstop is now also overridable viaASMTEST_DF_STEP_BACKSTOP(mirroringASMTEST_DF_ENTRY_WAIT_MS), which is what makes the bound reachable in a test at all — a real 2^20-step fixture was exactly the kind of “never exercised, needs 1M hits” gap this codebase already flags elsewhere. A new attach-based test (test_callout_step_backstop, needs the attach path’s exactpre_positionedentry so the trip point is deterministic rather than a coin flip on the fork prologue’s step parity) proves it: mutation (reverting the split) turns the check back into ETRACE.test_branchsnap.c’s multi-exit test assertedcovered(t, 0)as its entry evidence — vacuously.amd_replayappends block 0 unconditionally (amd_backend.c:267), socovered(t, 0)is always true by construction; the check was carried entirely by theni > 0conjunct beside it, same fact the Phase 9 tail-jmptests in the same file had already found and correctly stopped relying on.snap_default_runnow asserts the PATH-SPECIFIC block instead (covered(want_off) && !covered(other_off), the two exits’ ownmovblocks) — real evidence that the default-on snapshot captured the exit that actually ran, not just “some” data regardless of which path executed.test_dataflow_blockstep.cre-declaredasmtest_blockstep_info_twith no layout guard. The tier ships no header by design (keeps the producer off the public ABI), so the suite hand-copies the struct — exactly the skew that cost F6’s sibling telemetry struct 3 green checks before asizeof+offsetofguard caught it.asmtest_dataflow_blockstep_info_layout()(mirroringasmtest_dataflow_ptrace_win_info_layout) now lets the suite check its copy against the producer’s real layout before trusting anyinfo.*field.
All three were filed as open follow-ups (2026-07-17-dataflow-tier-open-followups.md) after the same day’s F1/F2/F6/F7 batch landed, deliberately deferred out of that diff to avoid scope creep. Verified:
make docker-dataflow-attach(126+118 checks, 0 skips, 0 failures) andmake dataflow-blockstep-testnatively on the Zen 5 dev box (119/119).test_branchsnap.c’s live leg needs the BPF toolchain (clang/libbpf-dev), absent on this host and gated behind asudopassword this session could not supply — verified by compilation + the ENOSYS stub path only.asmspy --dataflowon a symbol that is not running HUNG instead of erroring. The producer’s step backstop counts single-steps, and a region that never arrives burns zero steps — so the blocking wait never advanced (measured: rc=124, killed bytimeout, where--traceon the identical target answered “never executed” in 4 s). The entry wait is now bounded by aCLOCK_MONOTONICdeadline (default 10 s,ASMTEST_DF_ENTRY_WAIT_MSoverrides, 0 restores the old unbounded wait) and reports “not seen entering in pid N (waited M ms)” — an outcome, not a failure: the symbol resolved and the tracer worked; the code just is not being called right now. The unwind re-establishes the all-running invariant (restore the entry byte, rewind a thread stopped atbase+1, continue), which also fixes a latent hang on the never-exercised step-backstop disarm path.asmspy --dataflow --max=<n>failed for everynbelow the region’s step count — and blamed ptrace for it. The truncation branch did everything right (partial trace appended,truncated:true) and then returned the generic ptrace-failure code, so a valid cap surfaced as “ptrace/attach failure (permission? ptrace_scope?…)”. The flag worked only when it did nothing (measured:--max=3rc=1,--max=200rc=0 on an 83-step region). It now returns OK with the truncated partial trace; the smoke asserts the EXACT step count per cap, so “cap ignored” cannot pass either.Three
asmspy --tracefidelity defects: a bound that wasn’t, a diagnosis thrown away, and a documented self-skip that never happened. (1) The entry wait’s idle window reset on EVERY waitpid event — a target that stops more often than the window (a 1 Hz timer, a chatty clone) reset the budget forever and--traceblocked indefinitely; a 30 sCLOCK_MONOTONICwall bound now sits alongside the idle rule, checked unconditionally. (2) The entry race’s four distinct outcomes were bare integers collapsed by a barebreak, so “the region never ran”, “the target exited”, and “the entry could not be armed” all rendered as “never executed”; the outcomes are now named and each maps to its own answer (“pid N exited beforewas seen executing” for an exit, an attach/ETRACE report for an unarmable entry). (3) That fix makes asmspy.h’s promised W^X/JIT self-skip real: an entry page refusing the breakpoint now reports “possibly a W^X JIT page refusing the entry breakpoint” instead of the confidently-wrong “never executed” — verified against an unmappable explicit range, the same failure shape a genuinely W^X page produces.asmtest_trace_call_autocould report a window-overflowing AMD-LBR capture as complete, so escalation never fired (a real Zen 5 silicon finding). The Tier-A completeness check inasmtest_amd_ring_parse_decode— trust a single sampled window as complete only if it contains the region-exit branch — was gated behind!asmtest_amd_freeze_available()and thus skipped on freeze-capable parts (Zen 5). Withsample_period=1the capture picks the richest-in-region window, often an arbitrary mid-run fragment that never held the exit, so a 25-back-edge loop reconstructed a 4-edge fragment and reportedtruncated=0—trace_call_autoreturned it as complete instead of escalating to block-step (andtest_call_autopassed vacuously). The exit-presence requirement now runs on every part; combined with the existing overflow flag it is the airtight “complete iff a non-overflowed exit-anchored window exists” invariant. Verified deterministic across 16 privileged AMD runs;test_call_autocase (b) hardened to fail hard on a fragment-reported-complete. Surfaced only because the newdocker-hwtrace-privilegedlane runs the exact AMD paths live.The shared
libasmtest_hwtraceshipped with an undefinedasmtest_ibs_window_end. The Zen-2 F6 IBS survey fallback madehwtrace.ccall the IBS window primitives, but the shared-lib link recipe never includedibs_backend.o— every binding’s dlopen failed on every host (the static test binaries linkHWTRACE_OBJS, which carries it, masking the gap inhwtrace-test).pic/ibs_backend.ois now compiled and linked, and tracked by the knob-flip rebuild sentinel.A corrupt/huge
nrin an AMD branch-stack sample could drive the ring parse out of bounds (review F5/F7, followup Phase 2). The sampled-branch count from the perf ring is now clamped (nr <= 64, comfortably above the 32-deep hardware maximum) before thenr * sizeof(perf_branch_entry)size check can wrap, and a short-tail sample too small to hold the 8-bytenritself is rejected instead of read. Exercised by the new synthetic-ring tests on every host.asmtest_trace_call_autocould returnASMTEST_HW_OKwith an empty trace (review F24). Each escalation rung’s reset discards the prior rung’s partial capture, butrankept reading 1 from that earlier rung — so a rung that reset and then failed at runtime (seccomp/ptrace_scope, ENOMEM) reported a successful empty trace.ranis now cleared at every reset site and re-earned only when a rung actually commits; the legitimate truncated-but-OK partial (no downstream rung runs) is preserved.asmspy’s single-step engines could kill a V8/Node target seconds after a clean detach. V8 worker threads park in blocking futex syscalls; aPTRACE_SINGLESTEPthat completes across a syscall defers its#DBdebug exception until the syscall returns, so a parked worker carried a queued trap through detach — when its futex later woke, the trap fired with no tracer attached and terminated the whole process (reproduced: ~1 detach in 2–6 fatal on an 11-thread V8 target). Two prior defenses missed it: the two-phase detach orders resumes, and the trap-flag clear was gated on a read-back TF bit that a kernel-forced TF masks out ofGETREGS.detach_threadsnow clears TF unconditionally and drains the pending step — each stopped thread is single-stepped once more to consume its queued#DBwhile we are still the tracer (skipping threads poised on a syscall instruction so the drain cannot block); thePTRACE_SYSCALLengines skip both phases, since draining them would inject step state into a target that had none. 30 + 25 consecutive attach/trace/detach cycles on the 11-thread V8 target now survive.asmspyswallowed a target’s ownint3breakpoints (and could have killed it re-injecting them). The single-step engines treated everySIGTRAPstop as their own step, so an application-executedint3(a JIT/debugger breakpoint, e.g. V8’sIMMEDIATE_CRASH) was mis-decoded and silently dropped, breaking the app’s own breakpoint logic. Stops are now split bysi_code(PTRACE_GETSIGINFO): onlySI_KERNEL(an executedint3on x86) andTRAP_HWBKPTare delivered back to the target — viaPTRACE_CONT, neverSINGLESTEP, because re-arming the trap flag fires a#DBinside the (SIGTRAP-masked) handler and the kernel force-kills the target. Everything else (TRAP_TRACE,TRAP_BRKPTfrom a step completing across a syscall, theSI_USERexec trap) is still absorbed. Newint3_victim+ smoke prove a self-breakpointing target survives tracing with its handler intact.Fork-based tracers aborted the whole trace when an unrelated signal interrupted the post-fork handshake.
trace_call,trace_call_blockstep,trace_window_call, andtrace_call_descendwaited for the child’s initialraise(SIGSTOP)with a barewaitpidthat treatedEINTRas a failed handshake (rc=ETRACE, zero frames). A host runtime’s repeating timer — or the descent stale-alarm test’s deliberate 200 µsSIGALRMstorm, which failed ~70 % of runs on a fast box — could land in that window. All four handshake sites now retry acrossEINTRexactly as the step loop always has; a genuine child death still surfaces. The stale-alarm test passes 20/20.Four defects found by a deep multi-agent audit of the whole tree (each survived double adversarial verification; the newest subsystem —
asmspy— and the AMD/LBR, PT, single-step, orchestration and FFI-binding layers came back clean). (1) Reused-handle determinism leak in two emulator guests. The x86 and arm64 setups zero the GP + vector register file before every call so a routine that reads a register the caller did not set gets a deterministic0; the RISC-V (emu_riscv_setup) and ARM32 (emu_arm_setup) setups omitted it, so a long-lived handle (how every binding holds it) leaked the previous call’s callee-saved / FP-lane state into the next call and returned a stale result withok=true. Both now zero registers (x1..x31 + f0..f31 / r0..r12 + d0..d31 + condition flags) like the other two guests. (2) In-process stealth stepper could busy-hang forever.asmtest_hwtrace_stealth_trace’swhile (!sc->ready)spin only checked for early helper death underif (use_exec); in the in-process fork fallback (common under the ptrace-restricted container/CI posture this project targets) a helper killed by seccomp/OOM/watchdog before publishingreadyleft the caller spinning at 100% CPU. Thewaitpid(WNOHANG)death check now runs unconditionally, matching the two windowed spins. (3) Block-step tracer leaked its owned tracee on overflow.asmtest_ptrace_trace_call_blockstepbroke out on ablockstep_reconstructfailure (stream full / undecodable insn / no in-region terminator) withrcstill OK, so the post-loop cleanup — which only reaps onrc != OK— left the forked child alive, ptrace-stopped and unreaped; repeated calls could exhaust PIDs. It nowkill+waitpids on that path like the other overflow breaks. (4) DynamoRIO recording stack could pop a live region on deep nesting. The client’son_beginpushed only whiledepth < MAX_DEPTHbuton_endalways decremented, so nesting past 16 distinctly-named regions desynced the per-thread stack and silently dropped coverage withtruncatedleft0.on_beginnow tracks the true nesting depth unconditionally (matching the app side) and flags the tracetruncatedwhen a region falls outside the storable window — upholding the never-present-a- partial-trace-as-complete invariant.Native-trace “fidelity” gaps — three places a partial capture could escape without its
truncatedflag. The framework’s core invariant is that an incomplete trace is never presented as complete; a review found three leaks and they are now closed. (1) The whole-window single-step loops (asmtest_ptrace_trace_attached_windowedand the fork-internalasmtest_ptrace_trace_window_call) treated any loop exit as clean — so a window whose tracee died orexit()ed before reaching the return address was reported complete; they now flag the stream truncated unless the one clean terminator (pc == win_ret, or the async*stop) was reached. (2) The AMD MSR-direct LBR path (asmtest_amd_msr_trace) returned a partial branch stack as complete when a mid-stack MSR read failed; a short read now setstruncated. Both err toward false-truncated over false-complete.asmtest_hwtrace_arm_tid()reported a stale thread id after a single-step scope closed. The accessor’s documented contract is “the OS thread id that armed the active capture, or-1when none is active”, but the single-stepend()path (the default, most-portable backend, freshly wired into eight bindings) returned without clearing it — so it kept reading the last arming tid instead of-1. It now resets like the PT / AMD / whole-window paths already do; a regression assertion covers it.asmspy --tracesilently produced nothing for a function that runs only on a worker thread. The region engine attaches only the thread-group leader (unlike the whole-process syscall/stream engines, which SEIZE every thread), so a function executing on another thread was never single-stepped and the command exited cleanly with zero output. It now reports the region was never observed executing and points at--stream(which follows all threads).asmspyptrace-lifecycle hardening. A job-control group-stop (^Z/SIGSTOP/ tty stop) is now handled withPTRACE_LISTENinstead of being resumed, so a traced target can actually be suspended while watched. On OOM while seizing threads, an already-seized thread is now detached rather than left stranded seize-stopped. The ELF section-header walk in the symbol resolver now strides bye_shentsize(notsizeof(Elf64_Shdr)), so a non-standard object with a larger entsize resolves correctly instead of reading misaligned headers.Latent AMD-LBR test flake in nine bindings’ auto-select hwtrace test. Each binding mirrors the C reference
test_auto_resolve_traces_live: pickauto(BEST), trace a tiny five-instruction routine, assert the result. On a privileged AMD Zen 3+ hostautopicks AMD LBR, which faithfully truncates a too-fast-to-sample single-shot routine (socovered(0)is false) — the C reference and .NET already assertedcovered(0) || truncated, but the fix was never ported, so rust/cpp/python/lua/ruby/zig/node/java/go still asserted onlycovered(0)and would fail on such a host. All nine now assert the faithful invariant.The hwtrace options struct under-allocated the AMD-LBR fields in all seven FFI-mirroring bindings (8-byte OOB read in
asmtest_hwtrace_init). Whenlbr_period/branch_filterwere appended toasmtest_hwtrace_options_t(Zen 4/5 LBR work), only the .NET binding’s struct was updated; every other binding that hand-mirrors the struct still described the old 40-byte layout — Nodekoffi.struct, JavaOPTIONS_LAYOUT, Pythonctypes.Structure, Rust#[repr(C)], Ruby’s Fiddle packer, Go’s cgo typedef, and Lua’sffi.cdef.HwTrace.initpassed a 40-byte buffer, andasmtest_hwtrace_init’sg_opts = *optscopies the full 48 bytes — reading 8 bytes past the buffer on every init. Harmless for the SINGLESTEP backend (it ignores those fields), but for anAMD_LBRinit the garbage read could seed a spurious sample period / reduced branch filter, silently altering capture. All seven now mirror the 48-byte C layout (verified:koffi.sizeof == 48,ctypes.sizeof == 48; the rust/ruby/go/luadocker-hwtrace-<lang>lanes green). C++ (#include "asmtest_hwtrace.h") and Zig (@cImport) use the real header and were never affected. Surfaced by the adversarial review of the whole-window attribution work.Node binding: 64-bit trace-call results above 2^53 were silently rounded. Every fork/attach/stealth trace entry in the Node binding read the routine’s return (its RAX at the
ret) asNumber(readBigInt64LE(...)), which rounds any value pastNumber.MAX_SAFE_INTEGERthrough the double mantissa — so a routine returning a full 64-bit hash/id/pointer came back wrong, contradicting the documented “BigInt out of safe range” contract and the OOP capture forms’ exact-result guarantee. Added a_safeInthelper (Number when it fits the safe-integer range, else the exact BigInt) and applied it to all twelve result reads (callScoped,stealthTrace,windowCall,stealthWindow,traceCall/traceCallBlockstep/traceCallEx, and thetraceAttached*family). Surfaced by the adversarial review of the whole-window work; regression-tested with a leaf returning0x0102030405060708.Stealth stepper seized the wrong thread on a managed runtime (
getpid→SYS_gettid).asmtest_hwtrace_stealth_tracereverse-attached the helper togetpid()(the process leader), but on HotSpot the thread invoking the region is a JVM-created thread whose tid ≠ pid — so the helper single-stepped the wrong (idle primordial) thread and therun_tobreakpoint fired on the untraced calling thread, killing the JVM with a fatal SIGTRAP (exit 133). Node and CoreCLR were unaffected only because their calling thread happens to be the leader (tid == pid). Fixed to seize(pid_t)syscall(SYS_gettid)— the calling thread — matching what the windowed variantasmtest_hwtrace_stealth_trace_windowedalready did. Surfaced while adding the JavastealthTracewrapper; after the fix Java captures a complete, exact stealth trace.bindings-parity gate restored to green. The block-step / whole-window / snapshot commits added eight tier symbols wrapped only in the .NET binding, leaving the
check-bindings-parityCI gate failing with 75 missing (binding, symbol) pairs. The BTF block-step pair (asmtest_ptrace_blockstep_available,asmtest_ptrace_trace_call_blockstep) — siblings of the universally-wrappedasmtest_ptrace_trace_call— is now genuinely wrapped in all ten bindings, each with a self-skipping parity test asserting the block-step stream is byte-identical to the single-step stream. The managed-tier / C-level symbols (the §Z1 whole-window trio, §3.1(c)attribute_window, §D3trace_attached_windowed, and the AMD boundary snapshot) carry reasoned allow-list exemptions following the file’s existing conventions (the .NET tier keeps its real window-trio wraps).test_descent_stale_alarm_flagno longer flakes on loaded CI runners. The test spams the tracer with a 200 µsSIGALRMstorm to prove a stale L3 watchdog flag + EINTRs cannot abort a healthy L2 descent — but it left the descent’s real-time deadline at the 2 s default, which a loaded 2-core runner can legitimately exceed under 5000 interrupts/sec (a correct truncation, misread as the regression). The descent under test now carries an explicit 60 s deadline, so only the stale-flag bug it guards can fail it; the EINTR pressure is unchanged. (asmtest_hwtrace_call_scopedalso joins the parity allow-list under the same dotnet-only posture as the window trio, restoring the gate the lazy-arm commit tripped.)The scoped/windowed data-flow producers resolve r8d–r15b sub-register aliases.
gp_value(the register-file value reader in bothsrc/dataflow_ptrace.candsrc/dataflow_blockstep.c) anddfp_alias_shape(the F6 gap barrier’s alias-slice classifier) had cases foldingeax/ax/al/ahetc. to their 64-bit container but none forr8d/r8w/r8b..r15d/r15w/r15b— a step that wrote one of those aliases produced a def-use record with no captured value (value_validstayed false), and the gap barrier could not decide whether glue at risk had changed such a location (truncated). Both now fold every GP sub-register alias Capstone can emit on x86-64 to its container, exactly like the existing high-byte/32/16/8-bit cases.The block-step tier no longer compares or records architecturally undefined EFLAGS bits as if silicon defined them. A new explicit mnemonic(+count)-keyed table in
src/dataflow_blockstep.c(dfb_undef_flags) masks the undefined bits an instruction leaves out of both the coherence canary (regs_coherent, accumulated per replayed instruction and reset per block) and every captured EFLAGS write record (finalize_step), on both the single-step oracle and the block-step+replay paths — preserving their byte-identical property by construction, since both flow through the same shared classification. Coversand/or/xor/test(AF),mul/imul(SF/ZF/AF/PF),div/idiv(all six),bsf/bsr(CF/OF/SF/AF/PF), count-dependentshl/shr/sal/sarandrol/ror/rcl/rcr, andbt/bts/btr/btc(OF/SF/AF/PF); an instruction outside the table that touches flags at all is treated as fully flag-defining, matching every other x86 arithmetic instruction. New test hooksno_undef_mask(disables both mask sites — the negative control) andinject_flag_bit(forces a chosen bit to disagree right before the canary) land with the tier’s first opts-struct layout guard (asmtest_dataflow_blockstep_opts_layout). A dedicatedxor eax,eaxfixture — the AF-undefined case the tier’s primary oracle fixture deliberately avoided — proves AF reads 0 in every post-xor EFLAGS record on both paths while the trace stays byte-identical, and the canary discrimination checks prove the mask, not luck, is what tolerates it (an injected AF divergence is tolerated; the same injection withno_undef_maskset is caught).
1.1.0 — 2026-07-06¶
Fixed¶
Review-driven defect sweep (2026-07-02). Resolved the full backlog from the code-level review (54 findings) and the still-open 2026-07-01 / 2026-07-02 repo-review items, with a per-batch implementation note under
docs/summaries/. Highlights: AArch64 callee-savedd8–d15ABI checking + a correctedvm.s/structparam.s;SKIP()in SETUP/TEARDOWN reported as skip; JUnit XML made well-formed and no longer preceded by test stdout; hardware-trace truncation contract honored across the single-step / AMD-LBR / Intel-PT / code-image backends (block partition matches Unicorn/PT/DR); emulator SysV/AArch64 stack-and-register argument marshaling and a deterministic register reset per call on a reused handle; ptrace signal-forwarding, jitdump-truncation and tracee-reaping fixes; memory-safety and 64-bit-precision fixes across the Rust/Python/Node/Lua/C++/Java bindings; and Win64 runner teardown/DF/watchdog fixes. Build/CI: knob-aware object identity (SAN/COV/ASM_SYNTAX), header-prerequisite and PIC-object completeness, acheck-version+ third-party-version CI gate, publish tokens scoped to their step, the GPL corresponding-source release step, and a self-sufficient BTF-less eBPF fallback header. Added.mailmap.
Added¶
Scoped in-process tracing for .NET — the managed tier (§Z0–§Z5, §D0). The zero-config scope construct over the single-step hardware-trace tier:
using (new AsmTrace()) { … }captures whatever the thread executes (no region, noHwTrace.Init— the ctor auto-inits the portable backend and self-skips with a faithfulSkipReasonwhere it cannot run).byMethod: truelabels the captured window by managed method via an in-processMethodLoadVerboselistener (JitMethodMap), andwithRundown: truealso names warm + ReadyToRun BCL methods through a dependency-freeDOTNET_IPC_V1jitdump rundown over the runtime’s own diagnostics socket (no NuGet package, no launch knob). Results are data-first (Addresses,Methods,Disassembly,AsmMethod.Assembly/.Tier), withrenderPath: trueas the rendered opt-in. Labelling decodes against the code-image version live in the window (the map feedsasmtest_codeimage_trackper method load), so bodies that re-tier/move after the scope still render the bytes that ran.Named-method form —
AsmTrace.Method(delegate)(§D0.3). Trace one managed method’s own JIT’d body: resolution viaPrepareMethod+ the listener (jitdump rundown fallback for warm/R2R bodies), a region + step-over capture with exact offsets, andInvoke(args…)as the library-owned non-inlinable call site.outOfProcess: true(§D3) routesInvokethrough the concealed ptrace-stealth stepper — a bundled helper reverse-attaches and steps the body out of band, so the calling thread is never armed withEFLAGS.TF.Faithful-degradation surface.
HwTrace.DegradationNote()composes the tier ladder (Intel PT → AMD LBR → single-step → CoreSight, each with its skip reason, plus the ptrace fallback); cross-thread closes and overflows flagTruncated(native OS-tid assert + a complementary managed-thread guard);Disas.IsCall/IsBranch/IsRet/TryCallTargetclassify live instructions structurally. The packableAsmTestNuGet now shipsAsmTraceand the whole hwtrace wrapper alongside the bundled native payload.Eleven runnable .NET examples under
examples/dotnet/(whole-window, region, methods, rundown, assemblies, annotated, tiers, hotspots, coverage, callgraph, ptrace_native — plus the out-of-processptrace_dotnetattach demo), each split Program/Report, wired intomake hwtrace-dotnet-exampleandmake dev-dotnet. Validated on .NET 8 and .NET 9 (no diagnostics-IPC orMethodLoadVerbosedrift).Single-step native-trace tier: macOS-Intel front-end. The exact, unprivileged EFLAGS.TF (
#DB→SIGTRAP) single-step backend now runs in-process on x86-64 macOS, not just Linux — the first Phase-5 front-end a Linux CI host (or Docker-on-Mac, whose containers are Linux) cannot exercise. XNU delivers the single-step trap as a BSDSIGTRAP, so re-assertingTFin the saved thread state re-arms stepping acrosssigreturnexactly as on Linux; the only platform deltas are the feature-test macro (_DARWIN_C_SOURCE) and the mcontext field access (uc_mcontext->__ss.__rip/__rflagsvs.gregs[REG_RIP]/[REG_EFL]), both isolated behind shims insrc/ss_backend.c.asmtest_hwtrace_available(SINGLESTEP)now returns 1 on x86-64 Darwin and the wholeasmtest_hwtrace_*facade (region table,init/register,begin/end/begin_scope/render_scope) drives it; thesrc/hwtrace.cgateHWTRACE_LIFECYCLEis a superset of__linux__, so the Linux path is unchanged (verified:make hwtrace-test61 pass on macOS;make docker-hwtrace178 pass on Linux). The binding-facing W^X executable-memory helper (asmtest_hwtrace_exec_alloc/_exec_free,src/hwtrace.c) now runs on x86-64 Darwin too — itsPROT_NONE→RW→RXmmap/mprotectpath is plain POSIX and identical to the Linux one — so the per-binding single-step lanes are reachable natively on macOS, not just the C suite (verified on this host:make hwtrace-{python,cpp,ruby}-testpass, with the Linux-only ptrace/codeimage backends self-skipping). Off-platform hosts self-skip with “single-step backend is x86-64 Linux/macOS only (Windows/AArch64 planned)”.Scoped in-process tracing (the
using/RAII/withmodel). A cooperative, developer-ergonomics face of the tracing machinery: bracket a region of a program’s own code with a scope construct —using (new AsmTrace())in C#, RAII in C++/Rust,within Python,deferin Go/Zig, a block/try-with-resources elsewhere — and get back the assembly that executed inside it, rendered on scope close. Implemented across all ten language bindings over a small shared C/decode core (error-returningasmtest_hwtrace_try_begin, arming-thread assert,asmtest_hwtrace_render, idempotent-by-name region registration, per-thread single-step state, a recorder-backed image adapter, symbolize-and-bucket, and theasmtest_hwtrace_stitchasync-hop merge core). Linux-only; self-skips to a recorded no-op where no faithful backend is available. See docs/scoped-tracing-implementation.md and docs/internal/archive/plans/scoped-inprocess-tracing-plan.md.§Z0/§Z1 the aspirational empty-ctor form —
using (new AsmTrace()). A region-free whole-window scope with noNativeCodeand no[base,len): new C entry pointsasmtest_hwtrace_begin_window/_end_window/_render_windowover a whole-window frame mode inasmtest_ss_begin_window(the single-step handler records ABSOLUTE RIPs into the bounded ring, overflow →truncated), rendered from live self memory. The .NET reference shim gains the parameterlessnew AsmTrace()ctor +SkipReason(transparent self-skip). This is the single-step WEAK tier — native-leaf only, on any x86-64 Linux (test_wholewindow_singlestep,make docker-hwtrace→ 201/0;.NETmake docker-hwtrace-dotnet→ 33/0). The STRONG whole-window PT / AMD LBR tiers, arbitrary-managed-method capture, and the other nine binding shims remain forward-look. See docs/internal/plans/scoped-tracing-zeroconfig-plan.md.§D3 concealed ptrace-stealth stepper — now a bundled standalone binary. The hardware-free scope path (Zen 2 / Docker-on-Mac) reverse-attaches a helper to the caller (
PR_SET_PTRACER+PTRACE_SEIZE) and single-steps the region out of band. Its stepping body + discovery moved tosrc/stealth_helper.cso the same code runs either as an in-process forked child (the fallback) or as the standaloneasmtest-stealth-helperbinary — a real separate process the managed packages can ship — which the caller discovers via a dladdr-sibling lookup (mirroring the DynamoRIO payload) or theASMTEST_STEALTH_HELPERoverride, handing the shared trace over a memfd. New$(BUILD)/asmtest-stealth-helperbuild target +install-stealth-helper;test_ptrace_scoped_stealthasserts both paths reconstruct byte-identical offsets on any ptrace-capable Linux. The helper is bundled into every managed package payload (NuGetruntimes/<rid>/native, npm/Maven/gem/rock, the Python wheel_libs/) besidelibasmtest_hwtrace,$ORIGIN-rpath’d so it resolves the co-vendored Capstone in-package, and asserted present + rpath’d (and not leaked into a darwin slot) by a fail-closedpackage-libs-verifygate. Only the live-JIT cross-process address channel (needs a running managed runtime) remains forward-look.
Call descent for the out-of-process ptrace tracer. The single-step tracer (
asmtest_ptrace.h) can now optionally FOLLOW the calls a traced region makes instead of only stepping over them, at four opt-in levels (asmtest_descent_t):OFF(today’s behaviour),RECORD_EDGES(record eachcall-site → calleeedge, still step over),DESCEND_KNOWN(single-step into resolvable callees — an allow-set of method regions or an optional resolver callback — stepping over the rest), andDESCEND_ALL(into everything, default off, denylist + instruction-budget + real-time-watchdog gated). The flatasmtest_trace_tis unchanged — it is always frame 0, byte-identical across all levels; descent records into a separate opaque handle read through scalar accessors (edges + nested per-callee frames), soasmtest_trace_tstays ABI-frozen and every binding adds accessor calls, not a struct layout. New entry pointsasmtest_ptrace_trace_call_ex/_trace_attached_ex/_trace_attached_versioned_exthread the handle through the existing loops; the non-_exsymbols are unchanged (descent == NULL).The descender is a return-address shadow stack with an exact pop predicate (
PC == ret_addr && SP == caller-pre-call-SP && the just-stepped insn is a return) plus an SP-sweep for non-local exits (longjmp/unwind/sigreturn), same-region recursion as a distinct frame (with a recursion +max_depthcap), per-instruction byte windows viaprocess_vm_readv, benign-signal forwarding on the live path, and a backend-ownedITIMER_REAL/SIGALRMwatchdog so a blocked syscall in a descended callee self-truncates rather than hanging. AArch64 gained aNT_ARM_HW_BREAKhardware-breakpoint step-over path (the W^X JIT-heap fallback x86-64 already had). L3 is documented as best-effort / expected-to-perturb on a live managed runtime (the cross-thread lock-inversion deadlock vector is not fully mitigable) — see analysis/jit-runtime-tracing.md.Surfaced in all ten language bindings (a
Descentwrapper + descendingtrace_call_ex, with idempotent free and the per-FFI address/upcall hazards handled), pinned by a newptrace_descentconformance-corpus tier and the header-grep parity gate; the resolver callback ships to the six upcall-safe FFIs (Python/Go/Node/Java/.NET/Lua) and Rust/Ruby/ C++/Zig expose the allow-set only. Newjit_trace *-descend/*-descend-alldemo lanes (make docker-hwtrace-jit-dotnet-bcl-descend, …). See docs/native-tracing.md (“Call descent levels”) and docs/internal/archive/plans/call-descent-plan.md.
Clean-room install test — every bundled binding, on Linux and macOS, in CI.
make clean-room-test(any host) /make macos-clean-test(darwin alias) packages each binding that ships a native payload, installs it fresh into a throwaway prefix, loads it with everyASMTEST_*/DYLD_*/LD_*override scrubbed and the cwd outside the checkout, then asserts the native library it actually resolved lives under that fresh install — never a leaked devbuild/tree, a Homebrew dylib, or/usr/local. So “install fresh, noASMTEST_LIB” is proven, not trusted: the prior per-binding smokes only checked a tier was available, which a leakedbuild/or Homebrew dylib also satisfies. Bindings whose toolchain is absent self-skip; a real leak fails the run.All six dlopen bindings are covered — Python, Ruby, Node, Java, Lua, and .NET. Each core loader gained a resolved-path accessor:
library_path(Ruby/Lua),libraryPath()(Node/Java),Emu.LibraryPath(.NET, viaProcess.Modulesso it reports the real loaded path however P/Invoke resolved the name), and Python’s existingpython -m asmtest --where. The link bindings (C++/Rust/Go/Zig) ship source and linklibasmtestthemselves — no bundled payload to leak-check — so they are intentionally out of scope.Verified in Docker per language:
make docker-clean-<lang>builds the binding’s isolated image and runs the clean-room test in it withCLEANROOM_ONLY=<lang>— so a self-skip fails the lane (a missing toolchain can’t pass vacuously);make docker-clean-roomruns the set. A newclean-roomCI job (matrix over Ruby/Node/Java/.NET/Lua) gates every push, complementing the conformancebindingsjob (which loads the devbuild/tree). Python’s clean-room stays in the existing release.yml python job (which asserts on the repaired wheel — self-containing the wheel needsauditwheel/buildthe lean test image omits).New reusable pieces:
scripts/clean-env.sh— a sourceable env scrubber that pinsDYLD_FALLBACK_LIBRARY_PATHto/usr/librather than unsetting it (unsetting reverts to a dyld default that includes/usr/local/lib, where a Homebrew copy could still satisfy a bare-leaf load);scripts/assert-clean-path.sh— the leak guard (rejects the checkout,/opt/homebrew,$HOMEBREW_PREFIX,/usr/local; allows a temp extraction, e.g. the jar’s); andscripts/clean-room-test.sh— the cross-platform per-binding orchestrator (the first reusable local one; the release.yml smokes can call it next, per the plan’s Track E). (macOS clean-test plan, Track A)
The native-trace tiers now ship inside the packages. Both optional tiers — DynamoRIO (
libasmtest_drapp+libasmtest_drclient+ the pinnedlibdynamorio) and hardware trace (libasmtest_hwtrace) — are staged into the Linux payload slots bymake package-libs, so a freshpip install/ gem / npm / nupkg / jar / rock runsNativeTrace/HwTraceon a capable host with no manualmake shared-*and noDYNAMORIO_HOME, exactly as the emulator/Keystone/Capstone tiers already do. drtrace islinux-x86_64only (DynamoRIO auto-fetched viascripts/fetch-dynamorio.sh); hwtrace bundles on every Linux slot (single-step + ptrace always; the Intel PT / AMD / CoreSight decoders self-skip off the hardware they need). macOS/arm64 slots simply omit the Linux-only tier and the wrapper self-skips (available()→ false) — no API oravailable()behavior change, bundling only removes the build step.Every binding’s
drtrace/hwtraceloader learned a bundled-package candidate (env override → bundled slot → devbuild/→ system) and alibrary_path()self-report (python -m asmtest --where, and the equivalent accessor in the Go / Rust / Ruby / Node / Java / .NET / Lua / Zig wrappers) so a clean-room test can assert the tier resolved from the package, not a leaked checkout.A package-bundled
libdynamorioself-locates next tolibasmtest_drapp(viadladdr), so the DynamoRIO tier works with zero configuration —dlopendoes not consult a library’s ownRUNPATH, so drapp finds its sibling explicitly.Licensing unchanged in character: DynamoRIO (BSD-3-Clause core), and the “full” hwtrace’s libipt/OpenCSD/libbpf, are all permissive —
collect-licenses.shemits each only when the lib is actually staged, adding no copyleft beyond the existing Unicorn/Keystone GPL-2.0. The four source-distributed bindings (Rust/Zig/C++/Go) ship no binary payload, so their consumers buildshared-drtrace/shared-hwtracethemselves (documented, not bundled). (bundle-native-trace-tiers plan)
Native runtime tracing (two optional tiers). A third execution tier that traces code running natively, in-process, complementing the Unicorn emulator trace. Both fill the same engine-neutral
asmtest_trace_tshape (now extracted intoinclude/asmtest_trace.h+src/trace.c, shared by all backends) and the Capstone annotation layer renders any backend’s offsets. (Native runtime tracing)DynamoRIO in-process tier (
asmtest_drtrace.h,libasmtest_drapp+ CMake-builtlibasmtest_drclient.so):dr_app_*in-process attach with an enforced lifecycle state machine, begin/end region markers, basic-block and instruction coverage, and host-native W^X executable-memory allocation (asmtest_exec_alloc/asmtest_asm_exec_native). Uses DynamoRIO’s BSD core API only — no drmgr/drwrap, so no LGPL-2.1 obligation. Native-trace wrappers for every language binding — Python (asmtest.drtrace), C++, Rust, Go, Node, Java, .NET, Ruby, Lua, and Zig — each exposing the sameNativeTrace/NativeCodesurface and dlopen-loadinglibasmtest_drappat run time, so the core binding never link-depends on DynamoRIO and each wrapper self-skips (available()→ false) when the tier is absent. Targetsdrtrace-test,shared-drtrace,drtrace-client,drtrace-<lang>-test,drtrace-bindings-test,docker-drtrace, anddocker-drtrace-bindings(container lanes with DynamoRIO installed). Gated onDYNAMORIO_HOME; self-skips when absent. All wrappers are verified against a real in-process DynamoRIO in Docker: C++/Ruby/Java/Lua/Zig/Rust/Go trace live; Node and .NET self-skip there (in-process DynamoRIO can’t take over a JIT/GC runtime’s threads — the managed-runtime limitation, where Intel PT is the recommended backend). Linux x86-64.Hardware-trace tier (
asmtest_hwtrace.h,libasmtest_hwtrace): four backends behind one API, oneavailable()gating chain, and oneasmtest_trace_tsink. Intel PT capture viaperf_event_open+ libipt decode with branch-boundary block normalization; AMD LBR (Zen 3 BRS / Zen 4 LbrExtV2, 16-deep, exact within window thentruncated); ARM CoreSight (OpenCSD) scaffold; and single-step (EFLAGS.TF→#DB/SIGTRAP), the portable backend that records the same exact/complete offsets on any x86-64 Linux host (Intel, any-Zen AMD, VM, CI, plain container) with no PMU, perf_event, privilege, or decoder library.asmtest_hwtrace_available()encodes the full detect-and-skip chain; the PT/AMD/CoreSight backends self-skip off the bare-metal hardware they need (the common case). Targetshwtrace-test,shared-hwtrace,hwtrace-<lang>-test,hwtrace-bindings-test,docker-hwtrace, anddocker-hwtrace-bindings(plain unprivileged container lanes); auto-detects libipt/OpenCSD via pkg-config.Hardware-tier backend auto-selection.
asmtest_hwtrace_resolve(policy, out, cap)returns the host’s available backends most-faithful first (Intel PT > AMD LBR > single-step > CoreSight);asmtest_hwtrace_auto(policy)returns the single best pick ready toinit(orASMTEST_HW_EUNAVAIL).policyisASMTEST_HWTRACE_BESTorASMTEST_HWTRACE_CEILING_FREE(drops the one fixed-window backend, AMD LBR — what a caller re-resolves under after a trace comes backtruncated). On any x86-64 Linux host the cascade is non-empty (single-step is the floor), soauto()never fails there. Exposed through the C API and every language wrapper — Python, C++, Rust, Go, Node, Java, .NET, Ruby, Lua, Zig — each surfacingresolve/auto(C++ usesauto_select, sinceautois a keyword) withBEST/CEILING_FREEpolicy constants, plus a per-binding self-test of the selection invariants and a live auto-picked trace. Scope is the hardware tier’s own backends; a cross-tier fall to DynamoRIO/the emulator stays a deliberate, fidelity-aware caller decision.Cross-tier trace orchestration.
asmtest_trace_resolve(policy, out, cap)/asmtest_trace_auto(policy, &choice)(asmtest_trace_auto.h,src/trace_auto.c) are the front-end over all three tiers, not just the hardware backends: they walk the full descending-fidelity cascade — Intel PT → AMD LBR → DynamoRIO → single-step → CoreSight → emulator (DynamoRIO ranks above single-step because its code cache runs at native speed while single-step pays a per-instruction kernel round-trip) — and returnasmtest_trace_choice_tdescriptors{tier, backend, fidelity}. It callsasmtest_hwtrace_available()directly and dlopen-probeslibasmtest_drapp(via$ASMTEST_DRAPP_LIB) for the DynamoRIO tier, so it hard-links neither the DynamoRIO nor the emulator library — the three stay decoupled. Thepolicybitmask composesASMTEST_TRACE_BEST,ASMTEST_TRACE_CEILING_FREE(drop AMD LBR; re-resolve under it aftertruncated), andASMTEST_TRACE_NATIVE_ONLY— the flag that forbids the native→emulator fidelity crossing: under it the emulator floor is dropped, so a host with no native tier resolves toASMTEST_HW_EUNAVAILrather than silently downgrading real-CPU execution to an isolated guest. Shipped inlibasmtest_hwtraceand exposed through every language wrapper (Python/Rust/ Go/Lua/Rubyresolve_tiers/auto_tier, camelCaseresolveTiers/autoTierfor C++/Node/Java/.NET/Zig,ResolveTiers/AutoTierfor Go), each with a per-binding self-test of the cross-tier invariants. This is the cross-tier front-end the trace parity matrix flagged as the remaining gap.Out-of-process single-step backend (W2).
asmtest_ptrace_trace_call(code, len, args, nargs, &result, trace)(asmtest_ptrace.h,src/ptrace_backend.c) is the out-of-process sibling of the in-processEFLAGS.TFstepper: a tracer parentPTRACE_SINGLESTEPs a forked tracee that runs the registered routine, reads the program counter from the child’s register file at each stop, and reconstructs the same exact offsets in the parent — ordered in-region instruction offsets and the identical single-entry/ends-at-branch block partition the in-process stepper, Unicorn, DynamoRIO, and Intel PT produce — with no shared memory (the parent observes every step) and no library or privilege beyond ptrace of one’s own child. Because it touches none of the tracee’s signal disposition or code cache, it is the exact path for a JIT/GC managed runtime (JVM/.NET/Node) and the recommended managed-runtime backend on AMD (no Intel PT), and is the only single-step form possible on AArch64 (whoseMDSCR_EL1.SSis kernel-only). It runs on Linux x86-64 and AArch64 off one body: the AArch64 arm reads the program counter + integer return register viaPTRACE_GETREGSET/NT_PRSTATUS(AArch64 has noPTRACE_GETREGS) and decodes block lengths withASMTEST_ARCH_ARM64Capstone, while the fork/SIGSTOP/step/wait flow and the SysV/AAPCS64 register-arg call are shared. Built intolibasmtest_hwtrace;make hwtrace-testexercises it live — including in a plain unprivileged container — asserting byte-for-byte parity with the in-process stepper plus a 62-instruction loop (no depth ceiling). This lands the Linux x86-64 front of the Zen 2 single-step plan Phase 5 (W2).asmtest_ptrace_trace_attached(pid, base, len, &result, trace)extends it to the foreign-process case — the building block for tracing a managed runtime: it traces a region in a separate, already-running process you have attached to externally (the caller owns thePTRACE_ATTACH/DETACHpolicy), single-stepping the target from its current stop and reading the region bytes from the target viaprocess_vm_readv(so the tracer does not share the target’s memory). A live test attaches to a child that never calledPTRACE_TRACEMEand reconstructs the same[0,3,6,c,11]stream out of band. Two region resolvers turn the attach primitive into “point it at a running process”:asmtest_proc_region_by_addr(pid, addr, &base, &len)finds the executable mapping containingaddrin/proc/<pid>/maps(one interior address → the whole region to trace), andasmtest_proc_perfmap_symbol(pid, name, &base, &len)parses/tmp/perf-<pid>.map— the text format V8/Node, .NET, and OpenJDK (+perf-map-agent) write soperfcan symbolize generated code — to recover a JIT method’s(base, len)by name. A live test discovers a foreign process’s region from/proc/<pid>/mapsusing only an interior address and traces that region (no hardcoded base).asmtest_ptrace_run_to(pid, addr)closes the uncontrolled-timing gap that the rest of the flow left open:trace_attachedneeds the target stopped at the method entry, but a real managed runtime calls a JIT method on its own schedule, so you cannot attach at the right instant.run_toplants a software breakpoint ataddr(PTRACE_POKETEXT—int3on x86-64,brkon AArch64 — which patches an r-x text page the way a debugger does),PTRACE_CONTs until the program itself next calls in, then removes the breakpoint and rewinds the PC, leaving the target stopped exactly ataddrfortrace_attached(also hardened to record the entry instruction from either stop convention). A live test now drives the complete real-JIT flow with no cooperative go-flag — a child publishes a perf-map and calls its routine in a loop, the tracer resolves it by name, attaches,run_tos the entry, and traces that invocation to the exact[0,3,6,c,11]stream — completing the managed-runtime flow: resolve → attach →run_to→trace_attached→ detach. Call-depth awareness removes the last “leaf only” restriction:trace_callandtrace_attachedpreviously treated the first step out of the region as the return, so a routine that called a runtime helper (GC barrier, allocation, PLT stub — the norm for real JIT methods) truncated at the first call. The stepper now decodes the region-exit instruction (asmtest_disas_is_call, CapstoneCS_GRP_CALL) and runs a call-out to its return address at native speed (a breakpoint-cont over the callee, not a per-instruction step), resuming recording after it; only a genuine return ends the trace. The region’s own instructions are recorded, the helper skipped, the real return still found (test_ptrace_callout); without Capstone it falls back to the prior leaf-only behaviour. Real-runtime validation lanes point the whole pipeline at a live JIT (not a fixture) via one argv-driven harness (examples/jit_trace.c):make docker-hwtrace-jittraces Node.js (V8) —node --perf-basic-prof --no-turbo-inliningon a hot function, resolving the method from V8’s real perf-map, attaching to the live multi-threaded GC’d runtime,run_toing the entry, and single-stepping one invocation to recover the actual TurboFan code for(a+b)|0;make docker-hwtrace-jit-dotnettraces .NET (CoreCLR) —Program::Add→lea eax,[rdi+rsi]; ret(withDOTNET_TieredCompilation=0for a stable address), tracing .NET’s W^X code heap as-shipped via the hardware- breakpoint fallback below; andmake docker-hwtrace-jit-javatraces OpenJDK (HotSpot) —Hot.asmtjit→lea eax,[rsi+rdx]inside the real C2 nmethod (entry barrier and stack-bang and all), JIT’d once via-XX:-TieredCompilationand kept a standalone callable body with-XX:CompileCommand=dontinline. HotSpot needs two wrinkles the others don’t, both handled in the harness: it does not stream a perf-map, so the lane drivesjcmd <pid> Compiler.perfmapto materialize one for the live process; and thejavalauncher runs Javamain()on a secondary OS thread (not the primordial one V8/CoreCLR use), so the harness picks the spinning loop thread by CPU delta andPTRACE_ATTACHes exactly it — a software-breakpoint trap on an un-traced thread is fatal. A watchdog makes a re-tiered/moved address self-skip rather than hang, so the lanes never flake (resolve + attach are asserted against the runtime’s real output; the trace is asserted-or-skipped). Hardware-breakpointrun_to:run_until(behindrun_toand the call-out step-over) defaults to a softwareint3but transparently falls back to an x86-64 hardware execution breakpoint (DR0/DR7 viaPTRACE_POKEUSER) whenPTRACE_POKETEXTis refused — i.e. on a W^X JIT code heap whose executable page is not writable (.NET’s default double-maps it; POKETEXT failsEIO). The hardware breakpoint writes no code (so it traces W^X as-shipped) and is per-thread (so it never traps a sibling runtime thread, unlike a process-wideint3); software stays the default andASMTEST_PTRACE_HW_BPforces the hardware path (used to validate it deterministically on ordinary memory —test_ptrace_callout). AArch64 hardware breakpoints (NT_ARM_HW_BKPT) are a follow-on. A third lane,make docker-hwtrace-jit-jitdump, validates the binary jitdump byte source against real output:node --perf-profwrites a realjit-<pid>.dump, and the lane recovers a method’s recorded code bytes withasmtest_jitdump_findand checks them three ways — the address agrees with V8’s own perf-map, the bytes disassemble to real x86-64, and they match the live code at that address (jitdump’s temporal-capture guarantee) — the first validation ofasmtest_jitdump_findagainst a real jitdump rather than a synthetic fixture. A second jitdump producer lane,make docker-hwtrace-jit-java-jitdump, validates the reader against a jitdump from a different runtime and encoder: OpenJDK HotSpot has no native jitdump, so the lane loads the perf project’s JVMTI agent (libperf-jvmti.so, fromlinux-tools) with-agentpath, which records every C2 method to a real jitdump. It names methods in JVM descriptor form (LHot;asmtjit(II)I, unlike V8’s symbol) and interleaves debug/unwinding records the reader must skip, so it exercisesasmtest_jitdump_findon a genuinely independent encoder; the recovered bytes are checked against the live code and against HotSpot’s ownjcmd Compiler.perfmapaddress (two independent HotSpot outputs). A third jitdump producer lane,make docker-hwtrace-jit-dotnet-jitdump, covers .NET CoreCLR — which, unlike HotSpot, writes a real/tmp/jit-<pid>.dumpnatively (no agent) underDOTNET_PerfMapEnabled=1, naming the method identically in the perf-map and the jitdump. So it shares the sametrace_jitdumppath as V8 (one routine parameterized by runtime): recoverProgram::Add’s recorded bytes (lea eax,[rdi+rsi]; ret) and validate them four ways — disassemble, match the live code, and agree with CoreCLR’s own perf-map address/size. The binary jitdump is now validated against all three managed runtimes (V8, HotSpot, CoreCLR).asmtest_jitdump_find(path, pid, name, &entry, bytes, cap, &len)reads the richer binary jitdump image (jit-<pid>.dump— CoreCLR, HotSpot, V8; whatperf inject --jitconsumes), resolving a method to its(code_addr, code_size)and its recorded native code bytes (which the text perf-map cannot give). Because eachJIT_CODE_LOADrecord is timestamped, a method re-emitted at a reused address (tiered/OSR recompilation) resolves to the latest body — the temporal same-address-different-bytes problem; endianness is auto-detected and non-LOADrecords skipped. (The text perf-map remains the portable lowest common denominator for JITs that only emit symbols.) The full foreign-process toolkit —available/skip_reason,trace_call,trace_attached,run_to,region_by_addr,perfmap_symbol,jitdump_find— is exposed through every language wrapper (aPtraceclass / module surfacing the same methods, idiomatic per language), each with a per-binding self-test of the live-testable subset (out-of-processtrace_callparity,/proc/mapsand perf-map resolution, and a binary-jitdump round-trip), andasmtest_ptrace.his now covered by the binding function-surface parity gate. The/proc+ jitdump code-region readers (asmtest_proc_*,asmtest_jitdump_find) are pure file parsing, so they build and run on any Linux arch and are validated live on AArch64 (in alinux/arm64container, where the perfmap and jitdump tests pass). The single-step trace capture on AArch64 awaits a real host:asmtest_ptrace_available()is a cached, hang-proof self-probe (boundedWNOHANGpolling) that returns 0 under qemu-user — which does not emulate the ptrace tracer/tracee relationship at all — so the stepper self-skips on emulation just as the PT/CoreSight tiers self-skip off their hardware; the AArch64 fixtures are decode- and execute-validated under qemu, with only the live single-step stream pending Apple-Silicon / Linux-ARM64 / Windows-on-ARM hardware.Time-aware code-image recorder (
asmtest_codeimage). A userspacePERF_RECORD_TEXT_POKE: it records a timestamped timeline of a process’s code regions soasmtest_codeimage_bytes_at(img, addr, when, …)returns the bytes that were live at trace-positionwhen— the correct answer for a JIT whose code is patched, freed, or has its address reused mid-trace, where a single lateprocess_vm_readvsnapshot returns the wrong bytes (the temporal problem the JIT-runtime-tracing analysis calls “the innovative, buildable core”, approach #2). Change detection is soft-dirty (/proc/<pid>/clear_refsto arm + the soft-dirty PTE bit to detect, read via thePAGEMAP_SCANioctl where available, else by parsing/proc/<pid>/pagemap), which works cross-process — the foreign-JIT case — needing only permission to read the target. The W2 stepper consumes it:asmtest_ptrace_trace_attached_versioned(pid, base, len, img, when, &result, trace)decodes a foreign region’s blocks against the time-correct bytes instead of a live snapshot (the existingtrace_attachedis theimg == NULLcase, unchanged). An optional eBPF emission detector (a CO-RE program onmprotect/mmap/memfd_create, filtered to the target PID namespace viabpf_get_ns_current_pid_tgid, events drained from abpf_ringbuf) tells the recorder when code appears so it snapshots on thePROT_EXECedge instead of polling; it is built only whenclang+libbpf+bpftoolare present (-DASMTEST_HAVE_LIBBPF) and self-skips otherwise, with the userspace soft-dirty path as the always-available fallback. Validated live: the same-address-different-bytes temporal proof and the versioned W2 trace run inmake hwtrace-test/make codeimage-teston any x86-64 Linux host (no privilege); the eBPF detector is validated inmake docker-hwtrace-codeimage(a--cap-add=BPF,PERFMONcontainer — not privileged), observing a realmprotect(PROT_EXEC)emission edge. Exposed across all ten language bindings (aCodeImagewrapper) and covered by the binding function-surface parity gate. (asmtest_codeimage.h,src/codeimage.c,bpf/codeimage.bpf.c)ARM CoreSight reconstruction core (host-validated).
src/cs_backend.cis now split like the AMD backend: its decoder-independent reconstruction coreasmtest_cs_reconstruct(arch, ranges, n, base, len, trace)turns the ordered instruction ranges an ETM/ETE decoder emits (OCSD_GEN_TRC_ELEM_INSTR_RANGE) into the same instruction-offset stream and single-entry/ends-at-branch block partition the Intel PT backend produces. It is host-validated without a CoreSight board (examples/test_hwtrace.ctest_cs_reconstruction, the analogue of the AMD synthetic-branch-stack test), asserting byte-for-byte parity with the PT/AMD/single-step backends over the shared fixture. The remaining half — the live OpenCSD decode tree (ocsd_create_dcd_tree+ ETMv4/ETE decoder + memory accessor feeding ranges to the core) — needs libopencsd and a real AArch64 CoreSight board to write and validate, so per the project’s no-untested- hardware-code rule it is not yet implemented;asmtest_cs_decoder_present()still returns 0 and the tier self-skips on every host, but the half the board glue will feed is now proven (CoreSight advances from bare scaffold to validated reconstruction).AMD LBR Tier-B stitching (host-validated). AMD’s branch stack is 16 deep, so a single-snapshot (Tier-A) reconstruction sets
truncatedpast 16 taken branches. Tier-B lifts that ceiling:asmtest_amd_stitch(samples, nrs, n, out, cap, &gap)splices the overlapping windowssample_period=1emits (one per taken branch, consecutive windows overlapping by 15 edges) into one gapless taken-branch sequence — for each window it takes the smallest shift that still overlaps the accumulated tail and appends only the new edges, so a loop’s repeated identical edges stitch correctly; lost overlap (≥ a full window dropped to throttling) sets*gap.asmtest_amd_decode_stitched(...)replays the stitched sequence through the sharedamd_replayloop (factored out ofasmtest_amd_decode) without the 16-entry overflow flag. Host-validated without hardware (like the Tier-A reconstruction):test_amd_stitchsynthesizes an 18-iteration loop’s windows, stitches them, and reconstructs the complete trace (55 instructions, two blocks, not truncated) where a single Tier-A 16-window truncates — plus gap detection. Now wired into the live capture and Zen 5-validated:hwtrace_end_amdcollects every branch-stack sample in the perf data ring (time order) and, when the richest single window overflowed (best_nr >= 16), stitches them and decodes past the ceiling; the small-routine path (best_nr < 16) is unchanged. Completeness is gated on the precise loss signals — a stitch gap OR aPERF_RECORD_LOST/PERF_RECORD_THROTTLErecord (the non-overwrite ring drops the newest samples on overflow and emitsLOST, the signal the gaplessly-stitching survivors cannot otherwise reveal) → faithfullytruncated. On a Zen 5 (Ryzen 9 9950X,make docker-hwtrace-amd) a 20000-trip loop reconstructs ~290 instructions (≈95 stitched branches, far past one 16-deep window’s ~49) and stays truncated, as the perf ring size andsample_period=1throttling require; the live path is complete only for runs that fit the ring and survive throttling, beyond which DynamoRIO (no ceiling) remains the answer. (AMD LBR plan, Phase 5.)
Win64 wide-vector (AVX2 256-bit) capture. The Win64 capture trampoline topped out at 128-bit (
xmm); a routine’s fullymmresult under the Microsoft x64 ABI couldn’t be inspected past its low 128 bits. Newasm_call_capture_vec256_win64is the Win64 analog of the SysVasm_call_capture_vec256: it marshals four 256-bit args intoymm0..3, calls the routine, and captures the wholeymm0..15file into avec256_t[16], saving/restoring the callee-saved low 128 ofxmm6..15(the upper 128 is volatile per the ABI) andvzeroupper-ing on exit. Awin64_vaddpd_ymmroutinea
test_capture_win64case assert the full 256-bit return (all four doubles, exercising the upper-128 lanes the 128-bit path can’t see), self-skipping off-AVX2 via a local CPUID/XCR0 probe. Verified on both Win64 lanes — the nativems_abilane and the PE/Wine lane (Wine runs PE instructions on the host CPU, so it is real AVX2). Track D of the post-v1.0 expansion plan (the “Win64 wide path” follow-on); AArch64 SVE remains staged, hardware-gated on a runner that can execute it.
AVX-512 512-bit (
zmm) capture — across the core, Win64, and all ten bindings. The wide-vector path now reaches 512 bits, validated on real AVX-512 silicon (a Zen 5 / Ryzen 9 9950X). A newvec512_t(64 bytes) andasm_call_capture_vec512marshal eightzmmargs intozmm0..7and capture the fullzmm0..31file — AVX-512 doubles the register count as well as the width, so the capture is avec512_t[32](vsvec256_t[16]for AVX2) — using the EVEX-encodedvmovdqu64(required to reachzmm16..31). It is gated on a realasmtest_cpu_has_avx512f()(CPUID +XCR00xe6: opmask +ZMM_Hi256+Hi16_ZMM); theASM_VCALL512*macros and every binding wrapper self-skip where AVX-512 is absent, so the same suite runs everywhere. Shipped in both the GAS and NASM trampolines (src/capture.{s,asm}) and the Win64 path (asm_call_capture_vec512_win64, Microsoft x64 ABI, low-128xmm6..15saved/restored), withASSERT_VEC512_EQ+asmtest_assert_vec512_eqfor the 64-byte lane compare. Avec_add8dcorpus routine (vaddpd zmm, 8 packed doubles) andwin64_vaddpd_zmmassert the full 512-bit return — the 8th double lane proves the bits neither the 128- nor 256-bit path can see. Exposed across all ten language bindings (Python, C++, Rust, Go, Node, Java, .NET, Ruby, Lua, Zig) ascapture_vec512/cpu_has_avx512fanalogs with per-binding parity tests. Closes the AVX-512 half of gap #4 in the post-v1.0 expansion plan (its “no AVX-512 silicon” caveat is now lifted on this host); AArch64 SVE remains the staged remainder.Publishable packages (Track A): library-exposing artifacts, self-locating native libs, and a dry-run release workflow. The packaging scaffolding now produces artifacts a registry could ship. Each
make <lang>-packageexposes the reusable library module rather than the conformance test runner (the Ruby gem shipsasmtest.rb, npmasmtest.js, the rockasmtest.lua, the JAR theAsmtestclasses, and the NuGet packageAsmTest.dllvia a new SDK-styleasmtest-lib.csproj/dotnet pack), and bundles one native slot per platform present inbuild/dist/native/(the host slot locally; all four when a release has the CInative-all). The dlopen bindings self-locate their bundled native lib whenASMTEST_LIBis unset — Node/Ruby/Lua/Java fall back tonative/<os>-<arch>/next to the module (Java extracts the jar resource to a temp file; .NET resolves via theruntimes/<rid>/native/RID layout; Python already used_libs/) — so an installed package works out of the box. A newrelease.ymlbuilds the cross-platformnative-all, then per binding packages → installs the artifact fresh → smoke-tests the bundled-native load (ASMTEST_LIBunset) → dry-run publishes (twine check,npm publish --dry-run,cargo publish --dry-run); the live push is gated behind per-ecosystem token secrets, so it runs end to end with no credentials. Python wheels are built per platform: asetup.pytags the wheelpy3-none-<platform>(platlib, since it bundles a native lib), and the workflow repairs each into a self-contained manylinux / macOS wheel —auditwheel/delocatevendoring libunicorn — sopip installpulls no system libs. Every package + fresh-install + bundled-load smoke was verified in the per-language Docker images (the manylinux wheel checked to load with system libunicorn removed). Track A of the post-v1.0 expansion plan; see docs/packaging.md.Binding parity, round 2: the new emulator/capture capabilities reach all ten bindings. The Track F mid-execution guards, Track E coverage-guided fuzzing / mutation testing, and Track D AVX2 256-bit capture were C-core-only; they now have a binding ABI and a wrapper in every language. The C side adds opaque-handle FFI in
src/ffi.c(emu_watch_t/emu_reg_guard_tand the fuzz/mutation stat structs — alloc + by-field accessors; the arming/driver functions take plain pointers, so a binding calls them directly), andfuzz.onow ships in the emulator shared lib soemu_fuzz_cover1/emu_mutation_test1are reachable. Each binding gained the ergonomic surface in its own idiom — Python (ctypes), C++ (header structs), Ruby (Fiddle), Lua (LuaJITffi), Node (koffi), Go (cgo), Rust (#[repr(C)]+extern), Zig (@cImport), Java (FFM/Panama), .NET (P/Invoke) — e.g.Emulator.watch_writes/guard_reg/fuzz_cover/mutation_testand acapture_vec256+cpu_has_avx2gate, with the vector path self-skipping where AVX2 is absent. Done by hand (a binding-FFI codegen PoC was evaluated and reverted — it only covers the mechanical ~20%); each binding’s conformance runner gained native checks over the same byte-literal routines, verified on the host (Python/C++/Ruby) or the Docker matrix (the other seven).Track C disassembly reaches all ten bindings (via one superset lib). Capstone disassembly was C-core-only — binding it naïvely would pull Capstone into every binding lib, or spawn a combinatorial lib matrix. Instead a single
libasmtest_emuis the superset:make shared-emulinks the emulator (-lunicorn) plus both optional native tiers (Keystone assembler and Capstone disassembler) into that one lib, so any binding that pointsASMTEST_LIBat it gets disassembly and the assembler with no extra flag. Every binding gaineddisas/disas_availablein its own idiom — Python (ctypes), C++ (header, gatedASMTEST_ENABLE_DISAS), Ruby (Fiddle), Lua (LuaJITffi), Node (koffi), Go (cgodlsym), Rust (dlsym+extern), Zig (@cImport, free), Java (FFM/Panama), .NET (P/Invoke) — wrappingemu_disas(decode one instruction at an offset to"mnemonic operands") with a probe that self-skips against an older lib that lacks the tier. The per-binding*-asm-testchecks now drivelibasmtest_emu, so a single run exercises CallAsm and disas; thebindings-asmbase image andinstall-deps --asmgained Capstone. Each binding’s conformance decodes known x86-64 bytes (xor rax, rax/ret/nop), verified on the host (Python/C++/Ruby) and the Docker matrix (the other seven). Track C of the post-v1.0 expansion plan.Win64 runner parity — per-test isolation,
-jN, and benchmarks. The Win64 tier ran every test in one process (--no-fork); the POSIX runner’s fork-based per-test isolation,-jNpool, and benchmark mode were gated off. All three now work on Win64. With nofork(), the runner re-execs itself per test — a hidden--asmtest-child=<index>runs exactly one test and writes its result to a temp file the parent reads back — driven through the existingasmtest_win32_run/_run_poolprimitives (CreateProcess+WaitForSingleObject/WaitForMultipleObjects). Isolation is now the default (matching POSIX);--no-forkselects the in-process facility. A crash is contained in the child (caught there, or backstopped by the child’s death); a hang is killed by the parent’s deadline.--bench(rdtsc cycles per call) runs on Win64 too (aBENCHbody is trusted, so it runs unguarded).tests/win64/suite_win64.cgained aBENCH, andmake win64-runner-testexercises all four modes (isolation,-jN,--no-fork,--bench) under Wine; an optionalwindows-latestCI job signs the same suite off on a genuine Windows host with no Wine. Track B of the post-v1.0 expansion plan. See docs/win64.md.Wide-vector capture — AVX2 256-bit (
ymm). Vector capture was strictly 128-bit (vec128_t/ASM_VCALLn); a routine’symmresult couldn’t be inspected past its low 128 bits. Newvec256_t(the 256-bit analog union) andasm_call_capture_vec256marshalymm0..7args and capture the wholeymmfile into avec256_t[16](out[0]= return), withASM_VCALL256n/ASSERT_VEC256_EQand the existingASSERT_DEQ/FEQover the doubled lane counts. A runtime CPUID probe (asmtest_cpu_has_avx2/asmtest_cpu_has_avx512f, checking both the feature bit and OSXCR0enablement) makes the path self-skip (SKIP) on a host without the feature instead of executing an unsupported instruction. The trampoline is in both backends (src/capture.sGAS +src/capture.asmNASM,vzeroupperon exit),vec256_tis pinned in the manifest and by a_Static_assert, and anvec_add4dAVX2 example +test_simdcase assert the full 256-bit result (including the upper-128 lane) on both backends. Track D of the post-v1.0 expansion plan; AVX-512 (zmm), AArch64 SVE, a Win64 wide path, and binding parity are staged follow-ons — and the emulator wide path self-skips because its bundled Unicorn exposes YMM/ZMM but does not execute AVX (UC_ERR_INSN_INVALID). See docs/floating-point-simd.md.Coverage-guided fuzzing & mutation testing in the emulator. The emulator already recorded basic-block coverage but only ever fed a report; it now feeds input generation, and a mutation tester proves an input set actually catches a perturbed routine. Both run a one-int-arg routine inside the emulator (the instruction cap + fault hooks contain a pathological input or a broken mutant), are seedable for reproducibility, and live in a new dependency-free
src/fuzz.c:emu_fuzz_cover1— coverage-guided generation: keeps inputs that grow the block-coverage union, drawing candidates fresh or by mutating a corpus member (the feedback), so it reaches blocks fixed vectors miss (for theclassifyexample, 5 blocks vs a positive vector’s 3).emu_mutation_test1— mutation testing: flips bits of the routine, runs each mutant and the original on an input set, and counts mutants the set fails to distinguish (survivors = test-gap). A weak suite overclassifyleaves 100 of 192 mutants alive; a path-covering suite leaves only the ~16 equivalent mutants — a stronger input set demonstrably kills more. Reuses the framework’s seedable splitmix64 RNG (asmtest.h) and the emulator’s coverage trace; no new dependency. Track E of the post-v1.0 expansion plan. See docs/emulator.md.
Mid-execution guards in the emulator (watchpoints + register invariants). Assert properties while a routine runs, not just on its result — introspection no ABI-boundary tool can do. Armed on the emu handle and persisting across
emu_call_*until cleared, recording the first violation as data (no host crash), x86-64 guest. Memory-write watchpoints (emu_watch_writeswithEMU_WATCH_ONLY/EMU_WATCH_NEVER,ASSERT_NO_WRITE_VIOLATION/ASSERT_WRITE_VIOLATION) catch a logical scribble into mapped memory that does not fault — where a guard page sees nothing — and name the offending store (emu_watch_describereuses the Track C disassembler:write to 0x400800 (8 bytes): mov qword ptr [rdi + 0x800], rax (@0x3)). Register invariants (emu_guard_reg,ASSERT_REG_INVARIANT) assert a register holds a value at every basic-block entry — a callee-saved / stack-pointer guard that catches mid-routine corruption even when the value is restored by return (which ABI capture cannot see). Step-bounded assertions need no new API: run withmax_insns=Nand inspectout->regs. New hooks (UC_HOOK_MEM_WRITE, a secondUC_HOOK_BLOCK) insrc/emu.c; types, arming functions, and assertions ininclude/asmtest_emu.h. Track F of the post-v1.0 expansion plan. See docs/emulator.md.Disassembly in emulator diagnostics (Capstone). The emulator records faults, traces, and coverage as raw byte offsets —
@0x2f, never the instruction. With Capstone linked (the disassembler counterpart to the Keystone in-line assembler) those offsets now carry the instruction at them, across all four guests (x86-64, AArch64, RISC-V, ARM32). New helpers ininclude/asmtest_emu.h/src/disasm.c:emu_disas(one instruction at an offset, with PC-relative targets resolved to absolute),emu_fault_describe(a fault line that names the offending instruction —read fault accessing 0xdead0000: mov rax, qword ptr [rdi] (@0x0)), and disassembling counterparts to the reporters —emu_trace_disasm,emu_trace_report_disasm,emu_coverage_uncovered_disasm(turninguncovered: 0x2fintouncovered: 0x2f cmp rax, 0). Optional and auto-detected (pkg-config --exists capstone): every helper degrades to bare offsets when Capstone is absent (emu_disas_available()reports which), so the same call works either way and the core library / shared libs / binding images stay Capstone-free — onlybuild/test_emulinks it, exactly as the assembler tier keeps Keystone in its own object.make deps DEPS_ARGS=--emunow installslibcapstone-dev, so the CIemujob exercises the annotated diagnostics on every matrix OS. RISC-V disassembly needs Capstone ≥ 5 and self-skips on older builds. This is Track C of the post-v1.0 expansion plan. See docs/emulator.md.CI builds the cross-platform native payloads for the bindings.
make package-libsonly ever staged the build host’s shared libs, so a release shipped a single-platform payload. A newpayloadsCI matrix runs the native staging on each{x86-64, AArch64} × {Linux, macOS}runner (the Intel-macOS corner nightly, astest-macos-x86does) and uploads eachbuild/dist/native/<os>-<arch>/as an artifact; apayloads (collect + verify)job merges them into one tree, runs the newmake package-libs-verifyto assert every platform slot carries both the core and thelibasmtest_emulib, and re-uploads the combined set as a singlenative-allartifact a publish step would consume. This is the “multi-platform native payloads” step the packaging scaffolding stopped short of — no registry credentials or extra hardware needed. See docs/packaging.md.Static Mach-O verification of the macOS payloads, on Linux.
make package-libs-verify-macho(scripts/verify-macho.sh, folded intopackage-libs-verify) catches the most common macOS packaging regressions — a wrong/missing arch slice or a leaked absolute install-name/dependency — at build time on the Linux release collector, with no Mac needed, viallvm-otool/llvm-lipo. For everybuild/dist/native/darwin-*/.dylibit asserts: the slot’s arch is present (llvm-lipo -archs), the install-name (LC_ID_DYLIB) is@rpath/@loader_path-relative and neither it nor any dependency bakes in/Users,/opt/homebrew, or/usr/local(a dev-build or Homebrew leak; system/usr/liband/Systemare fine), and a min-OS load command is present (and<= MACOS_MIN_FLOORwhen that var is set). It self-skips where the llvm tools are absent (a dev host), sopackage-libs-verifystays green everywhere; thepackage-libs-collectCI job installsllvmso it runs there for real. This is Track B of the macOS clean-test plan — the independent cross-check thatscripts/package-native.sh’s macOS-side install-name rewrites actually produced correct Mach-O.The full emulator surface reaches every binding. A review found four core emulator capabilities that no binding could reach (the FFI lacked an opaque-handle wrapper), plus an assembler tier the corpus did not anchor. All are now exposed across all ten bindings, driven by a widened binding ABI in
src/ffi.c/src/emu.c:Cross-arch emulator guests — run raw AArch64 / RISC-V / ARM32 machine-code bytes on any host (
Guest/GuestEmulator+ per-arch register reads throughasmtest_emu_{arm64,riscv,arm}_reg), not just the x86-64 guest.Emulator FP / vector args and >2 integer args —
call_fp/call_vec/call_bytesover raw bytes, beyond the old two-integercall2.Execution trace / basic-block coverage — an opaque
Tracehandle (asmtest_emu_trace_*) recorded bycall_traced, withcovered(off).Win64 calling convention —
call_win64, to test a Win64 routine on a System V host. Anchored in the shared conformance corpus: newemu_bytes/emu_tracecases (cross-arch int, x86 wide/FP/vector, Win64, two-block coverage) run on every host via checked-in pre-assembled byte literals, and the assembler tier is now emitted intocorpus.jsonand executed by a newmake conformance-asmbuild. The C++ binding also gains the previously missingsum_via_rbx/clear_carrycases. See the binding-parity plan.
In-line assembler tier (Keystone). Pass a routine as an assembly string and run it, instead of only as pre-assembled object code.
asmtest_assemble()(in the newinclude/asmtest_assemble.h) turns text into machine code for the emulator’s guest set — x86-64 (Intel or AT&T syntax), AArch64, ARM32, and RISC-V where the linked Keystone supports it — with errors reported as data and output the caller frees viaasmtest_asm_free(). Bridge wrappersemu_call_asm/emu_arm64_call_asm/emu_riscv_call_asm/emu_arm_call_asmassemble at the emulator’s load base (so PC-relative and branch targets resolve) and run through the matchingemu_*_callin one call. Optional and pkg-config gated like the emulator tier, and folded into the supersetlibasmtest_emu(built bymake shared-emu, which linkslibkeystone+libcapstone+libunicorn):make asm-testbuilds the standalone in-line-assembler suite, withmake docker-asmand a CIasmjob on both x86-64 and arm64. Keystone has no Linux distro package, somake deps DEPS_ARGS=--asmpoints atscripts/build-keystone.sh(a pinned source build the CI job and Docker image use). RISC-V in-line assembly self-skips until a Keystone release ships a RISC-V backend (none does yet). See the implementation plan.In-line assembler reaches every binding, with a widened shim. All ten bindings now expose the assembler — the original five (.NET, Ruby, Lua, Node, Java) plus Python, Go, Rust, C++, and Zig — bound optionally so they self-skip against an older lib that lacks the assembler and pay no cost in the normal binding images. The dlopen bindings probe the symbol; Go and Rust resolve it through the libc dynamic loader (they statically link the plain lib); C++ and Zig link the assembler-carrying
libasmtest_emudirectly. The opaque-handle shim is widened from the original Intel-only, two-integer-arg, error-blindasmtest_emu_call_asm: the newasmtest_emu_call_asm6takes Intel or AT&T syntax, up to six integer args, and an instruction cap (max_insns);asmtest_asm_last_error()surfaces the Keystone diagnostic so a failed assemble reports why instead of a bare false; andasmtest_asm_bytes()exposes multi-arch text→bytes (x86-64/AArch64/RISC-V/ARM32) so a binding can assemble guests its x86-only emulator handle can’t run. Each binding presents this as acallAsm/assemblepair with a uniform failure contract — an assemble error raises/returns the diagnostic (it is never a silent miss). Thebindings-asmCI matrix grows from five to all ten (make <lang>-asm-test), each case now also covering the failure path and a multi-arch assemble; the Cmake asm-testsuite addsasmtest_emu_call_asm6/asmtest_asm_last_error/asmtest_asm_bytescoverage. The originalasmtest_emu_call_asmstays as a thin compatibility wrapper.Native Win64 tier (capture). A Microsoft x64 (“Win64”) capture trampoline (
src/capture_win64.asm) mirrors all eight System Vasm_call_capture*variants on real x86-64 silicon — integer/FP/vector args, the 32-byte shadow space, struct return and by-reference struct args, and ABI-preservation over the larger Win64 callee-saved set (rdi/rsiplus the callee-savedxmm6–15). The captured state has a first-classregs_tlayout ininclude/asmtest.h(selected by-DASMTEST_ABI_WIN64, LLP64-correct, with_Static_assertoffset pins) and a machine-readable manifest (make manifest-win64→asmtest_abi_win64.json). It runs with no Windows host, two ways: the native lane via GCC/Clang__attribute__((ms_abi))(make win64-msabi-test), and a real Windows PE built withnasm -f win64+ MinGW-w64 and run under Wine in an isolated image (Dockerfile.win64,make docker-win64). A new CIwin64job runs both on every push; the capture suite doubles as the native Win64 conformance check. This is the capture tier (suite runs--no-fork); the Win32 runner port is now underway (see below). See docs/win64.md and the implementation plan.Native Win64 tier — runner port. The framework’s process-level guarantees now have Win32 equivalents for the Win64 tier, each in
src/platform_win32.c(plus the platform-neutralsrc/glob_match.c), compiled only for the Win64 target and verified under Wine: per-test isolation + timeout viaCreateProcess/WaitForSingleObject/TerminateProcess(asmtest_win32_run, classifying OK / CRASH-as-NTSTATUS / TIMEOUT), the-jNparallel pool viaWaitForMultipleObjects(asmtest_win32_run_pool), the guard-page allocator viaVirtualAlloc+VirtualProtect(PAGE_NOACCESS), in-process crash-to-failure via a vectored exception handler +__builtin_longjmp(asmtest_win32_guard, no SEH unwinding), and a portable--filterglob matcher (*,?,[...]classes,\escaping) replacing MinGW’s missingfnmatch. Newmake win64-{guard,isolate,pool,filter,seh}-testtargets exercise each under Wine and joinmake win64-check/ the CIwin64job. A thin platform seam (src/platform.h,ASMTEST_FNMATCH) wires the--filterand guard-page paths intosrc/asmtest.cwith no POSIX regression. The runner itself is then built for Win64: a Win32run_one(the per-test facility’s vectored handler + watchdog, mapping the recovery reason to fail/skip/crash/timeout),main()running--no-forkwith the fork/pipe/poll isolation, parallel pool, signal handlers, and SysV-trampoline helpers gated to POSIX.make win64-runner-testbuildssrc/asmtest.cwith MinGW and runs a realTEST()suite (tests/win64/suite_win64.c) under Wine: the runner discovers and runs the suite, asserts real Win64 captures, and contains a crashing and a hanging test as reported failures while surviving. Still POSIX-only: a forked/-jNmode on Win64 and benchmarks. See docs/win64.md.Packaging scaffolding for all ten bindings. Each binding now has a publish-ready registry manifest and a
make <lang>-packagetarget that assembles a distributable bundling the host’s prebuilt native libs:asmtest.gemspec(RubyGems),asmtest-1.0.0-1.rockspec(LuaRocks),pom.xml(Maven),asmtest.nuspec(NuGet),CMakeLists.txt(afind_package-able C++ INTERFACE target),build.zig.zon(Zig package), plus upgradedpyproject.toml(wheelpackage-dataover a bundledasmtest/_libs/),Cargo.toml(crates.io metadata), andpackage.json(npmfiles).make package-libsstages the shared libs intobuild/dist/native/<plat>/; the dlopen bindings (Python/Ruby/Lua/Node/Java/.NET) bundlelibasmtest_emu, while the link bindings (Rust/Zig/C++/Go) ship as source. A new docs/packaging.md is the release guide (native-lib split, version pinning, per-language commands, the multi-platform caveat). Scaffolding only — no registry credentials or cross-OS build matrices.Go binding (Track G). A
cgowrapper inbindings/go/over the opaque-handle FFI layer — no struct layout mirrored: it declares the binding-ABI entry points (asmtest_corpus_routine,asmtest_capture6/_fp2+asmtest_regs_*,asmtest_check_abi,asmtest_emu_call2+ accessors) and links the prebuilt shared libs. ExposesRegs(capture / ABI / flags / FP),Emu+EmuResult(faults as data), and Tier-2Assert*helpers over a smallTBinterface that*testing.Tsatisfies (so the helpers are themselves testable — the suite proves each one bites).make go-testrunsgo test;conformance_test.goreplays the corpus, built + run in its ownasmtest-goimage (make docker-go) and thebindingsCI matrix. This closes the last language track — all ten bindings (Python, Rust, C++, Zig, Node, Java, .NET, Ruby, Lua, Go) now ship Tier 1 + Tier 2.Tier-2 idiomatic assertions (all ten bindings). Optional assertion layers over the Tier-1 result objects, with legible failure messages, idiomatic to each language: Python (
asmtest.assertions, raisingAssertionError), Rust (methods onRegs/EmuResult, panicking), C++ (asmtest::assert_*throwingassertion_error, for GoogleTest/Catch2), Zig (error-union helpers overstd.testing), Node/Ruby/Lua/Java/.NET (throwing/raisingassert_*helpers in the conformance runner), and Go (Assert*helpers failing a*testing.T). Each covers both the pass paths and the failure paths (the assertion fails when it should — pytestraises, Rustshould_panic, ZigexpectError, a recordingTBstub in Go, try/catch elsewhere).assert_ret,assert_abi_preserved,assert_flag,assert_fp,assert_no_fault,assert_reg, ….Node, Java, .NET, Ruby & Lua bindings (Tracks N/J/D/C). Five more language wrappers, all over a new opaque-handle FFI layer (
src/ffi.c+ emu helpers inemu.c):asmtest_regs_new+asmtest_capture6/_fp2+asmtest_regs_*accessors for the capture tier,asmtest_emu_call2+asmtest_emu_*accessors for the emulator, andasmtest_corpus_routine(name)for routine addresses — so a dynamic binding needs no C struct layout. Bindings: Node (koffi), Java (FFM/Panama), .NET (P/Invoke), Ruby (stdlibFiddle), Lua (LuaJITffi); each replays the conformance corpus (make node-test/java-test/dotnet-test/ruby-test/lua-test).Isolated per-language Docker images. Each wrapper is built and tested in its own image (
bindings/<lang>/Dockerfileon a sharedDockerfile.bindings-base), so toolchains never mix.make docker-<lang>builds + runs one language;make docker-bindingsdoes all ten. The CIbindingsjob is now a per-language matrix runningmake docker-<lang>.Zig binding (Track Z). The lowest-ceremony wrapper:
bindings/zig/consumes the C headers directly via@cImport— no separate binding layer — and replays the conformance corpus (make zig-test→zig build test,build.zigtargets Zig 0.13.x). Added to the Docker bindings image and thebindingsCI job.Rust binding (Track R). A no-crates-io crate in
bindings/rust/:#[repr(C)]mirrors ofregs_tand the emulator structs (arch-selected viacfg) plusextern "C"declarations of the binding-ABI entry points, linked against the prebuilt shared libs bybuild.rs. Exposescapture/capture_fp/capture_vec→Regs,abi_preserved(native verdict shim), and anEmulatorwhoseEmuResultcarries faults as data.make rust-testrunscargo test;tests/conformance.rsreplays the conformance corpus.C++ binding (Track X). The C headers now carry
extern "C"guards (and a portableASMTEST_STATIC_ASSERT), so a C++ TU both compiles and links against the framework.bindings/cpp/asmtest.hppadds an RAIIEmu, initializer-listcapture*, vector-lane helpers, andabi_preserved/flag_setpredicates;make cpp-testruns an example suite that drives the framework from C++. NewASMTEST_NO_MAINknob omits the runtime’smain()for embedding.Docker per-language wrapper testing.
Dockerfile.bindingsbundles the Python, C++, and Rust toolchains plus libunicorn;make docker-bindings(anddocker-python/docker-cpp/docker-rust) build and test every wrapper in one reproducible image — verifying a binding on any host, including a language not installed locally. AbindingsCI job runs the same tests natively on x86-64 and arm64 Linux.Python binding (Track P). A pure-ctypes package in
bindings/python/(nocffi/compile step) loads the shared library and theasmtest_abi.jsonmanifest and exposescapture()/capture_fp()/capture_vec()(returning aRegssnapshot withret,flags,fret, vector lanes,abi_preserved, andflag_set) plus anEmulatorcontext manager whoseEmuResultsurfaces faults as data. Struct layout is read from the manifest, so the binding is correct for whatever architecture the library was built for.make python-testbuilds the shared libs, manifest, corpus, and a routine fixture lib, then runs pytest; the suite replays the samecorpus.jsonthe C reference emits and reproduces every case. A newbindings-pythonCI job (x86-64 + arm64 Linux) runs it — the reusable per-language CI template (bindings plan 0.5), which completes Track 0.Shared libraries + ABI manifest (Track 0). The first slice of the multi-language bindings substrate.
make sharedbuildslibasmtest.{so,dylib}(framework runtime + capture trampoline, from-fPICobjects in a separatebuild/pic/tree) andmake shared-emubuildslibasmtest_emu.{so,dylib}(addsemu.o, links-lunicorn), both with platform-correct versioned filenames, soname/install-name, and dev symlinks;make install-shared/install-shared-emuinstall them plus a newasmtest-emu.pc.make manifestemitsasmtest_abi.json— a machine-readable struct layout (sizes, field offsets, host arch, sentinels, flag masks) compiled from the real headers viascripts/gen-manifest.c— so FFI bindings consume offsets instead of hand-transcribing them._Static_asserts inasmtest.h/asmtest_emu.hpinregs_tand the emulator register structs tooffsetof, preventing the headers, the trampoline’s stores, and the manifest from drifting apart.make install(static + headers) is unchanged. See docs/internal/archive/plans/multi-language-bindings-plan.md.Binding ABI + conformance corpus (Track 0). Non-jumping verdict shims
asmtest_check_abi/asmtest_check_flagreturn a verdict + reason instead oflongjmp-ing into the runner, so an FFI binding can validate a capture with no C runner present (the existingASSERT_ABI_PRESERVED/ASSERT_FLAG_*now delegate to them).ASMTEST_NO_MAINbuilds the runtime without itsmain()for embedding.make conformancerunsbindings/conformance/conformance.c— the C reference for a fixed corpus of canonical routines (int / FP / SIMD / flags / ABI capture + an x86-64 emulator case), checked against expected literals — and emitscorpus.json, the portable expected-results table every language binding must reproduce. The binding-ABI contract symbols are designated in the API reference.Parallel execution (Track E).
-jN/--jobs=Nruns up to N tests concurrently as forked children (a pool over the existing per-test fork model), while output stays in registration order regardless of finish order. Per-test timeout and crash containment are unchanged;--no-forkforces serial. Newexpect.shself-tests pin the ordering, failure reporting, and crash containment under-j4.libc-callback example (Track E).
examples/callback.s/.asmwithexamples/test_callback.c:sum_map(arr, n, fn)andcount_if(arr, n, pred)call a C function pointer per element, demonstrating an assembly routine calling back into C with correct callee-saved/stack-alignment discipline.Valgrind story (Track E).
make valgrindruns the example suites under memcheck (--no-fork) to catch bugs in the routine under test, complementing the always-on guard-page allocator;make docker-valgrindand the--valgrindflag ofscripts/install-deps.shround it out. Documented alongside the guard-page approach in the README.Emulator FP/SIMD (Track C). The x86-64 emulator guest marshals
doubleargs (emu_call_fp) and 128-bit vector args (emu_call_vec) into xmm0..7 and captures the whole XMM file (emu_x86_regs_t.xmm[]). The AArch64 guest gains the same (emu_arm64_call_fp/emu_arm64_call_vec, NEONv[]); the RISC-V (emu_riscv_call_fp,f[]) and ARM32 (emu_arm_call_fp,q[]) guests gain scalar FP, with their FP units enabled at open (RISC-Vmstatus.FS, ARM32 CPACR + FPEXC). GenericASSERT_EMU_VEC128_EQworks across guests.Emulator assertions (Track C).
ASSERT_NO_FAULT,ASSERT_FAULT,ASSERT_FAULT_AT,ASSERT_EMU_REG_EQ,ASSERT_EMU_FP_EQ,ASSERT_EMU_VEC_EQ, and coverageASSERT_BLOCK_COVERED/ASSERT_BLOCKS_AT_LEAST.Coverage reporting (Track C).
emu_trace_report,emu_coverage_uncovered(lists the blocks a run missed against a universe trace),emu_trace_lcov(offset-level lcov export), and theemu_trace_coveredpredicate.Emulator vector parity & source-line coverage (Track C, leftovers).
emu_arm_call_vecmarshals 128-bit NEON vectors into ARM32q0..q3and captures the wholeq0..q15file, matching the x86-64/AArch64 vector path. Source-line coverage: a caller-suppliedemu_line_map_t(ascending(offset, line)rows, produced out-of-band) drivesemu_line_lookup,emu_trace_source_report, andemu_trace_lcov_source, which report block coverage against source lines (hit and missed) — no DWARF parsing, no new dependency. The RISC-V “V” extension has no counterpart: Unicorn’s RISC-V guest exposes no vector registers, so it stays scalar-FP (documented insrc/emu.canddocs/emulator.md). Closes Track C’s open C items.
Changed¶
Consolidated the AMD native-tracing docs (design-doc curation). The AMD tracing story was split across four overlapping files; the three genuinely-AMD ones — the AMD LBR snapshot backend plan, the improvement analysis, and the improvement implementation plan — are merged into a single
docs/internal/plans/amd-tracing-plan.mdwith three parts (shipped LBR backend / improvement analysis / improvement roadmap), collapsing two duplicated “governing constraint” + “implementation status” preambles into one. All ~9 references (thesrc/amd_backend.cheader comment and the sibling plans / parity matrix) repoint to the merged file; the vendor-neutral single-step (“Zen 2”) plan stays separate. Thedocs/internal/analysis/trace-parity-matrix.mdoverlap the review also flagged was assessed and left intact: its matrices are trace-specialized (not restatements ofdocs/features.md) and its “Matrix N” numbering is cited fromsrc/trace_auto.c/include/asmtest_trace_auto.h, so trimming would lose detail and dangle those code references. Addresses review finding #15.
Fixed¶
DRAPP_KEYSTONEis now part ofdrtrace_app.o’s build identity. The app-side object was compiled with$(DRAPP_KS_DEF)(-DASMTEST_HAVE_KEYSTONE, gated by theDRAPP_KEYSTONEknob) but that flag was not among the rule’s prerequisites, so flipping the knob between sub-makes in the samebuild/tree reused the stale object. This is not hypothetical: the per-binding native-trace lanes invoke$(MAKE) shared-drtrace … DRAPP_KEYSTONE=0precisely because a Keystone-enabled drapp.sohas unresolvedemu_*symbols and won’tdlopen— so a reused Keystone-on object silently breaks exactly those lanes (CI only escaped it by running each in a clean tree). Bothdrtrace_app.orules (static + PIC) now depend on a$(BUILD)/.drapp-flagssentinel that records the full compile-flag string and is rewritten only when it changes (cmpguard, so no spurious rebuilds), folding Keystone — and by extensionSAN/COVviaCFLAGS— into the object’s identity. Verified: no rebuild when nothing changes, a rebuild the moment the Keystone define or anyCFLAGSknob flips. (The relatedmake -jrace on the aggregatedrtrace-bindings-testtargets — concurrent sub-makes sharing onebuild/— is a separate concern and not addressed here.)Version sync now covers the C header, not just the binding manifests.
include/asmtest.hpins the version in four macros (ASMTEST_VERSION_MAJOR/MINOR/ PATCH+ theASMTEST_VERSIONstring), andscripts/amalgamate.shderives the single-header version from that header — butscripts/sync-version.sh/make check-versiononly touched the nine binding manifests. So a version bump plusmake sync-version && make check-versionpassed green whileASMTEST_VERSION,ASMTEST_VERSION_NUM, and the amalgamatedasmtest_single.hall stayed at the old version — an unchecked second source of truth.sync-version.shnow writes and checks the four header macros too (splittingVERSIONinto numeric MAJOR/MINOR/PATCH, and requiring an exact 3-part numeric semver). Verified by round-tripping a bump: the header,ASMTEST_VERSION_NUM, and the amalgamation all follow, andcheck-versionfails on a stale header.RNG
asmtest_rng_rangedivided by zero on ranges wider thanLONG_MAX.asmtest_rng_range(rng, LONG_MIN, LONG_MAX)— a natural draw in differential / property testing — computed the span as(uint64_t)(hi - lo) + 1, wherehi - losigned-overflowslong(UB) and the wrap makes(uint64_t)(hi-lo)+1evaluate to0, so the following% spanwas an integer division by zero → SIGFPE (a hard crash, or a spurious “crashed by signal 8” under--fork). The span is now computed entirely inuint64_t(wraps to 0 only for the full 2^64 width, which is special- cased to “every value is in range”), and the offset is added inuint64_ttoo so a large offset can’t overflow thelo + …back through signed. New self-testsposit.rng_range_full_width_no_sigfpe/_wide_in_bounds/_narrow_in_bounds.Signed-overflow UB in the floating-point ULP-distance helpers.
fp_ulp_distance/fp_ulp_distance_fcomputed the gap between the sign-magnitude-mapped keys with a signedia - ib, which overflowsint64_t/int32_tfor far-apart operands (e.g.ASSERT_DNEAR(-DBL_MAX, DBL_MAX, …), ~1.84e19 ULPs). It yielded the right magnitude on two’s-complement hardware but is UB — and halted this repo’s ownmake sanitize(UBSan,halt_on_error=1) lane, so it was self-inconsistent. The larger-minus- smaller gap is now taken inuint64_t/uint32_t(bit-identical result, no UB; the signed comparison that picks the larger key is unchanged, so-0.0/+0.0still collapse to a zero-ULP distance). New self-testposit.near_full_range_no_overflow.Out-of-process ptrace stepper leaked the tracee when the traced routine faulted. Tracing a routine that takes a real signal (SIGILL/SIGSEGV) — exactly what the out-of-process single-step tier exists to trace — hit a break in the step loop (
src/ptrace_backend.c) that returnedASMTEST_PTRACE_OKwithout reaping the forked tracee. BecausePTRACE_O_EXITKILLfires only when the tracer exits (not when the trace function returns), the child was left stopped in signal-delivery and unreaped, so a suite of faulting routines slowly exhausted PIDs. The non-SIGTRAP break nowkill(SIGKILL)+waitpids the tracee like every other exit path, still recording the partial trace astruncated. New live regressiontest_ptrace_faulting_no_leaktraces aud2fixture eight times and asserts each leaves no child to reap (waitpid(-1) → ECHILD); it fails on the old code and passes on the fix (make docker-hwtrace, 95 tests green).AMD LBR live capture (first verified on real hardware, Zen 5). Running the branch-record capture on an actual AMD LbrExtV2 host (Ryzen 9 9950X, Zen 5 — the project’s real dev box, long mis-documented as Zen 2) surfaced a capture bug:
hwtrace_end_amdkept the lastPERF_SAMPLE_BRANCH_STACKsample, which for a small routine is all post-routine glue branches — it decoded to an empty in-region trace yet reported it complete (truncated=0). It now keeps the sample richest in in-region branches (the one taken at/just after the routine, whose 16-deep window still holds its branches) and setstruncatedwhen none is found — the faithful dynamic-fallback signal. A branch-heavy loop now reconstructs exactly from the live LbrExtV2 stack; a tiny single-shot routine (too fast for an in-region PMU sample) faithfully truncates. New live regressiontest_amd_live+make docker-hwtrace-amdlane (hwtrace image run with--security-opt seccomp=unconfined --cap-add=PERFMON); the standarddocker-hwtracelane is unchanged (AMD self-skips without perf). Docs corrected across the trace-parity matrix, native-tracing, and the AMD-LBR / single-step plans (dev host is Zen 5 withamd_lbr_v2; AMD LBR live-verified, no longer “unverified on dev”).Emulator handle reuse. Unicorn’s translation-block cache is now flushed when new code is loaded, so reusing an
emu_t/guest handle for a different routine no longer re-runs the previous routine’s stale translation.
1.0.0 — 2026-06-24¶
First tagged release. Captures the complete framework plus the Track A self-test suite and Track B packaging.
Added¶
Core framework. Auto-discovered
TEST(...)cases, a providedmain(), per-suiteSETUP/TEARDOWN,SKIP(reason), and colored TAP reporting with a nonzero exit on failure.Assertions.
ASSERT_TRUE/FALSE, signedASSERT_EQ/NE/LT/LE/GT/GE, unsignedASSERT_UEQ/UNE/ULT/ULE/UGT/UGE,ASSERT_STREQ,ASSERT_MEM_EQ(hexdump diff),ASSERT_REG_EQ, FPASSERT_FP_EQ/NEAR+ laneASSERT_DEQ/DNEAR/FEQ/FNEAR, andASSERT_VEC_EQ.ABI capture. Register/flags capture via
ASM_CALLn,ASSERT_ABI_PRESERVED,ASSERT_FLAG_SET/CLEAR; full call model (ASM_CALLN,ASM_SRET,ASM_FCALLn,ASM_VCALLn, struct-by-value) across the System V integer/FP/vector paths.Differential / property testing.
ASSERT_MATCHES_REF{1,2,3}with a seedable splitmix64 RNG (ASMTEST_SEED).Robustness & CLI. Per-test
fork()isolation with analarm()timeout, crash/hang containment, and a runner CLI (--filter,--list,--shuffle/--seed,--timeout,--no-fork,--format=tap|junit).Benchmark mode.
BENCH(...)cases timed in cycles/call viardtsc/cntvct_el0, run under--bench.Portability. x86-64 and AArch64, Linux and macOS; GAS (default) and NASM (
ASM_SYNTAX=nasm, x86-64) backends. CI covers all four OS/arch combinations.Emulator tier (optional). Unicorn-backed x86-64, AArch64, RISC-V (RV64), and ARM32 guests; Windows x64 ABI on the x86-64 engine; instruction trace and basic-block coverage.
Framework self-tests (Track A).
tests/positive.c,tests/negative.c, and thetests/expect.shblack-box harness, run bymake checkand wired into CI.Packaging (Track B).
ASMTEST_VERSIONmacros;make libbuildslibasmtest.a;make install/uninstallhonoringPREFIX/DESTDIR; anasmtest.pcpkg-config file; andmake amalgamateproducing the single-headerasmtest_single.h.