Property / differential testing

Fixed example values catch the cases you thought of. Differential testing catches the ones you didn’t: you supply a C reference model of what the routine should compute, plus a generator of random inputs, and the framework fuzzes many inputs through the real ABI and asserts the assembly agrees with the model.

The assertion

ASSERT_MATCHES_REF1(fn, ref, gen, n);   // routine takes 1 long
ASSERT_MATCHES_REF2(fn, ref, gen, n);   // 2 longs
ASSERT_MATCHES_REF3(fn, ref, gen, n);   // 3 longs

There is one macro per integer arity because C cannot dispatch on a function pointer’s arity. n is the number of random inputs to try.

  • fn — the assembly routine under test (called through asm_call_capture_args, i.e. the real calling convention).

  • ref — a plain C function with the same signature that returns the expected value.

  • gen — a generator that fills an input tuple from the RNG.

On the first mismatch, the offending input, both results, and the seed are reported, and the test fails. Matching inputs are silent.

Before reporting, the failing tuple is shrunk: each position is greedily replaced with 0, 1, -1, LONG_MAX, or LONG_MIN (else halved toward zero) while the disagreement persists, so the report leads with the boundary value that actually triggers the bug — shrinks to [0, 1] — instead of a raw 64-bit draw. The shrink is deterministic, adds at most a few hundred extra calls, and the original input is still reported alongside it.

Note

Shrinking applies to the integer engine only — the FP variants below report the raw failing tuple. And the whole differential engine is POSIX-only: on the native Win64 tier the ASSERT_MATCHES_* family is compiled out.

The FP variants

ASSERT_MATCHES_FREF1(fn, ref, gen, n, ulps);  // 1 double
ASSERT_MATCHES_FREF2(fn, ref, gen, n, ulps);  // 2 doubles
ASSERT_MATCHES_FREF3(fn, ref, gen, n, ulps);  // 3 doubles

The double counterparts for the FP/SIMD surface, where rounding, NaN, and lane bugs live. fn takes 1–3 doubles in the FP register file (called through asm_call_capture_fp) and returns a double; ref is the C model; gen fills a double tuple:

typedef int (*asmtest_fgen_fn)(asmtest_rng_t *rng, double *args, int cap);

Agreement is judged by ULP distance — ulps = 0 demands bit-exactness, and NaN matches only NaN. On a mismatch the input tuple (printed as %.17g, so it round-trips), both results, the ULP distance, and the seed are reported.

The generator

A generator pulls from a seedable splitmix64 RNG and writes up to cap arguments:

typedef int (*asmtest_gen_fn)(asmtest_rng_t *rng, long *args, int cap);

It returns how many arguments it produced. Helpers draw values:

long asmtest_rng_long(asmtest_rng_t *rng);                 // full 64-bit
long asmtest_rng_range(asmtest_rng_t *rng, long lo, long hi);

A complete example

#include "asmtest.h"

extern long asm_abs(long x);     // routine under test

static long ref_abs(long x) {    // the model
    return x < 0 ? -x : x;
}

static int gen_one(asmtest_rng_t *rng, long *args, int cap) {
    (void)cap;
    args[0] = asmtest_rng_long(rng);
    return 1;
}

TEST(refmatch, abs_matches_model) {
    ASSERT_MATCHES_REF1(asm_abs, ref_abs, gen_one, 10000);
}

This calls asm_abs and ref_abs on 10,000 random inputs and fails at the first disagreement, printing the input and both outputs.

Reproducibility and CI

The RNG seed is fixed by default, so a failure reproduces exactly on the next run while you debug. Override it with the ASMTEST_SEED environment variable so CI can explore a different stream each run:

ASMTEST_SEED=12345 ./build/test_refmatch

When a failure is reported, the seed that produced it is printed — set ASMTEST_SEED to that value to replay it.

Pairs well with the emulator

Fuzzing a routine that might loop forever or fault is exactly where the emulator tier shines: its instruction cap and fault hooks mean a runaway input is reported rather than hanging or crashing the harness. The fork isolation in the native runner provides the same guarantee for ABI-level property tests.