Open Manual

TODO

Open work only. Keep regression fixtures after fixes. Do not weaken assertions

or classify a timeout as solved by extending its limit. Fixed work is recorded

in docs/CHANGELOG.md, not here.

Local test runs: use NYTRIX_TEST_JOBS=32 ./make test (96-core host); the

CI image stays at NYTRIX_TEST_JOBS=1 for scheduler determinism.

Focused correctness-suite command (results change as concurrent fixes land):

NYTRIX_TEST_JOBS=32 NYTRIX_TEST_NO_BENCH=1 NYTRIX_TRACE=1 NYTRIX_TRACE_CALLS=1 NYTRIX_TRACE_VALUES=1 NYTRIX_TRACE_VERBOSE=1 ./make test etc/tests/runtime/{language/{extensible,type,proof-logic}.ny,modules/import.ny}

---

Bug Group T — Trace-mode compile regression (blocks 8+ suite fixtures)

**Fixtures:** case.ny ("list last"), defer.ny ("Defer unwind order"),

attr.ny, control.ny, primitives.ny ("load8 0"), rewrite.ny,

resource.ny, memcall-get-inline.ny — all rc=134/1 in the suite, all

passing solo without trace env.

**Repro** (LLVM default backend):

NYTRIX_TRACE=1 ./make ny -run /dev/stdin <<'NY'
use std.core
mut l = []
mut i = 0
while i < 100 { l = l.append(i); i += 1 }
assert(l.get(99) == 99, "list last")
NY
# → assertion failed (passes without NYTRIX_TRACE)

**Facts:**

set at COMPILE time; the trace-compiled binary fails regardless of the

runtime env (the corruption is baked in).

byte-identical IR except for ct_rt_result trampolines —

call i64 inttoptr (i64 <compiler-process-address> to ptr)(...) baked by

ny_jit_define_runtime_trampoline (src/code/native/llvm/jit.c:1010).

Those addresses are only valid inside the compiler/JIT process; an AOT

ELF (-o / the -run harness) calling them executes arbitrary bytes.

regression arrived with the in-flight const-fold/arraytab work

(nyir_eval_with_calls resolving ADDR_SYMBOL through

ny_native_arraytab_data), which widens compile-time evaluation to

runtime-shaped expressions (loop-grown lists) whose results then leak

into emitted code.

**Fix direction:** the comptime evaluator must never (a) evaluate

expressions over runtime-constructed aggregates, or (b) leave host-address

inttoptr call targets in AOT output — fold to VALUES only, and emit a

real symbol call when a fold is impossible. Then re-run the 8 fixtures.

---

---

Bug Group B — any-typed dict/nil handle corrupted at len/get boundary

The original raw any boundary is fixed; remaining fixture failures are

tracked in the prioritised list below.

B5 — Variadic argument packing missing in native NYIR lower_call.h (blocks type.ny)

**Root cause isolated:**

In native NYIR, a function declared with variadic parameters fn foo(...args) has its variadic parameter compiled as a list in src/code/native/lower/lower_stmt.h:3039:

bool is_list = ... (fn->as.fn.is_variadic && i == fn->as.fn.params.len - 1);

Inside the callee, args expects the standard list 3-tuple ABI (value_ptr, len, tag) in (rdi, rsi, rdx).

However, call-site lowering in src/code/native/lower/lower_call.h completely ignores callee_fn->as.fn.is_variadic:

1. It does NOT pack variadic arguments into a list. Instead, it emits each passed argument as a separate register/stack argument.

2. In vec.Vector3(1.0, 2.0, 3.0):

3. In type.ny:699:

def ratio = b / a divides 0.0 / 0.0, triggering ZeroDivisionError: division by zero in rt_flt_div.

4. In calls with integer arguments like foo(1, 2, 3):

Caller attempts to compute rt_len(1) for rsi, which triggers PanicError: len expects a sequence, got int (directly linking to Bug Group B4).

**Minimal repro:**

fn foo(...args) {
   print("args len:", args.len)
   print("args get 0:", args.get(0, -1))
}
foo(10, 20, 30)
build/release/ny-full --no-progress -e '...'
# → args len: 0 (or PanicError on integer sequences)

**Comparison with legacy LLVM compiler:**

In src/code/native/llvm/legacy/gencall/init.c:2147-2175, legacy compiler explicitly packs variadic arguments:

} else if (has_sig && is_variadic && !native_variadic &&
           i == (size_t)sig_arity - 1) {
  // Allocates list with __list_new(var_count)
  // Loops over trailing arguments storing each into the list with __store64_idx
  // Passes the constructed list as the final parameter
}

**Current status:** Trailing arguments are packed into a descriptor list and

the variadic callee receives the correct pointer and length. The remaining

failure is a float representation/ABI mismatch after args.get: emitted IR

loads f64 bits from an i64 local, then calls rt_f64_bits(i64 %value) even

though that C function accepts double. A reduced Vector3(1.0, 2.0, 3.0)

currently prints (1, 1, 1). Fix float register representation through

locals and call lowering; do not add a vector-specific compiler exception.

---

---

Bug Group E — Syntax registry macro integer representation mismatch

**Fixture:** extensible.ny (rc=134)

**Symptom:** Nytrix assertion failed: merge_registry_in with overwrite should replace existing handlers.

**Root cause isolated:**

In __macro_double_plus1 with x = 5, the expression evaluated is x + x + 1:

1. x + x (dynamic + dynamic): src/code/native/lower/lower_arith.h:1032 matches (left_any || right_any) and emits rt_any_add(a, rhs). rt_any_add returns a tagged dynamic integer: rt_tag_v(10) = 21.

2. (x + x) + 1 (dynamic + raw int literal 1): lower_arith.h:1014-1023 matches e->as.binary.right->semantic.rep == NY_SEM_REP_RAW_INT. It unboxes (x + x) via rt_any_to_i64(21) = 10 and emits nyir_emit(NYIR_ADD_I64, 10, 1) = 11. It returns a raw machine i64.

3. The macro function __macro_double_plus1 has an untyped signature and returns raw 11.

4. The macro expander / caller expects an untyped return to be a dynamic tagged value. It untags the return value: 11 >> 1 = 5!

5. Assert expects 11, but receives 5, failing the overwrite test.

**Action Blueprint:**

In src/code/native/lower/lower_arith.h:1014-1024:

When the enclosing function return or expression context is dynamic/untyped, box/tag the result of NYIR_ADD_I64 via rt_tag_v / rt_tag, or ensure all dynamic arithmetic paths return uniformly tagged representations.

---

Bug Group F — Remaining Correctness Test Suite Issues

F1 — Zero-capture lambda dynamic argument unboxing (blocks collections.ny & iter.ny)

**Symptom:**

assert([1, 2, 3].map(fn(v) { v + 1 }) == [2, 3, 4], "list map method") returns [4, 6, 8].

**Root cause isolated:**

ABI mismatch between caller indirect dispatch and callee zero-capture lambda entry:

1. **Caller:** In iter.ny:map(xs, fn1), fn1(xs[i]) is an indirect callable. In src/code/native/lower/lower_call.h:6302:

bool raw_scalar_parameter = callee_fn && !indirect_callable && ...;

Because indirect_callable is true, raw_scalar_parameter evaluates to false. The caller does NOT unbox xs[i] and passes tagged dynamic integer 3 (representing 1).

2. **Callee:** In src/code/native/lower/lower_stmt.h:3137-3143:

c

if (lambda_entry && lambda_entry->capture_count > 0 && fn->as.fn.name &&

strncmp(fn->as.fn.name, "__ny_lambda_", 12) == 0 &&

!param->is_any && !param->is_list && !param->is_cstr &&

!param->is_f64 && !param->is_f32) {

arg_val = ny_native_nir_emit_runtime_call(

&b, "rt_any_to_i64", arg_val, -1, -1, 1, 0);

}

Because lambda_entry->capture_count == 0 for fn(v) { v + 1 }, this check evaluates to false!

The callee does NOT unbox v either.

3. v remains tagged as 3. The lambda executes 3 + 1 = 4, 5 + 1 = 6, 7 + 1 = 8, producing [4, 6, 8].

**Action Blueprint:**

In src/code/native/lower/lower_stmt.h:3137:

Unbox untyped/scalar parameters with rt_any_to_i64 for all anonymous lambdas (__ny_lambda_) regardless of capture_count, or when called through dynamic function pointers.

---

F3 — Native BigInt arithmetic return ABI (blocks import.ny)

The JIT-symbol and HM import-resolution failures are fixed: import.ny now

compiles and runs through its earlier all(...) import assertion. The next

reproducer is independent of re-exports: bigint_cmp(bigint(-3), bigint(0))

returns -1, while bigint_neg(bigint(-3)) returns a malformed value and

bigint_abs(bigint(-3)) consequently returns zero. Trace the native ABI of

__bigint_sub / bigint_neg from typed bigint parameters through its return

value; do not special-case bigint_abs or the import resolver.

---

Bug Group G — Live probe failures (2026-09-13)

A fresh hands-on probe against the current tree (build/release/ny -run) found

the following failures. The typed scalar core works (fib(20)=6765, enum ADT

match, 2^10, proof<P>/prove, assert_compile, native -o ELF). Every

failure below sits on the dynamic/any boundary or a value-return ABI, matching

the F-group and return-ABI families already tracked here. Re-probe with the

single repro file command given per item; expected value in parentheses.

G1 — .map over a list literal still broken (F1 confirmed live)

**Symptom:** the documented F1 repro still asserts live.

use std.core
assert([1, 2, 3].map(fn(v) { v + 1 }) == [2, 3, 4], "list map method")

Nytrix assertion failed: list map method (signal 6). The zero-capture lambda

parameter is not unboxed (F1). Variant manifestation:

print([1, 2, 3, 4].map(fn(v){ v * 2 }))   ; → [11, 19, 27, 35]  (want [2, 4, 6, 8])

The output differs from naive tagging ((2v+1)*2 = 4v+2), so this variant is

worth tracing past the F1 fix to confirm the whole map/read chain is raw after

F1 lands.

G2 — reduce returns nil

sum([1, 2, 3, 4].map(fn(v){ v * 2 }))     ; or reduce(fn(a,b){ a + b }, 0)
print(sum)                                  ; → nil (want int)

reduce/sum over a mapped list evaluates to nil, consistent with the

untyped/any accumulator return ABI losing the value at the boundary.

G5 — comptime list map returns nil

use std.core
def base = comptime{ 2^5 }
def shifted = comptime{ range(4).map(fn(i){ i + base }) }
print(base)                                 ; → 32 (works)
print(to_str(shifted))                      ; → nil (want "[32, 33, 34, 35]")

This is the README's own comptime example (README.md:99-104). Scalar comptime

fold works; range(4).map(...) crossing the comptime evaluator boundary loses

its value. The README example is currently non-functional in the tree.

---

Benchmark Performance Blueprint — Crushing C & LLVM Across All Benchmarks

1. Benchmark Reality & Progress Log

Recent optimizations completed:

Current Benchmark Ratios (O3 peak profile, 62 fixtures):

BenchmarkBaseline RatioCurrent Ny nativeCurrent Ny LLVMC (host)Current Ratio vs CStatus / Key Mechanism
loop-unswitch0.07x139µs115µs2.09ms**0.06x****16x FASTER than C**
mixed-static-island0.66x19µs17µs31µs**0.56x****Beats C**
sieve0.91x89µs79µs69µs**1.14x**Near parity with C
vector534.73x471µs470µs305µs**1.54x**Massive win (was 534x!)
dgemm~2.5x413µs414µs259µs**1.59x**Near C
intops~2.0x923µs903µs534µs**1.69x**Near C
n-body~2.2x3.93ms3.92ms2.29ms**1.72x**Near C
binary~2.5x3.11ms3.08ms1.63ms**1.89x**Near C
nqueens~2.5x73.2ms71.8ms37.9ms**1.90x**Near C
spectral1.28x101µs101µs50µs**2.01x**Near C
mandelbrot~2.5x491µs493µs241µs**2.04x**Near C
calls~2.5x229µs227µs103µs**2.20x**Near C
splay~3.0x8.34ms8.49ms3.66ms**2.28x**Near C
fft~3.0x60µs51µs22µs**2.38x**Near C
json-parser~3.5x42.4ms41.7ms16.9ms**2.46x**Fast JSON tokenizer
sha256139.77x3.50ms3.51ms915µs**2.26x - 3.83x**Huge win (was 139x!)
linpack48.09x460µs829µs98µs - 182µs**2.53x - 5.72x**Huge win (was 48x!)
matrix799.14x130µs130µs22µs - 49µs**2.68x - 6.00x**Huge win (was 799x!)
havlak32.77x6.77ms6.79ms2.40ms**2.82x**Huge win (was 32x!)
fannkuch104.06x726ms725ms209ms - 216ms**3.35x - 3.65x**Huge win (was 104x!)
cmov-sort978.90x3.32ms3.14ms949µs - 1.02ms**3.31x - 4.26x**Huge win (was 978x!)
heapsort420.87x19.1ms19.1ms4.84ms**3.79x - 3.99x**Huge win (was 420x!)
sor52.59x3.76ms3.76ms663µs - 822µs**4.58x - 5.68x**Huge win (was 52x!)
revcomp589.16x183µs183µs12µs**14.7x**Down from 589x
iter188.39x10.9ms13.3ms325µs - 355µs**31.9x - 37.5x**Down from 188x
list272.81x3.14ms6.44ms78µs - 163µs**19.3x - 58.5x**Down from 272x
pbkdf2538.49x368ms371ms1.16ms**318x - 342x****PRIMARY BOTTLENECK**

---

2. Comprehensive Findings & Actionable Blueprints for Open Bottlenecks

Bottleneck 1: pbkdf2 (318x vs C) — Inner-Loop Buffer Churn & Dynamic Probing

1. **Allocation in inner loop (etc/tests/bench/pbkdf2.nshape:58):**

ny

while off < padded {

def w = zeros(64) ;; <--- ALLOCATES A NEW 64-ELEMENT BUFFER EVERY BLOCK!

...

zeros(64) calls i64buf_new(64) -> rt_tbuf_new_raw(64, 8).

Each block invocation performs calloc(1, 32 + 64*8), rt_map_oracle_add, and rt_tbuf_register_handle.

Over 1000 PBKDF2 iterations with 2 SHA256 passes per HMAC round, this executes **thousands of heap allocations** in the hot loop. In C (pbkdf2.c), uint32_t w[64] is an unallocated stack array reused across all blocks.

2. **rt_any_to_i64 -> rt_is_str header readability probe overhead:**

perf record shows ~4% of total runtime spent inside rt_addr_mapped and rt_header_readable_cached because dynamic dispatch checks string-ness by probing page tables.

3. **Leaf inlining of zeros() and rotr():**

rotr(int x, int n) is a user-defined function. nyir_func_is_inline_candidate (in src/code/ir/advanced.c:967) restricts inlining. If rotr is not inlined, every round executes 64 function calls.

---

Bottleneck 2: list (19x-58x vs C) & iter (31x-37x vs C) — Append Reallocation & Runtime Boundary

1. **20,000 external C calls:** lst = append(lst, i) in a loop calls rt_tbuf_append_i64_raw on every single iteration. Even when capacity exists, it crosses the JIT-to-C boundary, reads tbuf magic, reads count, reads elem_size, reads capacity, branches, writes, and returns.

2. **Geometric realloc churn:** mut lst = [] starts at capacity 4, then reallocs to 8, 16, 32, ... 32768 (~15 reallocations + rt_tbuf_replace_handle hash lookups). C allocates malloc(20000 * 8) once.

3. **The LLVM Inlining Hazard (CRITICAL FINDING):**

When we inlined rt_tbuf_append_i64_raw directly into LLVM IR via inttoptr GEPs without alias domain metadata, sor regressed from **3.76ms to 7.70ms** because LLVM's basic alias analysis assumed integer-derived pointers alias all loop memory, destroying loop vectorization across unrelated benchmarks.

ny

mut lst = []

mut i = 0

while i < N {

lst = append(lst, ...)

i += 1

}

When N is known or bounded, rewrite mut lst = [] to mut lst = list(N) (using rt_list_new_raw(N)) so the buffer is pre-allocated with capacity N. This eliminates all 15 reallocations and memory moves.

asm

mov rcx, [r_buf - 24] ; load count

cmp rcx, [r_buf - 8] ; cmp count, capacity

jae .slow_realloc

mov [r_buf + rcx*8], r_val ; direct store element

inc qword ptr [r_buf - 24] ; bump count

jmp .cont

.slow_realloc:

call rt_tbuf_append_i64_raw

.cont:

This avoids LLVM pointer aliasing problems entirely while giving C-level append speed in the native JIT tier.

---

Bottleneck 3: fasta Checksum Mismatch — Proved & Solved Root Cause

c

if (is_int(value)) return rt_untag_v(value);

In src/code/native/lower/lower_call.h (around NY_E_MEMCALL for "get"):

When target is proven to be an unboxed list/tbuf (elem_size == 8 or is_list && !dynamic_elements), emit rt_tbuf_index_read_raw and **DO NOT** emit rt_any_to_i64 on the returned value. The value is already an unboxed machine integer.

---

Bottleneck 4: dict Checksum Mismatch — Proved & Solved Root Cause

In dict.nshape, acc += d.get(to_str(i), 0).

The loop addition is lowered through rt_any_add(%acc, %v).

rt_any_add tags its result. On each loop iteration, the tag bit is doubled: (acc << 1) | 1. After 50,000 iterations, the tag bit shifts into the sign bit and overflows to -100003.

When acc is known integer, lower += directly to NYIR_ADD_I64 / LLVM add i64, completely bypassing rt_any_add. Ensure dict.get(..., default_int) unboxes the return value at the call site if the dictionary returns dynamic values.

---

---

Bottleneck 6: Small String Optimization (SSO) & Swiss-Table Dicts (dict: 7x-9x vs C, json-parser: 2.5x vs C)

1. **SSO (Small String Optimization):** In src/code/runtime/core.c: Strings <= 14 bytes can be stored directly inside the 16-byte object word (1 byte length, 1 byte tag, 14 bytes payload). This completely eliminates malloc() for to_str(0..50000).

2. **Swiss-Table SIMD Hash Map:** Replace the bucket-linked dictionary in core.c with flat 16-byte metadata control bytes and SIMD SSE2 _mm_cmpeq_epi8 probing.

---

Bottleneck 7: Repeated Call Sequences & Inline Caching (11,906 Instances in lib/)

Chained .get() or repeated load32/store32 operations go through full dynamic dispatch and boundary type checking on every step.

1. **Monomorphic Inline Caching (MIC):** In src/code/native/lower/lower_call.h: At call sites with shape-stable receivers, cache the resolved struct offset or dictionary index after the first lookup.

2. **Memory Coalescing:** Coalesce consecutive store32 / load32 into single 64-bit or 128-bit vector moves (movq / movaps).

---

Bottleneck 8: Preallocated Lists vs Dynamic Growth Churn (191 Instances in lib/)

191 instances across lib/ where a preallocated or empty list is iteratively filled using out = append(out, value) inside a known-bound loop:

append() performs capacity checks, potential buffer reallocations, and creates unnecessary temporary slice handles instead of writing directly into preallocated storage.

1. Compiler optimization: When mut lst = list(N) or mut lst = [] is filled in a loop while i < N, transform append(lst, v) to direct indexed write lst[i] = v with __list_set_len(lst, N).

2. Fast path: Emit native inline assembly for append in x64.c that checks cap > len in 3 instructions without C runtime call overhead.

---

Bottleneck 9: Repeated Dynamic Dictionaries to Typed Struct Layouts (80 Instances)

80 locations in lib/ construct record-like dynamic dictionaries with constant keys, e.g. {"x": ..., "y": ...}, {"ok": true, "tag": ...}.

Allocating a hash table, hashing string keys, and probing buckets for 2-4 static fields is 20x slower than accessing a 16-byte or 32-byte C struct.

Introduce implicit anonymous struct shapes in src/code/typing/: When dictionary literals have compile-time known string literal keys, lower them to fixed-stride typed buffer records with direct 8-byte field offsets.

---

Bottleneck 10: Type Signature Propagation (312 Untyped Params, 116 Missing Returns)

312 untyped parameters and 116 missing return annotations in lib/ force lowering to assume any representation, injecting rt_any_to_i64 unboxing and rt_tag_v dynamic tagging on every call.

1. Run Hindley-Milner intra-procedural type inference in src/code/typing/pipeline/hm.c to infer concrete scalar types (int, float, bool) from usage within function bodies.

2. Propagate inferred signatures into ny_native_nir_find_user_function to eliminate dynamic boxing trampolines.

---

Open items (prioritised)

A. Correctness & Test Suite Unblocking

---

Cross-cutting: compiler invariant / assertion hardening

always-on, actionable assertions at parser/AST ownership boundaries,

semantic-resolution output, HM representation changes, NYIR instruction

construction/CFG joins, native ABI argument and return shapes, runtime

handle validation, and backend emission. Every assertion should report the

source span, function/module, invariant name, and the relevant type/rep/ABI

state; convert recoverable compiler inconsistencies into structured

diagnostics instead of crashes or silent zero/nil values. Add a dedicated

compiler-assertion test matrix covering malformed ASTs, impossible semantic

reps, invalid NYIR operands/labels, mismatched call arity, stale handles,

and divergent interpreter/LLVM/JIT/native results.

---

C. Performance Epics to 1000x the Language