AI systems engineering and low-level ML.
Shown as I actually split my time — nothing here is a flagship, everything is a peer.
Shown as I actually split my time — nothing here is a flagship, everything is a peer.
The hybrid system beats both a tuned deterministic pre-pass and the model alone, on every metric, on every checkpoint tested — field F1 0.6432–0.6827 against 0.2911 for rules alone and 0.4461 for the model alone. It also has a dominant, unsolved failure: silent omission, 55–62% of every failure breakdown, immune to four independent fixes. And a non-deterministic teacher labeller quietly invalidated part of the project's own early evidence before an audit caught it.
V3 targets the one number V2 could not move: omission at 55–62% of failures. Four iterations in, the best configuration reaches 52.3% — a new low, and still not below half. No V3 checkpoint beats the V2 release candidate on the comparable eval. This is an interim account of work in progress, not a result.