I build software that leaves evidence.
Developer tools and AI reliability systems that expose hidden behavior, enforce deterministic gates, and make regressions impossible to ignore.
Portfolio · Technical writing · Email · LinkedIn
Most software says “it worked.”
I care about the harder questions: What happened? What changed? What can we prove?
I build the layer that answers those questions.
| Failure mode | Instrument | Evidence emitted |
|---|---|---|
| Mock drift | msw-inspector |
Coverage report + CI gate |
| Agent trajectory regression | agent-reliability-harness |
Policy findings + baseline diff |
| Retrieval without evidence | docagent-studio |
Citations + offline evaluation |
| Opaque coding-agent quality | agent-fight-club |
Scored, replayable bouts |
| Unproven product assumptions | TypeJung · source | A live, paid user journey |
A demo shows that something worked once. A trace helps explain why it will keep working.
01 Evidence over adjectives.
02 Deterministic checks in CI. Model judgment offline.
03 Narrow interfaces. Reproducible failures. Useful artifacts.
| Record | Public evidence |
|---|---|
PACKAGE / NPM |
msw-inspector-cli v0.3.2 + GitHub Marketplace Action |
PACKAGE / PYPI |
agent-reliability-harness v0.2.2 |
UPSTREAM |
14 merged pull requests across 7 repositories / 5 organizations |
COMMUNITY |
5 outside human contributors in msw-inspector |
PRODUCT |
TypeJung, a live full-stack paid product |
WRITING |
Deterministic checks beat model-judged evals in CI |
Dated public proof snapshot, verified 2026-08-01. Every count also links to its live public record.
Have a failure mode you cannot see yet? Send me the trace.
felmon.tech · felmonon@gmail.com · all repositories
CALGARY / CANADA · TYPESCRIPT / PYTHON · PROFILE DATA · NOT A RÉSUMÉ. A TRACE.



