AI red-team analyses, adversarial ML findings, and the methodology behind making published security results reproducible and honest.
One person, a phone, a Dell, a few SDRs, and a fleet of Claude agents that turn the radio spectrum from a firehose into a readable story
I hardened a network intrusion detector against adversarial attacks for almost nothing. The catch: you pay up front, and you can't bolt it on later
The feature importances you ship to build trust are the same features an attacker reads to break your model. I measured it
Raw radio is noise to a person and structure to an agent. I built a pipeline that turns captures into named entities and flags what's new
A passing test suite proves your code ran. It can't prove your science is right. Here's the gate I built to tell them apart
I stood up a misconfigured local chatbot, scored how it leaks and obeys attackers, then hardened it from 3 residual failures to 0
An ICS replay attack hides in plain sight: every packet valid, every distribution normal. The fix isn't more ML — it's physics
Most jailbreak claims are screenshots. I rebuilt mine as a scorecard with a pass/fail signal, a before/after delta, and external scanners checking each other
Why a receive-only red team that never transmits is both more realistic and more trustworthy than one that can
The radio spectrum is a massive recon and attack surface, and almost nobody is pointing AI agents at it. I am building a receive-only fleet to change that
Most safety in an open-weights model is a vibe nobody scores. I'm building a reproducible number for how much a fine-tune erodes
A prompt guardrail is a classifier that holds until the adversary pushes past it. Architecture is the training that actually holds
In forensics a hallucination is a false accusation. I stopped one with architecture, not a better prompt — and the refutals are the proof
Most of an adversarial evasion score is bought with packets that could never exist. Constrain the attacker to real physics and it collapses