Defensive AI Security & Governance
Building the discipline of autonomous defense — detecting drift, governing intent, and containing AI threats at machine speed, with the human kept where judgment matters.
Research-grade analysis of AI security — grounded in public, authoritative sources. Watch the narrated walkthrough of every topic.
Building the discipline of autonomous defense — detecting drift, governing intent, and containing AI threats at machine speed, with the human kept where judgment matters.
Attack & defend AI systems — threat labs, primitives and chains, every offensive step paired with its defense.
Most security education recycles the same dozen attacks. This threat lab covers the ones that break an assumption you never knew you were making — starting with the deepest of all: that two systems agree on what the same bytes mean. When a guardrail, a language model, and a tool each parse one input differently, the exploit lives in the disagreement, not in any single component. Each class is taught by the assumption it violates, the mechanism that makes it work, and the assumption-free defense — grounded in the seminal paper that named it, and tied back to the AI-agent stack every time.
The post-quantum transition beyond the lattice — the code-based and multivariate alternatives, what they cost, and where the standards stand.
A computer-use agent perceives a screen and acts by clicking and typing — and the moment the display it reads is attacker-influenced, the pixels on screen become an untrusted input that drives real actions. This threat lab maps the perceive-ground-decide-act-observe loop as a trust boundary, showing where hostile display content enters and what it can make an agent do. Each exposure — UI-grounding decoys, action hijack and irreversibility, the environment as adversary — is taught by the assumption it breaks, the mechanism that makes it work, and the containment control that limits the blast radius. Product-agnostic, grounded in the GUI-agent and computer-use attack literature, and tied back to the AI-agent stack every time.
A benchmark score is supposed to tell you what a model can do. This measurement series argues that for dangerous capabilities the number is only a lower bound, and a fragile one: the true capability can be hidden by a weak elicitation method, deliberately suppressed by a model that sandbags, or distorted by a model that notices it is being tested. Each article treats one threat to measurement validity — capability-versus-behavior, the elicitation gap, strategic under-performance, and evaluation awareness — with the data-scientist's tools of provenance, controls, and threats-to-validity, and closes with a defensible standard that pairs every capability claim with what it can and cannot prove. Product-agnostic and grounded in the primary dangerous-capability-evaluation and sandbagging literature.
The algorithms behind a model's answer — attention, the KV cache, speculative decoding, quantization, and vector recall — derived properly, to where each one breaks.
Feature dictionaries, sparse autoencoders, and activation probes are making model internals partially legible at runtime — and that raises a defensive question: can reading the internals detect deception, jailbreak states, or unsafe behavior before it reaches the output? This paper treats interpretability as a monitoring control rather than a research curiosity, and asks the harder question every detector must answer: what are its false-negative limits under an adaptive adversary? The organizing contribution is a Probe Assurance Ledger that separates 'evidence' from 'guarantee' for each internal-signal detector. Product-agnostic, grounded in the primary probing and feature-learning literature, and tied back to the AI-agent stack every time.
When one engineer holds the full context, plan, and intent for a feature, driving it end-to-end with an AI coding assistant can beat splitting the work across a team. This series builds the argument from evidence: why shared development pays a coordination tax that grows with team size, why context and cognitive load — not typing speed — are the real bottleneck, what AI assistants actually change by compressing the design-implement-test-refactor loop into a single continuous flow, when the single-conductor model delivers faster and more consistent results, and where genuine collaboration still wins. Grounded in Brooks' communication-overhead law, Conway's Law, cognitive-load and flow research, the GitHub Copilot productivity randomized trial, and the DORA delivery-performance findings.
Static analysis can find real defects in an agent's software substrate, but a clean report cannot prove that model-mediated intent, framework semantics, delegated authority, and multi-step behavior are secure. This series measures that boundary and develops an honest assurance envelope for reporting what was checked, what was missed, and what remains unknown.