Agentic Threat Modeling1Why Classic Threat Modeling Fails on AI AgentsWatch2Threat Modeling AI Agents by Capability and ProvenanceWatch3Modeling Agent Identity and Delegated AuthorityWatch4Threat Modeling Multi-Agent and Tool SurfacesWatch
Alignment & Fine-Tuning Attacks1The Alignment Threat Model: Where Safety Lives and What Each Layer GuaranteesWatch2Harmful Fine-Tuning: Dose-Response and the Capability-Safety DecouplingWatch3Refusal Ablation and Activation Steering: The Refusal DirectionWatch4Reward Hacking in RLHF: Specification Gaming and Proxy-Reward DivergenceWatch5Defenses: Safety-Preserving Fine-Tuning, Tamper-Resistance, and Adaptive EvaluationWatch
Data Poisoning & Backdoors1The Poisoning Threat Model: Availability, Integrity, and Backdoor Goals Across the PipelineWatch2Backdoor Triggers: BadNets, Clean-Label Attacks, and the Stealth-Success Trade-offWatch3Poisoning Web-Scale Corpora: Split-View, Frontrunning, and Expiring-Domain EconomicsWatch4RAG and Fine-Tune Poisoning: Dose-Response of Corrupting an Index or Instruction SetWatch5Detection and Provenance: Spectral Signatures, Dataset Signing, and Trigger Reverse-EngineeringWatch
Embedding & Retrieval Security1Embeddings Are Not Anonymized: Text-Embedding Inversion and How Much a Vector RevealsWatch2Cross-Tenant and Neighbor Leakage: The Similarity Oracle and Multi-Tenant Isolation FailuresWatch3Retrieval Corruption: The Geometry of Poisoning a Neighborhood to Dominate a QueryWatch4Index-Level Attacks: Exploiting Approximate-Index Knobs for Denial and EvictionWatch5Hardening Retrieval: Per-Tenant Indexes, Embedding-Space Access Control, and ProvenanceWatch
Extraction Attacks1The Query-Access Threat Model: What a Black-Box Attacker Can and Cannot LearnWatch2Stealing the Last Layer: Recovering a Production Model's Final Projection from LogitsWatch3Training-Data Extraction and Memorization: Eidetic Memorization and the Extraction RateWatch4Membership Inference: Shadow Models, Loss Thresholds, and Calibrated AUC as a Privacy MetricWatch5Defenses and Their Costs: Truncation, Noise, Rate Limits, and Differential-Privacy AccountingWatch
Inference Side Channels1The Shared-Serving Threat Model: Co-Tenancy, Batching, and the Attacker's ObservablesWatch2KV-Cache and Prompt-Cache Leakage: Prefix Caching as a Membership OracleWatch3Timing Side Channels in Decoding: Speculative Acceptance and Bits-per-QueryWatch4Denial and Resource Amplification: Sponge Inputs and Denial-of-WalletWatch5Isolation and Mitigations: Cache Partitioning, Quotas, and the Throughput CostWatch
Jailbreaks as Optimization1The Adversarial Objective: Formalizing a Jailbreak as a Loss over the InputWatch2Greedy Coordinate Search: Why Token-Level Search Beats Naive Gradient DescentWatch3Universality and Transfer: Shared Representation Geometry as the Reason Suffixes MoveWatch4A Taxonomy of Jailbreak Families: One Mechanism Behind Obfuscation, Role-Play, and Many-ShotWatch5Robustness and Its Limits: Adversarial Training, Classifiers, and Why Provable Bounds Stay SmallWatch
Model Supply-Chain Attacks1The Artifact Threat Model: What a Downloaded Checkpoint Can Do at Load TimeWatch2Deserialization and Loader Attacks: Pickle Code Execution and Why safetensors HelpsWatch3Weight and Adapter Tampering: Malicious Merges and Detecting Behavioral DriftWatch4Distribution and Typosquatting: Model-Hub Trust and Dependency ConfusionWatch5Signing and Provenance: Model Signatures, SLSA Attestations, and Verification at LoadWatch
Multimodal Injection1The Cross-Modal Threat Model: Where Non-Text Inputs Become Trusted ContextWatch2Instructions Hidden in Images: Perturbation and Steganographic Injection Robust to ResizingWatch3Audio and Transcription Attacks: Adversarial Speech and the ASR-to-Prompt BoundaryWatch4Document and OCR Channels: Invisible Text, Layout Tricks, and Metadata as Injection SurfacesWatch5Provenance and Sanitization for Non-Text: Trust-Tagging Perception and Modality IsolationWatch
Red-Team Measurement1What Is Attack-Success-Rate, Really? The Unit of Analysis, the Population, and the JudgeWatch2The Judge Problem: LLM-as-Judge Bias, Human-Label Reliability, and Calibrating a GraderWatch3Sampling and Confidence: Wilson Intervals, Per-Family Estimation, and Sequential TestingWatch4Transfer and Generalization: Measuring Whether an Attack Holds Across Models, Prompts, and TimeWatch5A Reporting Standard: Datasets, Seeds, Judge, Confidence Intervals, and Threats to ValidityWatch
Watermark & Provenance Evasion1The Detection Threat Model: Evasion, Spoofing, and Scrubbing GoalsWatch2How Text Watermarks Work: Green-List Schemes and the Robustness-Quality Trade-offWatch3Evasion and Spoofing: Paraphrase Attacks, Watermark Stealing, and ForgeryWatch4Content Provenance: Cryptographic Signing Versus Statistical WatermarksWatch5What Detection Can and Cannot Promise: Impossibility-Leaning ResultsWatch
Autonomous SOC: Adversarial1Telemetry Injection: When the Log File Is the WeaponWatch2Behavioral Mimicry Against ML Detectors: The Statistical Camouflage ProblemWatch3Alert Economy Attacks: Flooding, Suppression, and Prioritization GamingWatch4False-Positive Weaponization: Forcing the Defender to Be the AttackerWatch5Evidence Manipulation in AI-Mediated InvestigationWatch
Indirect Prompt Injection1Indirect Prompt Injection: Attack Tree & HardeningWatch2RAG Poisoning PrimitivesWatch3Tool Abuse & the Confused DeputyWatch