Research

Watch · narrated whiteboard episodesL3

Multimodal Injection: Payloads in Pixels and Audio

The instruction that hijacks an agent need not be text. Images, audio, and documents are first-class injection channels, and the model's own perception is the vulnerability. This threat lab extends the indirect-injection threat model to non-text modalities: adversarial and steganographic image payloads, audio that transcribes to attacker instructions, invisible-text and OCR document channels, and the provenance and sanitization controls that trust-tag perceptual inputs — each paired with a hardening. Grounded in the primary cross-modal injection and adversarial-audio literature.

Murali Chillakuru·5 episodes
  1. 18 min Episode 1The Cross-Modal Threat Model: Where Non-Text Inputs Become Trusted ContextA moderator and a security expert lay out why multimodal models face a structurally larger attack surface than text-only models — every image, sound, and document that becomes text is a door an attacker can slip an instruction through.
  2. 17 min Episode 2Instructions Hidden in Images: Perturbation and Steganographic Injection Robust to ResizingA moderator and a security expert examine how text and adversarial patterns hidden in ordinary-looking images bypass human inspection entirely and reach vision-language models as executable instructions — and what actually stops them.
  3. 17 min Episode 3Audio and Transcription Attacks: Adversarial Speech and the ASR-to-Prompt BoundaryA moderator and a security expert examine how adversarial audio, ultrasonic commands, and silent injections in recordings turn speech-recognition pipelines into an entry point for instructions no human in the room ever hears.
  4. 16 min Episode 4Document and OCR Channels: Invisible Text, Layout Tricks, and Metadata as Injection SurfacesA moderator and a security expert show how PDFs, office files, and scanned documents become injection vectors when extraction pipelines pull their text and hand it to a language model — and why what a viewer shows is not what a parser reads.
  5. 18 min Episode 5Provenance and Sanitization for Non-Text: Trust-Tagging Perception and Modality IsolationA moderator and a security expert close the series by assembling the defense architecture — provenance tracking, sanitization gates, and scope reduction — that together contain cross-modal injection from any modality, image, audio, or document.