Research

Watch · narrated whiteboard episodesL3

The Static-Analysis Confidence Gap in AI-Agent Software

Static analysis can find real defects in an agent's software substrate, but a clean report cannot prove that model-mediated intent, framework semantics, delegated authority, and multi-step behavior are secure. This series measures that boundary and develops an honest assurance envelope for reporting what was checked, what was missed, and what remains unknown.

Murali Chillakuru·10 episodes
  1. 19 min Episode 1What a Clean Static-Analysis Report Actually ProvesA moderator and a principal engineer trace, at the whiteboard, why an empty findings list is bounded evidence rather than a safety certificate — and why that boundary is unusually narrow for AI agents.
  2. 18 min Episode 2Measuring Static Analysis on AI-Agent VulnerabilitiesA moderator and a principal engineer build, at the whiteboard, the measuring instrument you need before any recall number about an AI agent can be believed — paired cases, per-class metrics, an ablation, and honest caveats.
  3. 18 min Episode 3The Agent-Framework Modeling GapA moderator and a principal engineer trace, at the whiteboard, why a scanner loses an agent's source-to-sink path in ordinary framework glue rather than the model — and how one declarative row of data reconnects it.
  4. 18 min Episode 4Model-Mediated Taint TrackingA moderator and a principal engineer derive, at the whiteboard, the rules taint must obey as it crosses a language model — preserve it, watch it escalate from data into instruction, and never trust the model to remove it.
  5. 20 min Episode 5Modeling Agent-Specific Sources, Sinks, and Trust BoundariesA moderator and a principal engineer draw, at the whiteboard, the endpoint catalog an agent needs — its real sources, sinks, and trust boundaries — before any clean report can mean a thing.
  6. 21 min Episode 6False Positives and Function-Breaking RemediationWhy a static-analysis finding is a hypothesis rather than a verdict, and how the fix that satisfies the scanner can be the fix that breaks the agent.
  7. 18 min Episode 7Agentic Risks Beyond Source-to-Sink AnalysisWhy a clean taint scan covers only the code-detectable slice of an AI agent's risk, and how a detectability taxonomy maps the rest to the instruments that can actually see it.
  8. 17 min Episode 8When Safe Operations Compose into an Unsafe PlanWhy every step an agent takes can pass its own check while the sequence they form is an exfiltration — and the trajectory-level defenses that actually catch it.
  9. 17 min Episode 9Static Analysis as a Gate for AI-Generated CodeWhy a scanner promoted to a merge-blocking gate for AI-written code becomes load-bearing without becoming more competent — and how to keep it honest and hard to game.
  10. 16 min Episode 10An Honest Assurance Report for Agent SoftwareWhy the deliverable of an agent security review should be a coverage manifest, not a green checkmark — a structured account of what was modeled, assumed, excluded, and still needs a runtime test.