Research

Watch · narrated whiteboard episodesL2

Runtime Observability for Non-Deterministic AI

Traditional APM assumes deterministic services; autonomous AI systems are not. This series builds runtime observability for non-deterministic agents from the ground up: why classic monitoring fails, traces + metrics + evals as first-class telemetry, semantic capture of reasoning and tool calls (OpenTelemetry GenAI), detecting regressions without ground truth, and a cost-, privacy-, and sampling-aware reference architecture. Grounded in NIST AI RMF, NIST SP 800-137, OWASP LLM Top 10, the OpenTelemetry GenAI conventions, and the evaluation-methodology literature.

Murali Chillakuru·5 episodes
  1. 13 min Episode 1The Observability Gap: Why Traditional APM Fails on Non-Deterministic AI SystemsA moderator and a staff engineer work through why request-response monitoring goes blind on autonomous AI, and what has to replace it.
  2. 11 min Episode 2The Pillars for AI: Traces, Metrics, and Evals as First-Class TelemetryA moderator and a staff engineer rework the three pillars of observability for autonomous AI, and promote evals to a first-class pillar.
  3. 10 min Episode 3Semantic Telemetry: Capturing Reasoning, Tool Calls, and Decision Provenance with OpenTelemetry GenAIA moderator and a staff engineer specify a semantic trace: a decision-span schema, provenance that binds claims to evidence, and recorded tool authority.
  4. 11 min Episode 4Detecting Regressions Without Ground Truth: Reference-Free Evals and Drift SignalsA moderator and a staff engineer build regression detection with no answer key: reference-free evals, change detection, a false-alarm budget, and a cost-ordered ladder.
  5. 10 min Episode 5An Observability Reference Architecture: Sampling, Cost, Privacy, and the Feedback LoopA moderator and a staff engineer assemble the pieces into one system: a pipeline with a feedback loop, a sampling strategy, a cost model, a privacy plane, and an adoption ladder.