Mikhail Kuznetsov

Senior Applied Scientist, Amazon AWS · AI Security and Observability.

New York, NY

mikkuzne [at] gmail.com

My current focus is securing AI agents at AWS, along four connected threads: detection models over agent runtime traces (Amazon Bedrock AgentCore) that surface intent deviation and policy violations (prompt injection being one driver), conditioned on an agent’s historical behavior and context; extending detections below the application layer with OS signals (eBPF, system logs), verifying what an agent actually did rather than what its trace claims (kernel-level evidence for agent security); auto-detecting agentic workloads in dynamic environments, the idea behind ClawGuard; and attacking the whole thing on purpose, via co-evolutionary red/blue-team auditing of multi-agent systems (PurpleAudit, NeurIPS 2026), where attacks and defenses evolve against each other over execution traces. Before that I tech-led an embedding model for audit logs, now in production for Amazon GuardDuty: learning representations of security entities (IPs, APIs, usernames) from dynamic audit data and cutting customer-facing false positives by 20–30%.

A single thread connects what I do: LLM workflows that learn from dynamic environments: observe outcomes, update a world model, act on calibrated belief. A co-evolutionary red/blue loop is that same idea pointed at an adversary. It shows up across my research (STARS alignment work accepted at SPIGM @ ICML 2026, earlier robust SSL for tabular data, extreme classification at NeurIPS) and in the open-source projects I build on the side.

PhD from MIPT in Computer Science (2016); earlier work at Yahoo! Research on extreme multi-label classification and multimodal retrieval for ads. See publications for the full list or grab the CV.


Built with Claude

  • ClawGuard · github.com/mikkuzne/clawguard: AI workload observer. Watches what local agents are doing on your machine (processes, network, parent chains) and never blocks or modifies anything itself. A small LLM loop maintains a live world model in place of hand-written rules.
  • apparty · github.com/mikkuzne/apparty: Telegram bot that builds sandboxed web apps on demand. An Anthropic tool-use agent writes the code; bwrap isolates it; /why <id> has a second model narrate what the agent did, in plain English.
  • pitchclaw · github.com/mikkuzne/pitchclaw: LLM-curated calibrated priors for football match outcomes. Claude maintains a weekly-rewritten team-strength model; a downstream mechanical filter flags actionable outcomes from calibrated probabilities.

news

Sep 30, 2026 Efficient Hierarchical Transformers for Representing Log Data (formerly HLogformer) accepted to the NeurIPS 2026 Workshop on Long-Context Foundation Models: a memory-efficient architecture that exploits the nested structure of log entries.
Sep 28, 2026 PurpleAudit accepted to NeurIPS 2026 (Evaluations & Datasets Track): co-evolutionary red/blue-team auditing of multi-agent systems against task hijacking.
Sep 24, 2026 New preprint: On the Effectiveness of Kernel-Level Evidence for Agent Security. We introduce ACE, a 4,047-session corpus pairing kernel syscall traces with agent transcripts, and show that kernel evidence is discriminative on its own and complements application-layer signals.
May 27, 2026 Released ClawGuard — an AI workload observer that watches what local agents are doing on your machine (processes, network, parent chains) and never blocks or modifies anything itself.
May 25, 2026 STARS accepted to the SPIGM Workshop at ICML 2026 — synchronous token alignment for robust supervision in LLMs.