Mikhail Kuznetsov
Senior Applied Scientist, Amazon AWS · AI Security and Observability.
New York, NY
mikkuzne [at] gmail.com
My current focus is securing AI agents at AWS, along four connected threads: detection models over agent runtime traces (Amazon Bedrock AgentCore) that surface intent deviation and policy violations (prompt injection being one driver), conditioned on an agent’s historical behavior and context; extending detections below the application layer with OS signals (eBPF, system logs), verifying what an agent actually did rather than what its trace claims (kernel-level evidence for agent security); auto-detecting agentic workloads in dynamic environments, the idea behind ClawGuard; and attacking the whole thing on purpose, via co-evolutionary red/blue-team auditing of multi-agent systems (PurpleAudit, NeurIPS 2026), where attacks and defenses evolve against each other over execution traces. Before that I tech-led an embedding model for audit logs, now in production for Amazon GuardDuty: learning representations of security entities (IPs, APIs, usernames) from dynamic audit data and cutting customer-facing false positives by 20–30%.
A single thread connects what I do: LLM workflows that learn from dynamic environments: observe outcomes, update a world model, act on calibrated belief. A co-evolutionary red/blue loop is that same idea pointed at an adversary. It shows up across my research (STARS alignment work accepted at SPIGM @ ICML 2026, earlier robust SSL for tabular data, extreme classification at NeurIPS) and in the open-source projects I build on the side.
PhD from MIPT in Computer Science (2016); earlier work at Yahoo! Research on extreme multi-label classification and multimodal retrieval for ads. See publications for the full list or grab the CV.
Built with Claude
- ClawGuard · github.com/mikkuzne/clawguard: AI workload observer. Watches what local agents are doing on your machine (processes, network, parent chains) and never blocks or modifies anything itself. A small LLM loop maintains a live world model in place of hand-written rules.
- apparty · github.com/mikkuzne/apparty: Telegram bot that builds sandboxed web apps on demand. An Anthropic tool-use agent writes the code;
bwrapisolates it;/why <id>has a second model narrate what the agent did, in plain English. - pitchclaw · github.com/mikkuzne/pitchclaw: LLM-curated calibrated priors for football match outcomes. Claude maintains a weekly-rewritten team-strength model; a downstream mechanical filter flags actionable outcomes from calibrated probabilities.
news
| Sep 30, 2026 | Efficient Hierarchical Transformers for Representing Log Data (formerly HLogformer) accepted to the NeurIPS 2026 Workshop on Long-Context Foundation Models: a memory-efficient architecture that exploits the nested structure of log entries. |
|---|---|
| Sep 28, 2026 | PurpleAudit accepted to NeurIPS 2026 (Evaluations & Datasets Track): co-evolutionary red/blue-team auditing of multi-agent systems against task hijacking. |
| Sep 24, 2026 | New preprint: On the Effectiveness of Kernel-Level Evidence for Agent Security. We introduce ACE, a 4,047-session corpus pairing kernel syscall traces with agent transcripts, and show that kernel evidence is discriminative on its own and complements application-layer signals. |
| May 27, 2026 | Released ClawGuard — an AI workload observer that watches what local agents are doing on your machine (processes, network, parent chains) and never blocks or modifies anything itself. |
| May 25, 2026 | STARS accepted to the SPIGM Workshop at ICML 2026 — synchronous token alignment for robust supervision in LLMs. |