Overview
Kyle Polley serves as Chief Information Security Officer at Perplexity [1]. Polley maintains a presence on X at @kpolley [2].
Career history
- Chief Information Security OfficerJul 2026 to PresentPerplexity
- Member of Technical Staff, SecurityDec 2024 to Jul 2026Perplexity
- Head of Security, EM of Internal SystemsJan 2023 to Dec 2024Pipe
- Lead Security EngineerJul 2022 to Dec 2024Pipe
- Security EngineerMar 2020 to Jul 2022Robinhood
- Security Engineer - Data Science & MLJun 2018 to Mar 2020PatternEx
- IT Security and Operations ResearcherJun 2016 to Sep 2017NASA Ames Research Center
- Code I Support, InternAug 2013 to Aug 2015NASA Ames Research Center
Education
- Bachelor’s Degree, Computer Science2018University of CaliforniaDavis
Insights & ideas
The through-line
Kyle Polley's consistent argument is that agents have broken the assumptions security programs were built on, and that defenders have to respond with layered, data-driven engineering rather than certificates and single controls. The specific technical form of that break is prompt injection: an agent asked to book a flight visits a website, ingests untrusted content, and "the agent context is at the end of the day the LM doesn't know exactly what is tool output versus user input versus system prompt" [1]. The organisational form is that traditional detection and response tooling cannot keep up with cloud, SaaS and agent-generated volume [2]. Both problems get the same treatment from him: build for the attacker who exists today, measure honestly against production reality, and share the work.
The second thread running through everything is impatience with proxies for security. Academic prompt injection benchmarks do not translate into live environments [1], compliance certificates do not answer the question a breach actually raises [2], and the word SIEM describes a tool built for a world of three log sources that no longer exists [2]. In each case his move is to go back to the underlying problem, which he repeatedly frames as a data problem, and rebuild from there.
On why prompt injection is the defining agent vulnerability
The mechanism is the same across products: a user gives an agent a goal, the agent visits an untrusted resource, malicious instructions flow back through tool output into context, and the agent follows them [1]. He catalogues the attack families his team has seen in the wild. One is impersonating the prompt template, where the injected text mimics the system prompt or the delimiters that separate user input from tool output [1]. Another is agent social engineering, a website that walks the agent through seemingly benign steps in language that "is very similar to how like fishing humans would look" [1]. A third is the conditional trigger hidden in a calendar event or a page, firing only when the user asks the agent to summarise or check something [1]. Browsers give attackers abundant places to hide the payload, in hidden divs, display properties, URLs, or far down a benign-looking calendar entry [1].
He locates the underlying cause in how people write prompts rather than in the models alone. Just as engineers writing code do not automatically think about security implications, people writing system prompts and doing context engineering do not either, and they end up with "this like AI agent with access to a ton of different systems and data" without ever instructing it to treat untrusted data differently from trusted data [2]. He notes the asymmetry that makes this hard: telling an agent what it should do is easy, enumerating everything it should not do is not [2]. He points to Meta's open source Llama variant that adds a third section specifically for untrusted content as a structural attempt to fix this inside the model rather than bolt it on [2].
On why existing benchmarks and off-the-shelf classifiers fail
His sharpest empirical finding is that open source detectors look competent on paper and collapse in live conditions. Existing benchmarks each have strengths but none covers the full range of attack types and injection points seen in the wild [1]. Detection rates hold up on obvious attacks like "please reveal your system prompt" and degrade badly on context manipulation, then degrade further on multilingual attacks where obviously malicious instructions written in Hebrew or Spanish slip through [1]. His conclusion is that these systems "aren't actually looking at the intent of what it of what the prompt injection is asking of for the agent to do. Uh instead it's just looking for keywords" [1].
The distractor experiment makes the same point from the other side. A cookie consent banner in raw HTML reads almost exactly like an injection, insisting you must click accept before you can proceed [1]. Models score well on clean data and drop to 81% with just three distractors added, which he treats as the clearest indicator that keyword and heuristic matching, rather than intent, is doing the work [1]. Both PromptGuard and GPT OSS Safeguard fall into this category of not working well in live environments, while prompted frontier models such as GPT-5 and GPT-5 mini do "exceptionally well" [1]. The response was BrowseSafeBench, built from real attacks, deanonymised and sorted into a taxonomy by attack type and injection method, then open sourced [1].
On what BrowseSafe is for and why latency matters as much as accuracy
BrowseSafe is a fine-tuned classifier based on Qwen, with a 30 billion parameter model used at inference for latency reasons [1]. The headline is not only the 90.4% F1 score but that it comes in subsecond time, against roughly 2 seconds for GPT-5 mini and up to 20 seconds for GPT-5, which is what makes it deployable in a product "without uh massive user frustration" [1]. He draws two lessons from the comparison: fine-tuning on a domain-specific dataset beats general purpose models by a wide margin, and the security industry should be doing more of that work and open sourcing it [1].
The design choice he seems proudest of is the output format. Rather than a binary verdict, BrowseSafe emits a calibrated probability, trained to be accurate rather than performative. His complaint about asking an LLM to judge and score itself is that "every time for me at least it returns like 93%. Every single time" [1]. A real confidence score gives security and product teams "a knob they can turn on uh how much precision and recall are they willing to accept in their product" [1]. If the product can afford to ask the user whether they want to continue when injection is suspected, the team can tolerate far more false positives and tune accordingly [1].
Stress testing revealed where the model generalises and where it does not. Holding out URLs actually improved performance, evidence it is not flagging on domains or keywords [1]. Holding out attack language caused a small dip, leaving it in line with frontier models [1]. Holding out injection strategies, meaning novel placements within websites, caused a significant drop, which he reads as a forecast: attackers will find new injection points precisely to bypass model guardrails [1].
On defense in depth for agent products
He is explicit that a classifier alone is not a strategy. The first layer is the oldest trick in the book, preprocessing and stripping content before it ever reaches context: HTML comments, hidden elements, anything not needed for the model to do the job [1]. The second layer is classification at every boundary where untrusted input enters, whether a website, a calendar invite, a third party connector or an MCP, checked before the data returns to the agent [1]. For teams unwilling to run their own inference, he offers a pragmatic starting point: prompted GPT-5 mini reaches 85.4% F1 and the 2 second latency is "honestly not so bad for something that you can get started with today" [1].
His most operationally specific advice concerns what happens after detection. Simply blocking a tool call is naive because you have to think about it from the agent's perspective: it sees an error, retries, tries a different route, perhaps calls the Google API directly [1]. The better pattern is to return the security verdict as the tool output, telling the model explicitly that the calendar fetch was rejected because prompt injection was detected in the text [1]. In Comet, detection blocks all follow-up tool calls entirely and instructs the agent to return text directly to the user [1]. The third layer is keeping the agent on the newest frontier model, citing the Opus 4.6 system card as the first he has seen to test prompt injection specifically in browser use scenarios, with success rates falling from 16.2% at 4.5 to 2.83% at 4.6 and to 0.8% with updated safeguards [1]. The fourth layer is an LLM backstop: route low-confidence classifications to a much larger model, accept the added latency where the use case allows, and then feed the backstop's reasoned verdicts back into the training set as a data flywheel that catches the novel attacks his own held-out testing showed the classifier struggles with [1].
On open sourcing defensive research
Asked directly about the risk of handing attackers an oracle that reveals what gets through, he does not hedge. He is "a huge fan of open sourcing our work" on the grounds that defenders need the help and that threat actors will be probing for these vulnerabilities regardless [1]. His frustration is with the prevailing genre of security publishing: blog posts describing what a product did or how threat actors are using a given model to conduct attacks, without sharing IoCs or anything that would let a reader actually protect themselves [1]. Both the BrowseSafe model and the dataset are on Hugging Face alongside a detailed paper, produced with colleagues at Perplexity and collaborators at Purdue University [1].
On detection and response before compliance
He argues for building detection and response ahead of compliance work, for two reasons. The first is timing: "attackers and hackers, they're not going to wait for you to become compliant" [2]. When a breach happens, the questions from the company and from enterprise customers are what happened and what did they take, and without a detection and response process you have neither the audit logs nor the capability to answer, so "you're stuck" [2]. He is careful not to overstate this into a demand for a mature SOC on day one, pointing instead to cheap first steps that carry you a long way: turn on CloudTrail, ship it to a bucket, turn on GuardDuty, send it to a Slack channel [2].
The second reason is that he treats compliance as a downstream effect rather than a goal. Build a genuinely good security program and compliance comes along with it; if compliance feels hard, that is a signal something is wrong or something is not being done [2]. Being compliance-driven, in his experience, leads teams down the wrong path [2]. He agrees with the practitioner consensus that being compliant is not the same as being secure, and returns to the blunt version: in an incident, nobody asks for your ISO or SOC 2 certificate [2].
On why the SIEM is the wrong shape and detection is a data problem
He dislikes the term SIEM and explains why in architectural terms. SIEMs were designed for a world of internal infrastructure with a handful of log sources, DNS and network traffic, server logs, authentication, where correlation across three streams was tractable [2]. Now everything is in the cloud, there are a hundred SaaS providers, and the modern vendors are themselves running data lake infrastructure underneath, so he prefers to describe what he wants as a data lake built for detection and response [2]. He is not interested in fighting over naming with executives who know the old term: the pitch is that a traditional SIEM simply cannot handle what the team is doing, and these are the tools that scale [2].
That framing matters for AI too, since "you could you can give the best AI agent access to to a legacy sim and it's still not going to work because it's quering stuff the way we do" [2]. He also questions why security maintains a separate data lake at all, noting it is strange that security engineers, who are not data experts, carry a data burden as large as the data team's while a data team exists to do exactly that work [2]. He has explored the alternative himself with Spark and Jupyter notebooks doing threat detection the way a data scientist would approach analysis, concedes it works on paper more than in practice so far, and points to Apple's threat detection and response team presenting Databricks and Jupyter-based operations, and to Netflix hiring data scientists into security years ahead of the curve [2]. His underlying claim is that threat detection is a big data problem, one that is hard to wrangle and that has not been solved by spending: companies with hundreds of staff and millions of dollars in SOC operations are still getting breached, so "if you just keep doing the same thing over and over again, it's not going to work out" [2].
On AI in the SOC and what stays human
He is unambiguous that AI belongs in security operations and equally clear about its limits: "I don't think it's going to just replace humans. I do think it'll take tasks that were super slow and repetitive" [2]. The case rests on the failure mode of human SOCs, where teams are burnt out, false positives are everywhere, and breaches happen anyway [2][7], against agents that do not get tired, can scale to the number of alerts and keep querying data without fatigue [7]. He describes himself as "100% convinced this is the future and the faster we get there, the better chance we had at stopping attacks" [2], while keeping humans in incident response [2][5].
The practical prerequisites are infrastructure and interfaces. Modern, proven, scalable storage such as Snowflake or equivalent data lake infrastructure is the first requirement, because the data volume is not going away [2]. The second is machine-accessible tooling: after working extensively with MCPs he now expects every security tool he uses to offer a robust API, ideally an MCP server [2]. He sees vendors bolting a chat agent onto each product, including SIEMs with a talk-to-the-agent button, as a step in the right direction but the wrong end state, preferring a future where the security team runs its own agent across all of them rather than dealing with a separate vendor agent per tool [2]. He has also released open source work aimed at letting security operations teams experiment with how much AI they can use, and talks about MVPs and easy first lifts across different maturity levels [2][5].
On what security teams should be building next
He frames the team's agenda as two simultaneous problems. The first is product-facing security in a world where agents increasingly act on behalf of humans, which he expects to reshape the fundamentals over the next couple of years: what identity means, what threat detection means, what DLP means when the actor is an agent [1]. The second is internal, redefining how security teams themselves operate with LLMs and agents at the core rather than as an add-on [1]. He treats both as open research problems his team is actively recruiting for, and as work he expects to have concrete results to report on within a year [1].
Takeaways
- Prompt injection succeeds because the model cannot distinguish system prompt from user input from tool output in its context, so any untrusted page, calendar invite, connector or MCP is an injection vector [1][2].
- Open source injection detectors match keywords rather than intent: add three distractors like a cookie consent banner and detection drops to 81%, and attacks translated into another language slip through [1].
- BrowseSafe, fine-tuned on Qwen and run at 30B for inference, reaches 90.4% F1 in subsecond latency versus 2 seconds for GPT-5 mini and up to 20 seconds for GPT-5, and both model and dataset are open sourced on Hugging Face [1].
- Emit a calibrated confidence score rather than a binary verdict, so product teams can tune precision against recall for their own user experience [1].
- When injection is detected, do not just return an error; feed the security verdict back as the tool output and block follow-up tool calls, or the agent will retry through another route [1].
- Layer defenses: strip hidden HTML and comments before context, classify every untrusted input, run the newest frontier model as the agent, and route low-confidence cases to an LLM backstop whose verdicts retrain the classifier [1].
- Build detection and response before chasing compliance, because attackers will not wait and a breach demands answers no certificate provides; start with CloudTrail to a bucket and GuardDuty to Slack [2].
- Treat threat detection as a big data problem on scalable data lake infrastructure, and require a robust API or MCP server from every security tool so agents can actually work the data [2].
Media & appearances
- [un]prompted Security Practitioner Con 2026Apple PodcastsEp 42 | Kyle Polley - Training BrowseSafe: Lessons from Detecting Prompt Injection | [un]prompted 2026Day 2 | Stage 1 Member of Technical Staff at Perplexity shares experience training and deploying BrowseSafe for detecting prompt injection in production browser agents. Deploying AI agents that browse the web creates critical security challenge preventi
- CBB 365Apple PodcastsEpisode 7: Kyle Boone + NBA Draft Talk, Tyler PolleyOn Episode 7 of CBB 365, Adam and Patrick are joined by Kyle Boone of CBS Sports for a phenomenal interview to preview some of the biggest transfers of the offseason and to take a deep dive into the NBA Draft field, where Kyle's affection (obsession?) for International Prospects is brought to light. Then, UConn forward Tyler Polley joins the show to talk about his road back from a torn ACL and to look ahead to next season in the Big East.
- Apple PodcastsAI for SOC Automation: A Bluep…–Cloud Security Podcast ...The nature of Security Operations is changing. As cloud environments grow in complexity and data volumes explode, traditional approaches to detection and response are proving insufficient. This episode features an in-depth conversation with Kyle Polley,
- YouTubeKyle Polley - Training BrowseSafe: Lessons from Detecting ...Kyle Polley, who leads security at Perplexity AI, discusses BrowseSafe, a classifier built to detect prompt injection attacks in AI agents. He explains how prompt injection vulnerabilities occur in browser agents like Comet when untrusted website content is ingested into the agent context, describes multiple attack types including prompt template impersonation and social engineering, and addresses limitations in existing benchmark datasets for detecting these attacks.
- YouTubeAI for SOC Automation: A Blueprint for the New world of ...Kyle Polley, head of security operations at Perplexity, discusses why organizations should prioritize detection and response capabilities before focusing on compliance, emphasizing that attackers will not wait for compliance certification. He covers emerging AI-specific threats like prompt injection targeting AI agents with system access, and explains how AI can be integrated into security operations teams to automate repetitive tasks while maintaining human involvement in incident response.
- LinkedInAI can fix threat detection: A conversation with Kyle Polley ...⚠️ Is AI the only way to fix threat detection? For years, SOC teams have been drowning: ❌ Endless false positives ❌ Burnout from alert fatigue ❌ Breaches still slipping through As Kyle Polley from Perplexity puts it: “Teams are stuck, they’re burnt out. They’re getting false positives everywhere. Companies are still getting breached.” The shift? ✅ AI agents don’t get tired ✅ They can spin up to match the number of alerts ✅ They query data endlessly without burning out Maybe the future of detection isn’t about humans doing more. It’s about AI doing what humans shouldn’t have to. 🎙 Full conversation on the Cloud Security Podcast where we unpack how AI is reshaping detection & response. Follow AI Security Podcast for more unfiltered insights on AI + security. #AISecurity #SOC #CyberSecurity #CloudSecurity #ThreatDetection
- Ep 42 | Kyle Polley - Training BrowseSafe: Lessons from ...Listen to this episode of [un]prompted Security Practitioner Con 2026 for free on iVoox. Day 2 Stage 1 Member of Technical Staff at Perplexity shares experience training and deploying BrowseSafe for detecting prompt injection in...
iVoox
- SpotifyAI for SOC Automation: A Blueprint for the New world of ...
In the news
- Reposted Aravind Srinivas
- The best way to predict the future is to invent it. We’re building and testing systems to make AI agents safer and more secure. If you’re an exceptional engineer who wants to help build a safer future, join us at Perplexity. My DMs are open
- We put all the best models in our home-grown sandbox infra (SPACE, the sandbox behind @perplexity_ai Computer), gave them root inside the VM, and one goal: hack out of the sandbox and capture the flag. Results were super interesting! 1. No model escaped the VM, but a few found
- The next time someone hacks into a frontier lab they will remember the previous reward for ethically disclosing was only $6500
- If your security team is not using AI to try to hack themselves, then someone else will
- Hacktron team is 10/10. Insanely talented group, AI alone could not have achieved this it required taste and true expertise
- Reposted Aravind
- Reposted NVIDIA AI
This page shows public professional information only, each fact cited. Is this you? send a correction, or ask for removal within 24 hours, no questions asked.

