P

Dan Shiebler

Co-founder of Artemis, an AI-native security platform; previously led the Humio acquisition work at CrowdStrike

Overview

Dan Shiebler is co-founder at Artemis [1], an AI-native security platform [1]. Shiebler holds the position of Co-Founder & CTO at Artemis [2] and previously worked at Abnormal and Twitter [3]. Shiebler holds a PhD in machine learning from Oxford [3]. Prior to founding Artemis, Shiebler led acquisition work related to Humio at CrowdStrike.

Career history

  1. Co-Founder & CTOJul 2025 to presentArtemis
  2. Head of AI/MLJan 2022 to Jul 2025Abnormal AI
  3. ML Engineering Manager (Revenue Science)Aug 2020 to Jan 2022Twitter
  4. Staff Machine Learning Researcher (Cortex)Aug 2017 to Aug 2020Twitter
  5. Deep Learning Researcher (Serre Lab)Sep 2016 to Oct 2018Brown University
  6. Senior Data Scientist (Team Lead)Sep 2015 to Sep 2017TrueMotion (acquired by Cambridge Mobile Telematics)
  7. Data ScientistSep 2013 to May 2015John D. Rockefeller, Jr. Library

Education

  1. Ph.D, Artificial Intelligence TheoryUniversity of Oxford
  2. Sc.B, A.BMay 2015Brown University
  3. Millburn High School

Insights & ideas

The through-line

Across a decade of talking about machine learning in production, Dan Shiebler keeps returning to the same claim from different angles: the model is rarely the hard part or the differentiating part, and the work that decides whether a system succeeds sits underneath it, in features, data organisation, defaults and the human structures around deployment. At Twitter that argument took the form of technical debt and ecosystem effects, at Abnormal Security it became aggregate signals as the foundation for every detector, and at Artemis it has hardened into a thesis about data architecture. His current framing is that frontier models are effectively a level playing field, so "to a great extents the core models themselves are in many ways commoditized" and "every attacker and every defender has access to extremely powerful models" [1]. What remains contested is whether a system can put the right data in front of those models fast enough.

The shift over time is in ambition rather than principle. The early advice was to avoid building a machine learning model at all where possible, because of the maintenance burden [2]. The current position is that agents should mediate everything, but only if the substrate underneath them was designed for that from the start [1].

On why data foundations decide whether AI works

He argues that the limiting factor in agentic systems is not reasoning but retrieval and organisation. Given the right context, "if they're given all the right inputs, they'll produce the right answer better than an expert almost every time", and so "the limitation is giving them the right inputs" [1]. Getting data into a shape that models can slice, query and join is the unglamorous obstacle: it is "like earth moving", with enormous volumes that are expensive to move and organised in ways that require manual manipulation before they become accessible at speed [1]. The task is fusing behavioural telemetry with the semantic data held in registries, HR systems and IT management systems, so that questions like who this user is, what this device is, and how an IP address relates to a machine can actually be answered [1]. He places current practice as an evolution through earlier waves, from the enthusiasm around RAG to the success of putting agents on top of SQL engines, and uses a combination of these approaches [1].

This is continuous with how detection was built at Abnormal Security, where every model rested on aggregate signals, "the most important component of our approach towards cyber security", built by aggregating raw emails, sign-ins and other events at the level of individual entities to form a historical picture of what has happened until now [3]. Those aggregates were the foundation for heuristics, trained models and everything above them.

On what "AI-native" actually means

He draws a sharp line between embedding agents and layering them. At Artemis, "every element of data organization of decision making of customer interface is mediated by agents", with flows that would once have been heuristics plus a smattering of models instead designed around an agent driving the loop end to end, through normalisation, decisions and the customer interface [1]. The alternative fails structurally: "When you stick AI on top of a system, you're working with a very rigid foundation," which makes it much harder to exploit the freedom the capabilities offer [1].

He extends the same standard to how the company itself operates, claiming that "Nobody has written a line of code manually in the history of Artemis. Every single line of code has been AI generated from the beginning" [1], and arguing that building this way inculcates agent monitoring, evaluation and visibility into the culture, so that the product mirrors how it was made [1]. For buyers trying to tell the difference from the outside, he offers a smell test: "How quickly is the company able to turn around feature requests and improvement requests and modifications?" [1]. If a customer needs three or four data sources a product does not support, the right foundations should deliver them in days rather than pushing them to next quarter, and while startups are always faster, "AI native development changes the physics of software engineering" and native products are easier to build natively [1].

On attackers with agents

He treats AI-enabled attack as the latest turn of a cycle rather than a rupture, since attackers have always used whatever tooling exists, with the difference now being the sheer power of a general capability arriving from above rather than being built bottom up [1]. The previous wave was the explosion of scam and spam made easier by GPT, which he says defenders at Abnormal Security bore the brunt of for several years and are still fighting [1]. Agentic execution of attacks is genuinely new, only becoming possible in the last year and hockey-sticking more recently, and it means unsophisticated attackers can now run fast, cheap, sophisticated campaigns, because "the agent is able to see each stage in the attack sequence then take the next step without waiting for human confirmation and human approval" [1]. His summary is that the technology "really is changing the physics of what attackers are capable of", which changes both what defenders can do and what they will be forced to do [1].

Adversarial pressure was already the defining feature of his detection work. Attackers actively cloak their behaviour, making messages resemble safe business traffic and using infrastructure that hides identity, so detection has to lean on "what are the things that the attack might not know that we know" and be resilient to the modifications attackers will try [3].

On choosing the right model for the job

He describes a three-tier hierarchy that ascends in cost and latency and descends in speed. First, heuristics and rules on top of aggregate signals, which he considers heavily data driven already, since the aggregates are themselves simple probabilistic models and rules over them resemble a Bayesian network of conditionals and derived probabilities [3]. Second, trained models: logistic regression, XGBoost, feed-forward networks, and deep and cross networks as the architecture of choice, borrowed from ad tech, where explicit cross layers learn multiplicative combinations of raw features [3]. He justifies that as an inductive bias, comparing it to how a convolutional network is less expressive than a feed-forward network yet performs better on images, and applies it where a frequency signal and a presence signal only mean something in combination [3]. Third, large language models, both out-of-the-box OpenAI APIs and fine-tuned variants of Falcon and llama used as classifiers for attack identification, attacker objectives and triage of a phishing mailbox product [3].

The tiers differ in how easily they absorb change: rules and prompts can be tweaked, while the middle tier needs retraining to incorporate new information, and fine-tuning a large model costs some of its generalisability [3]. Cost dictates placement. Scanning every sign-in or every email with a model of that size is prohibitive, so LLMs belong where a human was already in the loop and the burden can be reduced [3]. Underneath all of this sits an auto-retraining pipeline covering the core models on different cadences [3].

On errors that break trust

Asked which side of the confusion matrix hurts more, he refuses the easy answer. Missing a serious attack is the existential failure, so false negatives are worse in the limit, but a high false positive rate is "equally bad" because businesses cannot operate if security stops normal work, and the consequence is that "customers will end up putting overrides and ignoring remediation criteria", exposing themselves to exactly the catastrophic misses with no ability to control for them [3].

On resilience: defaults and fallbacks

He defines resilient machine learning as systems "capable of maintaining efficacy and maintaining good performance even when catastrophically bad things happen", such as the feature service that populates user attributes going down [4]. The first mechanism is defaults: add an indicator feature marking whether a block of features is present, clone the training data into a version with the features present and a version without, and let the model learn to behave in both regimes, so features can be dropped in production without collapse [4]. The second is fallbacks: build one lightweight model that predicts from the bare object alone, plus richer models that fire as additional data arrives, and iterate on the initial prediction, which keeps the system useful under high feature latency or outright service failure [4].

On technical debt and the true cost of a model

He defines technical debt as "making somehow long-term poor technical decisions in search of a short win", the figurative duct tape that accumulates when a team is flying by the seat of its pants shipping and maintaining models until nobody knows which model does what, who owns it, or where it runs [2]. His prescription starts earlier than most: production machine learning is the entire lifecycle, beginning with the question of whether a model is needed at all, and whenever the answer can be no it probably should be, because "the machine learning model is an enormous enormous engineering drain and burden" [2]. Once a model is unavoidable, the goal is to minimise the long-term drain, and he puts hard numbers on the trade-off, noting that a ten percent gain in model performance can cost a three to four hundred percent increase in engineering resources, which is rarely worth it [2].

He is equally insistent that models never live in isolation. Feature hydration must stay in concordance with training data, complex features like embeddings create long dependency chains where a small upstream change can badly damage downstream models, and a model that shifts the distribution of what gets served, such as a new candidate source on the timeline, can break assumptions in other teams' models while improving its own metric [2]. So the discipline is to err toward simplicity, understand the landscape, aim at actual business metrics rather than what seems cool, and "always be raising the question of whether or not a model is still useful and still driving value", since a product change can leave a model no longer worth its maintenance cost [2]. Attribution runs through a well-developed A/B testing framework, with dedicated teams owning that infrastructure, while responsibility for whether a model keeps moving product metrics sits with the product team or the modelling team partnering with it [2]. His organisational conclusion is blunt: "it almost never works for a team to separate the people who are building a model" from the product deploying it, except where the problem is extremely well understood, and the newer and faster-changing the problem, the closer those people need to be [2]. He also notes the difference organisational scale makes, since in a small startup a simple approach may be the first thing ever built for a given detection task, whereas a much larger company demands different work [2].

On what security teams should fix first

For anyone modernising a SOC or evaluating a SIEM, his first recommendation is coverage, because attacker activity shows up in the logs of the systems already in the organisation, so those logs need to be available, queryable and accessible [1]. The second is contextual data, the device and user information required to interpret an activity at all, brought together with the telemetry [1]. Only once those two foundations exist does it make sense to build a detection and response programme on top, and to start reasoning about remediation actions and playbooks, which depend on that visibility and accessibility [1].

Takeaways

  • Frontier models are commoditised across attackers and defenders, so the differentiator is data foundations and integrations rather than the model itself [1].
  • LLMs will beat experts when given the right context, so the real engineering problem is retrieval and organisation, which is expensive because moving and reshaping enterprise data is "like earth moving" [1].
  • Judge whether a product is genuinely AI-native by how fast the vendor turns around a new data source or feature request; days signals the right foundation, quarters signals AI layered on rigid architecture [1].
  • Agentic attack execution lets unsophisticated attackers run fast, cheap, sophisticated campaigns because no step waits for human approval [1].
  • Build detection on aggregate signals over raw events per entity, then choose deliberately between heuristics, trained models such as deep and cross networks, and LLMs, since each has different cost, latency and adaptability [3].
  • Reserve large language models for workflows that already involve a human, because scanning every email or sign-in with them is cost prohibitive [3].
  • A high false positive rate is as dangerous as a miss, because customers respond with overrides that expose them to the worst false negatives [3].
  • Make models resilient with indicator features trained on cloned present/absent data, and with lightweight fallback models that predict from the bare object [4].
  • Ask first whether a machine learning model is needed at all; a ten percent performance gain can cost three to four hundred percent more engineering [2].
  • Do not separate the people building a model from the people deploying it unless the problem is extremely well understood [2].

Media & appearances

  • ConferoYouTube
    What Is an AI-Native SIEM? Rethinking Security OperationsDan Shiebler, co-founder and CTO of Artemis, discusses how AI is changing the threat landscape and why data architecture is becoming critical in modern security operations. He explains that while large language models and frontier AI models are now commoditized across defenders and attackers, the key differentiator is building the right data foundations and integrations rather than simply layering AI on top of existing security tools, which is what most new AI security companies are doing.
  • Super Data ScienceYouTube
    630: Resilient Machine Learning — with Dan ShieblerML & AI Podcast with Jon Krohn: #MachineLearning #ResilientMachineLearning #FiveMinuteFriday@JonKrohnLearns sits with Dr. Dan Shiebler at the Open Data Science Conference (ODSC) to dive int...
  • Super Data ScienceYouTube
    717: Overcoming Adversaries with A.I. for Cybersecurity — with Dr. Dan ShieblerML & AI Podcast with Jon Krohn: #AIForCybersecurity #Cybersecurity #ResilientMachineLearningDr. Dan Shiebler, Head of ML at Abnormal Security, joins @JonKrohnLearns this week and unveils th...
  • Banana Data PodcastYouTube
    Life After Production: A Tale of Technical Debt Feat. Dan Shiebler, Sr. ML Engineer at TwitterDan Shiebler discusses technical debt in machine learning production systems, defining it as making suboptimal technical decisions in pursuit of short-term wins. He explores how technical debt accumulates in organizations as teams rapidly ship and maintain models without proper processes, and emphasizes the importance of quality assurance and productionalization practices, drawing parallels to manufacturing processes like McDonald's french fries production.

This page shows public professional information only, each fact cited. Is this you? send a correction, or ask for removal within 24 hours, no questions asked.