Julien Chaumond

Co-founder and CTO of Hugging Face, founded in New York City in 2016

Overview

Julien Chaumond is co-founder and CTO of Hugging Face[1], a company founded in New York City in 2016[1]. Chaumond holds the position of CTO at the organization[3] and maintains a professional presence on LinkedIn[2] and X[4].

Career history

  1. Co-founderJul 2016 to PresentHugging Face
  2. Advisor to the Deputy Minister for Digital AffairsJan 2016 to Jul 2016Ministère de l’Economie et des Finances, de l’Action et des Comptes publics
  3. Software EngineerJan 2015 to Dec 2015Stupeflix
  4. Co-founder and CTOJan 2013 to Jan 2015Glose
  5. Research AssistantMar 2007 to May 2008Stanford University
  6. InternApr 2006 to Aug 2006Infosys Technologies Ltd

Education

  1. M.Sc., Electrical Engineering / Computer Science2006 - 2007Stanford University
  2. Diplôme d'Ingénieur (M.Sc.), Applied Maths2003 - 2006École Polytechnique

Insights & ideas

The through-line

Everything Julien Chaumond says circles back to one conviction: natural language is the largest and hardest surface of machine learning, conversational AI is its summit, and the honest route to that summit is incremental, measurable, and open. Hugging Face began with an entertainment-oriented conversational product that reached roughly half a million monthly active users, and the founding lesson of those first years was that the technology was not ready for the assistant they actually had in mind [1]. Rather than keep polishing a product on top of insufficient science, they stepped back and pushed on the science itself, a decision he frames as deliberate rather than defeatist: the long-term goal is still a conversational AI so fluid it displaces how people use Google or Facebook, and he is careful to add that "we don't want to replace humans we like humans" [1].

The second half of the through-line is method. He repeatedly contrasts an aspirational, discontinuous vision of progress with the pragmatic one he prefers, tracking benchmarks and tasks so that improvement is "a bit more concrete, a bit more measurable over time" [2]. That preference explains the rest: open-source libraries that anyone can read, demos that put the state of the art in a browser, distilled models small enough to run in production, and infrastructure that turns published research into something a working developer can call [1][2][6].

On conversational AI as the holy grail

He treats conversational AI as the top of a difficulty pyramid rather than as one task among many. "I think conversational artificial intelligence is probably the hardest to solve because it essentially requires solving all the other NLP tasks" [2]: a usable system has to classify the user's input, extract entities, resolve what product or concept is being referred to and with what confidence, then retrieve and respond. His verdict on the industry is blunt, that "today none of the big companies have really managed to do anything in conversational interfaces that actually works well" [2], and he expects several more years before the quality is there for genuinely industrial needs. That difficulty is a recurring subject for him, from the early product experience through to later discussions of why conversation remains hard [4]. He is more optimistic about the lower layers of the pyramid, where he thinks the field is getting steadily better [2].

He also reads the field's history as validating the bet. By 2016 the first neural conversation systems, mostly RNN and LSTM based, were producing poor conversations but proving that end-to-end learning could replace hand-coded rules, and that trajectory was enough to justify starting a company on it [1]. What surprised him was the acceleration that followed [1].

On consumer science and refusing early monetization

He describes the company's shape as "consumer science": push the science of machine learning and simultaneously look for consumer usages it unlocks [1]. That posture only works with investors who accept it, and he attributes much of the freedom to having found people aligned with the premise, "let's not try to find monetized business application just yet let's just focus on the growth" and on pushing the state of the art, with monetization to follow later in a way consistent with an ethical view of NLP as a whole [1]. He is explicit about the oddity of a venture-funded company producing public research, and about the time that funding buys to keep doing it [1]. The partnership behind it he explains in equally unsentimental terms: he and Clément met around ten years earlier in Paris while working in the same startup space, stayed in touch through other projects, and started the company mainly on personal fit and a shared desire to work together [1].

On transfer learning and what large pretrained models changed

His account of the BERT and GPT-2 era is a two-phase story: train a very large language model on enormous, essentially unsupervised text corpora such as large slices of the web, a phase that is expensive and done once, then stack small, shallow networks on top and fine-tune them for specific tasks [2]. The payoff is that classic NLP problems, document and sentence classification, entity extraction, become much cheaper and much more accurate, and that a single base model supports many successive fine-tunings for very varied tasks [2]. He grounds this in ordinary usage: customers routing support email, deciding which department an incoming message belongs to, a problem that a little fine-tuning on the same base model solves very well [2].

He is equally clear that the architectural shift was not gradual. Before transformers the field used RNNs and CNNs that were difficult to train at scale; the transformers papers changed the state of the art in a way he insists was qualitative, "it's not incremental, it's a discovery" [2]. He expects the ecosystem of variants to keep growing, with RoBERTa and DistilBERT as examples of teams pressing the same ideas further, and while he thinks the field may need fewer attention heads than it first assumed, attention as a fundamental building block for discrete sequential tasks is, in his view, here to stay [2].

On building the transformers library

The library began as a practical problem. BERT and GPT-2 were released in TensorFlow, the team's stack was PyTorch, and Thomas Wolf took a week to sprint on porting BERT, which he describes as a genuine feat [1]. The lesson he draws from that port is that framework translation is never mechanical, and the difficulty lies where people do not expect it: "it's usually not the model itself but it's all the different parts around", the initialization, the inference path, the training and evaluation setup, all the places where different researchers do things differently [1]. What started as a port became a repository of large transformer models and, quickly, a widely used library, past ten thousand GitHub stars and with tens of thousands of model downloads [1].

The design philosophy is legibility. Models are kept self-contained, one architecture implemented in a few files and a few hundred lines, precisely "so you don't have to read through like hundreds of pages of documentation or hundreds of files of code to understand what's happening" [1]. The point of that constraint is not elegance for its own sake: he wants users to be able to onboard, run a model, then start tweaking training parameters and chase the state of the art themselves, which is why he stresses that researchers use the library to discover new methods rather than only to consume finished ones [1].

On demos, and putting the state of the art in front of people

He treats public demos as a serious output, not marketing. The one that impressed him most personally is Write with Transformer, a text completion tool built on GPT-2 that behaves like Gmail's suggestions but far more creative, generating much longer continuations and taking far more liberties, with model size and hyperparameters such as temperature exposed for the user to play with [2]. He notes the demo was built by porting the models to PyTorch and packaging them into a web application, and that people found many unexpected uses for it [1]. The same instinct runs through the coreference and conversational demos in the browser, and through deploying models onto iOS devices [1]. The conversational work also had a competitive proof point: at the ConvAI challenge organized at NeurIPS with support from Facebook AI, the team applied a GPT model to end-to-end conversation trained on the PERSONA-CHAT dataset and won the automated evaluation [1].

On what "solving NLP" means, and on benchmarks

Asked what the company's mission statement actually means, he concedes it is "intentionally a bit vague" [2], then gives it a concrete reading. Benchmarks like GLUE bundle paraphrase detection, sentiment analysis and natural language inference, and scores have reached or passed human level on them; whether that constitutes emerging intelligence or simply a model that captures the statistics better than a human doing a different kind of analysis is, he says, an open research question that touches the definition of intelligence and has no answer today [2]. What matters practically is the loop: as models solve a benchmark, harder ones such as SuperGLUE appear, and he calls this "a pretty healthy evolution where the difficulty of the dataset and the complexity of the models are going at roughly the same pace" [2]. Advancing NLP, for him, is participating in that iterative improvement over years, a stance he positions as complementary to and in tension with an AGI-framed programme, since there is no benchmark for AGI: "we want something a bit more pragmatic" [2]. He is also sympathetic to the argument that intelligence is multidimensional and that machines may develop along directions unlike the human one [2].

On data, compute, and who gets to do research

Presented with the claim that the big leaps have come from new data rather than new models, he splits the difference. In NLP specifically both are true, and the encouraging part is that the data driving progress is unsupervised or weakly supervised, so building pipelines that ingest larger parts of the web does not demand years of an annotator's life, meaning "we can build more intelligent systems just by training on larger parts of the web or larger datasets" [2]. But architecture matters too, and the transformer is his example of one family being strictly superior to another [2].

He is alert to the cost of that regime. If only labs with compute can train the largest models, genuinely interesting smaller ideas, including alternatives to attention, risk never being evaluated because they lose on the leaderboard [2]. His answer is disclosure norms already gaining ground: researchers reporting parameter counts, training budgets and the compute cost of the research process, which differ from the cost of the final run by at least an order of magnitude [2].

On making big models usable in production

The production answer he keeps returning to is compression. DistilBERT retains roughly half the parameters of BERT base, a reduction of between forty and fifty percent achieved mainly by cutting layers, while holding close to the same performance on benchmarks like GLUE and running more than twice as fast at inference, which makes it small enough to deploy, even on a server [2]. Getting large language models into production at millisecond inference latency is a problem he has discussed at length, alongside the argument for something like a CERN for machine learning and the BigScience collaboration [5]. The other half of the answer is infrastructure rather than modelling: hosting models on a model hub, inference widgets that let companies run models on CPUs and GPUs, automated training through Auto NLP, and a broader move toward machine learning offered as infrastructure as a service [6]. The consistent motivation, as his own description of the work has it, is democratizing state-of-the-art AI and machine learning for everyone [6].

On teams, autonomy, and learning the field from scratch

He hires against a combination he considers unusual: people strong in machine learning and in software engineering at once, because the field is full of people who are excellent scientists but not equipped to deploy and maintain systems in production. "we try to build like a kind of homogeneous team of people" who are good at both [1]. The structural consequence is that "everyone is pretty much autonomous" and self-driven, which frees him to move between projects every few weeks or months rather than own one, a way of working he says suits him and keeps him contributing across most of what the company ships [1]. He was proud that a team of ten, heading toward fifteen or twenty, had become visible and was moving the state of the art [2].

That team taught itself. He describes spending a large part of the first year catching up through online courses, singling out a Stanford NLP course whose playlists he rates highly and whose professor, after they told him how much it had helped them start, became an angel investor and has reinvested in every round since [2]. He is a straightforward advocate for that kind of educational content as the way past the field's steep learning curve [2]. Part of why the curve was steep for him is cultural: French engineering schools are led by theoretical computer scientists interested in proving an algorithm optimal, which he calls the exact opposite of the machine learning attitude where you "don't really know why it works but at the end we are going to have something that kind of works" [1]. In 2005 machine learning was seen as a toy set of brute-force methods, the compute and the techniques were not there, and his own attempts to apply small neural nets to robotics competitions did not work well [1]. He came back to it in 2015 specifically because he wanted to see what had changed in the intervening years, and found the answer exciting enough to build a company on [1].

Takeaways

  • Conversational AI is the hardest NLP problem because it subsumes all the others, classification, entity extraction, retrieval, and no large company has yet built a conversational interface that genuinely works well [2].
  • Hugging Face pivoted away from polishing a consumer conversational app with around half a million monthly active users once it became clear the science was not ready, and reorganized around what he calls "consumer science" [1].
  • Porting models between frameworks is hard not because of the model code but because of everything surrounding it: initialization, inference, training and evaluation conventions differ between researchers [1].
  • The transformers library is deliberately built so each architecture is self-contained in a few files and a few hundred lines, so that users can read, run and then modify models rather than only consume them [1].
  • DistilBERT cuts BERT base by forty to fifty percent of its parameters, mainly by removing layers, keeps close to the same GLUE performance and runs more than twice as fast, making it deployable in production [2].
  • Progress should be judged by an iterative loop of harder benchmarks such as GLUE then SuperGLUE, a pragmatic and measurable alternative to an AGI-framed research programme with no benchmark [2].
  • Concentration of compute risks burying good small ideas, which is why he supports researchers publishing parameter counts, training budgets and total research cost [2].
  • Hire people who are strong at both machine learning and software engineering, keep the team small and autonomous, and expect to rotate across projects rather than own one [1].

Media & appearances

  • MLOps.communityApple Podcasts
    Tour of Upcoming Features on the Hugging Face Model Hub // Julien Chaumond // MLOps Coffee Sessions #48Coffee Sessions #48 with Julien Chaumond, Tour of Upcoming Features on the Hugging Face Model Hub. Join the Community: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://go.mlops.community/YTJoinIn⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ Get the newsletter: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://go.mlops.community/YTNewsletter⁠⁠⁠⁠⁠ //Abstract Julien Chaumond’s Tour of Upcoming Features on the Hugging Face Model Hub. Our MLOps community guest in this episode is Julien Chaumond, the CTO of Hugging Face - every data scientist’s favorite NLP Swiss army knife. Julien, David, and Demetrios spoke about many topics, including:Infra for hosting models/model hubsInference widgets for companies with CPUs & GPUs (for companies)Auto NLP, which trains models“Infrastructure as a service” // BioJulien Chaumond is Chief Technical Officer at Hugging Face, a Brooklyn and Paris-based startup working on Machine learning and Natural Language Processing, and is passionate about democratizing state-of-the-art AI/ML for everyone.
  • The MLOps PodcastApple Podcasts
    🤗 Large ML models in production with HuggingFace CTO Julien ChaumondIn this episode, I'm speaking with Julien Chaumond from 🤗 HuggingFace, about how they got started, getting large language models to production in millisecond inference times, and the CERN for machine learning. Join our Discord community: https://discord.gg/tEYvqxwhah --- Timestamps: --- Relevant Links: FastDS: https://github.com/DAGsHub/fds BigScience: https://bigscience.huggingface.co
  • Chai Time Data ScienceApple Podcasts
    Hugging Face, Transformers | NLP Research and Open Source | Interview with Julien ChaumondSubscribe here to the newsletter: https://tinyletter.com/sanyambhutani In this episode, Sanyam Bhutani interviews the CTO of Hugging Face, Julien Chaumond. In this interview, they talk all about his journey into the field and his path as a CTO of Hugging Face. They discuss all the amazing work being done at Hugging Face, research and open source as well as about their team. Links: Transformers Repo: https://github.com/huggingface/transformers Follow: Julien Chaumond: Hugging Face: Sanyam Bhutani: Blog: https://medium.com/@init_27 About: A show for Interviews with Practitioners, Kagglers & Researchers and all things Data Science hosted by Sanyam Bhutani. You can expect weekly episodes every available as Video, Podcast, and blogposts. If you'd like to support the podcast: https://www.patreon.com/chaitimedatascience
  • YouTube
    Hugging Face, Transformers | NLP Research and Open Source ...Julien Chaumond, CTO of Hugging Face, discusses the company's founding three years prior with co-founder Clement, their initial focus on entertainment-oriented conversational AI with hundreds of thousands of monthly active users, and their pivot toward open-source NLP research and technology. He covers Hugging Face's work on transformer models, the transformers library for fine-tuning on NLP tasks, web demos including a text generation tool at transformer.hugging face.com, and deployment of models to iOS devices.
  • Listen Notes
    Julien Chaumond - Top podcast episodes - Listen NotesOctober 24, 2017. Conversation AI is hard. Tour of Upcoming Features on the Hugging Face Model Hub // Julien Chaumond // MLOps Coffee Sessions #48. Datasets: …
  • YouTube
    State of the art in NLP - Julien Chaumond, CTO Hugging Face ...Julien Chaumond discusses Hugging Face's founding with co-founders Clément and Thomas Wolff in Paris three years prior, his role handling the technical side, and the company's achievements in advancing machine learning and NLP. He highlights the team's work on democratizing NLP and making transformer-based models more accessible to people entering the field, emphasizing the importance of educational resources and reducing the steep learning curve in machine learning.
  • Listen Notes
    Hugging Face, Transformers | NLP Research and Open Source ...
  • Amazon Music UnlimitedAmazon Music
    Hugging Face, Transformers | NLP Research and Open Source ...

In the news

This page shows public professional information only, each fact cited. Is this you? send a correction, or ask for removal within 24 hours, no questions asked.