George Sivulka

Founder and CEO of Hebbia, the AI platform for finance and legal research

Overview

George Sivulka is Founder and CEO of Hebbia[1][2], a position Sivulka has held since August 2020[3]. Sivulka holds a Doctor of Philosophy in Electrical Engineering from Stanford University[4], along with a Master of Science in Applied Physics from Stanford University[6] and a Bachelor of Science in Mathematics from Stanford University[5]. Sivulka maintains a presence on X at @gsivulka[7].

Career history

  1. Founder & CEOAug 2020 to PresentHebbia.AI
  2. FounderHebbia
  3. Founder & CEOHebbia

Education

  1. Doctor of Philosophy - PhD, Electrical EngineeringStanford University
  2. Bachelor of Science - BS, MathematicsStanford University
  3. Master of Science - MS, Applied PhysicsStanford University

Insights & ideas

The through-line

Sivulka's fixed idea is that a technological revolution is worthless until someone builds the thing that captures it. He frames artificial intelligence, or more precisely metalearning, as the seventh major technological revolution in human history and possibly one of the last, then immediately adds the qualifier that matters: "technology is not a product" [2]. The historical rhyme he reaches for is blunt: "Fire wasn't important without the torch. The wheel wasn't really that important without the chariot" [2], and even computing stayed a hobby for nerds until Excel and the graphical user interface arrived [2]. Hebbia is his bet on what the capturing product looks like, and his answer to what it is not is equally blunt: "It's not going to be a chatbot" [2].

The second, older strand is a grievance about wasted human capacity. Teaching Stanford's math 51 course, he watched brilliant former students go to Morgan Stanley and Goldman Sachs and turn miserable, and concluded he "had never seen like such smart people going through so much hell like doing such stupid things over and over and over again" [3]. The original mission was to stop smart people doing stupid tasks; the founding insight was that "you had a lot of the smartest people in the world doing the stupidest tasks" and that there was an arbitrage in reclaiming their time [1]. What has shifted is scope. Where the company began with a Jupyter notebook running an early neural information retrieval model over 400-page DEFM14A filings for former students in Morgan Stanley's Menlo Park office [3], the stated ambition is now "building capable AI platform for a billion people" [1], horizontal rather than vertical, and he has said plainly that if the company stays in financial services he will have failed [3].

On why chat is the wrong interface for serious work

He treats chat as a demo of capability rather than a working shape. Chat is "maybe the first iteration of the technology, but it's definitely not the last" [2], analogous to a TI-84 calculator: show one to someone in the 1900s and they would have a heart attack, but it did not realise the promise of computing [2]. The limitation is that chat takes a single step. Real professional work involves collaboration, editing, iteration and back-and-forth in shared materials, so what is needed is "a much richer, deeper, more processoriented interface" [2]. Investment research, drug discovery, even scouting geological sites for new mines are multi-step, multi-hop problems that a single prompt cannot prosecute [2].

The design answer is orchestration made legible to humans. The spreadsheet form, where every cell is an agent and the cell shows the output of running that agent over some data, is the first iteration, but he describes grids and other human interfaces meant to portray the activity of many agents, agents of agents, and let people build pipelines of data and data processing without coding [2]. He has also framed the same idea through the image of the AI orchestra conductor [9][12]. The contrast with chatbots is where he plants the flag on usage: with most AI platforms "you're having a bit of a transactional relationship with the AI. You ask a question, you get a response", whereas here you hand over complex tasks that churn through vast quantities of data, so "It's much more of an agent than a chatbot" [1]. The volume metric he tracks is unstructured pages processed, from roughly 100 million pages last year to four or five billion on track this year, somewhere around 50,000 years of human reading [1].

On generalisation over specialisation

He argues against the vertical AI instinct on empirical grounds drawn from his customers: the best investors read genuinely strange data sources with nothing to do with finance, and the best lawyers have backgrounds far outside law [1]. His summary of the lesson is "You don't want specialization. You want generalization" [1]. That belief shows up in the product as a refusal to fine-tune. Working with many of the largest asset managers, many banks and the US government, he says no part of the software has needed to be personalised per customer or per vertical, which he compares to the absurdity of fine-tuning Microsoft Word [3]. It also shows up in deliberately hard product calls: he has passed on the simpler cut that would have served the initial customer profile in order to build something broader [3]. The platform now aims to let anyone, whether they can code, cannot code, or can only code with Cursor, build whatever knowledge work application or agent they want [1]. The question of generalised versus specialised foundation models, and whether they are commoditised, is territory he has covered elsewhere too [4].

On inference-time compute and the context window

His technical bet, made early, was that accuracy comes from spending more compute at runtime. He claims the inference-scaling paradigm was pioneered at Hebbia: two years ago with the early Matrix product "we're one of the first people to say, 'Hey, you get way better accuracy from using more large language model calls um at runtime'", and the company built infrastructure to scale that up and an agents team before the word agents was in use [1]. With training scaling laws feeling like they have slowed and no recent paradigm shift, he considers test-time compute the most interesting research direction going forward [1], a view that sits alongside his broader case that reasoning models and agents can supercharge knowledge-worker productivity and the global economy [7].

Alongside that he calls context length the central obstacle: "the number one problem in all of AI is elongating the context window" [1]. His objection to retrieval-only approaches is that you can jam the window with as many search results as you like but the model will not reason over that data, it will only find things that exist; a long enough window lets you connect dots and notice what is missing as well as what is there [1]. Hebbia's research has gone into artificially extending the window, which he describes as mostly an infrastructure problem, a patchwork or convolution problem of applying today's maximum window repeatedly, plus an information-theory problem of making each run's output both maximally informative and maximally compressed so each subsequent iteration receives exactly the relevant material and nothing else [1]. He is watching multimodal research and how context windows tie into the wider foundation model ecosystem for text applications [1].

On hallucination, accuracy and trust

The method he describes is an inversion of the usual order: rather than generate text and then hunt for citations, "we find the citations first and then generate the text" [2]. Every answer is corroborated by a highlight before it is generated, and there is a retrieval step even before creation [2]. He calls this a Copernican inversion and credits it with much higher accuracy [2], and the same discipline underpins the RAG approach that keeps outputs inside company-sanctioned data [11]. He is candid about the trade: "we leave creativity to the humans" [2], which means the system is less suggestive, and users have to drive it rather than trust it blindly [2]. He expects hallucination to fade from the discourse as bigger models, cite-before-you-generate techniques and tool use converge, on the principle that a model bad at maths should just use a calculator and get 100% accuracy every time [2].

Accuracy matters more than it looks because errors compound across steps. Models are "good at first order tasks" and can get 90% there, but chain three 90% systems and the output is the product, not the sum, so "you're going to have something that is you know 10 good" depending on how many steps you have [3]. That is why he insists the answer is not a prompt, not LangChain logical operators, not the right Pinecone database, but getting each step from 90% to 99.9% [3]. He was also early to say the models themselves would eat generic tooling work, citing OpenAI shipping function calling and wiping out startups built around tool-form implementations, and describes his obsession as building input, output and flexibility layers that make the model better rather than duplicating it [3]. Trust is the commercial constraint: people do not by default trust an LLM output, and when someone is allocating a great deal of money or making a mission-critical decision, being right matters [3]. He has argued more broadly that transparency is what makes or breaks trust in AI [11].

On agent employees and organisational design

He frames agents through the precedent of remote work. The internet decoupled the output of labour from location and democratised access to talent; AI decouples output from personhood, and from whether you can pay salaries or have humans doing the jobs at all [1][2]. What follows is a spectrum of organisational forms: fully human companies, fully agentic ones such as the one-person billion dollar startup or things that are effectively just APIs, and, most commonly, hybrids where agent and human employees work alongside each other [1][2]. Once you treat the agent employee as another node in the org chart, the practical consequences are mundane and specific: "It will probably have to have an email. It'll probably have to have a Slack. It'll probably be doing things the wrong way, and you'll have to manage it in the right direction" [1]. The mix, whether you are 0%, 10%, 50% or 90% agents, will define what the organisation can produce, in the same way Amazon's decision to give every AWS offering its own GM shaped the products and whether they work together [1].

The corollary is a change in what management means. He argues we are already AI managers, just with agents that take a single step, and that how good you are at prompting is how good you are at managing [1]. As agents roll out, the people who are best at prompting and best at defining a process will be the best managers and the most effective at extending their agenda inside an organisation: "prompting is managing and uh it will all blur pretty soon" [1]. Much of his interest in interface design comes back to this, since collaborative multi-agent frameworks are, in his reading, just different interfaces for managing agent employees [2].

On jobs, adoption and the economy

His public prediction is that AI agents will create more economic value than human workers within ten years [9][10][12], and he expects intelligence to become too cheap to meter [1]. But he is deliberately unromantic about the timeline: everything is backloaded. Chatbot-form AI will still be getting deployed in ten years, because organisational and technological change take time and effort, and "there's actually still plenty of businesses in New York that don't take credit cards" [1]. The counterweight is that when something genuinely works rather than being an experiment, adoption is fast, and some AI companies have gone from zero to double-digit percentages of their addressable market [1]. He accepts there will always be a long tail of stragglers [1].

On displacement he separates two questions. On sequencing, he concedes that for better or worse AI is currently better at knowledge work than at physical labour, so white collar work is more exposed first, but he expects metalearners that generalise over the physical world too, so the delineation will fade: "I'm not sure we'll all become plumbers and chefs", though a world where people realise childhood dreams might raise humanity's net happiness score [2]. On whether jobs disappear in aggregate he is optimistic, and reaches for the story of a Morgan Stanley figure who claims to have invented the analyst: the partners had a room of computers nobody could use, hired new grads from Columbia to operate them, and the analyst programme at every major investment bank grew from there [2]. His own framing of what he is building keeps humans central: "humans are really really important and making humans better is actually the ultimate goal of what I do" [3], and he compares it to bookkeeping moving out of physical spreadsheet books, where the technology changed jobs rather than erasing them [3]. Knock on wood, he said at the time, not a single job had been replaced by the product [3]. He has also spoken to the ethical challenges of AI [11] and, on the market side, argued that AI companies are collectively undervalued [4].

On the Excel analogy, the wedge, and building the company

The mental model for the product is Excel in 1985: it took a powerful technology, the SQL database, made it wranglable by normal humans in a WYSIWYG application, started in a niche in fixed income and ended up running governments [3]. He calls Excel the most important software ever made and expects archaeologists to be reading Excel files in 2000 years [3]. His version of the same move is a large language model native productivity tool where anyone with no technical training can wrangle models and documents into whatever output they want, an ETL pipeline, a search exposed to users, a chatbot spun up for an enterprise in seconds, or extracted data [3]. The proprietary work sits in three buckets, input, output and flexibility: integrations, document parsing and multiple index types to get good information in; programmatic wrangling to get controllable outputs back; and a product people can understand well enough to trust [3]. The platform is model agnostic by design, triaging calls to OpenAI, Anthropic, Google and others according to what the user wants, on the logic that "no one really cares what the back end of excel looks like it just works" [3][2].

The go-to-market wedge was chosen by looking for the equivalent of fixed income: a market the tool could take fast and then be leveraged elsewhere [3]. The insight was that these models are very good at mapping from one ontology to another, and private equity diligence is exactly that, a data room of unstructured company files on one side and a hundred or four hundred question tracker on the other, filled in by the highest paid people who hate their lives the most, deciding whether hundreds of millions of dollars get deployed [3]. Around that sit features like neural search, chatbots, semantic comparison of two documents, question lists, table extraction and alerts, the last of which he uses himself to be notified of the counterparty, senior stakeholder, fees and customer concentration every time a new contract is signed [3]. On the company itself he emphasises the team above metrics: a hundred people in the New York office five days a week and sometimes six, a London office already open and San Francisco opening, with a goal of 300 to 400 employees by year end, deployment at 40 to 50% of the world's largest asset managers and tier one investment banks, and a legal segment growing at an exponential clip [1][6].

He is B2B first by choice rather than by default. His framework is that consumer products map to flaws in human psychology, that "consumer products have to map to the seven deadly sins" or they fail, whereas B2B lets you appeal to knowledge, discovery and creation instead of laziness, gluttony, virality or echo chambers [2]. The long-term ambition is nonetheless universal, a piece of software that goes out to everyone like the Microsoft productivity suite [2], and his personal measure of success is that when his children take sixth grade computer science, the system tray they learn contains Google Chrome, Microsoft Excel and Hebbia [3].

The origin story he keeps returning to is the moment of surrender to a better result. He was building the world's first metalearner in his PhD, convinced it would be the most important technology ever invented, when OpenAI published GPT-3 under the title large language models are metalearners [1][2]. He sat back in his chair, realised the paper had beaten him to the punch, and concluded that if he could not create the fundamental technology he wanted to apply it, abandoning a fully funded programme and millions of dollars of research grants [1][2]. Before that there had been three separate moments of freaking out at early Transformers in 2016 and 2017, at Google Translate getting good, and at friends applying neural methods to information retrieval [3]. He has also discussed the coral reef origin of the company's name, working at NASA as a teenager, meeting Peter Thiel, and being a disappointment to his athlete parents [9][10][12], as well as going from being unable to afford rent to raising $130M [4].

Takeaways

  • A technological revolution needs a capturing product: "Fire wasn't important without the torch. The wheel wasn't really that important without the chariot", and for AI that product "is not going to be a chatbot" [2].
  • Chat is a single-step interface; serious work is multi-hop and collaborative, which is why the design goal is orchestrating many agents through human-legible surfaces like grids and spreadsheets where each cell is an agent [2].
  • Hallucination is addressed by inverting the pipeline: "we find the citations first and then generate the text", accepting that the system is less creative because "we leave creativity to the humans" [2].
  • Chained 90% systems multiply into roughly 10% reliability, so the work is getting each step to 99.9% rather than building logical operator trees [3].
  • Accuracy comes from spending more compute at runtime, a bet made two years before the industry named it, and the biggest remaining bottleneck is context length: "the number one problem in all of AI is elongating the context window" [1].
  • Vertical specialisation is the wrong instinct because the best investors and lawyers are generalists, and the software has needed no fine-tuning across asset managers, banks and the US government [1][3].
  • Treat the agent as a node in the org chart with an email and a Slack that will do things wrong and need managing, and accept that "prompting is managing" [1].
  • Agents will create more economic value than humans within a decade, but adoption is backloaded, since plenty of New York businesses still take only cash [1][9][10][12].

Media & appearances

  • AI and the Future of WorkApple Podcasts
    SE 20: The Founders’ Playbook: How to Build AI Companies That Last (Special Episode)Artificial Intelligence in the Workplace, Business, Ethics, HR, and IT for AI Enthusiasts, Leaders: Send a text In this special February compilation episode of AI and the Future of Work, we explore what it truly takes to build AI companies designed to last. While AI innovation moves fast, enduring companies are built on fundamentals. Clear problem sel
  • The MAD Podcast with Matt TurckApple Podcasts
    AI That Ends Busy Work — Hebbia CEO on “Agent Employees”What if the smartest people in finance and law never had to do “stupid tasks” again? In this episode, we sit down with George Sivulka, founder of Hebbia, the AI company quietly powering 50% of the world’s largest asset managers and some of the fas
  • AI and the Future of WorkApple Podcasts
    335: How AI is Changing Academia with Dave Marchick, Dean of the Kogod School of BusinessArtificial Intelligence in the Workplace, Business, Ethics, HR, and IT for AI Enthusiasts, Leaders: Dave Marchick is the Dean of American University’s Kogod School of Business and a seasoned leader with experience across the private, public, and nonprofit sectors. He spent over a decade as Managing Director at The Carlyle Group, where he served on t
  • The Innovation Civilization PodcastApple Podcasts
    #36 - George Sivulka : Knowledge Work 2.0: The Company Creating The Multi-Agent FutureWe’re joined by George Sivulka, Founder and CEO of Hebbia, for a conversation on how does the future of white collar work look like with multi-agents. Hebbia is backed by some of the most legendary technology investors of our generation including Peter Thiel (early investor: Paypal, Facebook), Marc Andreesen (early investor: Airbnb, Github, Coinbase), Eric Schmidt (ex-CEO Google), Jerry Yang (Co-founder Yahoo). George’s background from Stanford’s PhD program, combined with his work at the cutting edge of AI meta-learning, has led him to a bold mission: to build Hebbia into a generationally important company that captures the full power of the AI revolution, not through chatbots, but through entirely new interfaces for serious, complex work. We dive into: What does the future of white collar/knowledge work look like What the future UX/UI of Agentic AI might be (beyond chatbots). How Hebbia uses multi-agent orchestration to tackle tasks like investment research, drug discovery, and complex analysis. How Hebbia solves hallucination by "citing first, generating second." Why George believes AI won't eliminate jobs, but will transform how we work—and why humans will always find new ways to create value. The lessons George has learned from investors like Peter Thiel and Eric Schmidt about building great companies.
  • AI + a16zApple Podcasts
    Reasoning Models Are Remaking Professional ServicesHebbia founder and CEO George Sivulka discusses the potential for reasoning models and AI agents to supercharge knowledge-worker productivity — and the global economy along with it.
  • AI and the Future of WorkApple Podcasts
    322: How AI is Revolutionizing Finance and Decision-Making: Insights from wunderkind George Sivulka, Hebbia CEOArtificial Intelligence in the Workplace, Business, Ethics, HR, and IT for AI Enthusiasts, Leaders: Send us Fan Mail George Sivulka, hailed as a wunderkind by Peter Thiel, is the founder of Hebbia, a groundbreaking AI company revolutionizing financial research. Having worked at NASA as a teenager, he later earned a math degree from Stanford in just 2.5 years. George has secured $130M in funding for Hebbia at a $700M valuation, backed by Andreessen Horowitz, Index Ventures, Google Ventures, and Thiel himself. Hebbia’s AI platform, trusted by firms like Centerview Partners and Charlesbank, has driven a 15x revenue growth in 18 months by pioneering RAG techniques that ensure LLM outputs remain within company-sanctioned data. In this conversation, we discuss: How Hebbia is reshaping financial research with retrieval-augmented generation (RAG).The leap from academia to entrepreneurship—how George Sivulka turned research into real-world impact.The expanding role of AI in high-stakes industries where precision is critical.The ethical challenges of AI and how transparency can make or break trust.Why AI agents will change knowledge work—but not in the way you might think.The future of AI-powered research and its impact on decision-making.Resources: Subscribe to the AI & The Future of Work Newsletter: https://aiandwork.beehiiv.com/subscribe Connect with Sivulka on LinkedIn: https://www.linkedin.com/in/sivulka/ AI fun fact article:…
  • The Twenty Minute VC (20VC)Apple Podcasts
    20VC: Why All AI Companies Are Under-Valued | The Future of Foundation Models: Scaling Laws, Generalised vs Specialised, Commoditised? | From Unable to Afford Rent to Raising $130M From Index and Peter Thiel with George Sivulka @ HebbiaVenture Capital | Startup Funding | The Pitch: George Sivulka is the founder and CEO of Hebbia, is one of the fastest-growing gen AI companies and they recently raised a $130M series B. Investors include the company include hailed names such as a16z, Peter Thiel, Index, GV and others. In...
  • The Times Tech PodcastApple Podcasts
    Hebbia’s George Sivulka: “Bots will be most of the economy within decade”Artificial intelligence “agents” will create more economic value than humans within ten years. Sound outlandish? That is the prediction of this week’s guest, George Sivulka, founder of AI startup Hebbia, who comes on to talk about building AI that actually works for business (3:20), AI orchestra conductors (9:15), coral reefs and why he called the company Hebbia (10:30), why he started the company (19:30), being a “disappointment” to his athlete parents (29:30), working at NASA as a teen (33:00), meeting Peter Thiel (36:45), and how AI is going to revolutionize the economy (42:00). Hosted on Acast. See acast.com/privacy for more information.
  • The MAD Podcast with Matt TurckApple Podcasts
    Hebbia: Saving Smart People from Mundane Tasks with CEO George SivulkaToday we sit down with George Sivulka, CEO of the AI productivity tool, Hebbia, for a conversation on generative AI in fintech & government how how Hebbia keeps "smart people from doing stupid tasks" in fintech & government.
  • Listen to Hebbia’s George Sivulka
    Hebbia's George Sivulka: "Bots will be most of the economy within ...“Bots will be most of the economy within decade” from Times Tech. Artificial intelligence “agents” will create more economic value than humans within ten years. Sound outlandish? That is the prediction of this week’s guest, George Sivulka, founder of AI startup Hebbia, who comes on to talk about building AI that actually works for business (3:20), AI orchestra conductors (9:15), coral reefs and why he called the company Hebbia (10:30), why he started the company (19:30), being a “disappointment” to his athlete parents (29:30), working at NASA as a teen (33:00), meeting Peter Thiel (36:45), and how AI is going to revolutionize the economy (42:00).
  • Artificial intelligence “agents” will create more economic value than humans within ten years. Sound outlandish? That is the prediction of this week’s guest, George Sivulka, founder of AI startup Hebbia, who comes on to talk about building AI that actually works for business (3:20), AI orchestra conductors (9:15), coral reefs and why he called the company Hebbia (10:30), why he started the company (19:30), being a “disappointment” to his athlete parents (29:30), working at NASA as a teen (33:00), meeting Peter Thiel (36:45), and how AI is going to revolutionize the economy (42:00). Hosted on Acast. See acast.com/privacy for more information.Apple Podcasts
    Hebbia's George Sivulka: "Bots-The Times Tech Podcast - Apple Podcasts
  • The Innovation Civilization PodcastYouTube
    This AI AGENT App Will TRANSFORM All White Collar JOBS - George SivulkaGeorge Sivulka, founder and CEO of Hebbia, discusses why he believes Hebbia is a generationally important company, explaining that while AI is a revolutionary technology, the key is building the right product to capture it—comparing it to how fire needed the torch and the wheel needed the chariot. He describes Hebbia's approach to orchestrating AI agents for complex tasks and details their hallucination prevention method: finding citations first before generating text rather than generating text and then searching for citations afterward.
  • The MAD Podcast with Matt TurckYouTube
    AI and the Future of Knowledge Work with Hebbia’s CEO George SivulkaGeorge Sivulka discusses Hebbia's founding journey, starting from his observations of brilliant Stanford PhD students becoming unhappy knowledge workers at finance firms like Morgan Stanley and Goldman Sachs. He describes creating an early neural information retrieval model in a Jupyter notebook to help analyze long financial documents like DEF 14A filings, which gained organic adoption among former students. Sivulka explains that Hebbia evolved from a search-focused tool into a general-purpose large language model productivity platform that allows non-technical users to wrangle LLMs and documents to generate various outputs including ETL pipelines, chatbots, and data extraction, positioning it as more general-purpose than competitor products like Glean.
  • George Sivulka discusses Hebbia's AI agent platform built for finance professionals, lawyers, and knowledge workers. He explains the company's evolution from focusing on Wall Street to building a horizontal, general-purpose workplace AI platform, and describes the founding story triggered by GPT-3's release in 2020, which led him to abandon his Stanford PhD research on metalearning.YouTube
    AI That Ends Busy Work — Hebbia CEO on "Agent Employees"
  • GREY Journal Daily News PodcastApple Podcasts
    Why Investors Are Flooding Hebbia with Nearly $100M
  • Spotify
    Saving Smart People from Mundane Tasks - A Conversation with George ...

In the news

This page shows public professional information only, each fact cited. Is this you? send a correction, or ask for removal within 24 hours, no questions asked.