P

Tarun Amasa

Founder + CEO at Endex

Overview

Tarun Amasa is the Founder and CEO at Endex [1]. Amasa maintains a presence on X (formerly Twitter) [2].

Career history

  1. Founder + CEOMay 2023 to PresentEndex.ai
  2. Thiel Fellow2024 to PresentThiel Capital
  3. Apple2022 to 2022
  4. Tesla2022 to 2022
  5. Founder + CEOEndex

Education

  1. Computer Science, Mathematics, StatisticsDuke University

Insights & ideas

The through-line

Amasa's whole company sits on one bet: the tools finance runs on have stood still while everything around them moved. "Excel has looked the same way for the last 30 years, right? Very few professions are like that and this tool is the backbone for so many parts of the economy" [2]. The name follows from the thesis, "that in the future we'll be able to end traditional Excel workflows" [1], and the ambition is to bring the ergonomics engineers have enjoyed for two decades to people doing financial work [2]. What keeps this from being hype is a second, equally insistent commitment: honesty about where the models actually are. "We don't think you can build an a tab analyst model today" [1], and the company's "DNA has always been to underpromise and overd deliver" [2].

Over time the framing has moved from model capability to what surrounds the model. Early on the question was whether reasoning would arrive at all; the conviction to start came from betting "really heavily on reasoning" and seeing where the models were headed [1]. Now the question is harness engineering, the layer that lets a general model do specific work: "our goal is to build the right harness, which is almost like an iron man suit for the LLM model to do work in Excel and PowerPoint directly for finance professionals" [2]. Alongside that, his sense of where value sits has shifted from generating artefacts to checking them.

On what the models can and cannot do today

Amasa is precise about the ceiling. "I don't think these models can pull 12 years historical financials with quarterly data resegmentations and changes in tax" [1]. Asked why not, his answer is all of the above: the reinforcement learning environments for these tasks "are just not as built up," and there is too much inconsistent data to hold at once, which is "like trying to have some equity research analysts remember this all in their head and then write it out" [1]. He expects improvement from models getting better at using a scratch pad or writing down notes, and from sub-agent delegation, orchestrating many agents at scale rather than asking one to carry everything [1].

The payoff of that honesty is commercial. Knowing "exactly where the models are and where they're going to be in the next 8 months allows us to build a product interface that is both trustworthy and useful for users today" [1], and he reports no clients churning from a pilot [1]. He also frames the current moment as an awkward one: agents "are not good enough to be a full replacement for a human, but are almost there. And those slight cracks are what really like causes petribation" [2].

On verification being worth more than generation

This is his strongest and most repeated claim. "The verifiability and being able to catch mistakes is often more valuable than just the generation of materials, right? Where a single mistake is incredibly costly and giving the driver seat to an AI agent to do the entire model maybe is a higher trust barrier than using it to doublech checkck your results" [2]. He reads the software analogy the same way: for engineers "the bottleneck is not generating more code to get created. It is reviewing that code and making sure that there are mistakes" [2]. Being the second set of eyes on a model, catching the one cell that is off, is where he has seen the biggest lever in rollouts [2].

Building that verification layer for finance is harder than for code, where a working app is its own proof [1]. His tactical answer combines several checks: working closely with financial data providers such as Visible Alpha, CB and FACED so that data points are in line, models that "cite every number from a filing or earnings call so that every number is grounded," native Excel checks like a balance check, and interfaces that make it easy for a human to review assumptions [1]. Even so, the models "sometimes, especially with non-GAAP numbers get really confused" [1].

On accuracy versus creativity

The philosophical version of verification is knowing when you want a citation and when you want judgment. Amasa poses it directly: for an extra quarter of earnings, "do you want exactly what the CEO is citing in guidance or do you want some level of creativity where you want the element to start to judge like based on peers how they're performing based on macroeconomic trends?" [1]. His view is that with more post-training on reinforcement learning, models "will probably get to higher levels of intelligence than an equity research report" given the full data set, including expert calls, company websites and quality checks [1]. He expects the industry to pass through two paradigms: first a phase where trust demands every number be cited, then an inflection after which people want more creativity in the answers, a tension he says has already started hitting the street [1].

The mechanism for resolving this, in his account, is alignment through preference data rather than one enforced default, "because everyone has different preferences and trying to demand that this is the way you expect it to work leads to kind of user dissatisfaction" [1]. He expects models to become stateful and to remember which consensus source or brokers a user wants, and for which line items they trust management versus doing their own back of the envelope math, moving "from assistant level intern level that has no awareness of my preferences to something that will remember my idiosyncrasies" [1].

On context, retrieval and why tools miss things that are plainly disclosed

Amasa treats retrieval as the older, unglamorous discipline that still decides outcomes. Endex spent two years on financial data retrieval, "how do you kind of do information retrieval at scale across complex financial data for non-sandardized metrics," and that research powers what he claims is state-of-the-art accuracy for financial services tasks [1]. Failures come from several compounding places. First, modality: investor presentations bury numbers in pie charts and tables, so getting that information into a model is itself a problem, and Endex has evaluated document parsing from large companies such as Amazon and Microsoft down to small startups and the models themselves [1]. Second, context utilisation: "traditional retrieval where you kind of chunk these documents up and put them in vector space is often insufficient for financial tasks," so the goal is to get "somewhat near an infinite context window even if more expensive" [1].

He is sceptical of advertised numbers and of the benchmarks that certify them. "None of these models really have 10 million context windows to be honest" [1]. Needle in a haystack, the test labs have hill-climbed on, "turns out that's not analogous to real world tasks at all" [1]. Domain specificity has improved the whole training loop, and he draws the parallel between tracing a linting error through a stack trace and finding one inconsistent number that corrupts a terminal value calculation [1]. Where models get overwhelmed by many data points, the term is context rot, and the discipline has moved "past prompt engineering to now what is called context engineering" [1]. The consumer economics also explain the frustration: those tools are optimising document processing cost for an average user, while "an end user in a very domain specific task who might be willing to pay a lot more for higher accuracy doesn't really get that trade-off yet," a gap he expects to close and treats as the company's entire thesis [1].

On depth over breadth, and harness engineering

Amasa's competitive read is that generalist workflow products are dead on arrival. "Tools that try to go like Caroline mentions a mile wide and an inch deep are probably just going to get replaced by a chatbt or clot" [2]. His alternative is narrowness: "to be heavily focused on a very narrow set of very valuable work" [2], with a team split between people from private equity and people from engineering, so that the product knows "how do you do an ARR cube correctly" and what a cohort retention looks like the way a private equity firm looks at it [2]. Financial services is harder to generalise than code because it is heterogeneous: "a real estate model looks very different than a canal model," which makes any single generalised reward hard to define [1]. So the sequencing is deliberate, focus on a subset of users, expand to other verticals only after delivering value there, on the view that "the pie is really big" and the goal is "making something that we find a small set of folks find really really useful" [1].

The concept holding this together is harness engineering, with Claude code as one of the first breakthrough harnesses, followed by co-work and other coding agent tools [2]. In practice the harness has to abstract away the specific rewards that matter in finance: reconciling a number correctly, rolling a model forward, handling a resegmentation [1].

On the model layer: multi-model orchestration and RLVR

He does not expect one model to win. "We're going to go towards a multimodel world," with Anthropic, OpenAI and Google models each strong on different use cases, and the work is "orchestrating all of those for very specific tasks across the financial services space" [1]. He narrates the paradigm shifts that got here: the GPT-3 moment of generalisation at scale, the RLHF turn that made models conversational, the shoddy early days of tool calling structured through JSON or XML tokens before the labs shipped proper APIs, and then reasoning, where models "start to think before they answer" and long-running tasks became tractable [1]. The current frontier is "RLVR or reinforcement learning with verifiable feedback," today confined to narrow domains such as small coding sandboxes, which is why models are so good at deep research and code generation and not much else [1]. His expectation is that those environments generalise into simulations of Excel, PowerPoint, coding and clicking around a CRM [1].

Labs getting better is good news, not a threat. Endex works with OpenAI and other major labs on data initiatives, and "we want these labs to get really good at these tasks because it helps us downstream" [1]. The precedent he points to is Claude 3.5 Sonnet as the inflection point for coding, the moment output went from finishing one line or one file to something where "you could truly be like holy I see the trend" [1].

On generation versus distillation, and who the incentives serve

Amasa is pointed about lab economics. "The way that these AI labs make most of their money is on token generation," which he calls big token, so they are incentivised for users "to spend as much as possible on the generation of outcomes" [2]. Agents amplify this: instead of one turn, fifty turns, minutes at a time, and Endex has seen its own agents run for hours without interruption, all of which costs more [2]. His counter-position is distillation, getting to "the key core instinct of what is the underlying set of principles" rather than maximising produced material [2]. This connects back to his scepticism about summarisation as a product: the point of a model or a memo is that a human understands the company well enough to argue about it.

On why the tools were the problem in the first place

The origin of the conviction is auditability. Working on state budgets and nonprofit funding, he found files dating to 2004 that looked much like current ones, with almost nobody ever opening or auditing them, and organisations reporting strange COGS and SG&A line items while lacking offices at all [2]. He went to see for himself, and the conclusion he drew was structural: "Excel and these tools just don't have the ability to allow you to audit the same way that you have for a codebase" [2]. Colorado's Taber law, meant to give taxpayers observability, had "kind of the reverse effect" through arduous approval processes [2]. He wrote a bill, it failed, and he concluded that entrepreneurship was the better expression of the same aim [2]. A second early influence was the Pygmalion effect and what expectations do to outcomes, which became the founding question of "could you bring together a group of exceptional people to create exceptional outcomes" [2].

On what changes in banking and private equity work

The near-term use cases he and users describe are unglamorous and concrete: acting as the doer of manual retrofitting work, bringing in source tabs and updating a firm's own LBO template, or acting as reviewer, confirming that everything was linked up correctly [2]. He also sees agents automating the information gathering that today runs as multi-day back-and-forth between bankers and management teams [2]. The broader shift he reports being asked about by some of the world's largest consulting firms is what a world after PowerPoint and Excel looks like, a conversation he says is only possible now because of how good the technology has become [2]. His underlying goal is framed as communication: "how do we make communication between humans and humans better but also humans and agents" [2].

Takeaways

  • Catching mistakes beats making artefacts: "the verifiability and being able to catch mistakes is often more valuable than just the generation of materials" [2], because a single error is costly and handing an agent the driver's seat is a higher trust barrier than using it to double check.
  • Do not promise an AI analyst. Models still cannot "pull 12 years historical financials with quarterly data resegmentations and changes in tax," and building around that reality is what makes a product trustworthy today [1].
  • Generalist workflow tools lose. Anything "a mile wide and an inch deep" gets replaced by ChatGPT or Claude; the defensible position is narrow, high-value work done exactly the way the vertical does it [2].
  • Verification in finance is assembled, not automatic: cross-checking against providers like Visible Alpha, citing every number to a filing or earnings call, and native checks such as balance checks [1].
  • The unresolved design question is when to cite and when to judge; he expects an inflection after which users want more creativity, and expects reinforcement learning to push models past the quality of an equity research report [1].
  • Benchmarks mislead. Needle in a haystack is "not analogous to real world tasks at all," and no model truly has a 10 million token context window [1].
  • Lab incentives favour token generation, or "big token," while the useful work is distillation down to the core principles [2].
  • Excel's real failure is auditability: it "just doesn't have the ability to allow you to audit the same way that you have for a codebase" [2].

Media & appearances

  • Tarun Amasa, CEO and founder of Endex, discusses how AI and modern software engineering tools can transform financial modeling and Excel workflows in private equity and banking. He explains Endex's mission to replace outdated spreadsheet practices with better tools, drawing on his early experience observing inefficiencies in government and nonprofit budget management, and describes how AI agents can automate information gathering between bankers and management teams rather than waiting days for manual back-and-forth.YouTube
    Financial Modeling Will Never be the Same (w/ Tarun Amasa of ...
  • Tarun Amasa, founder of Endex, discusses the company's origins and vision to replace traditional Excel workflows with AI agents. He explains the technical challenges of adapting foundation models to the niche task of Excel-based financial modeling, covering paradigm shifts in LLMs including tool calling and reasoning capabilities. Amasa outlines Endex's approach of orchestrating multiple models for specific financial services tasks and describes reinforcement learning with verifiable feedback as the frontier for generalizing these AI systems across applications like Excel, PowerPoint, and CRM platforms.YouTube
    Can AI Replace a Hedge Fund Analyst's Excel Model ... - YouTube
  • AI-powered analysis of Financial Modeling Will Never be the Same (w/ Tarun Amasa of Endex) from Private Equity Funcast. 4 articles covering key topics.
    Financial Modeling Will Never be the Same (w/ Tarun Amasa of ...
  • tapesearch.com
    Financial Modeling Will Never be the Same (w/ Tarun Amasa of ...
  • muckrack.com
    Private Equity FunCast (Podcast) - Muck Rack

In the news

This page shows public professional information only, each fact cited. Is this you? send a correction, or ask for removal within 24 hours, no questions asked.