Overview
Renen Hallak is founder and CEO of VAST Data[1][3]. Prior to founding VAST Data, Hallak led research and development at XtremIO as VP R&D, overseeing a team of over 200 engineers and directing the architecture and development of an all-flash array from inception through over a billion dollars in revenue[4]. Earlier in their career, Hallak developed a content distribution system at Intercast from inception to initial deployment while serving as Chief Architect, and was a member of the CTO team at Time to Know[4]. Hallak holds a BA degree[4].
Profile introduction
Prior to founding VAST, Renen led the architecture and development of an all flash array at XtremIO, from inception to over a billion dollars in revenue while acting as VP R&D and leading a team of over 200 engineers. XtremIO is the world's leading all flash array, with over 40% market share and unprecedented growth from zero to three billion dollars of revenue in under three years. Earlier Renen developed a content distribution system at Intercast, from inception to initial deployment and acted as Chief Architect. Renen was also a member of the CTO team at Time to Know. He holds a BA and an…
Career history
Insights & ideas
The through-line
Hallak's consistent argument is that the trade-offs everyone in infrastructure treats as laws of nature are actually artefacts of the hardware and software available when the systems were designed, and that a team willing to design against future components can eliminate them rather than manage them. The origin was empirical: after leaving XtremIO in 2015 he spent his time asking customers what they were missing "not specifically in storage but in infrastructure in general" [1], and heard one answer repeatedly. All-flash arrays were denser, easier to manage and more resilient because there were no moving parts, but they were "so expensive that they could only put a very small subset of their applications of their data on them" [1], leaving many tiers of legacy hard-drive storage underneath. The forward-looking customers had a sharper version of the problem: analytics, AI and machine learning "required fast access to the entirety of the data set and so the tiering didn't work anymore" [1].
From there the position extends outward. If storage stops forcing choices, customers stop thinking about storage, and then they start doing things they could not do before, which in his telling points toward a data centre that no longer serves data to humans but understands it [1]. The same instinct to eliminate rather than solve runs through his views on failure handling, scale, protocols and where computation should sit.
On breaking the trade-offs rather than balancing them
The founding brief given to his first team of architects was deliberately absurd on its face: build a system "that is faster than the fastest that exists today cheaper than the cheapest more scalable than the most scalable more resilient than the most resilient" [1]. The finding was that "those trade-offs were there for a reason that they couldn't be broken with existing hardware and software underlying technologies" [1], which turned the problem into a hardware timing question rather than a cleverness question. He locates the cost-performance conflict specifically in shared-nothing design, "where every node came with a certain amount of performance and a certain amount of capacity and that capacity cost more if you wanted it to perform better"; disaggregating the two allowed the use of the lowest-cost flash available and broke the link [1]. Scale and resilience come apart the same way: when nodes hold no state and never need to talk to each other, "you can grow up to 10 000 notes to 100 000 notes and it's still okay", half of them can fail simultaneously "and the other half doesn't even know about it", so "scale and resiliency are no longer at odds with each other they come together" [1].
He frames the customer payoff in two stages. Immediately, "they don't need to think about storage anymore it just works" [1]. Then, usually a few months in, they realise they can analyse the entirety of a data set rather than a subset, in real time rather than as an overnight batch, and that a hundred teams, quant teams in finance or researchers in life sciences, can work on the same data without stepping on each other because of workload isolation and because each team can get its own nodes [1].
On designing against roadmaps you cannot yet touch
VAST's architecture was committed to hardware that did not exist. Hallak went to vendors including Mellanox and Intel and asked "what is on your roadmap what can you tell us about the future that may enable us to build a new architecture" [1]. The answers put the company on NVMe over fabrics two years before it was ready, QLC flash almost three years early and 3D XPoint two years before he could get his hands on it, alongside Docker containers [1]. The consequence was building to specification sheets: "we architected around them based on spec sheets even though we weren't actually able to test with them" [1]. The bet is presented as the only route to breaking trade-offs, since existing components made those trade-offs unavoidable.
On eliminating the failure path instead of coding it
The most personal design conviction comes from a scar he owns directly. At XtremIO, "it took us about six months to build the good path what we call a happy path and then it took about three and a half years to build HA" [1]. His diagnosis is that failure handling is where the danger lives, because "whenever something bad happens when a hardware piece fails or anything that isn't supposed to happens that's when a lot of code needs to run and that code is by its nature less tested and has more bugs and there are more corner cases" [1]. So VAST decided from day one "not to have to do anything when a bad thing happened": everything is always in its right place, with no journals to replay and no battery backup units [1]. That made the happy path slower to complete and the failure scenarios almost free once it worked. Intel's 3D XPoint is the enabler, letting the system land everything on persistent media with no RAM caching, so all state sits below stateless nodes that can fail together without data loss [1]. He is explicit that the team arrived from XtremIO, Kaminario, Isilon, NetApp and XIV carrying both scars and things worth keeping, and that the new architecture let them "basically eliminate the problem rather than try and solve it" [1].
On the pendulum, containers and disaggregation
Hallak reads the data centre through pendulum swings, using consumer devices as the analogy: a decade of best-in-class standalone cameras and phones, then convergence into a smartphone where "the camera isn't as good as that standalone camera" [1]. Storage, security, networking, compute and databases were separate two decades ago, then hyperconverged infrastructure "took it too far because those individual pieces aren't as good as they used to be although it's a lot simpler to manage" [1]. His prediction for the next decade is both worlds at once, with containers as the mechanism: software-defined networking in one container, storage in another, a database in a third, best of breed per service without requiring a single vendor [1].
The obstacle he names is that "storage is very stateful and containers clash with that philosophy", since a container that can only run on the specific node holding the data cannot be spun up and down freely [1]. NVMe over fabrics and disaggregation dissolve that, because every container attached to an Ethernet or InfiniBand network can reach all the information. This is where he borrows explicitly from the hyperscalers: "storage is no longer pets each node doesn't need to be cared for in the same way", since no node has affinity to a specific set of drives [1].
On writing your own protocols and opening the internals
VAST wrote NFS, S3 and SMB itself rather than taking open source, for two reasons: native performance, and full advantage of the new architecture, so that failover problems that plagued SMB "don't exist in our architecture because everyone can see the same shared state" [1]. All protocols reach the same user data and metadata, which is what makes exposing the internal APIs worthwhile. User metadata can be searched and indexed and correlated with files, and he argues the storage system is the right place for it: in life sciences, files tagged with the sequencer used, the organism sequenced and its symptoms or attributes can be cross-correlated, and "there's no better place to do that than within the storage system because we have direct access to that metadata sitting on crosspoint it's the most efficient way to do it" [1]. The same logic applies to the data reduction engine, which "finds similarity across the entire data set" and reduces data that was previously unreduceable; that emergent understanding of the data can itself be surfaced to applications as higher-level semantics [1]. He describes working with select customers to define what they want their next-generation storage system to do, on the assumption that others will want the same [1].
On the data centre as a computer that understands data
His view of what comes next is that the future data centre "will not be running VMs and Oracle databases it will be understanding of data" [1]. He calls the AI wave a turning point, with applications that stop serving data to human beings and instead generate insight from it, which will make it "clearer and clearer that existing infrastructure was not built with these applications in mind" [1]. The shift he expects is toward the data centre as a supercomputer, and he notes that supercomputing ideas are being tried in this new world with mixed results, "some of them are a good match some of them are a very bad match" [1]. On hardware, new SSD form factors such as rulers will improve density and serviceability, but the three-to-four-year item that excites him most is AI chips, GPUs or IPUs, moving into the storage system close to the data, so that storage systems "will not just be able to serve read and write requests but actually understand the data and give you search like abilities and insight generation abilities from within the storage system" [1].
Takeaways
- Trade-offs in infrastructure are usually hardware-era artefacts; VAST's founding brief was to be faster, cheaper, more scalable and more resilient than the best existing systems at once, which required components that did not yet exist [1].
- Design to vendor roadmaps: VAST architected on spec sheets for NVMe over fabrics, QLC flash and 3D XPoint two to three years before it could test with any of them [1].
- Failure-path code is the least tested and buggiest code in a storage system, so the better strategy is to need no recovery action at all, with no journals and no battery backup units, even if the happy path takes longer to build [1].
- Stateless nodes with all state on persistent media let scale and resilience reinforce each other: tens of thousands of nodes, and half can fail without the other half noticing [1].
- Disaggregating capacity from performance breaks the price-performance link inherent to shared-nothing designs and allows use of the lowest-cost flash [1].
- Hyperconvergence went too far; the next decade brings best-of-breed services from different vendors in containers, made possible for storage by NVMe over fabrics removing node-to-data affinity [1].
- Metadata search, cross-correlation and similarity-based data reduction belong inside the storage system, because that is where direct access to metadata makes them most efficient [1].
- Expect AI chips, GPUs or IPUs, to be embedded in storage systems within three to four years so they generate insight rather than only serve reads and writes [1].
Media & appearances
- AHEADYouTubeLooking AHEAD Podcast: Episode 4 with VAST Data CEO Renen HallakRenen Hallak, CEO and founder of VAST Data, discusses his background in computer science and previous role at XtremIO before starting VAST in 2015. He explains how customer feedback revealed the pain point of expensive all-flash storage systems that forced companies to maintain multiple tiers of legacy hard drive storage, and how emerging applications in analytics, AI, and machine learning required fast access to complete datasets, making tiered storage obsolete. Hallak details how VAST architected a storage system based on future hardware roadmaps from vendors like Mellanox and Intel, including NVMe over fabrics, QLC flash, and 3D cross-point technology, designed to break traditional trade-offs between speed, cost, scalability, and resilience.
This page shows public professional information only, each fact cited. Is this you? send a correction, or ask for removal within 24 hours, no questions asked.