
Summary In this episode Ragnor Comerford talks about OmniGraph, a lakehouse-native graph storage layer designed around the needs of agentic systems. He explores how graphs are primarily a semantic model for representing the world, rather than just a specialized engine for traversal workloads, and how that perspective shaped OmniGraph’s design on top of object storage, Lance, Arrow, and DataFusion. Ragnor explained the motivation for combining graph semantics with Git-style branching and merging so that teams can manage probabilistic writers such as AI agents with stronger governance, shared context, and safer collaboration patterns. He also dug into the practical tra...
Transcript
Automatically generated from the audio. May contain errors.
Hello, and welcome to the Data Engineering Podcast, the show about modern data management. Your host is Tobias Macy, and today I'm interviewing Ragnor Comerford about OmniGraph, a Lakehouse native graph storage layer with Git semantics. So Ragnor, can you start by introducing yourself? Yeah, sure. Thanks, Tobias. Glad to be here. Yeah, so I'm Ragnor,
I'm Dutch, but grew up mostly in France and Germany and spent, I guess, most of my career, especially started my career really in kind of machine learning, computer science, because I guess the common thread was I was always obsessed with the idea of knowledge or information or information theory as a whole. And then kind of after sort of really working quite a bit in
machine learning, I kind of dove quite a bit into the life sciences or biotech. So I spent quite a bit of time working in San Francisco at Edison's company called QBio that built like a medical digital twin and built a lot of their internal data infrastructure and graphs. And it also spent quite a bit of time in protein design, essentially using large language models, trained on amino
acids to design sort of new enzymes and proteins. So yeah, kind of the intersection of really computer science, ML and life sciences. And do you remember how you first got started working in the data space? Yeah, I mean, as I mentioned,
so quite early on, I was really kind of interested in sort of information theory and generally the power
of like information knowledge. So to me, you know, every kind of major outcome with, you know,
both in research or in anything from like life science research to like hardware design
was kind of downstream of how to use and operationalize information in a better way.
And to me, kind of, that was also fundamentally an infrastructure problem. And so, yeah,
quite early on when I was working in the life science, I was also, you know,
realized essentially how fragmented all the sort of natural information was. When if you went into a lab, you had sort of a bunch of data and data set like on local computers. Then you have essentially these huge biological databases like Uniprod and PubChem and things
like that, that contain loads of amounts of biological data. And yeah, we kind of across this kind of fragmentation information and the kind of the infrastructure. And so I kind of really saw that as kind of the fundamental upstream problem and got really
interested in sort of building better data infra and systems to really make use of all this information. And that brings us now to OmniGraph. So I'm wondering if you can just start by giving a bit of an overview about what it is and how did it get started? Yes, as you mentioned, so OmniGraph is a Lakehouse native graph engine built on top of like the open data format lens.
And so there are kind of two lenses to look at it. And then one is actually working sort of backwards from essentially the new operator we were kind of optimizing for, and that is really the agent kind of workload. And so like obviously the observation we're making now essentially is that going forward,
most of the knowledge work is going to be performed by agents or most particularly also agents working in the background, not just kind of chat assistants. And then working backwards from essentially what the constraints that imposes on that kind of system.
And I think fundamentally kind of split into kind of three different components. One is like context. The second one is coordination. The third one is really governance.
So context, actually, when an agent is supposed to perform a piece of knowledge work, how to make sure that the agent actually has access to the right information or the right knowledge. And then also for multiple agents, it means do they actually have access to a shared world mode, shared understanding, essentially, of the world.
And the second one is coordination. Okay, once you go from one, two, to actually multiple agents running, how do you make sure that they all perform essentially the correct global behavior rather than at the individual agent level?
And then the third one is like, okay, when we have agents automating large parts of the knowledge work, how to make sure that humans are still there where judgment matters and where they can provide essentially a signal to where things are going right or wrong. And this is kind of where also this kind of Git start semantics that I missed earlier and the description of OmniGraph that became really important is like, okay, can you apply this sort of governance aspect of branching and merging also to the graph layer or to the structured data layer basically. And so, yeah, that was kind of really the story behind it. How do you build like the optimal substrate for a multi-agent system?
But at the same time, I think we talked about this earlier, like also looking at the existing ecosystem, obviously a lot of the technology or existing database that were designed for the human operator, but also because of kind of historical path dependence of how we kind of build things.
And so like the other sort of more bottom-up reasoning or story behind OmniGrav is actually how would you build, you know, a database essentially from first principles. is also assembling essentially all the building blocks with amazing technology that has evolved
fast forward to 2026. Examples of this are building a database on top of object storage, decoupling storage
and compute, these open data formats and then obviously the amazing query engines like DuckDB or Data Fusion and things like that. So really sort of
the top-down and bottom-up reasoning behind OmniGraph. And so you mentioned that one of the core challenges that you're focused on is this idea of multi-agent coordination,
which is definitely the major concern of the day as the number and capability of agents continues to expand and the patterns are definitely still very much in flux. I'm actually just starting to read the agentic mesh book from O'Reilly, so good timing here. And graphs have gained a lot of
This is the opening of the episode. Open the player for the full interactive transcript with clickable words, translation and flashcards.
Open full transcriptAudio belongs to its publisher and is played from their feed. Rights holders can request removal — copyright & takedown policy