Artwork for podcast Data Driven
Unlocking Enterprise Knowledge – AI, Document Comprehension, and the Future of RAG
Episode 820th July 2026 • Data Driven • Data Driven
00:00:00 00:49:49

Share Episode

Shownotes

In this episode, hosts Candace Gillhoolley and Frank La Vigne are joined by Neil Katz, Chief Product Officer at Valantor AI—a four-time Emmy winner whose unconventional journey spans from technology startups to award-winning journalism, and now to the forefront of enterprise AI innovation.

Neil shares his unique perspective on the evolution of AI, from the early days of digital design and machine learning, to building large-scale AI and document intelligence platforms for major organizations. The conversation explores the critical challenges of knowledge extraction, document comprehension, and securing sensitive data in today's era of sovereign AI. Together, they uncover the hidden complexities behind Retrieval-Augmented Generation (RAG), discuss the importance of hybrid search strategies, and reflect on where the field is heading as enterprise needs push the boundaries of what AI can do.

Whether you're a data professional, an AI enthusiast, or just curious about how language models are transforming how we understand information, you won’t want to miss this candid and thought-provoking deep dive into the future of AI and data integrity.

Links

Time Stamps

00:00 Early career in tech and journalism

04:01 Early consumer AI experiences

06:49 Early AI and machine learning developments

11:43 Anthropic's new findings on AI models

16:15 Data sovereignty in AI systems

19:27 Implementing open source AI models

22:49 Breaking down documents for models

25:45 Understanding the RAG system process

28:47 Challenges in AI data processing

31:22 Challenges in RAG with Insurance Data

36:35 Understanding and managing data security

39:35 Early days with OpenAI GPT

40:48 Explaining vector and similarity search

46:57 The evolution of computing models

47:29 Computing evolution to cloud and edge

Transcripts

Speaker:

Hello and welcome back to Data Driven, the podcast where we explore the emerging field

Speaker:

of data science, artificial intelligence, and of course, all the hype

Speaker:

around AI. All of it is impossible without data

Speaker:

engineering. However, my favoritest data engineer in the world can't make it

Speaker:

today, but I did bring the most curious person I know,

Speaker:

and that sounds really bad, quantum curious, but she's also data curious

Speaker:

too, Candice Cooley. How's it going, Candice? It's great. I'm very

Speaker:

excited about today. Our guest has a really exciting

Speaker:

background. Yeah, just looking at his, his

Speaker:

LinkedIn profile makes me ask a lot of questions. Our guest today is

Speaker:

Neil Katz, who is Chief Product Officer at Valantor AI

Speaker:

based in New York. In the virtual green room, we geeked out on some New

Speaker:

York stuff, but he's also a 4-time Emmy winner

Speaker:

and he knows that it's a non sequitur. So welcome to the show, Neil.

Speaker:

Thank you very much, guys. Great to be here. Much appreciated. Good to have you.

Speaker:

I have to ask first, the Emmys. Did you work in media? Did you work—

Speaker:

what, how did you get an Emmy? I, yeah, I've, I probably have— 4

Speaker:

times. 4 times, true. I probably have one of the, one of the stranger

Speaker:

backgrounds to be in leadership at an AI company these days, but I actually

Speaker:

started my career in technology. When I came outta college, I built one of the

Speaker:

first digital design companies outta New York City. This is Web

Speaker:

1.0, so kind of dating myself here, but After a couple years of doing

Speaker:

that, we sold that company and I had a soul-searching moment. I just realized I

Speaker:

had spent my early 20s just spending all my time in

Speaker:

dark rooms with computers till 4 in the morning. And I didn't

Speaker:

want to do that anymore, or at least not for my whole life. So I

Speaker:

retooled, became a journalist, and spent 20 years in

Speaker:

journalism, having the great fortune to be able to report in places

Speaker:

like Iran, Vietnam, India for quite a time,

Speaker:

southern Mexico, all over the United States. for places like the New York

Speaker:

Times and CBS News and, and then Weather Channel, where I was running the

Speaker:

digital news division at the Weather Channel. We actually built a digital news

Speaker:

operation from the ground up, which was really a cool opportunity. And

Speaker:

strangely enough, even though I mentioned these nice brands like the New York Times and

Speaker:

CBS News, the 4 Emmys actually come from our time at the Weather

Speaker:

Channel, where we built a documentary

Speaker:

unit and went all around the world, really. Wow.

Speaker:

Reporting on the relationship between climate change and extreme weather

Speaker:

and social issues in a way that you wouldn't expect.

Speaker:

So we reported everywhere from Iraq to

Speaker:

Ethiopia to Sudan to Central and South America,

Speaker:

and of course all across America. And then we're lucky enough to win

Speaker:

some Emmys for documentaries we did. The first one we did

Speaker:

was actually— that we won for— was about the southern border of the United States.

Speaker:

and how many immigrants were dying actually at that

Speaker:

border, actually inside Texas when they crossed over. Things had become

Speaker:

hotter, things had become drier, and the passage had become more dangerous.

Speaker:

It's not how it works today because right now everyone goes to the southern border

Speaker:

and basically says, hi, I'm here and I want asylum. But back then you snuck

Speaker:

in, and in Texas it had become very dangerous, and they would

Speaker:

find a couple bodies, corpses a week, out in kind of the badlands of

Speaker:

southeast Texas. So anyway, starting off on a weird note for an AI

Speaker:

podcast, I realize, but I spent a long time in the journalistic world

Speaker:

and happy to go into that if you want. And then something interesting happened.

Speaker:

IBM came along and purchased the Weather Channel in one of probably the

Speaker:

strangest acquisitions in the era. Yep. We

Speaker:

can say a lot about that if we want to, but then I got thrusted

Speaker:

right back into technology. I'd run from it. So I'm going to do this content

Speaker:

thing and then tell people's stories and got thrust right back

Speaker:

into tech. And IBM really I think to their credit, made us much

Speaker:

more of a technology company than a media company. That was actually probably pretty good

Speaker:

for us. And then I was running big product and AI teams

Speaker:

under Watson, IBM Watson Group, and met my,

Speaker:

my current partner, Ben Fletcher, who's our CTO. And we formed

Speaker:

a company called EyeLevel, which just got acquired by Valontour. We can tell that

Speaker:

story in a minute. Ah, very cool. But even back in,

Speaker:

I actually looked up the date right before this podcast, back in 2016, so about

Speaker:

6 years before the launch of ChatGPT. We were using

Speaker:

the Watson model to build some really popular AI experiences for

Speaker:

consumers. We built something called Watson Weather, and about 2

Speaker:

million people a day were chatting, if you will, a chatbot that was on

Speaker:

Facebook Messenger, of all things, which was the popular place to chat,

Speaker:

that gab back then. And 2 million people a day were chatting with

Speaker:

our AI bot and trying to get their weather or hurricane warnings.

Speaker:

And these models were not nearly as good as the GPT-class models.

Speaker:

But what we saw was really interesting. We were, we understood what you

Speaker:

wanted maybe about 35% of the time.

Speaker:

So pretty high failure rate. I was thinking of the early Siri experiences. You're like,

Speaker:

if this is artificial— At the time it was magical, right? At the time

Speaker:

it was magical. 'Cause I also worked lots in the, the

Speaker:

before thing. I think you're right. I think people are gonna remember the, at least

Speaker:

the early part of the 21st century as there was before

Speaker:

ChatGPT launched and then after ChatGPT launched. You're probably right. The before time. But, but

Speaker:

a lot of people don't realize there was natural

Speaker:

language processing. It's old. I remember being a kid on the Commodore

Speaker:

64 playing Zork, right? And again, that seemed

Speaker:

magical at the time. Really couldn't chat with it per se. Yeah. But

Speaker:

just the idea that you could type something in. And before that there was ELIZA,

Speaker:

which you've probably heard of. But, you know, but

Speaker:

I remember building out chatbots. I remember using chatbots on Facebook Messenger

Speaker:

where it would basically do a lot of linguistic processing. was

Speaker:

not nearly as coherent or, or as

Speaker:

good of responses as you get from ChatGPT. But

Speaker:

yeah, no, people forget. Hopefully none of our listeners, because a lot of ours

Speaker:

are very savvy. But a lot of people, a lot of the normies out there,

Speaker:

just assume there were no chatbots prior to— There was a lot of

Speaker:

hard work that probably won't get any glory because it

Speaker:

just wasn't as cool. But we saw something I think really

Speaker:

we saw a piece of the future when we built this system, which was that

Speaker:

when we got it right, when people were able to communicate effectively with

Speaker:

this chatbot, the engagement stick looks

Speaker:

basically the same as the ChatGPT engagement stick. It just was vertical. Right. So

Speaker:

if we got you, you were hooked. You'd come every single day, you'd talk with

Speaker:

us for many minutes a day trying to get your weather. If we didn't, of

Speaker:

course you were gone because we didn't get it right. But you could see already

Speaker:

what was possible if you could make this work in a more generalized way.

Speaker:

The stickiness of chat cannot be overstated.

Speaker:

You know, yeah, the way you've been talking about it, you, you know, you were

Speaker:

involved before ChatGPT. So my question is, how

Speaker:

has your definition of AI changed over the last 5

Speaker:

years? Oh, that's a good— that's

Speaker:

a really good way of putting it. Look, I think in the before times,

Speaker:

You would build— like, most of what was built was machine learning, if we're honest

Speaker:

with ourselves. Most of the things that you would practically create was more about

Speaker:

basically statistical algorithms to try to find efficiencies or try to find

Speaker:

things that were popular. So much of it was something that people

Speaker:

wouldn't even recognize today, actually, when they say something like AI. These were like internal

Speaker:

tools that you would use to improve the various performances in a company, maybe improve

Speaker:

your supply chain by 5%. But already you saw

Speaker:

the kernels, as I'm saying, of what you could possibly do. I think now obviously

Speaker:

we're in the era of much more generalized models, which to me feels like one

Speaker:

of the big breakthroughs of GPT that was very different from what was happening with

Speaker:

Watson. Watson in its day was pretty badass. I don't know if I can say

Speaker:

that word on your pod. I don't know what your PG rating is, but— We'll

Speaker:

figure it out. All right. You gotta bleep me, let me know. But

Speaker:

you remember this thing won Jeopardy, which was very hard to do a long time

Speaker:

ago, but things were much more specialized. You'd be building models that were

Speaker:

around one very specific business problem. You try to train it

Speaker:

on the specific business problem. And that kind of made sense,

Speaker:

especially from IBM's perspective, because they have business customers that all it's about— it's only

Speaker:

about private data. But it was very limiting. And I think that's what actually

Speaker:

the breakthrough with GPT was, like, wait a second, actually the right way to get

Speaker:

to something that's more useful even for specific tasks is to make the

Speaker:

models more generalized. And obviously now today, people, you

Speaker:

can almost ask it anything And it can do almost anything.

Speaker:

For me, the exciting part though is not even where we are at the moment.

Speaker:

Language models— this is something a lot of people in the public don't understand, your

Speaker:

audience will understand— language models are trained on language.

Speaker:

A lot of people don't understand that they're not math geniuses. We're trying to make

Speaker:

them better at math, but they're actually word geniuses. And

Speaker:

that unlocks a huge amount of human potential that computers never touched before.

Speaker:

Our knowledge really has been— it's been unable to

Speaker:

communicate with computers in that way. But even though that's been

Speaker:

really amazing, I think the next step, which we're not there yet, is getting the

Speaker:

models to do things that are not language, right? So people don't quite understand that

Speaker:

you can't build a rocket ship to go to Mars with a language model. We

Speaker:

don't— that's what the models do today. So we're still like a step away, I

Speaker:

think, from some very different innovations that probably don't look like the language models

Speaker:

today. So that was a long answer on that question, but— All

Speaker:

the math that these models are good at all go—

Speaker:

they're not good at math per se, like you said, but they are good

Speaker:

at math. But all the math goes into linguistics and

Speaker:

text, right? You can ask it— what's interesting is now if you ask it to

Speaker:

say how many Rs are in strawberry or what's 2 2, it'll

Speaker:

actually write code. It's good at writing the code that'll give you the answer,

Speaker:

right? But no, you're right. And it's an interesting—

Speaker:

it's very curious to see how

Speaker:

people misinterpret what AI is doing, right? People say, oh, it's

Speaker:

hallucinating. I was like, from a certain point of view, it's always hallucinating, right? Because

Speaker:

it's just guessing the next logical step. And I know there's

Speaker:

attention mechanisms and all that that do keep it honest. But I

Speaker:

was once at an event, and I hate to name-drop, but

Speaker:

this was Esther Dyson. Oh, yeah. Right? Yeah. Somehow I ended up in the same

Speaker:

room with her. I'll never figure out how, but, but she was, I

Speaker:

was on the same panel and she was saying like, this was 2023, so early

Speaker:

days of what we're, what we have today. Yeah. And she was saying, well, there's

Speaker:

no real difference between autocomplete on your phone or

Speaker:

guesses the next number and, and ChatGPT. And I was

Speaker:

like, far be it for me to tell her she's wrong. Sure. Certainly in a

Speaker:

public forum. But also too, she's right in a sense that the stealth

Speaker:

bomber is the same tech, technology as a paper airplane,

Speaker:

right? They're both doing a lot of the same things. Obviously one's way more

Speaker:

advanced than the other. What fascinates me— I won't explain on that one.

Speaker:

Yeah, I think both are radar invisible. But no, but I'm

Speaker:

actually— I often find myself surprised at how far we've taken

Speaker:

the transformer and LLM architecture. I didn't think— I thought

Speaker:

we would run out of steam after a couple years, right? But

Speaker:

things like things that have no business working, fine-tuning,

Speaker:

That was— for fine-tuning, that's harsh. But for things like

Speaker:

distillation, it shouldn't work as well as it does.

Speaker:

Reasoning on these things also works way better than I would have

Speaker:

bet on. I'm sorry, I cut you off. I'm agreeing with you. I think reasoning

Speaker:

and logic is something you would think maybe would be beyond guessing the next word

Speaker:

or the next token. But you can see

Speaker:

emerging properties in language models that seem to go far beyond what you think guessing

Speaker:

the next word would offer way

Speaker:

more. And I can only imagine that's probably the attention

Speaker:

mechanism. There's some kind of wisdom in the attention mechanism. But that—

Speaker:

I think it's going to sound weird. I almost think we don't

Speaker:

necessarily know. No, that's crazy. Yeah. There's the

Speaker:

new paper that came out from Anthropic, I think last week, that was talking about

Speaker:

that. It wasn't even totally new because this idea of a latent space inside

Speaker:

models has been around for a long time. This idea that there's some kind of

Speaker:

its own thinking space where it's doing things in its own language, in its own

Speaker:

quote-unquote mind, whatever that means for a model. But Anthropic reconfirmed that

Speaker:

and is finding that their models are now— they've been able to figure out where

Speaker:

inside the model this activity is happening. That's interesting because

Speaker:

obviously no model has been built to do that specifically. And so what is so

Speaker:

surprising is just by training them on the world and then

Speaker:

refining them, they have emergent properties that aren't specifically developed.

Speaker:

That's a whole new kind of thing in the history of human invention. There's no—

Speaker:

like when someone invented the fork, it didn't have— there was no way it would

Speaker:

have some emergent properties and suddenly become a spoon and do something else. So

Speaker:

something, something special is going on as these models get bigger,

Speaker:

smarter, better trained, and more refined. Oh, that's a great way to put it. And

Speaker:

it makes you wonder too, what— maybe language is the secret

Speaker:

sauce to— I don't want to say consciousness, but intelligence. Maybe language

Speaker:

is in itself the magic. Emerg— no,

Speaker:

the notion of emergence, emergent properties is not unique to AI,

Speaker:

right? You'll see birds will flock a certain way, right? And

Speaker:

all they're doing is following the bird in front of them, right? And they inadvertently

Speaker:

makes these patterns. And it seems

Speaker:

language is the gateway to some kind of emergent properties of

Speaker:

intelligence and reasoning. It could be. And I'm actually also interested in

Speaker:

languages that aren't our own. If I think about like DNA, a chemical

Speaker:

language. It's one of those areas where I think language models will eventually be pretty

Speaker:

good, as opposed to the pure physics of building a spaceship, which they

Speaker:

will be good at as well, but this is a different problem. But there's lots

Speaker:

of languages that humans are not naturally conversant in. We can't think in

Speaker:

DNA, but a language model perhaps can, and that can— You could find—

Speaker:

that's a great point, right? What sorts of things could this possibly unlock?

Speaker:

And mathematics is also, to your point about building a spaceship, mathematics is also a

Speaker:

is a language in itself. Most humans don't

Speaker:

think it, can't think of it natively, but there are some, right? It might be

Speaker:

the gateway to, to other ways of thinking about this.

Speaker:

Very much so. Yeah, it's— Candace, you kicked off an awesome question, by the way.

Speaker:

I think that you answered that question maybe 10 minutes ago and it kicked off

Speaker:

a great conversation. So thank you. No, it was great. I really, I'd like to

Speaker:

dig a little deeper into the idea of DNA as a chemical language.

Speaker:

So what do you mean by that? And how has thinking

Speaker:

about biology influenced the way you think about AI?

Speaker:

Boy, that's a good deep question. Look, in our practical work and the things we

Speaker:

do for customers every day, I wouldn't say that we think in a biological sense

Speaker:

or that DNA is part of that story, at least the work we deliver today.

Speaker:

But it excites me because when I try to think about what are the unique

Speaker:

things that— where are the places in the universe where language can create something new?

Speaker:

DNA literally is the language of life. English or

Speaker:

Spanish, human languages, are like the language of consciousness.

Speaker:

DNA is the language of life, the fundamental building blocks of creation,

Speaker:

at least on this planet. We don't know what it'll be like on some others.

Speaker:

That's powerful witchcraft. That's powerful stuff to be playing with. And so I—

Speaker:

and things that the human brain just isn't really built to speak that, to

Speaker:

think conceptually in DNA. So to me, that's going to be a very

Speaker:

powerful space when we can really get language models just to be native,

Speaker:

natively speaking the language of creation of life on Earth. That's

Speaker:

super interesting. It's nothing we do. I wish I could say I'm delivering that to

Speaker:

Air France and EDP, some of our customers today. We're not, sorry. But that's some

Speaker:

cool sci-fi where I think we're not that— we're not 20 years from that. Maybe

Speaker:

we're 10 or 5. Baby steps. Yeah. So one of the things that I

Speaker:

see on the Valintor website is

Speaker:

really piqued my interest for a number of reasons. It was, you mentioned

Speaker:

enterprise visual intelligence, which one, that's

Speaker:

very interesting. 2, for sovereign AI

Speaker:

environments. So tell me about what sovereign AI means

Speaker:

to you, because everyone has a slightly different take on it. I've written a book

Speaker:

on it. I have my own take on it. I'll send you a free copy

Speaker:

of my book and you can, you can, you can tell me if I'm full

Speaker:

of it or— Same, same. What? I'll sign it for you. Yeah, I'll

Speaker:

sign you, I'll sign you the PDF. But the, the sovereign AI

Speaker:

went from being this really weird niche topic to

Speaker:

now is the topic of the G7, right? And

Speaker:

obviously there's fallout for that. So what does sovereign AI mean

Speaker:

to you? For me, it means, I think for us as a

Speaker:

company, it means that we believe, look,

Speaker:

fundamentally, I think the world's most important

Speaker:

The world's most important knowledge, in a lot of sense, is not on the

Speaker:

public internet and never will be. So much of

Speaker:

the vital data in the world sits

Speaker:

behind firewalls, right? Sits behind corporate or government or military walls.

Speaker:

And if we want to liberate that with AI, with the

Speaker:

newest technologies, we think that's not going to come from companies or

Speaker:

organizations that have built 20, 30 years of digital security around the golden

Speaker:

egg of knowledge in their company. That's probably not going to be them opening up

Speaker:

the kimono and firing that off across the internet and sending it to ChatGPT

Speaker:

or Anthropic. There's a big— the next wave, I

Speaker:

think, is letting companies, governments,

Speaker:

countries even have sovereignty over their

Speaker:

entire AI stack, from the data that they've been protecting for

Speaker:

decades to getting the models and the

Speaker:

harnesses and the stacks to run right next to the data. So

Speaker:

you have a boundary around your world, not like a

Speaker:

militaristic border, but like a boundary around your space so you can

Speaker:

confidently and safely build AI that sits right next to the

Speaker:

data that you've been working on protecting for the last few decades. I know in

Speaker:

the political space, as you brought that in, obviously the

Speaker:

EU is thinking about this, the Middle Eastern countries are thinking about this. We

Speaker:

often get calls from, from the Middle East on this stuff. I think

Speaker:

they think it more from like a geopolitical perspective of Can we trust

Speaker:

America? Can we— increasingly, companies, countries feel like they can't.

Speaker:

Our relationships are fraught. So if we can't trust America, maybe

Speaker:

we can't trust the products coming out of America, and we need to have control

Speaker:

from the global perspective. We're not really thinking about it from

Speaker:

geopolitics, not what we do. But when we go to a company like

Speaker:

ADP, for example, obviously they have a tremendous amount of

Speaker:

enormously privileged data from their customers, the largest payroll provider in the

Speaker:

country, hundreds of millions of people's financial information is stored in that

Speaker:

company, a lot of it actually down in mainframes in the basement, old

Speaker:

IBM machines that we used to work for, old iron.

Speaker:

That's not going out to ChatGPT. They wanna build AI around that

Speaker:

stuff too, but they're not gonna use some public endpoint to throw that at

Speaker:

Fable, whoever it is. So they need solutions to bring really powerful AI

Speaker:

that can sit next to their best data. And that's really how we think about

Speaker:

it. Because as

Speaker:

organizations— I'm sorry, go ahead. No, go ahead, Candace. I've been hogging the mic because—

Speaker:

As organizations continue to adopt AI, how do they— how do you balance

Speaker:

innovation with the need to keep sensitive data

Speaker:

secure, compliant, and under your own control?

Speaker:

I think when people— when people say— I'm trying to guess what the word innovation

Speaker:

means here for you. I think what you might mean is the latest, greatest model

Speaker:

from the biggest companies. And, but the open source universe

Speaker:

of models is not far behind. It's generally, maybe it's 6 months

Speaker:

behind. There's a lot of debate about how robust the models are. But it's

Speaker:

getting shorter, right? The time span's getting shorter. And I think that's an interesting metric,

Speaker:

right? 'Cause it used to be, ah, they're about a year behind or, or early

Speaker:

in 2022, right? Ah, no one can ever touch OpenAI and where they

Speaker:

are. Look, the, look, the, those companies may also

Speaker:

make business decisions, the OpenAIs and the, the frontier

Speaker:

model companies, as we're calling them, they may eventually move to a model that's not

Speaker:

just SaaS, right? Today they've decided to be pure cloud. That doesn't

Speaker:

mean they have to be pure cloud in the future. But from a, like, how

Speaker:

do you help a company implement today standpoint, what we're really talking about

Speaker:

is helping them run open source models

Speaker:

inside their security perimeter. And our core technology, Ground X,

Speaker:

is an enterprise-grade RAG and document understanding system.

Speaker:

And the way that we've built it allows you to bring in all the documents

Speaker:

of your company, all the knowledge of your company. And we've done

Speaker:

some really special things on the first principles so that you don't even need a

Speaker:

frontier model to get the best performance out of the information

Speaker:

inside your company. We've actually built our own, trained our

Speaker:

own vision models to take apart your documents, break

Speaker:

them down into small pieces. Here's a table, here's a diagram, here's a text

Speaker:

block. You break things down into atoms, then you could feed those pieces,

Speaker:

small pieces of data one at a time into models and get the same kind

Speaker:

of performance you would get as if you were trying to work to a frontier

Speaker:

model. So we make it possible for companies to use smaller, faster, cheaper

Speaker:

models on-prem or in managed cloud

Speaker:

so they get all the benefits of the frontier on their private data.

Speaker:

Yeah, that's how we're approaching it today. Interesting. So do you— are you talking about

Speaker:

a chunking strategy? For RAG, or are you talking about

Speaker:

fine-tuning or all the above? It starts with a chunking

Speaker:

strategy. It's funny, our product really comes out of— we're talking now in a

Speaker:

philosophical way about AI, but our product really comes from really practical problems.

Speaker:

When we were at IBM, we'd see this problem. The same thing

Speaker:

would happen over and over again on big projects where you were

Speaker:

trying to get a company's information to work with a model. At that time,

Speaker:

Watson was the model. Now it's other models. But the problem is always the same.

Speaker:

What is a company's knowledge? Essentially, it's typically stored in documents,

Speaker:

and documents are visually complex. And fundamentally, these

Speaker:

documents confuse language models because they're

Speaker:

either really long, or they're very visually confusing,

Speaker:

or they're very dense in certain parts. The frontier models are getting

Speaker:

better at understanding these things, but when we started, it was really bad.

Speaker:

So our first problem was like, Company

Speaker:

X hands you 100,000 pages of

Speaker:

stuff. They say, this is what I know. This is what my department— this is

Speaker:

what my legal department knows. This is what my HR department knows. Okay, how do

Speaker:

you make that work with AI? At the time, you really

Speaker:

couldn't. And what we did was we

Speaker:

trained a vision model. We actually trained a video model because the document

Speaker:

models weren't good enough. The Detectron

Speaker:

2 is where we eventually landed. It's our most recent one we're working on from

Speaker:

Meta. It's a really great video model. So the purpose of that video models designed

Speaker:

to follow a soccer ball around the— around a video screen or

Speaker:

security footage and things like that, but they're really good at detecting

Speaker:

objects. We trained that vision model on a million pages of

Speaker:

corporate data, all kinds of stuff— legal, medical, financial, what have you.

Speaker:

And the purpose of that was because we realized the

Speaker:

problem of a document or a million documents was too big for

Speaker:

models to deal with. And this is still true because no matter what they

Speaker:

say the context window is, it's not really true. And so you— it

Speaker:

became obvious to us that to make models work with your documents and your

Speaker:

data, you had to break them into really small pieces. And so we did that

Speaker:

with a vision model that we trained to basically identify on every

Speaker:

single page where's the table, where's the

Speaker:

text block, and where's the graphic. And then we have an

Speaker:

agentic pipeline that essentially feeds those small

Speaker:

objects into a visual language model. Again, one that can be

Speaker:

smaller, cheaper to run because we're giving it small bits to work on,

Speaker:

and have it explain in great detail what those objects are

Speaker:

into text that's friendly to language models, or text that's

Speaker:

friendly to essentially search is the other aspect of this. Because when

Speaker:

you build these systems— I'm in the weeds now, and you tell me if you

Speaker:

want me to come back out to 30,000 feet, the engineering weeds a

Speaker:

bit— these systems are really— RAG

Speaker:

systems are document understanding is the first problem and

Speaker:

search is the second problem. Yes. The document

Speaker:

understanding problem, everyone thinks it's solved. It's actually not. It's still pretty hard.

Speaker:

And if you get it wrong,

Speaker:

the downstream consequences of it are extremely severe.

Speaker:

It's not so much that you have incorrect

Speaker:

information in your database. That's a problem. You ingested 100,000 documents or a

Speaker:

million documents. And let's imagine you're using a system that's not ours and you've got

Speaker:

15, 20% of that's incorrect when it gets turned into LLM-ready

Speaker:

data. Okay? Okay. That's a problem.

Speaker:

But the real problem is it's a silent failure.

Speaker:

So you'd never know which chunks are right and which chunks

Speaker:

are wrong. So now you've got this burning problem sitting in

Speaker:

your database, in the, in your AI representation of your data, which you've now

Speaker:

vectorized, you've turned it into numbers, you've done these things with it. But underlying, it's

Speaker:

not the same as the actual information that came out of your company. And you

Speaker:

don't know where it is. Kind of smolders just below the surface,

Speaker:

and you never know when it's going to cause trouble. I'm sorry, I cut you

Speaker:

off, but— No, that's right. And it— how does that represent to folks? It represents

Speaker:

in a funny way, because at first blush, it looks like we don't have a

Speaker:

problem at all. You produce an MVP, you're like, oh, it looks pretty

Speaker:

good. Most of these answers are right. Looks pretty good. Then the SMEs get

Speaker:

in there and really play, and they're like, wait a second, this is in— this

Speaker:

is In an intermittent way, which every engineer knows is the worst problem you can

Speaker:

have. You don't want an intermittent bug. In an intermittent way

Speaker:

that's nearly impossible to figure out why or track, this thing is wrong.

Speaker:

Now everyone thinks these are model hallucinations. We have

Speaker:

found that 90 to 95% of the errors in RAG-based systems are

Speaker:

not model hallucinations. They're incorrect—

Speaker:

it's like an information gap. It's a comprehension gap.

Speaker:

It's that the things that you have bad information— it's garbage in, garbage out. The

Speaker:

information that sits in your RAG database is not an

Speaker:

accurate representation of the world, of the documents you put in the first place.

Speaker:

Right. Now someone asks a question or an agent asks a question of your RAG

Speaker:

system, you've done a search. We can get into how you would do search and

Speaker:

why that's complicated. You basically search your RAG database for the 20 to

Speaker:

100 blocks of text most likely to contain the answer to this question.

Speaker:

And then the model synthesizes those together and produces a final answer.

Speaker:

This is RAG. Most of the errors are not because the

Speaker:

model just made some stuff up, because the whole point of RAG is that you

Speaker:

can only answer based on the content I've just given you. It's typically the

Speaker:

answer— the content I've given you is wrong, and it's hard to

Speaker:

figure out why or where. So I went down a rabbit hole here, but this

Speaker:

is the core of our technology, the problem we've been trying to solve for a

Speaker:

number of years, this is the problem we started solving at IBM. And

Speaker:

frankly, IBM wouldn't let us solve it, and so we left. From

Speaker:

IBM's perspective, it was a research

Speaker:

problem that was too middle

Speaker:

space. I couldn't commercialize it 6 years ago at IBM,

Speaker:

so IBM said, I can't make money today. And it wasn't a

Speaker:

$100 billion bet like quantum computing that they hope will make them into the next,

Speaker:

you know, bigger than SpaceX. It was this weird middle research problem.

Speaker:

And so we left it high level to build that. And that's the core of

Speaker:

the Ground X tool, that and really powerful search, which I can get into if

Speaker:

you want me to get into that. But that's the ultimate beginning data layer

Speaker:

of how you build knowledge, an AI empowered knowledge system inside a company,

Speaker:

whether you're using that for customer support or agent orchestration

Speaker:

or finding fraud in insurance claims, which a lot of our customers do.

Speaker:

You got to start with that, the engineering right to turn a company's knowledge

Speaker:

So you don't have a comprehension gap into something AI can understand.

Speaker:

No. And it's funny you mentioned IBM. Sorry, Candace, I cut you off.

Speaker:

But one of— I used to work at Red Hat up until a few days

Speaker:

ago. So a few days ago, literally. And so I'm quite familiar with that.

Speaker:

So, and then one of— I'm sorry, what? We're built on actually the first

Speaker:

Kubernetes, the first on-prem implementation we did was built on OpenShift. Oh, very cool. I

Speaker:

worked on OpenShift AIs. It was But

Speaker:

what's fascinating about this is that I, my last project was

Speaker:

a RAG, kind of like how to teach the field sales, like what RAG is,

Speaker:

right? Because there was this product called InstructLab. InstructLab said, hey, just

Speaker:

fine-tune your models. You'll be fine. Don't worry about RAG.

Speaker:

Didn't really work out that way. But so it was like, no, let's reintroduce

Speaker:

RAG into the conversation. So ultimately, kind of like the lab I

Speaker:

worked on, so I, I had the opportunity to, to deep dive into the space

Speaker:

for a couple months starting in January. Yeah. So it's a spot near and dear

Speaker:

to my heart because RAG RAG solutions are really going to be

Speaker:

the shortest path to enterprise ROI. I still

Speaker:

think so, obviously partially because we sell a product that does it, but

Speaker:

we had to— we could choose to pivot into something else, but we— and every

Speaker:

few months someone comes out and says RAG is dead. Why do they say that?

Speaker:

They say that because the models are bigger and better, but we've

Speaker:

always had this point of view that if you really think macro,

Speaker:

Okay. The world's— it would be—

Speaker:

it's very difficult to move all of the world's information,

Speaker:

which is currently stored on the cheapest medium possible, hard drives and flash

Speaker:

drives, move that into GPU, the world's most expensive

Speaker:

silicon. So unless there's a huge architecture

Speaker:

change on the hardware side, and I think eventually there probably will be,

Speaker:

but for where we are today, you're not going to move everything

Speaker:

from a hard drive onto a GPU. And that means by

Speaker:

definition, you need something in the middle that can figure out what to

Speaker:

feed the model when, depending on what kind of— what an agent wants

Speaker:

or what a human wants to do next. And it's like RAG, I think, is

Speaker:

that middle layer that's gonna be with us for quite, quite some time. Yeah, I

Speaker:

think at most it won't die. It might evolve. I know there's RaFT and then

Speaker:

there's something else. PEFT, I think, is the other one. No, RaFT

Speaker:

is the other one. RAG-assisted Retrieval augmented fine-tuning, right? But

Speaker:

I think the RAG pattern, whether or not it'll still be called RAG, is— you're

Speaker:

right. I think it's fundamental. I think it's a fundamental force of the AI

Speaker:

universe personally. But— Yeah, we do too. And you

Speaker:

mentioned fine-tuning models. That was when this whole craze started. I think it would go

Speaker:

into big enterprises and all of them were like,

Speaker:

yeah, RAG sounds interesting, but also boring to me. I want to go fine— I

Speaker:

want to fine-tune my own model. I want to do it because that looks like

Speaker:

a lot of fun, frankly. Then you find out, and it's

Speaker:

cool, at the other side of the engineering project inside a company, you're like, hey,

Speaker:

we have our own model inside the company, and you're a hero. And that's awesome.

Speaker:

Resume-driven development is probably what's behind it. Maybe.

Speaker:

But when you get in there and you're like, wait a second, how do you

Speaker:

actually fine-tune a model? You have to fundamentally— you need

Speaker:

question-answer pairs. You have to distill knowledge out of, again,

Speaker:

those millions of pages of documents we started with. You still have the same problem

Speaker:

in the beginning. Now you're going to try to distill that into

Speaker:

basically question-answer pairs that you can then go fine-tune a model with.

Speaker:

And I think everyone found out pretty quickly that is dirty, messy, difficult

Speaker:

business. And companies' knowledge is really

Speaker:

chaotic. And trying to distill out the perfect

Speaker:

tens of thousands of QA pairs out of that was actually pretty difficult and

Speaker:

time-consuming and out of date. As soon as new information happens, your

Speaker:

fine-tunes are out of date. So that being said, now

Speaker:

we're doing something that's a little bit in the middle. One of the things our

Speaker:

technology is really good at is extracting information from documents. And the

Speaker:

pattern that most of our customers did in the beginning was pure RAG. Okay.

Speaker:

And GroundX is an end-to-end RAG platform that

Speaker:

does the ingest, the processing, the parsing, the storage,

Speaker:

the search, the re-ranking, all of that. And you can set it up in 3

Speaker:

lines of code. Wow. So that's cool, you know, problem solved.

Speaker:

However, increasingly we're seeing is a different kind of pattern

Speaker:

where a lot of the information, a lot of the things people

Speaker:

ask are not actually good for RAG.

Speaker:

People would— and customers don't often understand this, right? So

Speaker:

people ask questions that typically merge structured

Speaker:

information and unstructured information and sometimes graphed

Speaker:

information, relationship between things. Okay, so let me give you

Speaker:

a— let me give you a perfect example from some of our customers in the

Speaker:

insurance space of a question that RAG is natively bad at. Okay,

Speaker:

we do a lot of work, we have a lot of customers, we have another

Speaker:

product called FraudX, and FraudX helps insurance carriers

Speaker:

find the potential evidence of fraud in their claim files

Speaker:

about 40 times faster than human review. Interesting. When you have

Speaker:

an accident, let's say an accident in a building or on a construction site, How

Speaker:

do you actually— what does fraud mean? Often someone's faked an accident,

Speaker:

or they've had a real accident and they've inflated their injuries so they can have

Speaker:

a much bigger settlement. Okay, the evidence for

Speaker:

that is sitting inside the documents. It's sitting inside 5,000 to

Speaker:

10,000 pages of documents. A human would take about 50

Speaker:

hours to read these 5,000 to 10,000 pages. Yeah, a well-trained

Speaker:

human will take about 50 hours to go try to find the inconsistencies in the

Speaker:

stories. I'm going to get to the point. A lot of that is

Speaker:

contained in the medical history, a medical chronology. Okay,

Speaker:

so this person had an accident, 2 years later they had back surgery.

Speaker:

Does the accident relate to the back surgery 2 years later? You could discover that

Speaker:

in the medical chronology of every single thing that happened to them medically over 2

Speaker:

years. An investigator might want to say something like,

Speaker:

how many times did this person go to physical therapy Before they had back

Speaker:

surgery. Because if you rush to back surgery,

Speaker:

if you haven't gotten all the physical therapy, maybe you're trying to get a settlement,

Speaker:

or maybe the doctors you're working with are trying to push you too fast.

Speaker:

Okay. RAG cannot answer that question. Why?

Speaker:

Because RAG can't count. So it's a judgment call, right?

Speaker:

I mean— Not that. Not that.

Speaker:

They're actually asking a very concrete question. They're asking a structured data question.

Speaker:

How many physical therapy appointments did this person have before they had

Speaker:

surgery? If you ask that question in SQL,

Speaker:

you could answer it instantly for a millionth of a penny.

Speaker:

Easiest thing in the world to do in an Excel spreadsheet. Or select

Speaker:

star from wherever. Yeah, easiest thing to do.

Speaker:

Weirdly, a RAG system cannot answer this question well.

Speaker:

Why? Because RAG is using unstructured data. It's

Speaker:

unstructured search is what RAG really is. Fundamentally what RAG is, is you've

Speaker:

taken all this doc, all these unstructured documents, you've stored them.

Speaker:

Someone asks a question, you then search that data

Speaker:

to try to find the 20 to 100 blocks of text most likely to answer

Speaker:

the question. Semantically most likely, right? That would— that's sure.

Speaker:

Okay. We use a mixture of techniques, but let's imagine it's just purely semantic for

Speaker:

the moment. What if you've had— let's

Speaker:

imagine you get that, you nail that search, but someone has had 200

Speaker:

physical therapy appointments. By definition, you're going to miss them.

Speaker:

The other problem is RAG doesn't actually know how many things, how many events

Speaker:

exist, because it doesn't— by its nature, it's searching, it's

Speaker:

filtering out information, searching and giving you a piece of it. So let me

Speaker:

get to the— sorry, let me get to the point. Okay, we're going to— Candace,

Speaker:

I see you're floating. I'm going to, I'm going to take you home. We'll get

Speaker:

it when we get there. Okay. We're gonna get right back to DNA in a

Speaker:

second, I promise you.

Speaker:

What we see now is that

Speaker:

customers don't understand the difference between— even you guys were a little like, wait a

Speaker:

second, there's a difference between a structured and unstructured question? You guys do this all

Speaker:

day long. They don't understand very upfront that

Speaker:

you might be asking a structured question— show me how many times a thing happened—

Speaker:

You might be asking an unstructured question like a judgment call,

Speaker:

Candace. Explain what happened in the ER visit.

Speaker:

That's a good unstructured question. You might be

Speaker:

asking a relationship question. How often did this doctor and lawyer work

Speaker:

together on a case? And so what we really need are

Speaker:

systems that when you ingest your documents, you can

Speaker:

extract the information and interpret it in multiple ways so that you could

Speaker:

store it in structured databases, unstructured

Speaker:

databases, which mostly is vector, and even graph

Speaker:

databases. And then when someone asks a question, they don't have to care

Speaker:

that whether it's structured, unstructured, or relationship-based like a graph.

Speaker:

So that's where we're going on the tactical level. We're seeing companies really need—

Speaker:

it's going beyond RAG, but it starts with that

Speaker:

document comprehension. If you don't get the documents right, if you can't

Speaker:

extract it, all is lost downstream. All right, let's talk about DNA. Let's get off

Speaker:

this topic, right? Let's get out of here. Let's talk about using language models to

Speaker:

talk to whales or something like that. Let me ask you though, so then

Speaker:

is knowledge extraction becoming the new competitive advantage?

Speaker:

We think it's 2 things, and this is where the conversations start. It's knowledge extraction.

Speaker:

It's, it's, can you accurately represent, we call it the comprehension gap. Can you

Speaker:

accurately represent the information inside this company in some way that AI can do

Speaker:

something with? And can you keep the information

Speaker:

safe, which gets into this sovereign AI? And can

Speaker:

you keep it inside the bound— whatever boundary a company is comfortable with? Are

Speaker:

they comfortable with our SOC 2 cloud? Smaller

Speaker:

companies are. Are they comfortable with a boundary that lives inside their

Speaker:

managed cloud on AWS? Some information is

Speaker:

okay there. Do they need the boundary to be in the basement? Do they need

Speaker:

to be in a data center they control machines? Do they need to be air-gapped?

Speaker:

The military needs things to be air-gapped. So we're built to run in all of

Speaker:

these environments, wherever the boundary is. You can natively install us right

Speaker:

next to the mainframes if you want to and still get that data comprehension

Speaker:

to work. So would better

Speaker:

documentation, better document

Speaker:

understanding be more valuable than simply deploying another

Speaker:

AI agent? The AI agents haven't—

Speaker:

what's an agent? An agent is just a, it's a bunch of prompts. To me,

Speaker:

an agent is like the future of the if-then statement. Okay. A very

Speaker:

sophisticated if-then statement, really. An agent is a bunch

Speaker:

of prompts that, you know, use a model for judgment to do a thing,

Speaker:

whatever that function is. But underneath— but you have to have the data underneath for

Speaker:

it to know what it's talking about. Ultimately, a company's knowledge is probably going to

Speaker:

be stored in documents and conversations, and that has to be turned into something that

Speaker:

agents can understand, or else

Speaker:

you've got a bunch of nothing at the end of it. You've got a bunch

Speaker:

of agents running around in circles talking to themselves.

Speaker:

That makes sense. That's an interesting way to put it. Think of it.

Speaker:

What, what do you think is next for

Speaker:

kind of the rag space, right? I think you and I are both in agreement

Speaker:

that it's here to stay. It's not going to die. There'll be many of clickbait,

Speaker:

ragebait articles saying that. But,

Speaker:

and the other thing too, actually, is you said there's multiple approaches,

Speaker:

right? Obviously semantic is the most obvious, but What are the other ones

Speaker:

that you find most effective? Yeah,

Speaker:

mechanically under the hood, search is a hard

Speaker:

problem, actually. And it's really core to making RAG work. It's obviously

Speaker:

fundamentally 2 things. Can I get the document— the documents right?

Speaker:

And when someone asks a question, can I find the right pieces inside the RAG

Speaker:

database? We come from a funny place on this, actually.

Speaker:

Because we started doing this before GPT launched. Once GPT

Speaker:

launched, let me get a little tinfoil hat for a second. That's

Speaker:

going to make it juicy for you. Okay, maybe a little tinfoil, make it juicy.

Speaker:

I'm not a conspiracy guy, I'm just going to give you a little bit. Okay,

Speaker:

juice it up for you. All right, so when—

Speaker:

a lot of people don't know this— OpenAI originally had a RAG

Speaker:

product. Okay, and

Speaker:

the RAG product actually was— No, that's right. I'm sorry, I cut you off.

Speaker:

I had a moment of realization. We were actually one of the first, first 20

Speaker:

testers of OpenAI. We were like one of the early guys who were given access

Speaker:

to the GPT models back in the day. And they had— they weren't really

Speaker:

commercializing anything just yet, but they had a RAG service. And

Speaker:

what a lot of people don't realize is it was not a vectorized service. They

Speaker:

were not vectorizing the data yet. They were actually— their first pass

Speaker:

at this, which is a bit more similar to what we do, was just storing

Speaker:

it as text. And then doing various kinds of text search

Speaker:

against the text to try to find the chunks that would eventually go

Speaker:

into the language model. Now, here's the tinfoil hat part. Okay,

Speaker:

when they launched ChatGPT publicly, they

Speaker:

removed that RAG service. They killed all the

Speaker:

web pages, so you could never even see that they had one unless you use

Speaker:

the Wayback Machine. We've got some emails that we're talking about it with them way

Speaker:

back, and then they told everybody, you know what? The better way to do

Speaker:

search actually is for you to vectorize all your data and do

Speaker:

vector search, not text search. And what do they wind up selling

Speaker:

the next day? An embedding service to vectorize your data.

Speaker:

So embedding models, that's the— and I'm not saying they're a tinfoil hat. They're an

Speaker:

awesome company. They have great products. We use them all the time. They have plenty

Speaker:

of detractors without us joining the— I'm not in the hater group at all.

Speaker:

I'm making it fun and funny, but We always

Speaker:

thought actually vectorization and vector

Speaker:

search, which you call— most people call similarity search,

Speaker:

is one of several techniques you need to blend together to

Speaker:

get to the answer when you do search. Why

Speaker:

vector search? When you say similarity, it's basically trying to

Speaker:

find things that are related to each other, that's more similar to each other, to

Speaker:

the question, right? Semantically similar to the question. Here's

Speaker:

the problem. When you have a lot of data,

Speaker:

you have too many things that are similar to each other. It's— if

Speaker:

you've got a jar of red marbles and you've got 10

Speaker:

red marbles, all different shades, and you say, find me the medium

Speaker:

red marble in a jar of 10 marbles, you can do it.

Speaker:

If you have 100 million marbles, all different shades of

Speaker:

red, find me medium red marble. It's really hard.

Speaker:

And so what happens in systems that are heavily vector-focused or

Speaker:

100% vector-focused, which is a lot of systems out there today,

Speaker:

the more information you add, the worse your RAG database performs.

Speaker:

We've done a lot of testing on this. Most vector systems will

Speaker:

start to fail or decline in performance in as few as 10,000 pages of

Speaker:

information. Wow. Remember, our world, we

Speaker:

do a lot of scale. But we have customers that have— an insurance customer

Speaker:

might have 20, 30 million pages of information

Speaker:

across their thousands of insurance claims.

Speaker:

So I can't go to a customer and say, after 10,000 pages, look out,

Speaker:

things are not gonna be so good. A lot of projects don't realize that's a

Speaker:

problem because a lot of things haven't scaled yet. A lot of projects are still

Speaker:

working on small knowledge bases. Yeah, they're all POCs. Yeah, there's

Speaker:

a lot of that. We're getting past that now. We're going beyond that now, but

Speaker:

still a lot of things are living in the small database world. When you scale,

Speaker:

you start to get problems. So our approach has always been really fundamentally

Speaker:

different and maybe a bit more similar to OpenAI's original idea.

Speaker:

We actually begin with the first search is a

Speaker:

bigram text search, taking advantage of 20 years

Speaker:

of good text search that's been like solid technology around a long time. We

Speaker:

use that to downsample to 1,000 potential retrievals out of the

Speaker:

system. Then We've, we have a

Speaker:

heavily tuned model that we designed that real-time

Speaker:

vectorizes those 1,000 results, then does the similarity,

Speaker:

and then does a re-rank at the same time. And then you get down to

Speaker:

the 20 to 100, let's say, much more likely

Speaker:

answers to your question. At the same time, we're

Speaker:

also doing something different with the data. When we

Speaker:

ingest your data we don't just represent it precisely as it

Speaker:

exists on the page. We're creating a lot of

Speaker:

metadata around every chunk so that

Speaker:

the chunks themselves are more differentiated than

Speaker:

any other system, so that when you do similarity, you wind up

Speaker:

with things that are more different than each other. So you don't run into that

Speaker:

marble problem we just discussed. So it's a more powerful way to do vector

Speaker:

search, more differentiated on the data. It's merged with other

Speaker:

techniques, and it's also merged with what we call like a

Speaker:

micrograph inside of it. So we're— every time we ingest documents, we're

Speaker:

creating keywords around all the chunks in there, and then we graph those.

Speaker:

That allows us to graph the chunks together so we can see relationships between things.

Speaker:

So basically, it's a 3— the marketing version would be like

Speaker:

3-way hybrid search, but now I just told you how you actually do it.

Speaker:

And well, yeah, that makes sense, right? Each one of these Each one of these

Speaker:

is gonna have their own weaknesses, right? So hopefully it's like the power system where

Speaker:

it's a 3-phase power, right? You know, when they all kind of, if you

Speaker:

do enough different approaches, they, where one is strong, one's gonna be weak

Speaker:

and vice versa. Yeah, I think that's part of it for sure. Yeah.

Speaker:

So we've always come at it from just like a first principles perspective, which is

Speaker:

sometimes good and sometimes bad cuz it's, you're, you're going against the grain of

Speaker:

maybe what everyone thinks is the right way to do it. But to your point,

Speaker:

what people think is the right thing— way to do something is

Speaker:

there's— everyone has an agenda, right? That's interesting. For sure.

Speaker:

Even OpenAI itself was a controversial bet. Now that the bet paid off, people don't

Speaker:

realize it was a bet, right? There was a lot of controversy a decade

Speaker:

ago of would this transformer architecture work

Speaker:

well enough? If I just add more data and more money and more training, does

Speaker:

this become something really interesting or not? Or do I need a different technique?

Speaker:

And OpenAI made this enormous financial bet that, nah, just keep working on it, keep

Speaker:

putting more cash, more training, and more data into this particular

Speaker:

model architecture. It will be something quite extraordinary.

Speaker:

So did that bet ever pay off? We don't know if it'll

Speaker:

pay off yet, right? They haven't— Oh, pay is the operative word. But I think

Speaker:

you're right. And I think, like we said earlier, like the transformer

Speaker:

architecture, we've gotten more mileage out of it than I originally thought.

Speaker:

And out of— than a lot of the original creators of it thought. I mean,

Speaker:

there's a reason why— I'm like, I'm not inside Google, but I assume there's a

Speaker:

reason why Google created— literally created it and was like, interesting,

Speaker:

but am I gonna throw $10 million at— $10 billion at that? Maybe not that

Speaker:

interesting just yet. Someone else proved it.

Speaker:

Sorry, Candace, I was hogging the mic because I was geeking out pretty hard.

Speaker:

No, I'm totally geeking out. I'm totally fascinated. And I'm wondering

Speaker:

if you think RAG is a long-term architectural

Speaker:

pattern, or do you think it'll eventually

Speaker:

evolve into something more sophisticated as enterprise AI

Speaker:

matures? In a way, my answer is—

Speaker:

you're gonna hate this answer— I was a journalist for 20 years, and I hate

Speaker:

when someone says both. I was like, pick a lane, man.

Speaker:

And here's why I say both. It would be silly to say that what we

Speaker:

have today is what we'll have tomorrow. Obviously, that's just not true. There'll be more

Speaker:

interesting patterns and techniques to do these things over time. And

Speaker:

we might be 10 years out, but the chipset will evolve. So maybe it can

Speaker:

take on— maybe it can hold more data inside of the

Speaker:

model than we can do today. But at the same time, if I think about

Speaker:

the history of computing, so much never changes. We're still—

Speaker:

take a really simple example, edge versus cloud. This

Speaker:

conversation has been happening since the beginning of computing, right? There was the

Speaker:

mainframe that was cloud, if you will. There was the

Speaker:

compute was in the center, and then the mini computer moved it to a

Speaker:

little bit of a rim around that, but it still was in the center. Then

Speaker:

PC revolution, massive. Everyone's forget it, compute's on the edge.

Speaker:

We have all this power on our desk. And then that happened

Speaker:

for 20 years. Then I was like, wait a second, maybe we should go back

Speaker:

to the cloud. And then the cloud came along, we start to put the compute

Speaker:

in the center. Then mobile came along, and now we've got these powerful supercomputers in

Speaker:

our phone. It's like, now it's at the edge. Now it's going to— same thing.

Speaker:

So just a lot of metaphors in computing just seem to be forever. And to

Speaker:

me, I think there's something in RAG that's pretty fundamental, which is

Speaker:

that the big fence, fancy, expensive thing in the middle

Speaker:

will probably never have all the information it needs to. And it's always going

Speaker:

to need something that has a larger database of knowledge than it

Speaker:

and filter it in real-time ways so that it can answer a question or perform

Speaker:

a function. And if you just think about what RAG is really fundamentally, it's that.

Speaker:

It's that the universe for information is really vast. It's out here.

Speaker:

It's hard for me to store all that and use it in a functional way

Speaker:

in real time inside whatever thing I'm building. And so you need something in the

Speaker:

middle that helps you search and sort and filter and deliver it in the right—

Speaker:

in real-time way. Sorry, that was not a straight-line

Speaker:

answer. That was like both sides of the mouth on that one, but I think

Speaker:

that's the truth. That makes a lot of sense.

Speaker:

We are— I would love to talk to you another hour or 2, but I

Speaker:

think we're at the end of time and I'm going to be respectful of your

Speaker:

time. So where can folks find out more about you and your

Speaker:

company? Yeah, so the company is Valantor

Speaker:

and you can get us at valantor.com.

Speaker:

V-A-L-A-N-T-O-R. Luckily it's spelled like it

Speaker:

sounds. I just spent a week and 2 weeks in Croatia and

Speaker:

I don't know what they're doing with the alphabet over there, but I'm trying to

Speaker:

make sense of it. It's an awesome country. I got nothing bad to say, but

Speaker:

they need some vowels. Okay. But Valontor is easy to

Speaker:

pronounce. So valontor.com, check us out. And for

Speaker:

me, look, if you want to see the documentaries, you could search— would you search

Speaker:

weather films or Weather Channel documentaries on YouTube and see some of the cool stuff

Speaker:

from my past? But check us out at valontor.com. And if

Speaker:

you're an enterprise that has the kind of problems we're talking about, large

Speaker:

amounts of data, scale, security, accuracy is what you need, say hi.

Speaker:

Cool. Excellent. That sounds great. That sounds great. Thank you so much for your time

Speaker:

today. It was fabulous. Thanks a lot, guys. I really appreciate it. And

Speaker:

next time we'll get back on the DNA, Candace. I'm sorry, it felt like we,

Speaker:

we ran outta time before we can get back there. It's okay. We'll have you,

Speaker:

we'll have you back and then we'll— Have you back. Yeah. That's

Speaker:

it. Thanks a lot. I'll let the outro music play.

Links

Chapters

Video

More from YouTube