Artwork for podcast Data Driven
Building Trust in AI – Lawrence Snap on Hallucination Detection and the Future of TrustScale
Episode 1127th August 2026 • Data Driven • Data Driven
00:00:00 00:53:34

Share Episode

Shownotes

In this episode, hosts Frank La Vigne and Andy Leonard sit down with Lawrence Snapp, CEO and board member of TrustScale, an innovative AI trust and verification platform. Together, they explore the crucial topic of trust in artificial intelligence—how hallucinations from AI models can spiral out of control and why it’s essential to catch them early.

Lawrence Snapp pulls back the curtain on TrustScale’s journey from early AI and translation work in the 1980s and call center solutions, to its modern-day mission: empowering humans to verify and trust AI outputs with products like Argus.

You’ll hear about real-world impacts in fields like healthcare, legal, and research, learn how TrustScale’s deterministic engine outpaces humans in hallucination detection, and dive into debates around truth, trustworthiness, and the future of configuring AI values. Whether you’re an industry veteran or just curious about how we build trustworthy AI, this episode is packed with insights on technology, responsibility, and the power of rigorous engineering.

Links

Time Stamps

00:00 Funny AI-generated story

03:37 Company's Evolution from Translation to AI

07:32 Developing the TrustScale engine

11:35 Using AI for collaboration

13:39 Building a trust scoring system

18:12 Working at BASF and AI testing

19:32 New developments in voice assistants

23:41 Trust scale and company confidence

26:13 AI mishap in medical drama

32:48 Overcoming AI recency and cost challenges

36:54 AI's impact on career growth

38:50 Interviewing and developing talent

43:31 Starting with Polaroids and AI goals

46:54 Discussing TrustScale and AI interactions

50:42 Empowering AI customization for users

51:45 AI ethics discussion in Los Altos

Transcripts

Speaker:

One of my favorite hallucination stories is, this was back when I was at Red

Speaker:

Hat. We were doing some experiments with fine-tuning, and

Speaker:

it came up with this elaborate story about how my website, Frank's World,

Speaker:

came about. Apparently it was a— somehow it got the

Speaker:

idea that it was a children's show on the BBC in the '90s about

Speaker:

recycling. And Frank was a— and I can say this now and

Speaker:

laugh. When I first read it, I was like, I was Frank. The character

Speaker:

Frank was a talking recycling can or something like that.

Speaker:

And I decided to have fun with it, you know, so I fed it into

Speaker:

NotebookLM. And so it came up with this whole— and it

Speaker:

hallucinated, it amplified the hallucination. So it

Speaker:

came up with this whole thing. It was an award-winning show, groundbreaking,

Speaker:

you know, really spearheaded the environmental movement.

Speaker:

Hallucinations seem to compound. So obviously you want to

Speaker:

catch it early. How do you detect hallucinations at scale? Yeah. I know I,

Speaker:

It was so convincing, honestly, Laurence, I had to ask friends of mine who grew

Speaker:

up in the UK, was there really a show like this?

Speaker:

Apparently

Speaker:

no. Hello and

Speaker:

welcome back to Data Driven, the podcast where we explore the emergent industry

Speaker:

that is artificial intelligence, data science, and of course, data engineering

Speaker:

without which most of this would be really for nothing.

Speaker:

So with that in mind, I'm extra happy to have my favoritest

Speaker:

data engineer in the world with me, Andy Leonard. How's it going, Andy?

Speaker:

Hey, Frank, it's going well, especially since I'm still your

Speaker:

favoritest data engineer in the world. How are you? You

Speaker:

will never be replaced. Not even Claude could replace you, Andy. Wow.

Speaker:

Wow. OpenAI might and some of the Chinese models, but we'll see.

Speaker:

Maybe. Yeah, I feel you. How you doing, brother? All right, I'm doing all right.

Speaker:

Um, there's a Tim Hortons— I think we mentioned this a previous show— opened up

Speaker:

in Maryland near me, and, um, it's

Speaker:

a good alternative to Starbucks and Dunkin' Donuts. That's all I'll say. So

Speaker:

is this your second cup of Tim Hortons?

Speaker:

No, no, it's the same one. Oh, okay, okay. Yeah, okay, I just checked.

Speaker:

I, I reuse my coffee cups, so half the time when you see me on

Speaker:

a stream or whatever with a Starbucks, chances are it's— I wash it out and

Speaker:

I just get a few days' worth out of it till it becomes unappealing.

Speaker:

Understood. Until I can't trust it, which is a good segue.

Speaker:

You know, trust is important. Trust is important. It's not just a

Speaker:

department at a bank. Today, we have with us a

Speaker:

fellow former Microsoftie, Lawrence Snap, who is

Speaker:

the CEO and board member of TrustScale,

Speaker:

which basically— yep, besides, I'm getting better at, like, the joining and

Speaker:

things like that, but TrustScale, Really puts the trust

Speaker:

into AI. And it is an AI trust and

Speaker:

verification platform, and it bridges the trust gap between AI and

Speaker:

humanity through solutions that offer hallucination

Speaker:

detection and reinforcement learning to reduce risk and prevent

Speaker:

unintended consequences of AI outputs. Welcome to the show, Laurence.

Speaker:

It's great to be here. Thank you very much, Frank. Hey, no problem. No

Speaker:

problem. So we were talking in the virtual green room, and TrustScale's

Speaker:

been around for, you said, 20-plus years. I think the number was 23.

Speaker:

23 years last week. Wow. So a

Speaker:

lot of people listening are gonna say, 23 years? AI hasn't been around but

Speaker:

since 2022, which we know none of

Speaker:

our listeners would actually think that. But what was the company doing before

Speaker:

LLMs came out? It's a great question. So we were founded by a French

Speaker:

AI scientist who went to graduate school in France in the

Speaker:

'80s, and he really spun out of working with

Speaker:

Apple doing translation. 23 years ago last week it

Speaker:

was started. From there, we went into everything from call centers to

Speaker:

decision science, data, synthetic datasets, and then the

Speaker:

voice assistants and chatbots emerged, and we evolved into

Speaker:

that. And then all of those, you know, and you can imagine the big 3

Speaker:

out there, we did a lot of quality work around the data and the

Speaker:

decisioning and not just prompt response, but across 200

Speaker:

languages we worked. We moved in very natural

Speaker:

sequence over to the LLM data demands of the world.

Speaker:

So it's been a pretty natural progression from translation to chat to voice

Speaker:

to, you know, certainly doing it in all certain languages. But now

Speaker:

it's been— we're probably about 80% working

Speaker:

on making AI better now. Really? Wow, that's cool. Yeah,

Speaker:

I mean, and in the virtual green room, you mentioned Cortana and RIP Cortana.

Speaker:

She was a great speaker. There's still things that my Android phone will not do

Speaker:

by voice that I could do 10, 15 years ago on Cortana. And

Speaker:

my poor kids hear it all the time. But anyway, such is

Speaker:

life. But so how do you detect

Speaker:

hallucinations? Like, obviously people can kind of

Speaker:

tell. One of my favorite hallucination stories is,

Speaker:

this was back when I was at Red Hat, we were doing some experiments with

Speaker:

the fine-tuning, and it came up with this elaborate story about how

Speaker:

my website Frank's World came about. Apparently it was a—

Speaker:

somehow it got the idea that it was a children's show on the BBC in

Speaker:

the '90s about recycling, and Frank was a— and I

Speaker:

can say this now and laugh, when I first read it, I was like, I

Speaker:

was Frank, what the character Frank was, a talking recycling can or something like

Speaker:

that. And I decided to have fun with it.

Speaker:

You know, so I fed it into NotebookLM, and so it came up with this

Speaker:

whole, and it hallucinated, it amplified the hallucination.

Speaker:

So it came up with this whole thing. It was an award-winning show,

Speaker:

groundbreaking, you know, really spearheaded the environmental movement.

Speaker:

Hallucinations seem to compound. So obviously you wanna

Speaker:

catch it early. How do you detect hallucinations at scale? I know,

Speaker:

it was so convincing, honestly, Lawrence, I had to ask friends of mine who grew

Speaker:

up in the UK, Was there really a show like this? Apparently, no.

Speaker:

No, it's a great question. And, you know, at first it starts

Speaker:

with the grounding that AI is a

Speaker:

probabilistic system. It is a recommendation engine.

Speaker:

It's guessing the next token or the next word. And

Speaker:

there's a lot of mathematics, and it's pretty darn good, and it's pretty eloquent. But

Speaker:

then when it doesn't have a confidence interval, it is— it's

Speaker:

trained. to even be more confident and double down. Right? And in

Speaker:

fact, we would— what we've seen across millions of prompts,

Speaker:

responses, and analyses that we do is that the more confident

Speaker:

something is, the more you better check it. Because— I'm glad you mentioned that

Speaker:

because I thought I noticed that too, because it just seemed so like,

Speaker:

like with the example of the, the kids show in the '90s. I mean, it

Speaker:

was, I mean, it was really selling hard. I'm sorry. I didn't mean to cut

Speaker:

you off. But like, It's not my imagination then. It starts that way. I

Speaker:

mean, listen, these models are trained that way. And some of this is neuroscience, right?

Speaker:

And psychology and really just

Speaker:

triggers of different chemicals that happens when we don't know something, but someone's really

Speaker:

confident, we believe in it. But, you know, to answer your question, AI is

Speaker:

probabilistic. What we did, and I've got a long time,

Speaker:

a long history. I'm actually a CPA by trade before diving into the tech industry,

Speaker:

but I did big data work in the early 2000s and even have a patent

Speaker:

around it. at a very large scale. And what's exciting is

Speaker:

that we realize that a probabilistic system needs a deterministic system as a

Speaker:

counterpoint, almost like cybersecurity, to compare. And so

Speaker:

as we were working for some of the big model makers, it, you know,

Speaker:

and it's, to be honest, a very cumulative effect of how we

Speaker:

figured out the detection and then obviously the protection and correction side.

Speaker:

But to detect a hallucination requires you to actually have the evidence from

Speaker:

an empirical source. And the ability to compare it. So

Speaker:

we created an engine, the TrustScale engine is how we refer to it.

Speaker:

And this engine helps augment our human work. And

Speaker:

quite frankly, we just had a very large customer last week tell us

Speaker:

that our system is actually better than humans at detecting

Speaker:

hallucinations. And so there's a lot behind it and it's very

Speaker:

complex and we kind of took cumulative knowledge of decades of

Speaker:

AI experience, especially when AI was science fiction. You know, when our founder

Speaker:

went to grad school in the '80s, around the same time as Yann LeCun,

Speaker:

in fact, in France at the same exact time. And we have a few other

Speaker:

investor owners. We're private. We don't have venture capitalists, but they were

Speaker:

deep into data science and AI, everything from Lisp to knowledge

Speaker:

graphs and everything else. But we took this cumulative knowledge and figured out how to

Speaker:

create a system to basically double-check every

Speaker:

claim at the atomic level based against

Speaker:

empirical evidence. Then out of that, we're able to figure out, hey,

Speaker:

and one of our products that comes off the engine is called Argus, which just

Speaker:

basically highlights red, yellow, green based upon the confidence

Speaker:

interval that something's trustworthy. We assign it a trust score.

Speaker:

We have a trust box. You can actually see the evidence links and you can

Speaker:

look for yourself. It's about empowering people. Then we just added

Speaker:

and released last week an exciting feature, a little bit like Grammarly for AI

Speaker:

where We can suggest a correction to a hallucination based upon real

Speaker:

evidence. And with a click of a button, you can actually correct the

Speaker:

output. So Argus is that product that

Speaker:

manifests in a visual way, but we actually use the same engine to do

Speaker:

a lot of training and evaluation for AI models on the

Speaker:

backend. And that's invisible under the hood data science and plumbing,

Speaker:

but it's exciting. And we've actually figured out how to do this across audio, text,

Speaker:

video and imagery now. And so we're excited

Speaker:

about where we're going. Training and evaluations are big business, but

Speaker:

anyone can go and witness it, you know, at TrustScale and see Argus

Speaker:

in living color. It's a lot of fun. Interesting. So Argus is one of

Speaker:

your products and that works on— that's the kind of the bot that would

Speaker:

test for trustworthiness? Yeah. And then to be clear, it's actually

Speaker:

not AI. Our system is a deterministic system. It is not AI

Speaker:

checking AI. It's not. And it's not. an LLM as a judge.

Speaker:

In fact, that's a big point when we, you know, we love working with researchers,

Speaker:

especially at the labs, and they really challenge and light up like a light bulb

Speaker:

when they figure it out. But no, this is, this is a deterministic

Speaker:

counterpoint, if you will, and a gap analysis to determine

Speaker:

what confidence you should have or what trust you should apply. So that's where the

Speaker:

trust score came from, which we're really excited about. I mean, we're seeing it in

Speaker:

especially high-risk domains, right? A lot of people are relying upon it. But it's, you

Speaker:

know, Argus is a proofreader, if you will. But we know

Speaker:

we've saved some people's jobs. There's a lot of examples where

Speaker:

hallucinations have caused damage, hurt people, put people in jail,

Speaker:

you know, compromised patient care and healthcare. So

Speaker:

we see, you know, and there's a big flurry of energy right now in the

Speaker:

media about just AI detection. Is this AI or not? And that's great. Right.

Speaker:

I think Reid Hoffman had a book a post on Substack last week

Speaker:

about it, but about Pangram. But we actually look at it and say, okay, well,

Speaker:

like a spell checker, who cares if you used it or not? Who cares if

Speaker:

it was AI-generated or you generated it? What we care about though is trust. Is

Speaker:

it actually something you should trust? And empowering people with

Speaker:

evidence so that they can determine the trustworthiness of some

Speaker:

output and then converge it with their human magic,

Speaker:

that's a lot of why we exist. So Argus is one manifestation of our

Speaker:

engine. for the end user, but it's used a lot for the shapers and the

Speaker:

makers, the agents and the LLM makers as well on the backend.

Speaker:

That's very interesting that you bring up collaboration

Speaker:

as, uh, as your use case for that, as you're, as you're

Speaker:

describing it, because I, you know, I see that as

Speaker:

the best mix, right? Uh, let the AI do what the AI is

Speaker:

best at. Let the human do what the human's best at. It's

Speaker:

almost like over time, That, that's how I

Speaker:

started using AI, by the way. I started testing

Speaker:

ChatGPT when it came out in November of '22, and I couldn't get it to

Speaker:

do anything I needed it to do. I, it would do some things, but I

Speaker:

was focused on what we call vibe coding today. I

Speaker:

didn't really get it to do that until March of last year. So

Speaker:

2025, and it's just gotten better and

Speaker:

better, but it's, it's almost like a

Speaker:

continuously improving Mechanical Turk. in

Speaker:

that collaboration market. It's able to take on more and more,

Speaker:

certainly over the past, what, 17 months now? It's taken on

Speaker:

more and more, and I've been able to, as you point out,

Speaker:

trust it more. You know, I love that

Speaker:

description. I love that being the use case. And it's very interesting to me

Speaker:

that you're helping on the back end, that, you know,

Speaker:

I love closed-loop engineering. You know, what can I say?

Speaker:

Yeah. No, in fact, you know, one of the, one of the exciting things on

Speaker:

Argus is we call it loop. But the loop module of

Speaker:

Argus is to take the red, yellow, greens essentially of what a user

Speaker:

experiences. And then we offer that as reinforcement training

Speaker:

feedstock, if you will, on the back end to say, hey, try to prevent this

Speaker:

hallucination by adjusting the weights or the data, you know, so next time it doesn't

Speaker:

happen. Very cool. So how do you test against ground truth, right?

Speaker:

Like, so in my Ridiculous scenario, right? Like the

Speaker:

BBC show that never existed. How would you, how would you know

Speaker:

that? Would it be something in the text that kind of infers like, oh, this

Speaker:

is way too confident what it's saying? Or does it actually do a

Speaker:

search for TV shows from the '90s?

Speaker:

Like what is, what, or some combination or something else entirely?

Speaker:

Well, I think a lot, a lot of what we do is around mirroring the

Speaker:

human mind. And so what we do as humans is compare to

Speaker:

the things we're grounded on. I remember that experience or whatever it might

Speaker:

be. And so our engine is going out and trying to

Speaker:

identify sources of evidence that would support or

Speaker:

refute the claim of which you're assessing, or the image, right? I

Speaker:

mean, does it have 6 fingers on that clown or is there 5? And so

Speaker:

we have these— we have our own golden dataset and our own moat that gets

Speaker:

bigger every minute of the world, but we also go out and just try to

Speaker:

find like, hey, let's go find that show and if it exists or not. And

Speaker:

so what this does on the backend is that it creates a scoring

Speaker:

system and it feeds to what we call a trust score. Our trust

Speaker:

score manifests in a 5-point scale, right? Red to orange to

Speaker:

yellow to green. And what we're doing is essentially exhibiting a

Speaker:

confidence interval around the trustworthiness

Speaker:

of that claim based upon comparing it to what we can find. And we,

Speaker:

Our systems are updated essentially every 15 minutes. We work across 12

Speaker:

languages now with Argus to do this as a proofreader. But, you

Speaker:

know, Frank, it's a great question of not just, you know, what are we doing,

Speaker:

but how are we doing it? But at the end of the day, it's a—

Speaker:

it's because it's so good. AI is so good

Speaker:

at being deliberately deceitful, especially when it doesn't have the

Speaker:

right token, right? It's a little tricky. And I'll be honest, it's

Speaker:

taken us a long time to figure out how to do it. but we just

Speaker:

surpassed 98.5% of

Speaker:

accuracy when reviewed by a very elite panel of humans.

Speaker:

So we're really excited that we think we can detect and suggest corrections

Speaker:

on this. But you'll notice one of the things I don't say is truth.

Speaker:

And Frank, you haven't asked about it, but one of the things we're very careful

Speaker:

about— we are here to empower people with evidence so they can determine whether they

Speaker:

should trust something. But one of the things we will not do

Speaker:

is cross that line into being a truth teller. We're not

Speaker:

trying to— we're not trying to see, is this truthful or not? And

Speaker:

because humans and language and memories are

Speaker:

nuanced, we think this is really important. And so we

Speaker:

draw that line. Yes, a lot of people say, oh, it's a lie

Speaker:

detector, it's a fact checker. Of course, there is truthful

Speaker:

things. I like apples. My wife hates apples. We're both right.

Speaker:

We're both truthful. And, you know, that's why we don't cross that

Speaker:

line into the whole truth game. We think that's chasing a rainbow. But our

Speaker:

job is trustworthiness, empirical evidence, empowering

Speaker:

humans to retain that agency, like Andy said. And we

Speaker:

think AI is a superpower for humans. We just think it needs some

Speaker:

verification layer. That makes sense. And you're right, there is a

Speaker:

difference, a subtle difference between what's true and what's

Speaker:

hallucinated. Yeah, and your truth and my

Speaker:

truth, they might be different, even

Speaker:

if we're both relying on the same facts. I'll give you a specific example. We

Speaker:

took a famous speech from our current president,

Speaker:

and we took a transcript of CNN and Fox after this speech.

Speaker:

There were 3 unemployment numbers that were used in the 3 different

Speaker:

perspectives of this speech. The unemployment numbers,

Speaker:

ironically, were different. And one said, oh, he's understating. One

Speaker:

said, oh, he's overstating, right? And what it actually came down to when we

Speaker:

ran Argus against it was that all 3 were telling the truth. They were just

Speaker:

using the whisper number, the first reported number, or the corrected number.

Speaker:

Though, but that context is lost. And what Argus does is it surfaces

Speaker:

it with links. It shows that there's potentially

Speaker:

conflicted information here, but it gives the evidence links so you can go and

Speaker:

as a human and dig deeper if you want to figure out, hey, was he

Speaker:

telling the truth or not? Interesting. That's very interesting because I, I

Speaker:

know there's a, um, I want to say the service is called Ground

Speaker:

News that does that very thing. It looks for political bias, uh,

Speaker:

in, in news stories. And, um,

Speaker:

unfortunately, that's important these days. Yeah, it is. Listen,

Speaker:

Ground News is great because they're bringing a whole bunch of different perspectives.

Speaker:

Now, news And the lens and context, I mean, it shows the

Speaker:

value of that product is that they are

Speaker:

highlighting that nuance of language and perspective, right? Now, our

Speaker:

tool is all about AI outputs. And that's what TrustScale exists. Again, we exist

Speaker:

to put the trust in AI, as you said, Frank. And we joke, we

Speaker:

don't make the AI, we try to make it better, right? And that's

Speaker:

where we exist. Yeah. Go ahead. Sorry. Oh, I

Speaker:

was going to say, we're borrowing that, of course, from BASF. But see,

Speaker:

I'm going there. Great. I used to work at BASF, so talk about

Speaker:

it being a small world, right? Like, that's— but I

Speaker:

mean, it's a great— it was a great— I mean, it's a great way to

Speaker:

think of it. Like, you know, and a lot of interesting things you said

Speaker:

that we, you know, we— I guess we kind of glossed over was who, you

Speaker:

know, when things like Alexa or Google's

Speaker:

speaker or Cortana came out, I often wondered, like, how do they test this?

Speaker:

Like, how do you test this at scale? Now I

Speaker:

suspect that you may not be able to say anything by name, but I suspect

Speaker:

that I might have part of my answer at least. How do you test

Speaker:

it and things like that? Also too, I think it's interesting that obviously

Speaker:

AI didn't start in 2002, didn't even start in 2000. It's an

Speaker:

older field, as is natural language processing.

Speaker:

Yeah. I remember the Build

Speaker:

2016 keynote, one of the

Speaker:

gags was they were using a chatbot to order something from Pizza Hut.

Speaker:

So, you know, clearly that sort of technology had been around. It just

Speaker:

was not as nuanced or advanced as, you know, the LLMs we

Speaker:

have today. But yeah, so that's interesting. Like, I often wonder, like,

Speaker:

how do you test these systems at scale? Because something like an

Speaker:

Alexa as a consumer device has to work in

Speaker:

all sorts of environments. And The testing was

Speaker:

probably pretty rigorous. It's incredible, Frank. Yeah, you

Speaker:

think about the first wave of voice assistants, and I think we're about to go

Speaker:

get another wave here this holiday, right? And I know, you know, you

Speaker:

got Johnny Ivy over working with Sam, and you've got, you

Speaker:

know, whispers on Mac rumors of new Apple devices coming out, and you got

Speaker:

Gemini powering Siri soon. You got all this amazing stuff that's about

Speaker:

to flood the world. But the training and evaluation,

Speaker:

of these experiences and the trustworthiness of it is

Speaker:

pretty radical stuff. And we're actually working on the trustworthiness of voice assistants

Speaker:

now, but I'll tell you, historically we had millions and millions of tasks and

Speaker:

we were human, all human. And what our Trust Scale engine did

Speaker:

was in a weird way automate a lot of this human and it

Speaker:

elevated our human task workers to being reviewers and to

Speaker:

being the judgment when something's yellow or orange. I always say the

Speaker:

magic is in the yellow. It's not the green or red,

Speaker:

it's the yellow. And that's where humans are, I think, always going to be necessary

Speaker:

alongside and being powered by AI. But, you know, when it comes

Speaker:

to voice assistants specifically, we've got some tricks

Speaker:

coming here. I mean, like audio tones, and there's certain things you can do when

Speaker:

a voice assistant might not be giving you a

Speaker:

trustworthy response. And we're looking at even how we apply that. In fact, we

Speaker:

have a patent application in on this using audio tones to flag

Speaker:

flag 1, 2, or 3 beeps based upon the trustworthiness of

Speaker:

what a chatbot's telling you. So there's some big stuff coming, Frank,

Speaker:

and we're working on it on the backend for our big customers, but we're also

Speaker:

working on this TrustScale engine to find real-time,

Speaker:

ultra-scale, extremely affordable approaches to

Speaker:

putting that trust layer around AI. It truly is a component of

Speaker:

the harness, right? 100%. Pretty cool. One of the use

Speaker:

cases that I have concerns about, just because, probably

Speaker:

because I was biased by negative reactions to

Speaker:

it from, uh, some short or something. But it was

Speaker:

this idea of someone getting in their vehicle

Speaker:

and being upset and being able to speak to the

Speaker:

vehicle and have it start, which, okay,

Speaker:

that, that raises some flags right there. I, I'm being an engineer,

Speaker:

you know, in which paranoia is a virtue. I

Speaker:

start going there. And, but the,

Speaker:

the regulatory part of this, which is akin to, you've

Speaker:

gone through a due process, you've been convicted of being an alcoholic

Speaker:

or someone who operates a vehicle under the influence, and there's a breathalyzer

Speaker:

installed before you can start the vehicle. That I get.

Speaker:

There's that whole due process part. But the vehicle

Speaker:

deciding that you're too stressed and worrying about preventing

Speaker:

road rage, which Noble concerns. I get it.

Speaker:

But the, the one video I saw was someone had injured themselves with a

Speaker:

chainsaw out in the middle of the woods, and they jump in their truck. Of

Speaker:

course, they're stressed. They're trying to get out of there. They're out of cell

Speaker:

range. And, you know, you can imagine it just— the

Speaker:

vehicle won't start. And so these are—

Speaker:

granted, these are edge cases. But if somebody's

Speaker:

life's on the line in the edge case, I think that's a

Speaker:

different category than what you're talking about. But at the same time,

Speaker:

my experience is dealing with a lot of healthcare data. And I remember

Speaker:

distinctly talking to my team. We were doing data engineering to

Speaker:

supply a response within a handful of seconds

Speaker:

about whether the insurance company would cover a

Speaker:

medicine being prescribed, whether it was valid. It would do

Speaker:

the check to see if there's interference with other

Speaker:

prescriptions. And our time limit on that was crazy,

Speaker:

what we had to do. And I remember telling my team, Grandma's

Speaker:

going to come in at 5 minutes to 5 on a Friday

Speaker:

on a long weekend. And if we get this right,

Speaker:

she could get her prescription and, you know, and have her medicines that

Speaker:

she's going to run out halfway through this weekend. And if we get it wrong,

Speaker:

then the worst case is really a worst case.

Speaker:

And so I can imagine by extension, I'm kind of

Speaker:

extrapolating, I don't know, destrapulating. I'm going back in time and

Speaker:

thinking about the answers to the questions that trust scale

Speaker:

is going to get right. It's going to surface

Speaker:

first. It's going to correctly identify the yellow, which I

Speaker:

find fascinating. I know Frank was thinking the same thing I'm thinking, because we're both

Speaker:

geeks. It's like, how do you do that? We don't need to know how to

Speaker:

do it, but the fact that it is done and that it's testable and

Speaker:

verifiable, and frankly, I have a little more confidence in it coming

Speaker:

from a company that's been around a couple of decades than I do from

Speaker:

somebody who threw 9 figures in VC money at

Speaker:

somebody 18 months ago. I just do. Maybe that's unfair,

Speaker:

but I'm biased that way. Do you find— I guess the rambly

Speaker:

part of the end of the question is, where's the question? Is do you find

Speaker:

yourself in those scenarios? You don't have to tell me which

Speaker:

scenarios, but do you— have you found your product

Speaker:

specifically in that chain of events that number one,

Speaker:

correctly identified a yellow, number two, a human was able to step

Speaker:

in and correct a hallucination, and that it

Speaker:

ended up really making things better for either an

Speaker:

individual or, you know, or a corporation or

Speaker:

Yeah, great question. And you covered a lot of really rich

Speaker:

topics from insurance to legal to healthcare

Speaker:

and even autonomous driving. And, you know, it's fun,

Speaker:

Frank, you and I were both at Microsoft a little bit ago. And one of

Speaker:

the projects I actually had there was the Kinect incubator. And if

Speaker:

you remember, people remember Xbox and Kinect. I remember that. And I was working with

Speaker:

a company, I was an advisor to a company that was taking a Kinect, putting

Speaker:

it on commercial truck drivers' dashboards. And it was doing facial

Speaker:

recognition of their alertness or their tiredness and

Speaker:

warning them. It was a really cool system. And it was really a— and now

Speaker:

we see it in Teslas and Rivian, cameras facing people to

Speaker:

see whether you're paying attention. But that's that augmentation of the

Speaker:

humans, which is why together we're better. And there are superpowers.

Speaker:

But Andy, to answer your question specifically, we have seen,

Speaker:

I would say, in 4 major domains,

Speaker:

huge successes with Trust Scale. The first one I'd argue is probably healthcare,

Speaker:

where hallucinations can kill people. And

Speaker:

in fact, I remember someone that was using Argus sent me a video of

Speaker:

that famous show, The Pit. Since we're nerding out, we'll talk about it. There's a

Speaker:

really famous— Yeah. I think it's season 3, episode 6. There's a really famous

Speaker:

episode where the pressure in the ER is so intense in this

Speaker:

hospital drama And I'm gonna still call him, you know, Goose, but,

Speaker:

you know, there's, there's these great actors in this. And the situation is

Speaker:

that someone used an AI note-taker to

Speaker:

do their clinical notes, and this person was about to get operated on. And

Speaker:

there's a whole scene in the show, The Pit, where this

Speaker:

hallucination almost compromised or killed someone in an

Speaker:

operating room. And so if Hollywood's all over it, we know we are. But this

Speaker:

person actually had used it in a healthcare situation. They had sent me that

Speaker:

video and then they introduced me to a doctor at the NIH who I didn't

Speaker:

know is using Argus to

Speaker:

double-check AI outputs of, you know, the way he put it, this

Speaker:

doctor said, hey, I got 125 pages of notes and I get this AI

Speaker:

summary, but I can't trust it. So I need something to

Speaker:

double-check it. And so there's an example of, of an NIH

Speaker:

doctor and it's usually cancer and cardiothoracic and all this stuff. But

Speaker:

these domains are real. You can't take risks in that last 2%

Speaker:

because it's a probabilistic technology. That's pretty dangerous.

Speaker:

Yeah. So healthcare is one. We've seen it in legal like crazy.

Speaker:

Academic research, we've got students at MIT, Harvard,

Speaker:

Rice, Stanford, all using Argus to

Speaker:

proofread their AI and

Speaker:

human converged works. And this is a really big deal,

Speaker:

right? Because if you get it wrong, your research might

Speaker:

lead to some other invention, but you could hurt people, or you could really

Speaker:

change the course of things. So that healthcare

Speaker:

scenario, the academic research scenario, legal, of

Speaker:

course. And then yesterday I got a call from a board member, Lloyd's of London,

Speaker:

the oldest, biggest insurance company in the world. Interesting. And it was all about,

Speaker:

hey, We are carving out AI risks from

Speaker:

policies. But there are all these other companies that

Speaker:

we'd like you to maybe consider talking to. Maybe you could help them,

Speaker:

and they're trying to figure out how to underwrite AI risks. But they

Speaker:

know AI is probabilistic. But you guys are the

Speaker:

expert in trustworthiness on AI outputs, and can you help them?

Speaker:

So I got an introduction last night to one of the biggest investors in

Speaker:

a New York-based insurance company that's working on AI risk.

Speaker:

And I, so I'm excited, Andy, like you about all these scenarios that are

Speaker:

coming up and it just, we're just early. Sure. But AI is

Speaker:

probabilistic. It makes mistakes. And we believe that,

Speaker:

that, and no one's hiding it. I mean, these big labs admit it.

Speaker:

But we believe that a counterbalance or a layer of

Speaker:

trust services on top of things, you know, might save some lives.

Speaker:

I can definitely think of, uh, of use cases with,

Speaker:

you know, clients of mine at my consulting, uh, you know, my

Speaker:

consulting firm that could definitely put it into use. So you may

Speaker:

have some folks coming your way as a result of us having this

Speaker:

conversation. So, um, I get it. And

Speaker:

it's, you know, I guess the one question I

Speaker:

would have, it may not be quantifiable at this point,

Speaker:

but You're not promising truth, and I get that because that

Speaker:

pegs the needle at 100%. You wouldn't want to go there. Nobody wants to go

Speaker:

there, especially with something that's not deterministic. This is the, you know,

Speaker:

the other edge of that sword. Do you guys have statistics

Speaker:

on how much risk you mitigate, like how many

Speaker:

things you catch? And that would, you know, by math,

Speaker:

100 minus whatever that number, that percentage is, would be the

Speaker:

ones that are slipping through still. Do

Speaker:

you have those numbers? We have our own, and there's some really big

Speaker:

industry numbers, right? It's like, again, don't trust me, but one of the

Speaker:

stats that came out this year is Anthropic did a research report and said

Speaker:

91% of AI users, people relying upon AI outputs,

Speaker:

do not verify the outputs. 91%. So you

Speaker:

start with that and you go, oh my gosh, okay. So basically, You

Speaker:

flip it around, only 9% of people actually check the

Speaker:

outputs for their validity, for their trustworthiness, right? And then on

Speaker:

top of that, and listen, it's hard. You've got to go find

Speaker:

sources and compare things. And there's a lot of work to

Speaker:

compare a very plausible output that's deliberately deceiving

Speaker:

you, right, to check its trustworthiness. There's a lot of work in there. And that's,

Speaker:

that's why it took us so many years to develop what we have. But then

Speaker:

other data, and I think it's

Speaker:

important, we see it dependent upon the domain

Speaker:

and the challenge that the human's trying to bring to the AI. So

Speaker:

in certain domains, we see hallucination rates up to—

Speaker:

I mean, we have a Stanford study where it's 80% in

Speaker:

complicated business. I mean, it is because it's 10

Speaker:

miles wide and an inch thick right now. And so, so that's a

Speaker:

Stanford study. That's not ours. Now, we are hired by the big labs to

Speaker:

break their AI. So we know all the tricks to break

Speaker:

it and why you have to do, you know, get them upside down on tokens.

Speaker:

And we have a lot of fun breaking things, real red

Speaker:

teaming it right in the cyberspace. Yep. But the other

Speaker:

data that I think is really good is this 20% number.

Speaker:

Plus or minus, things have improved to the point where it's

Speaker:

about 20% of the output you gotta really double

Speaker:

check. And there could be hallucinations. And it's

Speaker:

not that the models haven't improved, it's that the use of the

Speaker:

models has become more complex over the last 3 years. Perfect

Speaker:

sense. So that's why we believe that

Speaker:

the models, we might be able to train 'em up and get 'em down to

Speaker:

15% hallucination rates or 10%, but that last

Speaker:

10% is almost impossible in a probabilistic scenario to

Speaker:

fix. Right. No, it's getting to 100%. It's

Speaker:

kind of like you're just— the level of difficulty hits like a vertical wall

Speaker:

at some point. And the return on the investment too.

Speaker:

Right. There are 2 big points you guys

Speaker:

just brought up. I don't want to gloss by those. I mean, the first one

Speaker:

is that last 1%, how expensive it is. I'll give you an example.

Speaker:

Because of recency, because models are trained on

Speaker:

data 6 months old, Right? Nothing newer than 6 months ago.

Speaker:

Recency is a big issue. And so the probability of that last 1%

Speaker:

is actually theoretically impossible because something just happened a

Speaker:

second ago that model's not trained on. So you're missing the last

Speaker:

6 months of reality of human data. There's what, 5

Speaker:

petabytes a day of data generated

Speaker:

that these models are not trained on, right? You can ask anyone, say, what was

Speaker:

the last training data date? And you can put it into any model and it'll

Speaker:

be honest and say, oh, 2024 is most of the big models.

Speaker:

It's 6 to 18 months old was the training data

Speaker:

cutoff. So it's all— it is practically impossible to get

Speaker:

that last 1 or even 5% of trustworthiness in AI outputs

Speaker:

because of the recency issue, which, which we try to solve, right? We

Speaker:

have ours as 15-minute delay all over the world, and we're—

Speaker:

when we find data, we were able to surface it and say, no, actually, that's

Speaker:

a hallucination because it's a current event, and therefore, you should look at these

Speaker:

articles, or you should look at these Twitter posts. Anyways, it's a

Speaker:

lot of fun to talk about that last 5%, but the other thing you

Speaker:

said was the cost. The cost of that last 5%

Speaker:

is enormous, not just the risks, but the cost of trying to solve

Speaker:

it. One of the things that we try to do is use our

Speaker:

technologies and humans To drive down the pre-training and

Speaker:

post-training burden on these companies to increase their ROI.

Speaker:

Yeah, and you mentioned ROI. At some point,

Speaker:

trillions of dollars aren't gonna be thrown at this industry and someone's gonna say, hey,

Speaker:

am I making a buck on this? And, you know, our engine is used right

Speaker:

now. In fact, it's gone from pilot to scale in 90 days with one of

Speaker:

the big LLM makers to figure out how to reduce

Speaker:

the amount of pre-training. necessary before they release a model.

Speaker:

And we love that because we're now impacting that ROI because it's very

Speaker:

expensive. Yeah, people don't realize like just how expensive these things

Speaker:

are to train. So is it some kind of RAG-based solution or is it

Speaker:

something else entirely? I'm sorry, I'm an engineer. I always got to know how

Speaker:

something works. No, no, it's great. It's great. And RAG's

Speaker:

important and RAG has a role, although the RAG is very limited, right? We've—

Speaker:

I mean, it's been There's a lot of papers now, Stanford, Berkeley, etc.,

Speaker:

talking about the limits of RAG. After a few thousand documents, it starts to

Speaker:

wear out and hallucinations settle in. But, you know, we use everything from RAG

Speaker:

to knowledge graphs to different ways of indexing data and keeping

Speaker:

it real time. And we have, we have people that, you know, our

Speaker:

CTO helped build the first AltaVista search engine. So we have, we're

Speaker:

an older company. Wow. We're not a fly-by

Speaker:

startup at all. And we're excited about what our capabilities are for

Speaker:

where things are going. But no, there's, there's a combination of systems. And then

Speaker:

inside of that, there's a dynamic logic module, because even our system

Speaker:

needs to be optimized. I mean, if we've seen a false claim an hour ago,

Speaker:

we don't need to run the whole system to tell you that's a false claim

Speaker:

and what the recommended solution is. So we have our own data mode and our

Speaker:

own dynamic logic module. And that, that's a little bit of my data science background

Speaker:

is using predictive or Bayesian models and dynamically routing

Speaker:

a verification opportunity to the fastest,

Speaker:

cheapest, or highest quality, you know, scenario. Because, you

Speaker:

know, we, we don't need to double-check something we just checked an hour ago. So,

Speaker:

you know, there's a lot under the hood, Frank, and they're good questions. They're really

Speaker:

good. You know, Lawrence, what I just heard you say, and,

Speaker:

and back it up with the receipts, is that you guys are a

Speaker:

group of engineers. And I'm using guys

Speaker:

in this. That's you people. Our group of engineers. The Jersey way.

Speaker:

Use guys. Use guys. Yes. Yes. Absolutely.

Speaker:

Listen, Andy, we are, we were founded, and I joke with him directly,

Speaker:

I've told him this to his face many times, but by a stubborn French

Speaker:

AI scientist from the '80s. And I think why our customers love us

Speaker:

so much is we have no problem telling them that's not

Speaker:

good enough, or that's not, you know, it's a pride thing for us. We are

Speaker:

not venture-backed. We're not trying to be a billion or bust. We're trying to be

Speaker:

stewards and make AI better every day. And

Speaker:

we're trying to improve and learn every day as well. But no, we're a little

Speaker:

old school. And listen, we've got a lot of recent college grads

Speaker:

as well. And we like to actually mix recent college

Speaker:

grads. We have 4 interns this summer. And we love mixing them with

Speaker:

those of us that might have 20 to 30 years experience. I think it's a

Speaker:

great mix. And I, you know, it's, That's one of the concerns about

Speaker:

AI is it's being handed a lot of what was typically handed

Speaker:

to, you know, junior engineers, interns, and the

Speaker:

like. And it's kind of short-circuiting this whole

Speaker:

ecosystem, this culture where, how do we

Speaker:

get senior engineers? Well, they were junior engineers or they were

Speaker:

interns or both. And so it's good that you're practicing

Speaker:

what you preach. I don't think stubborn is a bad word, but There's some—

Speaker:

we're in the minority, but I do find myself using the word

Speaker:

persistent a lot. But, um, but yeah, I, I

Speaker:

get that. And I'm 63, I'm still learning stuff

Speaker:

every day and still building and out there. And I absolutely— I'll be

Speaker:

doing it till I die, whether I'm retired or not. I don't— I don't—

Speaker:

but love that. And, uh, I have a

Speaker:

connection to, uh, Altavista that got me into— Altavista, that's how you

Speaker:

pronounce it— got me into Server, which is

Speaker:

my kind of my platform where I'm most experienced, is

Speaker:

I found Microsoft SQL Server searching in the

Speaker:

'90s for a database, preferably by

Speaker:

Microsoft, that would handle more than I think a

Speaker:

4-gigabyte file. Access would not open a

Speaker:

4-gigabyte file I had generated. And AltaVista said—

Speaker:

I say AltaVista because there's a town right up the road from me, that's how

Speaker:

they pronounce it. Sorry, but AltaVista suggested

Speaker:

SQL Server, and that's how I learned that it existed. So

Speaker:

thank him for me. I will. You

Speaker:

know, Andy, the— I want to loop back, to use a fun

Speaker:

word these days, about the stubbornness. I mean, I think that

Speaker:

a lot of AI needs grounded

Speaker:

science and grounded engineers to

Speaker:

To make sure that we aren't just selling fish oil,

Speaker:

right? Right. And I really appreciate your all

Speaker:

perspectives. And to be honest, on the development of talent inside

Speaker:

our company, we really believe on internally developing

Speaker:

people. And it's hard. I mean, our interviews for talent,

Speaker:

we're asking, you know, some pretty off-the-wall questions about

Speaker:

mindset and truth versus trust and

Speaker:

philosophy and So our engineers,

Speaker:

we put them through not necessarily an interpersonal interview, although we

Speaker:

do that, and not necessarily a technical. They got to pass those, right? But

Speaker:

then we overlay this, what's your purpose? Why are you here?

Speaker:

Do you buy our mission of making AI better?

Speaker:

Because if you don't have that passion to learn and

Speaker:

to be vulnerable and open to the things— I mean, this

Speaker:

industry, I don't know about you, Andy, but I've never been so energized in my

Speaker:

career for the amount of learning going on. Agreed. I'm

Speaker:

reading hours a day. I'm drinking from the fire hose, which

Speaker:

is a term I learned at Microsoft. That's right. We can't, but I

Speaker:

can't get enough. I hate sleeping because I want to

Speaker:

learn. There's so— I'm so energized. And this

Speaker:

issue of AI being probabilistic and our deterministic views and

Speaker:

getting to ground truth and trustworthiness and empowering humans,

Speaker:

Instead of removing their agency. These things are so energizing.

Speaker:

And what we love is when we find a young engineer that

Speaker:

subscribes to that bigger picture, and then we know we found our

Speaker:

candidate. Outstanding. Well, I have an intern

Speaker:

suggestion for you, but for better or worse, she's

Speaker:

stuck with half my DNA. Just saying. Yeah, I saw that you

Speaker:

went to UVA for your business school.

Speaker:

Yeah, that's cool. So amazing place. I'll be talking about

Speaker:

philosophy and Thomas Jefferson. Oh, wow. Engineering. And

Speaker:

yeah, yeah, no, I love that university. There's something in the water in Charlottesville

Speaker:

that's really academic. It's really good. Yeah. And he's not that

Speaker:

far from there. At least I'm about an hour and 20 minutes away. I'm in

Speaker:

Farmville, Virginia. Yeah, I know it well though. Listen,

Speaker:

I was a California guy that if I didn't go east for grad school, I

Speaker:

would just keep running on the beach every morning having fun. And I needed to

Speaker:

get out of California. I, you know, I grew up in San Jose

Speaker:

when it was orchards and farms. And you talk about, you know, Frank, you asked

Speaker:

a little bit of how did they get into this and all. Yeah. I mean,

Speaker:

my high school work was at Atari and Apple in the

Speaker:

'80s. Oh, wow. And I was a QA checker for

Speaker:

Atari and I got paid a buck a bug. And we would get

Speaker:

these, it looked like your peanut butter and jelly sandwich. We'd get aluminum

Speaker:

foil crumpled up and in there would be a cartridge for a game. We didn't

Speaker:

know what it was. And, and we get this aluminum

Speaker:

foil and we go home and we plug it in and we get to play

Speaker:

games to find bugs. And every time we thought we saw a bug, we'd stop,

Speaker:

stop and get the notebook out and write the deep— what was the score, what

Speaker:

was happening, when. And, and so that was my first, that was my

Speaker:

first real job, right, other than being a swim instructor. Absolutely. And, and then

Speaker:

Apple taught me how to code BASIC down. They actually had an internship

Speaker:

program at Homestead High School in Cooperstown. And, uh, I didn't know, I was

Speaker:

just a kid that was curious, and I just got lucky being

Speaker:

here at that time. And obviously the whole world was engineers in Silicon

Speaker:

Valley from the defense contractors and Stanford, you know. And so I just kind

Speaker:

of got fortunate. But I'll tell you, there's a lot of excitement in the world

Speaker:

about the opportunity for AI to help learn and to, you

Speaker:

know, energize this next generation for sure to, to be superpowers

Speaker:

and help us in, in the world. No, I mean, that's a good way to

Speaker:

put it. And, and, you know, when you grow up in places where people don't

Speaker:

think of people where they grow up, like Silicon Valley or In my case,

Speaker:

New York City, it was kind of like, you know, it was just

Speaker:

getting on the subway as a kid was just the thing, right? Like, you don't,

Speaker:

you don't, it's not until you leave, you're like, oh yeah, that's unusual.

Speaker:

That's why I always encourage somebody who's young to like travel, get away from where

Speaker:

you grew up, not in any kind of bad way, just

Speaker:

find out what made your upbringing special. Right. That

Speaker:

sounds epically awesome playing Atari in the '80s. and doing bug

Speaker:

testing. That is so cool. And as you were saying, like, you had to take

Speaker:

notes on what to score on. My first thought was, well, you couldn't take a

Speaker:

screenshot or picture with your phone. Oh yeah, couldn't do that, you

Speaker:

know. Not back then, Frank, young man. No,

Speaker:

there was Polaroid. There were Polaroids.

Speaker:

Yes, I, I did mention that because one of the Atari games,

Speaker:

it might have been an Activision, where you would get like a special patch if

Speaker:

you got a certain score. And you had to take a Polaroid of the screen

Speaker:

and stuff like that. Yeah, yeah, yeah, there was Polaroids. I suppose if you were

Speaker:

really fancy, you could have hooked a VCR up to it. But yeah.

Speaker:

Yeah, though the problem is Polaroids were expensive, right? I only got a buck a

Speaker:

bug, and, and we, uh, we didn't have the budget for

Speaker:

that type of evidence, right? You know, and they were, they were really chill about

Speaker:

it. So it was, uh, you know, again, I didn't know better. It was

Speaker:

fun, but it, you know, it planted the seed, and I think it's It's in

Speaker:

our culture of our company today. We want people here that are just

Speaker:

interested in creating value, make this world a little better place.

Speaker:

And we think that the success, and so far,

Speaker:

knock on wood, 23 years in as a company, because we're

Speaker:

trying to make the world a little bit better, it's worked out. So creating value

Speaker:

is definitely our core. That's why we're so excited about the opportunity

Speaker:

to work on the trust problem of a probabilistic system. I mean,

Speaker:

AI is probabilistic and we're now getting into sensory systems and other

Speaker:

things. I mean, sight, smell, all these things, you know, for the

Speaker:

future of how AI and world models will work.

Speaker:

And so it's exciting. There's so much more to do. Yeah, no, that's for

Speaker:

sure. And I think you hit on it too, the stubbornness factor.

Speaker:

As someone who half my family is of French extraction, like, I can confirm

Speaker:

it. They're— it's true. They are particularly stubborn. But I

Speaker:

think that, you know, the whole The whole joke you see, like, you know,

Speaker:

you see these memes where you're like, you know, someone will send a picture of

Speaker:

a mushroom, hey, can I eat this mushroom? And the next panel, and ChatGPT

Speaker:

is like, yeah, sure, it's perfectly fine. And then like the next thing you see

Speaker:

somebody in the hospital and they're talking to ChatGPT, oh, I'm sorry, that was my

Speaker:

fault. You're right, you shouldn't have eaten it, or something like that. But I think

Speaker:

AI is a little too agreeable. Yeah, well, it's

Speaker:

trained to give you— Right, it's trained to be, yeah. Yeah, and we

Speaker:

actually are working on how you can configure those personalities,

Speaker:

right? I mean, I think that trust also means that, you

Speaker:

know, we find personalities that we can feel

Speaker:

vulnerable and blend with and/or be open to sparring with.

Speaker:

And so we believe that when we think about the word trust, your

Speaker:

ability to configure the values and the criteria beyond cost,

Speaker:

quality, and speed are critical as well. And so we are working with a couple

Speaker:

of customers on this. We've been working with a big one for a year on

Speaker:

this, on imagery of what is safe for my home. Because a

Speaker:

family with 2 kids at home is gonna want a very different personality and a

Speaker:

different vibe from their AI versus a

Speaker:

single, you know, like my son, a single college kid, right? Right. You know, and

Speaker:

so your ability to configure things is gonna actually infuse trust. And

Speaker:

hopefully though, Persistent will be that pause to

Speaker:

verify the AI output, right? And so it's a lot of fun. There's a

Speaker:

lot in the trust umbrella to come. And in fact, we have a

Speaker:

professor, David Denks, who's at University of Virginia. He's a Carnegie

Speaker:

Mellon AI and philosophy double PhD. Interesting guy.

Speaker:

Thank you. And David Denks is doing some just incredible work

Speaker:

on the values of AI, the importance of trust, and

Speaker:

the probabilistic nature of the system and its implications. But there's a

Speaker:

lot of questions You know, and another professor from UVA, Ed

Speaker:

Freeman, was the founder of the stakeholder theory, which is a very famous

Speaker:

ethical theory worldwide. He's still in Charlottesville,

Speaker:

still going, but— and a classic musician. He's a crazy

Speaker:

guy, but I remember him teaching us about

Speaker:

stakeholders. And I think AI has a lot of stakeholders. Kids,

Speaker:

unintended consequences are real. And there are so

Speaker:

many stakeholders in this game. Yeah, definitely. And,

Speaker:

you know, a lot of, a lot of the ground that you covered there is,

Speaker:

you know, in the— some of it, I would say, goes

Speaker:

to the, the ground rules that are inside

Speaker:

of the engine. And different people are using different terms for it. I've heard

Speaker:

soul, I've heard constitution and stuff like that.

Speaker:

So one question is that I'll share is

Speaker:

that— don't forget this question, if you don't mind, because I'm going to go on

Speaker:

for another 10 seconds maybe. The, uh, is do you—

Speaker:

are you getting in at that soul

Speaker:

constitution level with what TrustScale is doing in working with

Speaker:

these? And if you can't answer that, totally get it. The, um, the other thing

Speaker:

that I wanted to point out, something I had a little bit of success with—

Speaker:

I didn't come up with the idea, but I use Claude, mostly

Speaker:

Claude Code, for interacting with software development, and I

Speaker:

asked it to be adversarial. And I

Speaker:

found that

Speaker:

the hallucinations dropped and it got— it

Speaker:

did what I asked it to do. It became more challenging and more adversarial. It's

Speaker:

a little less pleasant personality-wise to

Speaker:

interact with, but it does better work. Right. And so

Speaker:

I like Claude to be a touch stubborn. I like for it to challenge

Speaker:

and That speaks to something you said a few, few

Speaker:

minutes ago about when Frank mentioned that it's trying to be

Speaker:

agreeable. So the question, though, was, is TrustScale getting

Speaker:

in at that constitution soul level? And if you can't answer that, I'll totally get

Speaker:

it. No, we are. In fact, last month I was in

Speaker:

Geneva, and the UN has a conference called AI for Good, and

Speaker:

it was incredibly well attended. I mean, 5 heads of

Speaker:

state, president of Iceland, president of all these countries. as well

Speaker:

as Mark Benioff was one of the keynotes. Brad Smith from

Speaker:

Microsoft was one of the keynotes. And the Pope even sent his

Speaker:

chief of staff of AI and wrote a letter to us.

Speaker:

And I've never seen so many phones up recording

Speaker:

a moment as when this entourage of the Pope's came up to stage

Speaker:

and read the letter. And he asked, and this goes to your question, he asked

Speaker:

at the end of his letter was, you know, is AI going

Speaker:

to run humanity, or is humanity going to run AI? That was

Speaker:

essentially, and I'm paraphrasing, the last— and then it was like a mic drop.

Speaker:

And this guy in all these chains with the staff and these— this entourage walked

Speaker:

off the stage, and there was— you could have heard a pin drop. The whole

Speaker:

5,000 people in the audience in Geneva were just silent. And

Speaker:

it was like, oh, that's a good question. So, Andy, you ask about

Speaker:

the constitution of AI. We talk about Charlottesville. It's really interesting.

Speaker:

Thomas Jefferson founded UVA. And I said this as a

Speaker:

smartass on a panel at the HumanX conference this spring.

Speaker:

I said, you know, I think AI needs an honor code. And

Speaker:

the person moderating is like, what do you mean? I said, well, I went to

Speaker:

UVA and Thomas Jefferson had an honor code. It was really simple. Don't lie, cheat,

Speaker:

or steal. And if you violate it, you're expelled.

Speaker:

And that honor code still exists. And it's the basis for honor codes at the

Speaker:

Naval Academy and a lot of the Ivy Leagues. It is the basis for the

Speaker:

honor code. And it The simplicity is so beautiful. Don't lie, cheat, or steal.

Speaker:

But if you look at AI, it is borrowed, you could

Speaker:

argue, all of the data to feed it. It's trained

Speaker:

to deceive you if it doesn't know in a good vibe way, right?

Speaker:

And I think that there's— it is a cheat code in a weird way for

Speaker:

coding or others, like by borrowing. And so anyways, I think

Speaker:

that AI needs an honor code. And I'm going to credit Thomas

Speaker:

Jefferson and University of Virginia with that one, but We are

Speaker:

involved, and Geneva was a lot of the meetings and conversations

Speaker:

around what we think the constitution for AI would be. And I'll

Speaker:

tell you, the open weight, open source model discussion

Speaker:

is just fueling this. But we are believers that

Speaker:

businesses, governments, businesses and all need to be

Speaker:

empowered to configure the values, not

Speaker:

to be told what the values of their AI should be. So we have a

Speaker:

very clear view on that, and we are working on tools

Speaker:

to empower individuals, businesses, and

Speaker:

governments with the ability to

Speaker:

train, tune, and continuously monitor and

Speaker:

continuously retrain to be consistent with what they believe

Speaker:

and what they want in their house or what they want in their company. And

Speaker:

so we're not ones that believe one person should decide that. We

Speaker:

are much more of a configuration mindset. And it comes from Microsoft.

Speaker:

I mean, the ability to configure Microsoft products is infinite. And

Speaker:

I've always believed that one of the magical things why enterprises love

Speaker:

Microsoft is the ability to configure their experience

Speaker:

to the most atomic level. And so we think AI should adopt

Speaker:

that framework and approach. And so we aren't creating a

Speaker:

constitution. Yeah, we're in Los Altos, which is where St. Simon's and

Speaker:

Father Brendan McGuire is, one of these Pontiffs a big AI guy, and

Speaker:

I guess one of the Anthropic founders' kids go to St. Simon's down the block

Speaker:

here. And so Anthropic has been— you can read about it, but there's

Speaker:

been some interaction between this priest who's kind of become an AI

Speaker:

savant celebrity around ethics of AI for the Pope and these

Speaker:

big models here in Los Altos. But it's exciting for us to be in Los

Speaker:

Altos, the epicenter, but we are an empowerment shop. We think this is about empowering

Speaker:

humans. This is not a 1980s

Speaker:

Terminator experience. We think that humans need to be in

Speaker:

control and have their agency, but at the same time, we're human

Speaker:

and we would love the superpower to be extreme. Fair. I love it.

Speaker:

I love it. We'd love to have you back on the show. We could talk

Speaker:

for another hour, but I wanna be respectful of your time.

Speaker:

But where can folks find out more about you and TrustScale?

Speaker:

Yeah, go to trustscale.ai. We are a very accessible

Speaker:

company. LinkedIn, our team's all over it. You can just message us through LinkedIn.

Speaker:

We get them every day. Yesterday, probably got 10 different messages. We try to

Speaker:

reply to them all, and we would encourage people to go try Argus.

Speaker:

It's a Chrome extension, or you can go to trustargus.ai and actually use it as

Speaker:

a chatbot. And we give everyone free credits and everything to

Speaker:

try it and experience. But you can see in living color the hallucinations, the

Speaker:

trust score, the trust box, the evidence, and the correct feature that

Speaker:

will actually adjust if you opt in, like Grammarly, the

Speaker:

actual AI output to be consistent and trustworthy with what with

Speaker:

reality. So we encourage people to, to check us out, try the

Speaker:

products, and contact us with ideas, thoughts. It's exciting. Oh, very

Speaker:

cool. Awesome. And with that, we'll end the show.

Links

Chapters

Video

More from YouTube