Artwork for podcast Cognitive Engineering
Good Judgement
Episode 50 • 10th April 2017 • Cognitive Engineering • Cognitive Engineering
00:00:00 00:30:12

Share Episode

Shownotes

Nick, Peter and Fraser speak with special guest Michael Story of Good Judgement Inc about the uptake of good analytical practice.

For more information on Aleph Insights visit our website https://alephinsights.com or to get in touch about our podcast email [email protected]

Transcripts

Speaker A:

Hello and welcome to the Cognitive Engineering Podcast produced by Tell Me Studios for Aleph Insights. In this series of podcasts we take a look at interesting topics and discuss what we think they tell us about analysis and decision making. I'm Fraser McGruer and I'm here with Nick Hare and Peter Coghill of Aleph Insights and this week we've got a special guest with us, Michael Story of Good Judgment Inc. And this week we're discussing the impact of the Good Judgment Project. So Michael, just starting us off, tell us about yourself.

Speaker B:

Well, so I currently work for Good Judgment Inc. which was the successor to the Good Judgment Project, which is something I joined as a grad student. So I was studying at the LSE and somebody forwarded me an email, he forwarded me an email, which was inviting people to participate in a research study looking at political forecasting. I clicked on that email, joined the project, collected a book token as thanks for participating and I just gradually got more and more and more involved in a project that was trying to investigate how you could make more accurate forecasts about the future and I ended up working there and now it's my full-time

Speaker A:

job. Okay. What was your grad degree in? What was it? Public policy research. Okay. Nick?

Speaker C:

Yeah. So the reason I've asked Michael to come along is just in the spirit of disclosure, we were both in the Good Judgment Project and we were both in the group that ended up being called super forecasters who were people who for using a variety of different techniques and approaches were able to just produce very accurate forecasts. But anyway, the Good Judgment, the findings of the Good Judgment Project have been pretty well covered. I mean, you can Google it. There's a book, Super Forecasting by Phil Tetlock and Dan Gardner, which I highly recommend. But actually, the reason I've got Michael in is because I think I'm interested in his firsthand experience of trying to get people to do something about the findings. The findings broadly give us the recipe for being a better forecaster and now Michael's job involves trying to get people to implement those findings. And I suppose the question is, what impact is it having? How are people receptive or resistant to the findings? How's the project to improve the world's forecasting going?

Speaker B:

That's the question you want me to answer? Oh yeah. How's it going? That's where you come in. Well, so yes, it's been very interesting. So the key discovery of the project is that forecasting is its own specialism, right? And we should regard that as a kind of independent skill that you can have that's separate to the expertise that you may have about the area that you're forecasting. So traditionally, most organisations that do forecasting will have people that are regional experts or subject matter experts on different things. And they will assume that what if you know lots and lots about say, China or contemporary Chinese politics, that you will be the best person to ask to deliver a forecast about what Chinese politics is going to look like next year. And you will have obviously some ability to do that. But the idea that that there's a kind of separate set of skills that you might need that will help you make that forecast be more accurate and actually that separate to your specific knowledge about China. And that's the bit that the Good Judgment Project was really concerned with is, well, how do you do that bit better? And what can we say about that? And what we can say is that forecasting is a kind of skill. It's something that everybody can get better at, everybody can improve. Some people are better at it than others. And it's a kind of independent, separate component to your subject matter expertise.

Speaker A:

Well, let me ask some questions. So who is it for?

Speaker B:

So originally, the project was funded by IARPA, which is the Intelligence Advanced Research Projects Activity, which is a sort of sister agency to DARPA. So they fund any researchers of interest and value to the intelligence community in the US. So initially, they're looking at obviously, that's a big branch of the state bureaucracy that produces forecasts. And so that was the primary interest. We now still have the US government as a client. But we're now able to do things outside of that, following the conclusion of a kind of exclusivity deal that we had with IARPA. So the last year or so, we've been able to go outside of that. So partly, that's foreign governments. So people, you know, in a similar sort of role, you might have interest in doing this a bit better. But it's also corporations, big organizations. And also, we're starting to look at humanitarian organizations as well. So we're building up, we've done a bit of work with a fantastic organization called Start Network, which is a conglomerate made up of contributions from lots of big humanitarian agencies like DFID as part of it, Red Cross, Save the Children, these big organizations. And they want to get better at forecasting disasters, the responses to disasters, how much they're going to need to build things together. So we we're branching out gradually. So it's governments, big organizations, humanitarian stuff, and then

Speaker A:

kind of eventually the public, hopefully. And so what sort of response have you had to your findings? And also, have you found that people have come to you already with a kind of a prejudice, thinking about what you will already think about them? If that makes sense?

Speaker B:

Yes. So that's been quite interesting. So our project got a lot of press, particularly last year when the book Superforecasting came out, which was a which kind of had quite a big splash. And that created some perceptions that were a little bit off compared to what we wanted, the lessons that we wanted people to learn. So the big one being that the crowd component kind of got lost a little bit. So we're saying that if you have lots of analysts doing something, there are systematic ways you can get them to be better forecasters individually. But there are also lots of systematic ways you can aggregate those forecasts to be more accurate, right? So you can take a simple unweighted average, and that's pretty good. But there are other things you can do about, you know, how you weight the averages and which people you pay more attention to and things like that, and the balance between those things. So, for example, more recent forecasts, you should give high weight to even if that person's a bit less reliable, because they've had more information, or if that person's got a stronger track record, you can give them a higher weighting generally. And that's a kind of a new a bit of a new finding that if someone's generally good at forecasting, they're generally good in all domains. So someone who's generally been right on China is more likely to be right on Egypt than perhaps an Egypt expert who's not got such a good track record. So those things are very interesting, both the individual and the collective. But the press coverage really focused on two things. One was the kind of individual stories. So they were very, very interested in how do I as a as a single person get better at forecasting, and the kind of aggregation side got a bit lost, which makes sense. I mean, very few people have access to big banks of analysts and how you, you know, structure aggregation between analysts is a little bit, you know, a little bit more on a big subject. But the other focus was really on this idea that the super forecasters who were the most accurate individual forecasters from the Good Judgment Project, these are people who signed up like I did. And like Nick did, you know, for a book token to volunteer to be in a study and make forecasts, the people who are most accurate and were most accurate across all domains, a lot of the coverage focused on them being like the regular Joes that beat the experts. So I when I started handling some of the press requests for Good Judgment, they would always say, who's the most ordinary super forecaster, that's who we want to meet. Now I understand why that is, because of course, it's a more interesting story. And we do have a few people with odd with odd type of careers. And but then that would tend to produce press that argued, well, you know, experts are all useless, expertise is not valuable. And look, these ordinary people turn out to be to be better than the expert. And there is a really good claim there, right. So there's a report which has not been released, but was leaked to the Washington Post, comparing the performance of Good Judgment Project against aggregated forecasts from the intelligence community in the United States. And the Good Judgment Project's forecasts were more accurate, about a 30% reduction in error. That's according to a leaked document, it's not official, but that was in the in the post as a according to a leak. So there is something about about that, right, there is this idea that open source stuff can be useful. But the idea that the super forecasters are themselves not experts, like most super forecasters are, are very well qualified in things half of the supers have a PhD of the remaining half, half of them have a master's, you know, I'm among the least educated of the of the super. So it's it's not really the case that that it's like, you know, a lot of them are in the type of things where you expect to find people with who know what they're

Speaker D:

doing. Okay. Nick, Peter, so the press, perhaps concentrating a bit on the individual underdog story, missing some of the key points of the study. Yeah, definitely. There's a obviously,

Speaker B:

everyone loves an underdog, everyone loves a maverick. So we had supers, there's one, I think, who runs a tractor dealership, there's one who's a sports coach, things like that were like, were big stories, right, that out of 20,000 people, you pull out the most accurate 150. And it turns out that some of them are kind of in wacky careers that you wouldn't expect to be making very accurate geopolitical forecasts. But they're unusual, most of the people are exactly who you'd expect to be good at making geopolitical forecasts. I think we should probably add that

Speaker D:

this document, comparing performance of professional intelligence analysts, and Good Judgment Project people wasn't leaked by the Good Judgment Project, or Good Judgment Inc, presumably was leaked by somebody in the client organisation. Yes, absolutely should make that

Speaker B:

clear. Yes. No, we don't know where that came from. It ended up in the Washington Post. But

Speaker C:

we don't know. Favorable, though it might have been a good judgment project. I suppose my I mean, I should sort of say, if people don't know, my background is from the intelligence community, I used to work in defence intelligence. And in fact, one of the reasons I found out about the Good Judgment Project was through I just I read an article by Phil Tetlock talking about forecasting at the time that I was running a team whose whose role was it was to try and implement better ways of doing intelligence analysis. So it's fantastically relevant as far as I was concerned. And I found it quite hard to really to make other people within the intelligence community see the relevance. I think partly there was there was a an unwillingness to believe the results. I mean, that the business model of intelligence analysis production all the way to the top is entirely based on knowledge on human knowledge, you know, okay, we got the guy who looks at Syria, and he's been looking at Syria for three years. So he's the guy we're going to get to write our intelligence reports. And that includes forecasts, you know, so there's a kind of you can see there's a sort of self serving element. Well, if this is true, we're going to have to change everything about the way we do business. But I wonder if I wonder if that's if there's something more fundamental, which is that actually, you know, I sometimes wonder, even though from a decision theory point of view, being right about your forecast is better than being wrong, whether actually being right in your forecasts is really what forecasting is about most of the time, or whether actually, what most organisations want is just to sort of do things and say things, and they're not too bothered about being right. And I just wondered, because you've had that experience with admittedly with people who are coming to see you, but whether or not you've you've experienced, you know, resistance and whether or not actually people really do they really care about getting

Speaker B:

it right? Well, that's a very interesting question. I think broadly, the people that come and see us do care about getting it right. But they may be at the end of their organisation that does care. So a typical first contact for good judgment will be from head of data, head of research, head of whatever, that's normally the title of the person that gets in contact. So they do tend to have a real interest in forecasting research, and they want to they want to get better. Broadly, I think a lot of a lot of senior decision makers do do want to get better. They don't necessarily care how it's done. They're not really that interested in legitimately say, right, they're not interested in the actual mechanics of it. If you unless you find the science of forecasting interesting, you just want information that you can act on. And I think that a lot of organisations are pretty receptive, right? If you say, okay, what do you currently have? Well, you have a tree of people who report to each other, with all of the biases that are thrown in from individuals, from the relationships that can throw things off. I mean, you know, you've got people telling each other what they want to hear, and all that sort of stuff, going up the chain. And that becomes a written report that says, here's a few reasons why we think this might happen. Here's a few reasons why we think it might not happen, decision maker, you decide. And we say, well, we can set something up where you can turn that into probabilities that you can be more confident in, and we can say we think we can rank these scenarios and how likely they are. And therefore, you can allocate your resources, it just it speeds up that decision time. And so they tend to be quite receptive. But it depends on the organisational structure. And that's another finding of the project that I think has really not been particularly well explored elsewhere, is that the the environment in which people make their forecasts, or the architecture really matters. So one of the key things is the ability to, for example, to be anonymous. So you want to be graded on how accurate your forecasts are, but you don't necessarily want everyone to know what your particular forecast on a particular topic is. So when you so that avoids this kind of yes man problem, you don't want to be seen to be agreeing with the boss. So you but you want at the end of the year, for example, all your forecasts to be treated in aggregate and say, well, overall, you are more accurate, but at the time, it may be very difficult for you, for kind of internal political reasons to give you a real view. So you need to kind of lop off these things that will cause problems. So there's individual level biases, but there's also kind of incentives that might cause you to, to not reveal your true opinion. And, and all of these things can go up the chain.

Speaker A:

And I know we're talking about forecasting, but forecasting what, what kind of stuff

Speaker B:

does it concentrate on? So in Good Judgment Project, it was anything that the intelligence community would be forecasting, right? So they want it to be as close as possible to, to what to replicate that yet what the client is paying for. So it could have been anything. So in it was, you know, is North Korea going to test this missile by this date? Will Ebola spread beyond this, these levels, bird flu, anything that that requires the government to act and that they would sort of need to know about to, to act on it. And we, the really fun part of the project was discovering that there's a quite a tight correlation between your ability to forecast pretty much anything. So people that are good at bird flu were good at North Korean missiles. So, so that kind of generalizability of those techniques means that that's, that's good news if you're going out to, to corporate clients or other organizations, and they've got their own things they want to forecast. The general structure of it, though, is quite particular, like you need to be careful to make sure that you're forecasting something that is, where the definition is very clear, because if the value is in aggregating the forecast, you want to include this diversity of opinion, you want to get the disagreements, you want to kind of use these variations from the median to, to lop off inaccuracies when you're aggregating from a group. What you don't want is for people to be giving different answers to the question, because they've interpreted the question differently. So you want to make sure that any disagreement among your analysts who have, is down to their having a different, likely scenario in mind, rather than that they've just read the question differently. There are some ways you can get around that. But, but this is the kind of key component. So it's not necessarily the subject matter. But the structure has is somewhat limited. And there are things you can't forecast,

Speaker C:

because of that. And, well, yeah, no, I think this is actually, I think, from taking part, the most interesting lesson, and actually something which is very practical, for use by organisations like intelligence analysis organisations, where their job is to try and form true beliefs about the world, was the was the rigor relevance trade off, which I think, you know, is has been talked about, but actually is really fundamental. You know, what the kinds of things people ask intelligence analysts are, you know, is, is China getting more aggressive? You know, is, is, is Iran going to try and be more pro Western? And actually, these things, it turns out are really, really hard to define. And, and, but when you try and define something very specific, it turns out, actually, a lot of the time, that's not necessarily going to be terribly relevant, you know, will, will Iran agree to some, you know, new nuclear treaty or something. And, and the other surprising thing, I think, is how, how you can really, you can feel like you've nailed a concept down, and then some weird edge case will always come along. And the famous examples, I think it was the example of the of the fatality in the South China Sea, where, you know, it was, it turned out to be, I think, a ship got rammed, and a fisherman fell overboard and died or something, whereas people were naturally thinking about a, you know, a firefight. And there was also the case about US intervention in Syria, I think one of the questions, where it really, really tight definition of, you know, not troops being somewhere and fatalities being caused. And, you know, and then it, we realized that none of that quite captured drone activity, you know, the use of drones, and you think, well, you know, it's really, really hard to be specific, it's almost like some kind of fundamental level of indeterminacy at the level of definitions. But, but I mean, that, to me says, well, we need to take that really seriously, as an organization, actually, just being more specific about what you mean, is can deliver really big benefits. Because you're taking away one of the one of the key problems, which is a disagreement about what we mean by conflict, or threat, you know? Yeah, if you can lock that off, if you can

Speaker B:

eliminate the those disagreements, and you start to be able to apply some of those techniques that you get from aggregating a large number of people, so you you you isolate what the disagreement is about, right? And there are a lot and a lot of the success of the project was about doing that. So it was about trying to, to find out sources of disagreement to share them to say, okay, if you have, you know, a dozen people, and two of them think this thing is very likely, and some of them don't think it's very likely, yes, if they're disagreeing about the meaning of the question, that's not useful. If they're disagreeing, because they've got different models and different theories about how that how they expect things to work in that region, or things in that domain to to function, then that that is useful, you want them to hear those things and kind of flush out

Speaker D:

what the disagreement is and discuss it. I'm interested, how much so how well do you think you're penetrating this potential market of people who would benefit from the advice and the services that you provide? Because I'm thinking, I'm thinking to Ted Locke's earlier book, Expert Political Judgment, where he would talk about the pundits, and that's a big market of people who, who get away with making pretty shoddy forecasts. And partly because the people, their clients don't care very much, because they just want a nice, sexy sounding sound soundbite to put on the news. But they, but, but so many of them were quite resistant, actively resistant to this concept of this something they could get better at, because they were in a, they're sitting pretty in a unique position where actually, they don't want anyone to challenge the way they

Speaker B:

do things. Yes. So in terms of our market, there aren't very many of us, right? We're a pretty small organisation. So we can't say yes to everybody that comes to us, which is a nice position to be in. But, but it's, that's just the reality, right? So we, we tried to pick projects that we feel are have a kind of long, you know, where it'd be a long term relationship, and also that it offers some interest and meets, you know, personal, you know, developmental goals, right? Like, there's a, there's a very well known article about Palantir, who, you know, provide forecasting for the US government. This, I think, started from the Air Force funded project. And they really had trouble retaining staff when they started to expand outside of their traditional background. So they'd hired people out of college saying, well, you, we're going to help try and catch bin Laden. So would you like to work on that? And of course, what new graduate isn't going to want to do that. But then you expand and you take on a commercial client and you say, okay, you know, we want you to try and flog this new brand of supermarket cola, can you give us some forecasts that are going to be helpful in that, and you don't quite get the same level of commitment from people. And I think from Good Judgment Projects, original recruitment of forecasters, that was a really big component, right, that you're helping make this thing better, you're helping the government make better decisions. And there is a real sense of mission there. And that's still there, right. And people really do still feel that that's important. You want there to be this better process of decision making, you need, you want to feel that that's happening. So there's more interest than we've got. Eventually, of course, we'd like it if everybody adopted those views. I think everybody would, right, we'd all be better off if there was a bit more accurate forecasting around. Sorry, I mean, this might have just been

Speaker A:

covered, but I just want to get a better sense of this. Has there been a lot of resistance to your

Speaker B:

findings? Yes, well, so I divide that, I divide the resistance to the project into two groups. One is where they've kind of slightly misinterpreted what the project is saying. So this idea that experts are all idiots and all that sort of stuff. So that's not true. We don't think that, but people who think we think that are resistant to us until they find out we don't. And so yes, and the second part of resistance to the project has come from people whose position is perhaps a bit threatened by this more rigorous approach, right? So the people who really benefit from the status quo of forecasting, which is not very accountable, not very rigorous, are people who are themselves perhaps not very accountable, not rigorous and benefit from that. But in

Speaker A:

institutional terms, you haven't had any resistance, it sounds like, unless it's to the former.

Speaker B:

Within an institution, you can have that, right? So, so we, I mean, I was with a colleague, and we gave a presentation to a very eminent institution that does forecasting. And we talked about some of the findings, in particular, some of the personality and so on components, right? And we said, one of the things we found is that there's no relationship between extroversion, social status, etc. And forecasting accuracy, you can be a shy, retiring, brilliant forecaster, or have a gregarious one. And, and a very high status person immediately busted in and said, what are you sure? I mean, because, you know, and so that so somebody in that position, who's perhaps blustered their way a little bit, or who's a little bit high status, and benefits from being taken very seriously, perhaps without any accountability is going to be threatened. And there are some good examples of that Robin Hanson talked about when he tried to apply more rigorous forecasting methods in for the US Missile Test Agency that test strategic missiles. And they had some problems into integrating some of these forecasting techniques into that organization. Because it's not very nice if you're a senior decision maker to find out that the proposed date by which you're going to complete some task is unlikely to succeed, right? And your staff are kind of anonymously telling you that I think that's very likely. So there are individuals within organizations rather than necessarily organizations themselves.

Speaker A:

Sounds very leveling your project or yeah. Look, I've got one question. I'm not gonna ask it now. But I've got one question that I want to round things off with. We don't have much time left. Peter.

Speaker D:

Yeah. So it sounds like there's this competing incentives. So you have the I'm thinking again, of the pundits who, who make a living of not being very rigorous, just coming up with the soundbite versus their customers, perhaps who should, who ultimately care more about how accurate those forecasts are, because they're making use of those to make decisions. So would you agree that there's, there's a strong to encourage general adoption of the good print, the principles explained by Good Judgment Project. There's a role for the customer to demand accountability, perhaps even demand a portfolio track record and not hire those pundits who don't keep a record of how good they are to actually flip the incentive the right way to get best

Speaker B:

forecast. Yeah. And for people to demand better, they have to know that it's possible, right? Because most of the time, if you request reports, or if you go to a consultancy and ask them to do something, you won't get back numbers. So it's kind of revolutionary. We're so used to it because it's hard bread and butter, right? That we don't even think about the fact that for a very large number of people, they commission an analysis of a situation, there won't be any numbers in it. They'll say, well, I want to know whether my investment in such and such a country is safe, they won't get something back saying, we think there's a 60% chance your oil field is going to be nationalised in the next five years, right? They'll say, oh, here's a list of reasons why that could happen. We think it's reasonably likely blah, blah. And so knowing that there's an alternative is the thing that probably drives it up, right? Once everyone's, you know, it's like in the 80s, when nobody had, you know, no one had extra virgin olive oil. But once it's there, if your supermarket doesn't provide it, you go somewhere else. So once you create the expectation that it is possible, and it should be normal that if you commission something, you get back something you can actually act on, right? You can say, 80% of my planning is going to go on this thing that's

Speaker C:

80% likely. Okay, quick question, quick answer. Well, more of an observation, really, which is this is exactly, I mean, this idea of having a way of characterising where a forecast comes from is a fantastic gift that the Good Judgment Project has given us. From my direct experience, one of the things that analysts struggle with is being able to justify a forecast, you know, they want to say, well, it's, you know, there's only a 10% chance of, you know, Assad stepping down, but they don't have a framework with which to express where that comes from. Normally, it's, you know, they run a simulation in their head, and it seems to happen about one in 10 times. And that's where 10% comes from. The great thing about what the Good Judgment Project has shown is that you can break it down into some very simple and intuitive components and say, you know, there's a base rate here, here's what my base rate is based on, here are the exceptional circumstances about this situation, which means that, you know, the probability is going to be a bit higher. And it gives you a way of justifying a forecast. So, you know, you no longer have to rely on your analysts as a black box with a number coming out, or indeed, no numbers at all, you can see what the provenance of those numbers is. And it's, you know, it's fantastically useful. I mean, it's a great tool.

Speaker A:

Okay, well, you may have preempted my question there. We don't have anything, time for anything beyond what Michael is about to say. Let's just go back to the original question, what has been the impact of the Good Judgment Project? So can you answer that question for me, Michael?

Speaker B:

A big, well, the Vote Leave campaign made everybody who worked on it read a copy of Superforecasting. So you could, to some extent, we could attribute Brexit to the Good Judgment Project. Pat on the back. And they were quite keen on that. I think the initial impact of the book was quite strong in the kind of popular sphere. But in the year since then, there's been a bit more of a trickle down of the ideas into institutions. So I think a lot of individuals read the book, it was fascinating, got a lot of press coverage. But it's this year that we've started to have more inquiries from institutions. I think there are a lot of institutions looking at this, as some of the kind of flashy bits about experts being useless and so on have kind of dissipated, and they've realised, okay, there are some practical things we can do. So lots of institutions are starting to say, is there some way that we can more formally improve what we're currently doing? And the fact that we're still going, that we're still producing research and so on, this kind of flash in the pan thing, didn't apply to us, and our project is a bit more rigorous and had good statistical power for the answers, is standing in good stead. So the impact broadly has been that people have started to realise forecasting is its own skill, and that there are systematic ways to make that better with whoever your staff are. It's not about throwing out your experts, and it's about using them better. Quick question. I mean, you can go

Speaker A:

off and study Middle East politics for a degree, right? Does it already exist where you can do a

Speaker B:

degree in forecasting? Well, there are master's programmes and things like structured analytics and things like that, where part of that is forecasting. So just to some extent, you sort of can, but the Good Judgment Project and Good Judgment Inc's current research stuff is pretty near the frontier. So it'll take a while to get through, I think, to back through academia.

Speaker D:

And I suspect that in these political-focused master's courses, there won't be a module on how to do forecasting well. It will just be subject matter, subject matter.

Speaker A:

There won't be the methodological part. Why? What? I mean, you mean now, right?

Speaker D:

Yeah, now. Now, but perhaps there should be. Perhaps there should be more consideration put

Speaker A:

on methodology. Sure. I mean, it sounds like, hopefully, one of the impacts of the Good Judgment Project, it sounds like, we can say with a reasonable amount of certainty, is better

Speaker B:

forecasting, right? That's what's going to happen. One hopes. Yeah, one hopes. Better decisions, better forecasts, more rational behaviour by states. That'd be nice. Excellent. Okay, I'm going to

Speaker A:

wrap it up there. So look, thank you very much, Nick Hare and Peter Coghill of Aleph Insights, and thank you especially to our special guest this week, to Michael Story of Good Judgment Inc. We've really enjoyed chatting with you. So thank you very much. Thank you to everyone. Thanks for listening. Until next time. Bye-bye.

Links

Chapters

Video

More from YouTube