Artwork for podcast Cognitive Engineering
A-Level Algorithms
Episode 216 • 30th September 2020 • Cognitive Engineering • Cognitive Engineering
00:00:00 00:31:11

Share Episode

Shownotes

What is a fair way to decide an exam result in the absence of being able to sit the exam?

In this podcast we discuss the background to the controversial A-Level algorithm debacle. We also touch on the concept of fairness in examinations and consider the essence of what we are trying to measure through an exam. Finally, we look at the application of algorithms to other areas of performance assessment, such as sport.

A few things we mentioned in this podcast:

For more information on Aleph Insights visit our website https://alephinsights.com or to get in touch about our podcast email [email protected]



This podcast uses the following third-party services for analysis:

Podtrac - https://analytics.podtrac.com/privacy-policy-gdrp

Transcripts

Speaker A:

Hello and welcome to the Cognitive Engineering Podcast produced by me, Fraser McGruer for Aleph Insights. In this series of podcasts we take a look at interesting topics and discuss what we think they tell us about analysis and decision making. I'm here with Chris Wragg and Nick Hare of Aleph Insights, and this week we're discussing the UK's exam fiasco. Should it be fiasci if there's more than one of them? Yeah. Fiasco, I presume it's an Italian word, right? Sounds like it, yeah. I mean they have plenty of fiascos. It's a good place to, yeah, yeah, yeah. Okay, Nick, lead us in. What are we talking about here?

Speaker B:

Well, I don't know if you're aware of this coronavirus that's been going around a bit, but one of the things, the sort of knock-on effects it had was very early on in March, April, the government announced that exams would not take place, that they weren't going to have the main exams that we sit at school in the UK, which are GCSEs when you're 16 and A-levels when you're 18. Which all takes place in about June, I think, if memory serves. Yes, so we have this big problem, which is that people need to get qualifications, among other things, because they will need to know whether they've got into the university of their choice. So Ofqual, who is the UK's exams regulator, essentially qualifications regulator, designed an algorithm which took account of various things, took account of the past performance of the school, took account of the individual pupil's past performance, and teacher predictions, etc., etc., and gave people grades based on this, based on the forecast. So your grade was a result of lots of information, but not actually an exam, right? As a result, about 40% of pupils got a lower grade than they were predicted to get at A-level GCSE. It caused a massive controversy, a huge outcry, can't get into the university I want to get into because this algorithm says no. So now the political outcry, everyone's, the Aleph Insights official MP, Gavin Williamson, said that there would be no U-turn, right? Of course, yeah, because, sorry, he's the education secretary, and I forgot, our favourite, as you say, Aleph. For some reason, he crops up on many podcasts. Anyway, he announced definitively there would be no U-turn, and a couple of days later, there was a U-turn, and they then said they were going to sort of, essentially, the moderation system had resulted in significant inconsistencies, had caused a lot of distress, etc., etc., and they were going to abandon the algorithm and more or less go with the prediction grades.

Speaker C:

It was more of an A-star turn than a U-turn.

Speaker B:

Now, had they gone with the algorithm, the results would still have been about 2% higher, you know, on average, across the board, than they were last year. They were building in a bit of grade inflation. However, there is now, as a result of abandoning that algorithm, we now have an increase of something like, from about 25% of people getting the top grades, it has gone up to 38%, which is the biggest increase in 20 years. So, and with his resignation, Sally Collier, who was the chief regulator of Ofqual, has resigned. Jonathan Slater, who was the permanent secretary of the Department of Education, has stood down.

Speaker A:

So, Gavin Williamson resigned as well, right?

Speaker B:

Yes, on principle, of course. Weirdly, no. So, that's the situation. Now, and I think this raises a whole load of questions. Was the algorithm good or bad? Should we use algorithms for this kind of thing? Were there other ways we could have approached it? You know, whose fault was it, if anyone's? And yeah, so lots of issues, lots of interesting issues to think about.

Speaker A:

It's quite a meaty one, this one, isn't it? Yeah, yeah. Well, that being the case, Chris, is there anything you immediately want to pick up on here?

Speaker C:

Yeah, I think the first thing to say is that this is a relative judgment that we're making. So, whether or not the algorithm was fair or not is itself a very thorny question. But the question is, is it fairer than what they ended up doing, which was to accept effectively unmoderated teacher assessments? And I think the answer to that is probably no, because the alternative is essentially a highly subjective human judgment. Yes, based on professional opinion, but there is no ability within that judgment to sort of moderate for the biases within the teacher themselves. And that's why you have these, you know, these sort of moderated processes. So, yeah, I think that in terms of, you know, whether it's fair or not, it's a relative judgment,

Speaker B:

I think. Well, I can actually, I've got quite a lot of fairly good evidence about the accuracy of teachers' predictions. Right, let's, yeah, let's start with that, yeah. The shit. Okay. So, teachers are incredibly bad at predicting. Presumably they over-predict. They over-predict quite badly. So, one of the most recent study, which looked at this comprehensively, looked at the total number of points that someone gets at A-level, right? And in A-level points is an A star is six, an A is five, a B is four, and so on, right? So, you know, if you got, if you predicted B, B, B, but you got A, B, C, that would be kind of, yeah, that's on point. That's the measure that's being used here, rather than looking at individual grades. One in six grades were accurately predicted, and 75% were over-predicted. Good Lord. Right? So, if you, whatever your results were, there's a 75% chance that you were predicted to get higher than your results, right? Under normal circumstances. There was a few centres that were looked at that put in judgments where they gave everyone A stars and A's, for example, in a particular subject where, you know, that had never happened in the past. And this ties in actually with previous research that has looked at individual grades as well. It's very highly consistent. Like, so this is not an unusual thing of the year they looked at it. It's, you know, it ties in with other research that's been done. The lower you get, the bigger the wrongness gets, the bigger the over-prediction. So, on average, right, someone who was predicted B, C, C got E, E, E. Wow. Someone who was predicted A, A, A, on average, got A, A, A. The lower you, the worse you are, the higher your over-prediction will be. So immediately we're thinking, well, probably a lot of people who should have done badly have done much better than they actually deserve to have done, right? And actually the very highest scores are always under-predicted. Well, that's not surprising. Like, you know, someone who gets three A stars, they can only be under-predicted. Independent schools are much more accurate, right? So about 20% are accurately predicted rather than 15% at state schools. And they are, like, they're only on average, like, a point off rather than two or three points off. The most disadvantaged students are more likely to have to be over-predicted. And ethnic minority pupils are more likely to be over-predicted. So they will have benefited from the U-turn. And this, I mean, to use another Italian word, this explains the furore when this all happened, right? Yeah, although, but it's not, I mean, the algorithm didn't look at ethnicity. It just so happens that in this study, ethnic minority peoples were more likely to be over-predicted. So that's all.

Speaker A:

Yeah, no, but that's what I'm saying is, I remember in the press, the big news was that, you know, the more impoverished parts of society are the ones that had done the worst. Well, that's because historically there's always this over-prediction for-

Speaker B:

Yeah, they've done much, they're more like, if you're disadvantaged, you go to, you know, or you go to a school which has got historically bad results, blah, blah, you will end up doing much more worse than your predictions. Yeah. But that's, you know, because teachers are bad, very bad at predicting.

Speaker C:

e, you know, the graduates of:

Speaker B:

Yeah, well, they are going to university and, you know, high ability students are going to be sitting alongside people who don't deserve to be there, who are taking up places and, you know-

Speaker C:

A number of university places, obviously.

Speaker B:

Well, I know, I think they've had to, like, introduce sort of bulge years and things, as I understand it, because of the fact that now they've had to offer people, or they, a lot of people are being offered deferred places, for example, but that's just punting the problem to next year.

Speaker C:

Yeah. The other thing is sort of looking at the algorithm itself and whether or not that was fair. And you mentioned the sort of, you know, the sort of penalizing schools that had underperformed. And of course, one of the probably the most controversial thing about it was that it effectively presented a distribution of grades based on the school. So, you know, your school and your subject were told you get, you know, two A stars, three A's, seven B's, you know, whatever. That was your profile. And effectively, the teacher grades, the teacher assessments, rank ordered students within that grade profile, right? So, if your school didn't have any A stars to give out, it didn't matter how good a student you were, if your school had traditionally not performed well in that subject, it made no account of outliers. So, that's where the ferrari came from, largely, you know, was the- Because you could have a one-off sort of brilliant student. You could have a brilliant student in a terrible school and there was nothing they could do to achieve the top grades. The other thing that sort of, that tripped them up about it was that, let's say you had a small sample of students in a, you know, doing music in a small private school, for example. For those schools, they realized that the algorithm wouldn't work, it wouldn't function because the sample was too small to produce a distribution curve for. So, they got to opt out. So, you had this double thing of sort of, you know, impoverished schools, not being able to get high grades and small, you know, subjects, not math subjects and in small schools, which tended to be in the independent sector, were allowed to opt out and use their predicted grades. So, you had this kind of something that was perceived to be socially inequitable going on, which is always, you know, a no-no in British society.

Speaker B:

Yeah, I mean, I think that's fair. There are a lot of criticisms you can make of the algorithm. I have a few myself, but I won't go into much detail. But I think, as Chris said, the question at the beginning is, you know, what's more unfair? Right. And the problem here is that you need an algorithm. Teachers are not doing their job. They're bad at forecasting. And they're not just bad, you know, because it's inherently hard to forecast, although it is. They're systematically bad. They're systematically over-predicting people's grades. And you know who's most likely to be under-predicted? The people who are most likely to be given worse grades than they actually get. The top students. It's high-ability disadvantaged students. So, they are, they are most, these are the people who suffer, who are going to suffer the most from cancelling the exams. And that's about 3,000 kids a year, apparently, fall into that category of people who, you know, they're in a sink school and the teacher's like, you know, everyone gets Ds, I'll just, I'll give them all Cs. Which is pretty much what the algorithm did. Yeah. Well, I mean, sort of. I mean, that's right. And, but those people are in, they're between a rock and a hard place in terms of, you know, algorithm versus teacher predictions. And I think, I think this is where- We're just not very, we're just not predicting people's grades very well. And this is, and now

Speaker C:

we're in this sort of problem. And I think, I think the question is, this is, this is where it's quite interesting and where the concept of an exam is an interesting one. That, you know, teachers, teachers are being asked to predict exam results. And I think what teachers are actually trying to do is they're trying to predict the person's competence in the, in the subject, right? Rather than predict their exam results, which is not the, not the same thing. You know, systematic, systemized exams, you know, were, were created to try and remove this type of, of bias. So, you know, it was really only in, in like the 17th century that the European education system started to take on board the idea of, you know, standardized exams. And prior to that, it had all been done orally. And of course that's much, you know, they got, they did away with that effectively, that the, the predicted results kind of process, because it was so subject to bias and, you know, corruption and so on.

Speaker A:

So look, we've described some of the, in quite depth, a lot of the pitfalls.

Speaker B:

Well, I do want to know, given that Chris has a very small amount of teaching experience. Yeah, that's what we've got. Why do you think that this, I mean, you, you know, why is it, I get that, you know, they might be predicting competence, but you'd think that over time they would learn, they would adjust to their forecasts to be more accurate. Like, why is that not happening? And why is it worse in state schools? I mean, why?

Speaker C:

Yeah, well, it depends. So I think, so I think grade predictions are one thing. Nobody, I mean, grade predictions are used in the university admission process, but by and large, that's their, that's their sort of sole function. So it's quite, it's quite an important one, but actually when it comes down to it, you know, there's the safety, the safety net of people's GCSE results and their actual performance in the exams, which are, you know, what they wait for until they confirm your position. So I think predictions are effectively a bit of a meaningless exercise, where they're more meaningful further down the schooling system in primary schools, where you've got SATS results and they basically, the teacher, not the second set of SATS results that people do, but earlier on, then they're not purely test-based, you know, they're based on, partly on teacher assessment and so on, and the marking of work and being able to demonstrate that somebody has done work at a particular level of competence. That process is much more broadly moderated. There is no moderation of exam predictions. A teacher says, a start, there's no, you know, there's no internal moderation. And that's probably why they're bad at predicting grades, because there's no objective checking, there's no inter-assessor reliability, there's no, you know, people removed from the student relationship assessing and saying

Speaker B:

you sure, where's your evidence for that and so on. Yeah, but I guess, you know, why is it that they're always over-predicting? I mean, the allegation is that it's supposed to motivate people, you know, oh, I know Jemima hasn't handed in any geography homework all year, but I had a chat with her last week and she said she's really going to turn things around and it's important to not lose faith in people. But I mean, you know, speaking for myself, being given bad predicted grades, that would be the motivator, not being given better predicted grades than I'm going to

Speaker C:

get. So I don't get it, I don't get how that works. I mean, obviously, mock exams are a much more objective measure of how you're likely to get on an exam, right? I mean, you know.

Speaker B:

Although, you know, the rational student sees them for what they are, which is totally irrelevant

Speaker C:

and doesn't really bother mugging up. No, that's right. But in terms of, you know, the human psychology of giving somebody a grade, you're effectively valuing that person and then you have to maintain a relationship with that person for the, you know, remainder of the year. So there's no incentive to under-predict, there's no incentive to accurately predict really other than your credibility. But that's fairly distantly removed from, you know, nobody comes back and say, hey, do you know what? You, teacher there, are massively over-predicting all your students and letting everyone else down. There's none of that sort of process. It's essentially a pretty meaningless, non-incentivized process. So why wouldn't you give them good grades?

Speaker A:

Okay. So look, we're close to needing to conclude this part of the podcast. So that being the case, where have we got to? We need to round it off a bit.

Speaker B:

Well, I think this is one of a general class of situations where you don't have the information that you would normally get and you've got to do something about it. You've got to classify something in the absence of the information that you need. You know, and you ask, well, what's the point of exams? Like, imagine that I had an algorithm that was so good. It was, it was correctly predicting people's exam results at A-level, you know, in January. Well, then you might say, well, what's the point of the exam? And the point of the exam is that it does get, it gives you information that you do not have, that isn't contained in any other piece of work that that student has done because, you know, it is designed to test your ability to perform in precisely the situation of a high stress, single incident of having to demonstrate stuff, which is going to affect the rest of your life. Nothing else you've done until then is analogous to that. So there are aspects of your character, which are only revealed by being in that situation. That is the whole point of it. So I think, you know, that what I would say is it's probably inherently impossible to predict, to have an algorithm like that, but, you know, it's designed to find those kids who, you know, maybe they're coasting, but then when push comes to shove, they step up and, you know, spend two weeks working really hard and ace the exam. And I, you know, we want to find people like that. They're useful. So, but it doesn't, I mean, like the only, another example of an algorithm used in precisely this sort of situation that I'm aware of is the Duckworth-Lewis rule. But I, what with Chris being a sports expert, I think you might have looked at that. But the interesting thing about that is that it is sort of seems to be widely accepted in cricket as far as I understand it, but I'm not sure.

Speaker C:

Yeah. But I mean, it was born out of controversy of the previous ways of measuring that. So it was only after a sort of a World Cup semi-final between England and South Africa where you had a ridiculous situation where I think rain basically meant that there was a place stoppage and South Africa had something like, I think it was, they had to get 23 runs off 13 balls or something like that. And by the time they came back to the game after a short rain shower, there was enough time for a last ball effectively. And the method they had been using to determine what the target should be left them with a situation of needing to score over 20 runs off the last ball, which if you know anything about cricket. It's fairly unlikely. It's pretty much impossible given six is generally the highest you can get. And so, yes, two kind of English statisticians, Frank Duckworth.

Speaker B:

Sorry, just because I don't really understand it. And I think there may be international listeners who don't really understand it. Why couldn't you in that situation just finish the game? Why could they not just get all of their overs?

Speaker C:

Well, often it's to do with light, the amount of time there is to play the game and so on.

Speaker B:

So the amount of, so it's not like there's some certain number of overs in this situation. You lose basically by the end of the day. If you haven't done it, you've lost.

Speaker C:

Yes. So the Cricket World Cup is one day, one day cricket. So you don't normally, you know, in a test match where you might be playing over several days, you have the capacity to sort of play a bit more, have a bit less tea or whatever it might be, or generally sort of make up time and so on. But in one day cricket or shorter fixed overs matches, that's not an option. And obviously in England in the summertime, you know, we're quite susceptible to rain stopping play. So it's no surprise that, yes, it was two English statisticians, Frank Duckworth and Tony Lewis, who if you imagine two English statisticians called by those names, just picture them in your head. That's what they look like. They came up with a system that is now being added to. It's now the Duckworth, Lewis and Stern method after the custodian of it is now a Professor Stern. But they essentially, what it did was the problem you have is you might have 50 overs, right, six balls in an over. You might, both sides might get 50 overs and it's time to, it's how many runs you get in that, within those overs and you add them up and the one with the most runs wins, right. That's the nature of cricket. But if one team loses half its overs because of rain, so the first team finished their 50 overs and get 200 runs, let's say, and the second team only have 25 overs to make their runs, you need to set them a target. Now, you could just say, well, give them half the number of runs. They've got half the number of overs, half the number of runs. But of course, that doesn't take into account the fact that they have still got all their wickets left. So, you know, they've still got 10 wickets effectively. And so they can play much more aggressively. They've got more, they can take more risk as they're playing. So effectively, what this algorithm does is it's just a giant lookup table and it assesses the number of overs you have left and the number of wickets you have left as your resource effectively.

Speaker B:

But how did it get so, I mean, it's accepted that this is how you decide the results of matches. What made it accepted? Why did people not say, well, that's not fair because you're not taking account that Johnson's into bat next and he's awesome?

Speaker C:

Yeah, well, it has a certain degree of face validity. It's sort of not, you know, whereas previous methods had created those ridiculous situations, this has created fewer, you know, ridiculous situations. So it's generally accepted to kind of work. And it still enables people to have agency over the result. What's adjusted is the target they're aiming for, you know, not they're told, all right, okay, we've just said you've got 150 runs you've lost. You know, they've still got some capacity to engage in the event.

Speaker A:

We need to finish. Okay, just so I've got a question I want to ask, but is there anything you want to finish off on?

Speaker C:

Well, I just thought what I was going to say is that effectively, you know, we're now moving into an era where algorithms are going to be being used to decide lots more outcomes about our lives, you know, things like car, you know, what your car insurance premium should be, what your recidivism rates might be as a somebody coming up for parole, so on and so forth. And those algorithms are only going to get more opaque as you know, they become more subject to from our perspective, black box AI kind of predictions. So, you know, we're going to have to get used to being judged by algorithms. And it may turn out that they end up being more effective than whatever the essence of what we were trying to measure was in the first place by our traditional means anyway.

Speaker A:

Nice. That being the case, I've got a question for you, which is, can you think back to a moment in your life? And I guess in the first instance here, I'm thinking about exams, but it doesn't have to be. Can you think about when there was a life changing moment for you? Or know an important period of your life when you would rather have had an algorithm deciding things for you than actually what was used as an assessment at the time. So just to kick us off, for example, had I been in late had this had the had Coronavirus happened when I was an A level student, it would have been the best thing that could have happened to me. Because I did. My predicted results were very high. My actual results were just an absolute bombshell. And I just performed terribly. And I would have just that would have, it would just be brilliant. I'd have just shut down if I had just shut down the exams anyway. And I was the stone cold opposite.

Speaker B:

Right. Okay, only ever my exams, I was aced exams and always got terrible predictions. Right. That implies... Partly to try and show my teachers up for the morons they were. But that implies that your teachers didn't like you. Yeah, that's probably fair.

Speaker A:

st one of those students from:

Speaker B:

I suppose this makes me think about having been... The example I think of is the Oxford interview process, which as you know, I failed and you inexplicably passed. And I can't help thinking actually, just the opposite of your example, you know, if you'd have got your real exam results, and an algorithm had decided whether I ought to go to Oxford, you know, then maybe I'd have done better than a, you know, this kind of sit down in a chat with a politics professor who asked me about those are things I've never heard of to do with the UK Constitution. And I, you know, I had no idea what I was doing. And yeah, whereas I think probably an exam and an algorithm, I will probably have done a little bit better. So essentially removing any kind of human contact. Take the humans out of it. Yeah, as long as humans don't have to have contact with me, I look pretty good. I look pretty good on paper. Yeah. In a spreadsheet,

Speaker C:

you know, it's just that's very interesting, though, because on the one hand, you know, what you're missing with an algorithm versus an exam is you're missing the sort of the crucible of the of the pressure and the not being able to refer to anything and what's actually in your your head, but also time pressure, and so on. And so that's something you lose from that process potentially unless you factor factor in. But what you what you lose from a from an interview is versus an algorithm is that human interaction, you know, are you are you a good egg? Well, not necessarily. But can you you know, can you converse? Do you rub people up the wrong way? You know, those kinds of things. So you're missing missing that information as well. And so it's quite interesting that, you know, what you two have gained or lost by by the real process versus what an algorithm may have come up with. But of course, algorithms can account for those things. I think for me, I would probably rather have had certainly from about sort of 14 to about 17. I think I would rather have had an algorithm take care of my attempts to sort of woo. I think that sort of, yeah, just just essentially lining me up with potentially suitable, you know, people and, and not forcing me to have that long sort of terrifying walk in a bar, you know, to another table, or across a, you know, a school disco or something. Sweaty palms. Yeah, exactly. So I think just having something pony up and say, hey, we think this person's compatible for you. Well, a lot of people seem to complain about Tinder, but I think, you know,

Speaker A:

it's got that going for it. Yeah. And also, I'm sure there's an algorithm out there or some kind of connection to people who do well in exams, and people who need help with wooing the opposite sex. It seems like a stereotype, but probably a fair one. Okay, I think we'll stop there. I really enjoyed that. Thank you, as always, for listening to the Cognitive Engineering Podcast. I'm Fraser McGruer. We've been here with Nick Hare and Chris Wragg of Aleph Insights. Until next time, goodbye.

Links

Chapters

Video

More from YouTube