My What If lecture on using Economics to analyse ODI cricket is now up on YouTube. In it I give an "Economics in 4 lessons" overview, and then used the same graphs to talk about cricket. Watching the playback, I wish I had sped through the intro to Econ a bit faster, and I rather wish the cameraman had kept the camera on the screen rather than panning back to my fidgeting, but overall I'm reasonably happy with how it came out.
Showing posts with label cricket. Show all posts
Showing posts with label cricket. Show all posts
Wednesday, May 8, 2013
Thursday, April 25, 2013
What If Lecture on Cricket
Last year, Eric presented a talk on alcohol in the University's What If Wednesday series. This coming Wednesday it will be my turn, talking about economics and cricket. The title is "What if...Economics could help cricket teams win matches?" I have blogged a bit about my work with students on various aspects of cricket in the past (click on the tag "cricket" to see all those posts).
On Wednesday, I am going to focus on how an economics-derived approach to the analysis of cricket can yield interesting analyses of on-field strategy. Specifically, I will be talking about an aspect of the work from Scott Brooker's doctoral thesis that I haven't blogged on previously--how we can estimate batter "production possibility frontiers" showing the tradeoff between risk and return for specific batsmen, and from that to suggest one explanation for why New Zealand has traditionally punched above its weight in ODI cricket.
The announcement and link to register for the lecture are here. I understand that the registration process can be a bit cumbersome as it seems like you are registering for a course, but it is free to attend and registering enables the University to ascertain likely numbers. (And to be honest, I haven't registered for the ones that I have attended.)
I hope to see as many loyal readers of Offsetting there as possible, conditional, of course, on your caring about cricket.
On Wednesday, I am going to focus on how an economics-derived approach to the analysis of cricket can yield interesting analyses of on-field strategy. Specifically, I will be talking about an aspect of the work from Scott Brooker's doctoral thesis that I haven't blogged on previously--how we can estimate batter "production possibility frontiers" showing the tradeoff between risk and return for specific batsmen, and from that to suggest one explanation for why New Zealand has traditionally punched above its weight in ODI cricket.
The announcement and link to register for the lecture are here. I understand that the registration process can be a bit cumbersome as it seems like you are registering for a course, but it is free to attend and registering enables the University to ascertain likely numbers. (And to be honest, I haven't registered for the ones that I have attended.)
I hope to see as many loyal readers of Offsetting there as possible, conditional, of course, on your caring about cricket.
Tuesday, March 5, 2013
Declarations and Nightwatchmen
A couple of years ago, I supervised an Honours research paper by Johnny Sharland (now at the RBNZ) on the efficacy of using a nightwatchman in cricket. For the uninitiated, a nightwatchman is a player (usually a bowler**) who is sent in ahead of a better batsman when a wicket is lost near the end of the day in a multiple-day cricket match. The idea is that a batsman is at most risk of getting out at the start of his innings, so the nightwatchman's job is to protect the better batsman from having to start his innings at the end of the day, and then make a fresh start the next morning.
We wanted to compare the cost to a team of changing the batting order to the benefit of a reduced probability of dismissal for a top-order batsman, to calculate when and if the benefits would outweigh the costs. As it turned out, the benefit-cost calculation turned out to be completely uninteresting for an unexpected reason: We could find no evidence in a database of 200 test matches that having to make a second start in any way increases the probability of a top-order batsman being dismissed early in his innings. That is, the key variable that determines a batsman's probability of being dismissed is how many balls he has faced in that innings, not how many he has faced that morning. The cost-benefit calculation then becomes irrelevant, as there is simply no benefit from using a nightwatchman to balance against the costs.
I thought this made the results uninteresting, and so did not seek to write it up for publication. The latest test between Australia and India, however, makes me think that maybe that decision was wrong. The Australian captain, Michael Clarke, has just earned the unenviable distinction of being the first ever captain to declare his first innings closed and then go on to lose by an innings.* Clarke declared the Australian innings closed when they were 9 wickets down with 5 overs remaining on the first day. It seems that his reasoning was that Australia's last two batsmen were probably not going to be able to score many runs anyway, and by declaring he would force the Indian batsmen to start their innings that night and make a fresh start the next morning. Johnny's results suggest that there was no expected benefit to this declaration, only a cost. Maybe a shortened version of his paper concentrating only on the estimate of the second-start effect would be interesting.
* Explanation for non-cricket fans. Each team bats until 10 of its 11 batsmen have been dismissed. The 10th dismissal signals the end of the "innings" after which the other team gets a chance to bat and score more runs. The two teams get two innings of up to 10 dismissals. A captain, however, can choose not to use it full allotment and declare its innings over before 10 batsmen have been dismissed. If a team scores fewer runs in its two innings combined than the opposition scores in its first innings, the opposition does not need to bat a second time and is said to have "won by an innings".
** For an example of a batsman being used as a nightwatchman, check out this match (which is also notable for featuring Bradman's first test century). Batting on a notorious Melbourne sticky wicket, England sent specialist batsman, Douglas Jardine, in ahead of his normal batting position in order to protect superstar batsman, Wally Hammond.
We wanted to compare the cost to a team of changing the batting order to the benefit of a reduced probability of dismissal for a top-order batsman, to calculate when and if the benefits would outweigh the costs. As it turned out, the benefit-cost calculation turned out to be completely uninteresting for an unexpected reason: We could find no evidence in a database of 200 test matches that having to make a second start in any way increases the probability of a top-order batsman being dismissed early in his innings. That is, the key variable that determines a batsman's probability of being dismissed is how many balls he has faced in that innings, not how many he has faced that morning. The cost-benefit calculation then becomes irrelevant, as there is simply no benefit from using a nightwatchman to balance against the costs.
I thought this made the results uninteresting, and so did not seek to write it up for publication. The latest test between Australia and India, however, makes me think that maybe that decision was wrong. The Australian captain, Michael Clarke, has just earned the unenviable distinction of being the first ever captain to declare his first innings closed and then go on to lose by an innings.* Clarke declared the Australian innings closed when they were 9 wickets down with 5 overs remaining on the first day. It seems that his reasoning was that Australia's last two batsmen were probably not going to be able to score many runs anyway, and by declaring he would force the Indian batsmen to start their innings that night and make a fresh start the next morning. Johnny's results suggest that there was no expected benefit to this declaration, only a cost. Maybe a shortened version of his paper concentrating only on the estimate of the second-start effect would be interesting.
* Explanation for non-cricket fans. Each team bats until 10 of its 11 batsmen have been dismissed. The 10th dismissal signals the end of the "innings" after which the other team gets a chance to bat and score more runs. The two teams get two innings of up to 10 dismissals. A captain, however, can choose not to use it full allotment and declare its innings over before 10 batsmen have been dismissed. If a team scores fewer runs in its two innings combined than the opposition scores in its first innings, the opposition does not need to bat a second time and is said to have "won by an innings".
** For an example of a batsman being used as a nightwatchman, check out this match (which is also notable for featuring Bradman's first test century). Batting on a notorious Melbourne sticky wicket, England sent specialist batsman, Douglas Jardine, in ahead of his normal batting position in order to protect superstar batsman, Wally Hammond.
Wednesday, January 30, 2013
Do Catches win Matches? (UPDATED)
The University put out a press release yesterday describing some summer research I did on this topic with a summer student, Marcus Downs. It has been picked up by a few electronic media outlets, including the Hearld, and the two of use were interviewed by One News, with the plan to include it on the the Six O Clock News tonight. (UPDATE 1: It appeared on Sunday 3 Feb. It will be here till the 9th. The clip starts at around 33:14.)
What we did was to go through the cricinfo commentary for 122 ODI matches played in 2011 and 2012, and identify every ball where there was an opportunity for a fielder to bring about a dismissal, either by taking a catch, effecting a run-out, or making a stumping. We then used the commentary to characterise the degree of difficulty of the opportunity (blinder, difficult, normal, absolute dolly), and find the probability that the dismissal would be made for each of these difficulty levels.
We then used the same analysis that produces the first-innings score predictor in the WASP (see my previous post on that here), to calculate how many runs a batting team's expected score in the first innings would increase by after each ball. Batters get credit (discredit) for all of that increase (decrease), whereas the credit or discredit is shared between bowlers and fielders in those cases where there is a fielding dismissal opportunity, with fielders getting more of the credit for a blinder, and more of a penalty for dropping a dolly.
We calculated a distribution across all batters, bowlers, and fielders in our datbase. What we found was that a batsman who is one standard deviation above the average contributes about 8 runs more to his team than an average batsman; a bowler who is one s.d. above average contributes about 6 runs more (that is, he restricts the opposition's score by about 6 runs more than an average bowler), but a one-s.d.-above-average fielder contributes less than 2 extra runs. 8 runs may not sound like much, but an additional 8 runs can make quite a difference to the chances of a first-innings score being successfully chased. (UPDATE 2: Scoring 8 runs more than par rather than par, pushes up the chance of winning from 50% to 56%.)
We still have some improvements to make to the analysis, but they are only going to further minimise the relative importance of fielding.
There are two main reasons for why catches and run-outs are not that important (notwithstanding the recent 2nd ODI between NZ and South Africa, where 5 run-outs tipped the balance in New Zealand's favour). The first is that a lot of the run-outs and catches in ODI games occur near the end of the innings where their impact on the score is not so great. The second is that most of the opportuntities that arrive are ones that are (or should be) straightforward for an international cricketer. We all recall moments of fielding brilliance, but those opportunities simply don't arrive often enough to make the contributions of great fielders worth a place in the team for that reason alone.
There are a couple of caveats to any coaches taking policy conclusions from this.
What we did was to go through the cricinfo commentary for 122 ODI matches played in 2011 and 2012, and identify every ball where there was an opportunity for a fielder to bring about a dismissal, either by taking a catch, effecting a run-out, or making a stumping. We then used the commentary to characterise the degree of difficulty of the opportunity (blinder, difficult, normal, absolute dolly), and find the probability that the dismissal would be made for each of these difficulty levels.
We then used the same analysis that produces the first-innings score predictor in the WASP (see my previous post on that here), to calculate how many runs a batting team's expected score in the first innings would increase by after each ball. Batters get credit (discredit) for all of that increase (decrease), whereas the credit or discredit is shared between bowlers and fielders in those cases where there is a fielding dismissal opportunity, with fielders getting more of the credit for a blinder, and more of a penalty for dropping a dolly.
We calculated a distribution across all batters, bowlers, and fielders in our datbase. What we found was that a batsman who is one standard deviation above the average contributes about 8 runs more to his team than an average batsman; a bowler who is one s.d. above average contributes about 6 runs more (that is, he restricts the opposition's score by about 6 runs more than an average bowler), but a one-s.d.-above-average fielder contributes less than 2 extra runs. 8 runs may not sound like much, but an additional 8 runs can make quite a difference to the chances of a first-innings score being successfully chased. (UPDATE 2: Scoring 8 runs more than par rather than par, pushes up the chance of winning from 50% to 56%.)
We still have some improvements to make to the analysis, but they are only going to further minimise the relative importance of fielding.
There are two main reasons for why catches and run-outs are not that important (notwithstanding the recent 2nd ODI between NZ and South Africa, where 5 run-outs tipped the balance in New Zealand's favour). The first is that a lot of the run-outs and catches in ODI games occur near the end of the innings where their impact on the score is not so great. The second is that most of the opportuntities that arrive are ones that are (or should be) straightforward for an international cricketer. We all recall moments of fielding brilliance, but those opportunities simply don't arrive often enough to make the contributions of great fielders worth a place in the team for that reason alone.
There are a couple of caveats to any coaches taking policy conclusions from this.
- We have only looked at dismissal chances. If we were able to get good data on ground fielding, it might make a difference. I suspect not, though.
- We have only looked at ODI cricket. I supsect the role of catching might be greater in test cricket. (I am showing my age here, but I continue to believe that Jeremy Coney should have been in the NZ team between in the 76-78 period, simply to make sure there was someone who could hold on to the slip catches that Richard Hadlee was generating and having continually dropped at that time.)
- It may be that fielding is more dependent on coaching and practice rather than natural talent, relative to batting and bowling, and so the reason that the better-than-average fielders are not that much better than average, is because coaches have correctly emphasised bringing all fielders up to a minimum standard, and have not selected players who don't meet that threshold.
Wednesday, November 21, 2012
Cricket and the Wasp: Shameless self promotion (Wonkish).
In their coverage of the Wellington-Auckland game in the HRV cup last Friday, Sky Sport introduced WASP—the “winning and score predictor” for use in limited-overs games, either 50-over or 20-20 format. In the first innings, the WASP gives a predicted score. In the second innings, it gives a probability of the batting team winning the match.
I am very happy about this as it is based on research by my former doctoral student, Scott Brooker, and me. Not surprisingly, the commentators didn’t go into any details about the way the predictions are calculated, so I thought I would explain the inner workings in a wonkish blog post.
The first thing to note is that the predictions are not forecasts that could be used to set TAB betting odds. Rather they are estimates about how well the average batting team would do against the average bowling team in the conditions under which the game is being played given the current state of the game. That is, the "predictions" are more a measure of how well the teams have done to that point, rather than forecasts of how well they will do from that point on. As an example, imagine that Zimbabwe were playing Australia and halfway through the second innings had done well enough to have their noses in front. WASP might give a winning probability for Zimbabwe of 55%, but, based on past performance, one would still favour Australia to win the game. That prediction, however, would be using prior information about the ability of the teams, and so is not interesting as a statement about how a specific match is unfolding. Also, the winning probabilities are rounded off to the nearest integer, so WASP will likely show a probability of winning of either 0% or 100% before the game actually finishes, even though the result is not literally certain at that point.
The models are based on a database of all non-shortened ODI and 20-20 games played between top-eight countries since late 2006 (slightly further back for 20-20 games). The first-innings model estimates the additional runs likely to be scored as a function of the number of balls and wickets remaining. The second innings model estimates the probability of winning as a function of balls and wickets remaining, runs scored to date, and the target score.
The estimates are constructed from a dynamic programme rather than just fitting curves through the data. To illustrate, in the first innings model to calculate the expected additional runs when a given number of balls and wickets remain, we could just average the additional runs scored in all matches when that situation arose. This would work fine for situations that have arisen a lot such as 1 wicket down after 10 overs, or 5 wickets down after 40 overs, etc.), but for rare situations like 5 wickets down after 10 overs or 1 wicket down after 40 it would be problematic, partly because of a lack of precision when sample sizes are small but more importantly because those rare situations will be overpopulated with games where there was a mismatch in skills between the two teams. Instead, what we do is estimate the expected runs and the probability of a wicket falling on the next ball only. Let V(b,w) be the expected additional runs for the rest of the innings when b (legitimate) balls have been bowled and w wickets have been lost, and let r(b,w) and p(b,w) be, respectively, the estimated expected runs and the probability of a wicket on the next ball in that situation. We can then write
Now many authors have applied dynamic programming to analyse sporting events including limited overs cricket (see my previous post on this here), although I don’t know of any previous uses of such models in providing real-time information to the viewing public. Scott’s and my main contribution, however, is in including in our models an adjustment for the ease of batting conditions. I have previously blogged about our model for estimating ground conditions, here. Without that adjustment, the models would overstate the advantage or disadvantage a team would have if they made a good or bad start, respectively, since those occurrences in the data would be correlated with ground conditions that apply to both teams. Using a novel technique we have developed, we have been able to estimate ground conditions from historical games and so control for that confounding effect in our estimated models.
In the games on Sky, a judgement is made on what the average first innings score would be for the average batting team playing the average bowling team in those conditions, and the models’ predictions are normalised around this information. At this stage, I believe this judgement is just a recent historical average for that ground, but the method of determining par may evolve.
I gather that the intention is to unveil more graphics around the use of WASP throughout the season, with the system fully up and running by the time of the international matches against England. It’s going to be interesting listening to what the commentators make of the WASP. Last Friday’s game wasn’t the best showcase, since when Auckland came to bat in the second innings, their probability of winning was already at 92% and quickly rose higher. It was fun, though, hearing the commentators ask Wellington captain, Grant Elliot, who was wired for sound while fielding, what he thought their chances were given that WASP had the Auckalnd Aces at 96% at that point. Grant's reply was lovely: "Sometimes even pocket aces lose". This is worth remembering when (as will inevitably happen), a team has a probability of winning in the 90s but still goes on to lose.
I am very happy about this as it is based on research by my former doctoral student, Scott Brooker, and me. Not surprisingly, the commentators didn’t go into any details about the way the predictions are calculated, so I thought I would explain the inner workings in a wonkish blog post.
The first thing to note is that the predictions are not forecasts that could be used to set TAB betting odds. Rather they are estimates about how well the average batting team would do against the average bowling team in the conditions under which the game is being played given the current state of the game. That is, the "predictions" are more a measure of how well the teams have done to that point, rather than forecasts of how well they will do from that point on. As an example, imagine that Zimbabwe were playing Australia and halfway through the second innings had done well enough to have their noses in front. WASP might give a winning probability for Zimbabwe of 55%, but, based on past performance, one would still favour Australia to win the game. That prediction, however, would be using prior information about the ability of the teams, and so is not interesting as a statement about how a specific match is unfolding. Also, the winning probabilities are rounded off to the nearest integer, so WASP will likely show a probability of winning of either 0% or 100% before the game actually finishes, even though the result is not literally certain at that point.
The models are based on a database of all non-shortened ODI and 20-20 games played between top-eight countries since late 2006 (slightly further back for 20-20 games). The first-innings model estimates the additional runs likely to be scored as a function of the number of balls and wickets remaining. The second innings model estimates the probability of winning as a function of balls and wickets remaining, runs scored to date, and the target score.
The estimates are constructed from a dynamic programme rather than just fitting curves through the data. To illustrate, in the first innings model to calculate the expected additional runs when a given number of balls and wickets remain, we could just average the additional runs scored in all matches when that situation arose. This would work fine for situations that have arisen a lot such as 1 wicket down after 10 overs, or 5 wickets down after 40 overs, etc.), but for rare situations like 5 wickets down after 10 overs or 1 wicket down after 40 it would be problematic, partly because of a lack of precision when sample sizes are small but more importantly because those rare situations will be overpopulated with games where there was a mismatch in skills between the two teams. Instead, what we do is estimate the expected runs and the probability of a wicket falling on the next ball only. Let V(b,w) be the expected additional runs for the rest of the innings when b (legitimate) balls have been bowled and w wickets have been lost, and let r(b,w) and p(b,w) be, respectively, the estimated expected runs and the probability of a wicket on the next ball in that situation. We can then write
V(b,w) =r(b,w) +p(b,w) V(b+1,w+1) +(1-p(b,w)))V(b+1,w)Since V(b*,w)=0 where b* equals the maximum number of legitimate deliveries allowed in the innings (300 in a 50 over game), we can solve the model backwards. This means that the estimates for V(b,w) in rare situations depends only slightly on the estimated runs and probability of a wicket on that ball, and mostly on the values of V(b+1,w) and V(b+1,w+1), which will be mostly determined by thick data points. The second innings model is a bit more complicated, but uses essentially the same logic.
Now many authors have applied dynamic programming to analyse sporting events including limited overs cricket (see my previous post on this here), although I don’t know of any previous uses of such models in providing real-time information to the viewing public. Scott’s and my main contribution, however, is in including in our models an adjustment for the ease of batting conditions. I have previously blogged about our model for estimating ground conditions, here. Without that adjustment, the models would overstate the advantage or disadvantage a team would have if they made a good or bad start, respectively, since those occurrences in the data would be correlated with ground conditions that apply to both teams. Using a novel technique we have developed, we have been able to estimate ground conditions from historical games and so control for that confounding effect in our estimated models.
In the games on Sky, a judgement is made on what the average first innings score would be for the average batting team playing the average bowling team in those conditions, and the models’ predictions are normalised around this information. At this stage, I believe this judgement is just a recent historical average for that ground, but the method of determining par may evolve.
I gather that the intention is to unveil more graphics around the use of WASP throughout the season, with the system fully up and running by the time of the international matches against England. It’s going to be interesting listening to what the commentators make of the WASP. Last Friday’s game wasn’t the best showcase, since when Auckland came to bat in the second innings, their probability of winning was already at 92% and quickly rose higher. It was fun, though, hearing the commentators ask Wellington captain, Grant Elliot, who was wired for sound while fielding, what he thought their chances were given that WASP had the Auckalnd Aces at 96% at that point. Grant's reply was lovely: "Sometimes even pocket aces lose". This is worth remembering when (as will inevitably happen), a team has a probability of winning in the 90s but still goes on to lose.
Subscribe to:
Posts (Atom)