08 March 2013

Distributions and Consistency, Part II


The first thing to look at is power vs OBP. And I would think that the casual fan would think that OBP is more consistent, because it will generate base-runners and ‘threats’ more often, by definition. But this is actually the opposite of what is actually true. A little bit of playing around with a good model which will show you distributions (again, I’d suggest http://www.tangotiger.net/markov_wes.html) will show you this.

And actually, a little bit of thought will let you realize this, too. In order for a team relying on OBP to score a run, it’s going to need 3-4 successes to get in. A team relying on power will only need 1-2. So for the OBP team, they need to take that OBP to the third or fourth power to get to the first run, whereas a home run, for example, only needs to hit once. Because typical OBPs are in the 1/3 range, this makes it very sketchy to have to tack on several of them in a row. Of course, if they’ve already achieved this, then every additional success is another run, and their higher OBP can start raking in a substantial number of runs extra – but this is the trade-off, OBP teams are less likely to score overall, but more likely to score more. Of course this makes them less consistent.

It’s definitely worth nothing that this effect is much more dramatic on smaller scales. There is a very big difference over a single inning, but much less of a difference over a nine inning game. Basically, the longer the game goes, the less of a consistency difference there is. (For the statistically inclined, this is the central limit theorem at work – in the limit as the number of innings goes to infinity, things start getting closer to the mean and more symmetric, approaching a normal distribution).

Now, you might think that this would mean that extra-inning games are going to favor the high OBP teams, as the longer the game is, the more they’re favored. But there’s a nuance of conditional probability that reverses this; namely, if the game is going into extra innings, we KNOW that it’s tied. Ergo, we essentially are getting a one-inning game, and if it’s tied after that, another one inning game, etc. etc. until there is a winner. Which means that the extra-inning games are much MORE in the favor of the consistent, high-power team.

What makes this effect more or less important? Well, it’s a continuous, changing thing that will fluctuate a good bit depending on the actual matchup you’re looking at. But there are a number of general factor which will change how relevant it is. First up, the bigger the difference in consistency is, the bigger impact that it will have on who wins. Moreover, close games will make the consistency effect increasingly important – i.e. if the expected average runs scored by the two teams are very close, consistency will start to play a big factor, but if they’re far enough apart, then while consistency might help the lower-scoring team, it won’t save them. (I should point out that this WILL save them if the averages are reasonably close – up to several percent different if one is enough more consistent than the other).

So there are two things that are going to drive this, usually. One is that the teams should be close-matched to each other. The second is that the scoring, overall, is low. So you’d expect pre-1969 teams and deadball teams to see this more, and actually, you should expect to see the difference rise now (and including the past couple of years) compared to how things were even 10 years ago, when offenses were much higher. Actually, there’s a feedback effect, too – the more consistent the both teams are, the more likely the game is to be close, and the more important consistency is.

Okay, but how big is this effect? Over a nine inning game, it’s pretty small. One thing you have to remember that most major league teams – player even – aren’t going to be all that different from each other. For instance, last year in the AL, the team with the best OBP was barely 4% different from the team with the worst, and the top gap in SLG was only about 85 points there. On top of this, in most games, one team is going to have a reasonably substantial lead on the other in terms of overall score rate, and especially when this is the team that is more consistent (i.e. more power – which is often the case, as they will have an advantage in both areas), the consistency just won’t move the needle much at all.

There are a couple of reasons why this effect can be important, anyway, though. One is in very particular games – say there are a couple of very very good starting pitchers, a pitching-friendly umpiring crew and ballpark, etc. and the teams are close. In this case, (and the pitching matchup is probably the big thing here – you really can get very low run-scoring environments in several dozen games a year), the consistency thing can actually start having some kinds of macroscopic effects. But the second reason is bigger, and that’s extra-inning games. Whereas the length of nine innings mitigates it out pretty well, once you get down to the one-inning affairs, this consistency thing actually starts being pretty important. Towards the extreme end of things, a team which scores as few as 60% as many runs as their opponents over the long run can still expect to win more extra-inning games than their opponents if they have enough more consistency. Realistically, you’re never going to see numbers this far off, but it wouldn’t be too ridiculous to see this figure in the 85% range. Or if you want to look at it the other way, when teams are scoring about the same number of runs (as you’d probably expect them to, given that they’re in extra innings), teams which are more consistent can win as much as 62% of the time theoretically, and actually in practical terms, as much as 57% of the time, which is pretty massive when you consider the average difference in MLB teams overall.

All of which brings me to the interesting point of last year’s Baltimore Orioles. They won 16 extra inning games in a row to end the season, after losing their first two. And the statisticians told us that this was so incredibly lucky. And they’re absolutely right – to have a 50% chance of winning that many in a row, your win% in each one would need to be over 95%, which is ridiculous. And even the 16 of 18 would need a win percentage of a little over 85%, which is way too high. But while the odds you’ll see quoted of this most often would give a 0.067% chance of the 16 of 18 happening, that’s based on a 50% win rate, and I think that’s absolutely unfair to the Orioles. For one thing, and this is the biggest factor, they had quite a good bullpen and a pretty good team overall. This is going to give them over a 50% chance to start with. But if you look at their offensive profile, they were one of the more power-dependent teams we’ve seen in the past several years, significantly more on the power side of things than the rest of the league was. The combination of the consistency that this power gave them with the strength of their bullpen would lead me to think that they would have a ‘true winning percentage’ of at least 60%. Indeed, it wouldn’t totally shock me to find out that this figure was all the way up at 2/3, though it would surprise me a little bit. Anyhow, even just a 60% winrate for extra innings would give them a 0.82% chance of getting 16 out of 18 or better, and 2/3 would raise that to a whopping 3.3% chance. So yes, it was definitely still really unlikely, I believe it was at least 10 times as probable as the quick 50-50 assumption would get you.

In the last installment of the series, we’ll take a look at why this consistency effect isn’t showing up in the studies that have been done, and why I still think it’s there anyway.

07 March 2013

Distributions and Consistency Part I


Most of the sabermetrics you’ll find out there deals in terms of runs. But who cares about runs? We want to win games. Well, the basic point is that generally if you increase your average run scoring (or decrease the average amount of runs you’re allowing), you’re going to win more games. But this is only generally correct, and it is imprecise.  More exactly, you win when you score more runs than the other team. What this means is that we don’t need the mean of the distribution or any other measure of its center, really, we need the whole distribution itself.

Fortunately, there are lots of a priori tools which can give you the whole distribution and not just a point estimate. Any simulator can give it to you (that doesn’t mean they all do, but it’s not too hard to program your own, and there’s probably at least one out there which already does this for you). The homogeneous model can do it very easily. And this http://www.tangotiger.net/markov_wes.html site, which I literally have bookmarked as Markov Gold, will do it for you, too. What do you see when you do this? Well, you get a characteristic shape that looks like this:


(Incidentally, if anyone knows how to get graphs in here at all better than this - which is pretty terrible, I can only figure out how to save it as a JPEG and then it's very difficult to resize... anyway, help here would be more than appreciated).

It’s worth noting that for this particular graph (a quick grab of AL 2012 stats, without much cleaning), the mean number of runs per game is 4.57, and the median is 4. Clearly, the mode here is at 3 runs per game. The exact numbers aren’t important, but what is important is the shape. Technically, it’s possible to get graphs that don’t look like this – but to do so, you need to use values which are way, way off from any we’ve ever seen in the ranks of competitive baseball (MLB, MiLB, NCAA, big-time softball, Little League, Japanese leagues, Olympics, WBC, I presume high schools… pretty much anything you can name).

And with a graph which more or less resembles this shape, the big thing to notice is that it is skewed. The inequality Mean>=Median>=Mode (and you only get equality for degenerate values or for median and mode because of the discrete nature of the distribution) is going to hold. The reason this is important is that it generally means that more offensive consistency is going to trump offensive inconsistency (and conversely, pitching + defensive would prefer to be inconsistent).

Next time, we’ll look at what leads to this kind of consistency and how big of an impact it can have.

28 February 2013

Homogenous vs other run modelers

In the comparison of the Homogeneous model to other run-scoring modelers out there, it actually performs pretty darn well. I want to look at a few of the most notable ones, which can beat it, and why that happens.

First of all, there's simulators. Now, simulators with the same assumptions and parameters as this model will always underperform it, but if you include extra parameters (like lineups and matchups and such), then it's not very difficult to make something which outperforms this model, given enough iterations of the simulator.

A Markov Chain Modeler - Actually, if you take the time to improve the homogeneous model by allowing advancement on outs (and include strikeouts separately, as this necessitates) and having differing values of advancement on singles and doubles depending on how many outs there are, they turn into the same thing. Now, the markov is a heckuvalot simpler to program, so that is a big plus going for it, and it also has those extra features, at least right now, which help it be more accurate as well. (Incidentally, both of them can also be improved by giving some set value for a double play if there's a guy on first with less than two outs, and then further refining this based on other factors like if there's a guy on third with nobody out and the defense is 'giving' the run; this is all a lot of work though). The Homogeneous model has some advantageous in terms of practical work, though. First of all, it's very easy to get a break-down of run-scoring with it - though this version of the Markov is a piece of solid gold, as it will do all that stuff for you as well. Still, the closed form, symbolic nature of the solution that the Homogeneous model gives you is still useful for some particular manipulations you might need to do - for isntance, you can take partial derivatives to get PRECISE linear weights - or if you just really like being able to do it in equation form, and not haivng to punch the numbers in.

Various linear guys - these will do well, essentially by taking the empirical advantages to the extreme. They get all the stops possible out of typical lineup construction, talent distribution, pitching usage, maybe even particular managers' or players' idiosyncracies, by pulling data from really recent previous seasons. They pay the price for this 'cheating', though, if you get them in the wrong spot, by changing the data source very much at all - i.e. if it's trained on 2008-2011 data, it may predict 2012 really well, but give in 1985 and it will fall flat; furthermore, if you want to look and see what value you should give things, forget about it - these are not very predictive at all.

BaseRuns - BaseRuns is the reigning champion, except maybe simulators, of the run modelers that are out there. But while a part of me really really wants to love it based on the simplicity of the formula and the intuitive rightness that is there, it still has some issue. Namely, the B value here is determined empirically, just like in those linear models, and while that does have many accuracy advantages, it can also lead to ruin if you're looking at something very different from your training set - even if it just happens to be a team or pitcher (or umpiring crew or ballpark or...) that is very different from the average of what you have been looking at. So watch out!

In any case, you are going to be very limited by your projections, which are almost always going to be less accurate than any of this stuff. In the limit as the projections approach more perfect accuracy, though, the Markov/Homogeneous version of things (which can by approximated by simulator with enough trials with less difficulty in construction) is going to win out, and the linear ones will fall into the dust.

Re-Introducing the Homogeneous Model

Some time ago, I introduced what I call the Homogoneous model of run-scoring. However, I don't think this is the best explanation, and together with launching this new blog, I wanted to give it a fresh introduction.

The essential assumptions of the model are that every plate appearance has the same probabilistic distribution of outcomes as every other plate appearance; there are never any outs made on basepaths (i.e. no double plays, pick-offs, caught stealing, etc.); there is never any play of any kind which happens in the middle of a plate appearance (so no steals, passed balls, wild pitches, etc.), and that runners don't advance on outs (this one can be eliminated, though this would require roughly tenfold the work as went into the entire model so far). The model accepts nine parameters - rates for walks, singles, doubles, triples, homers, and outs (these must be constrained so as to add to one); as well as the probabilities that a runner on first will advance to third on a single, such a runner will score on a double, or that a runner from second will score on a single. Technically, it is also assumed that all runners advance at least as much as the batter on singles, doubles, and triples, and whilst this isn't necessarily always the case, in almost every application, the assumption is quite a safe one to make.

The model works on an inning basis. The basic premise for the model is that the weakest positive offenisve event is a walk, and four walks scores a run, five walks scores two, six walks scores three, etc. We can then mathematically derive that every team which only makes walks or outs will have (given an on-base percentage p) a chance of p^(n)*(1-p)^3*(1/2)*(n^2+3n+2) of scoring n-3 runs (for n>3). (The trick in getting the combinatorics right is that you know for sure that the final event of the inning was one of the outs; statisticians may recognize this as an instance of a negative binomial distribution looking for 3 failures). We can then get an overall run expectancy by taking the product of each of these probabilities with the number of runs being scored, and then summing from one run to infinity, which eventually leads us to (6*p^6 - 18*p^5 + 15 * p^4)/(1-p), which is the expression for the number of runs you score in an inning if all your positive events are walks, or in other words, the minimum number of runs you'll score at a given OBP.

To make the leap from this to a full distribution, given all the different possible offensive events, you simply have to realize that only the last three successes matter in an inning with three or more successes, as the previous ones are going to have been taken care of by the distribution from walks, as derived above. So you make an adjustment for the last three successes in an inning, and adjustments for the runs you score if you only get one or two successes in an inning, and you end up with the very long and complicated formula for run scoring given in the link at the top of this article.

Okay, so what are the strengths and weaknesses here? Well, first off, the assumptions are fairly hampering on the model. Steals and various during-a-PA basepath information is actually not such a big effect, as it ends up mostly cancelling out. Double plays, however, will bring the runscoring down a fair bit. But the bigger issue is every PA being the same as every other. This is just wrong. There are three reasons. The first is that the lineup is constructed in a way that makes certain sequences of events more or less likely than others. This probably helps the lineup slightly, though it depends somewhat on the actual construction. The second is that in different situations, different events are more or less likely to happen. With the bases loaded, for instance, the batter will be more apt to take a walk, and the pitcher will be less likely. Defensive positioning, mental approach, kind of swing, kind of pitch, aggressiveness - all of these things can change, and they can all have potentially very significant impacts. Exactly which side benefits out of this is something that empirical research would need to be done on. If this is out there at all, I'd be grateful for you sending it to me.

The final reason for the difference is that the strength of opposition is not constant, but rather it varies. So, if you're looking at an offensive lineup, the opposing pitchers will have different strengths, and if you're looking at a pitcher, the opposing offenses will have different strengths. Defense, park and weather factors, umpiring, and time-throug-the-order effects all play here as well. You also have to watch out for the expectations of the given situation, e.g. with platoon splits. In any case, all of these variations will serve to increase the run-scoring in comparison with the average case. The reason is that run-scoring is not linear. In fact, it's VERY non-linear. You can actually see this just from the on-base terms, where you have factors of OBP to the fourth, fifth, and sixth powers, and much more importantly, you are dividing by (1-p), which makes enormous impacts in making the curves go nearly vertical even on logarithmic scales near OBPs of 0 and 1. But the general point is, for any given average, the time you score less, you score rather significantly less, but the times you score more, you score WAY significantly more. And the overall effect is that if you add it all up, the mores are going to always overway the lesses, and this effect increases the higher your variance from the average.

A Priori Sabermetrics: An Explanation


The kind of Sabermetrics most people focus on is what I call “a posteriori sabermetrics” (or “empirical sabermetrics”). It uses data in order to construct statistical models in order to measure and predict performances in the game of baseball (or other sports, though I’m going to focus on baseball). It is based on established data, and it is also therefore dependent upon this data. This approach has a number of advantages. Data is fairly plentiful, and a number of statistical tests and models are increasingly common and easily done with recent advances in technology. It also has an advantage in how it focuses – it does not necessarily need to model every piece of how a final measure comes into place in order to be able to predict it. For instance, if you are looking at runs scored, you can focus on only a few statistics – maybe just the various different kinds of hits, walks, and outs – and not worry about things like lineup construction, ballpark, etc., and even with simple estimates, you can come up with a pretty good estimate, as all of the other things you are ignoring get taken care of as hidden variables which you essentially just copy over from past to future. This is particularly useful in that you don’t always know what all the pertinent factors are.

The problems you can run into with a priori sabermetrics come when things start to change. When you aren’t explicitly modeling something, a change in it is going to disrupt your statistic, potentially very greatly. You also get problems when you try to take your model beyond what it was developed on – for example, if you are trying to predict run scoring from a model built on leagues where OBP runs from .280 to .370, you have to be incredibly leery of applying this model to a team which has OBP over .400 or under .250 (or even when you get near the edges of your bounds at all, in some cases).

 

A priori sabermetrics, on the other hand, don’t require any data to build at all. They instead rely only on a knowledge of the rules, as well as almost always some assumptions and/or parameters. There are a couple of ways of doing a priori work. My most preferred method, possibly due to my deductive, scientific background, is to try to develop purely mathematical models. The nicest thing about these is that they can be absolutely and perfectly accurate, constrained of course by any assumptions and parameters. They also take essentially no time to run once they are computed. Finally, the equations you get here, being ‘real’ equations, and not just random artifacts or constructs, can be manipulated and adapted for other uses that you might have for them down the line, provided that your assumptions still hold in that new frame of observation.

Another, more popular kind of a priori sabermetrics is simulations. The nice thing about these is that they are really not very difficult to program – you can get a pretty good one going with a computer, a compiler, and several free hours. They can also be run to arbitrary decision – if you need more accuracy, you just run the simulator for more trials. Because they’re so easy to program, they’re also pretty easy to adapt, and they can be used in quite a wide variety of applications.

In either case, a priori sabermetrics, not being derived from data, are neither enslaved to it – they are applicable to ALL situations in which the assumptions hold and the required parameters are available. They are good for all kinds of crazy and wacky run environments. Perhaps their biggest triumph is when you are looking at major changes between the data you have ample supply of to look at and what you are trying to adapt it to – if you are trying to run sabermetrics on your local little league, a priori metrics are going to be far friendlier to you than most of what is being currently developed for the majors. But perhaps on a more practical level (man, don’t apply saber stats to the little league…), if you are trying to make a concerted change, a priori sabermetrics are going to be much better to you than empirical ones. For instance, if you want to look at having batters changing their approach in such a way that home runs and strikeouts both go way up, to ridiculous proportions, but things like singles and doubles go way down, you aren’t going to have any empirical basis for this change, and you need to rely on more purely theoretical a priori stuff. Of course, the big drawback of these kinds of statistics are that they are a slave to their assumptions and parameters – or lack thereof. If you are looking for run scoring as a function of ten factors, you can PERFECTLY model it to those ten factors, but you have absolutely no feel for any other factors you aren’t modeling. Where the a posteriori stats can buffer for this by assuming the same hidden mechanics will persist into the future as they’ve been in the past, here you have nothing.

Ultimately, the answer to these problems is to simply account for those hidden variables. But as you always have these parameters floating around, you are going to need at least some a posteriori research – which can actually include non-statistical things like scouting as well – in order to give you these inputs. But the further down the line this is, the more robust of a system you are developing. If you try to model runs straight off, you aren’t going to get as far as if you theoretically develop a system of how they are produced based on things like the rates at which basic events happen (e.g. walks, doubles, home runs, strikeouts, etc.), and then try to predict those basic statistics based off of your empirical research.