How to Build an Effective Player Index
Choosing metrics and weights, and avoiding the risks that can skew your rankings and lead you to the wrong players
You have identified a hole in the squad: a deep-lying playmaker. But going from that need to the scouting team’s final recommendation is a long road. In this article, we will look at a key, foundational stage of this process: building a player index.
We will go through the necessary steps to build an effective player index, from defining the targeted profile, a specific subset of metrics and their assigned weights, through to the three risk areas to watch out for.
Where It Fits in the Workflow
In a standard scouting workflow, building an index sits at the player search stage. It is the moment you need to narrow down a large set of players to a shortlist of possible targets. A poor choice of index metrics or an unbalanced weight distribution across them can prematurely remove players from consideration or ensure that unsuitable ones make it through to the next stage.

What is an Index and How to Build One
An index is a single number built from the weighted sum of normalised metrics. Indices are extensively used in football to filter down a large set of players into a shortlist of potential options who match a targeted profile.
To build one, you need three things:
- A target profile (e.g. a "ball-playing centre-back").
- A chosen subset of metrics that map to the attributes required for the target profile.
- Assigned weights for each metric.
That sounds simple enough, at first. The problem is that index rankings are subject to mathematical significance, which does not always link to practical, football relevance.
An index is, at its core, mathematics. It takes numbers, normalises them, adds them up according to their assigned weights and spits out a ranking. While that exercise will always produce a neat result, it doesn’t necessarily mean that the difference between two players ranked far apart will be meaningful in practical football terms. You have to account for the choice of metrics and how they are distributed across the whole sample.
Below, we will explore three potential pitfalls to watch out for when index-building: why they matter, how they could skew your results and how to avoid them.
Risk # 1 - Single Metrics with Poor Discrimination
The first question to ask about any single metric before adding it to your index should be whether it actually separates players. That is what the Coefficient of Variation (CV) tells you.
The CV is a metric’s standard deviation divided by its mean and multiplied by 100. It allows you to evaluate the spread of a distribution.
A low CV means most players score close to the average, so the metric is barely separating players. In the example below, using Avg PSV 99 — a metric with a CV of 3.9% — a 43-position ranking difference turns out to be driven by an 86-millisecond difference in a 10-meter sprint – statistically real, but basically imperceptible on the pitch.
(The sample used in this article is comprised of midfielders with at least 5 games of 60 minutes in the Big Five leagues. Conclusions depend on the sample)
Alternatively, Sprint count p90, a metric with 0.81 correlation with PSV 99 and a CV of 37.6%, captures much of the same signal, but separates players better.
To further illustrate that point, you can see that while the charts below show a similar difference between the 25th (P25) and 75th (P75) percentiles after normalisation, the PSV 99 one indicates only a 67 ms difference in a 10-meter sprint (not football-relevant), while the sprint count points to a difference of 4.5 sprints per 90 minutes (fairly relevant).

While a higher CV generally indicates better separation, on some occasions it can go too far and be misleading. Take, for example, Line-breaking passes through the first line (per 30 TIP). Its CV of 100% stems from the metric's relationship to a very specific midfield role. It doesn't do a good job of separating players; it highlights only a small handful of specialist number 6s, while most players compress near zero. This would only be useful if you are analysing specifically number 6s that receive before the first line.
As you can see above, a difference of only 1.41 passes through the first line per 30 TIP brings about an 88-position swing, driven by a metric with an extreme distribution.
A more reliable alternative for a broader search is Line-breaking passes through the second line, with a more workable 58.9% CV and a 0.64 correlation to the first-line metric. After normalisation, Passes through the first line show only a 0.83 difference between P25 and P75, against 1.49 for Passes through the second line for the same rank gap.

Evaluating a metric's discriminating ability, both statistically (via CV) and in practical football terms, is a critical step before including it in an index. Selecting metrics with stronger discrimination power that rely less on specific tactical scenarios provides a more reliable foundation for effective player comparison.
Risk # 2 - Overly Correlated Metrics
Separation power isn’t the only thing to keep in mind when selecting your metrics. Your next question should be how closely two or more of your metrics are correlated to each other.
For example, High Intensity Count (p90) and High Intensity Distance (p90) correlate at 0.966. Effectively, it’s as if you were including the same metric twice. The impact of these correlated metrics on the results will be higher than intended, doubling the weight of one single underlying attribute.
If you build an index with, say, five metrics, each meant to account for 20% of the final score – a fifth for Possessions, a fifth for Off-ball runs, a fifth for Intensity, etc – and two of those metrics are ultimately measuring the same ability, you end up with an unbalanced result.
Inversely, negative correlation may cause a different type of issue. High intensity count (per 90) and Line-breaking passes through the second-last line (p30 TIP) correlate at -0.51. This means that, even amongst the top-ranked players, you are less likely to find players who score high on both metrics. If your target profile genuinely needs both qualities, you should expect your pool of realistic candidates to shrink significantly.

On the flip side, that scarcity could also be occasionally read as a strength. Finding profiles who score high on both can be a legitimately rare and valuable find, which could be worth focusing on.
Evaluating metric pairs for correlation, whether positive or negative, is, once again, a crucial step. A high positive correlation essentially duplicates a metric and doubles its weight, while a strong negative correlation may result in a search for combinations that are statistically rare or even unattainable.
Risk # 3 - Too Many Metrics
While it might be tempting to think that adding more metrics to an index would make it more complete, the data shows otherwise. In general, the more metrics you include, the more you lose understanding of what drives a player's high rank, and the more weak spots even top-ranked players will show.
The chart below displays the top 25 players on a series of random draws of metrics (uniform weights, 100 random sets per point), tracking what share of an index’s metrics falls below the 50th, 75th and 90th percentile as the number of metrics grows, from five to 20.

At five metrics, 18% of slots for a top-25 player fall below P50, 36% below P75, and 60% below P90. At 20 metrics, those figures climb to 29%, 54%, and 77% respectively.
Below, you can see what that looks like in a real-world situation. This index is very similar to one a prospect on trial built in our WebApp Lab, with 18 different metrics. We can surmise that it brings an elevated risk of detrimental CVs, correlations (counts and distances), and weak spots due to the high number of metrics.
And this is what that index produces: Rank 1 is elite across almost every metric, but has very visible weak spots. Rank 34 has no red flags at all: a steadier, more even and complete profile than the player ranked far above him. If the brief was to find the most well-rounded player, it's hard to argue that this would be the intended result.
The same disconnect shows up again with this request we received for a "#6 Playmaker".
- Technically composed, shows for the ball in possession and is calm under pressure. Highly press resistant. —> Passing Options in Create, Retained Ratio Under Intense Pressure.
- Progressive passer, highly efficient with through balls in all 3rds of the pitch. —> Line-breaking passes through all 3 lines.
- Tactically aware and capable of changing the point of the attack through switches of play.
- Game awareness to control tempo and balance retention with risk.
- Intelligent defender, efficient at screening, condensing space and intercepting opposition progressive passes.
- Physically capable of covering high distances to always be available in possession and to counterpress out of possession. —> High Intensity Distance, Counterpress Engagements.
Even before building an index, you can deduce that finding an all-round profile that can do all those things, ranked over P60 in his league for the metrics above, for example, is extremely difficult.
The result? Only one player out of 541 met that bar.
The alternative approach here would be to focus only on what matters most at the shortlist stage, with fewer requirements and the appropriate weighting, assessing other attributes further down the line in the scouting workflow.
Closing Thoughts
Building a player index is one of the most efficient ways of navigating through a large set of players at the search stage. It does come, though, with its own challenges. And as one of the foundational stages of a scouting process, it requires a special level of scrutiny. The wrong setup here can lead to a strong candidate being filtered out before anyone gets to watch him play.
It is crucial to understand what each metric in it is actually measuring, how they relate to each other, and what you are trading off each time you add one.
Go Further with SkillCorner data
If you’d like to find out more about SkillCorner data, our easy-to-navigate WebApp Lab platform, and explore how you could leverage it in your own scouting processes, get in touch today.
.webp)


