2025-07-21
Albert Cohen and Jimmy Risk
This lecture is taken from a presentation by Dr. Cohen on March 7, 2025 at the Simon Fraser University Sports Analytics Group Virtual Seminar Series. That recording can be found HERE and the companion code to the paper HERE
Player valuation in European football is a complex endeavor, requiring nuanced metrics that go beyond traditional sports statistics
Due to the dearth of goals in a game, other in-game events should be utilized to understand the contrbution of individual players to the team’s performance
Existing methods often fail to capture the dynamic, stochastic nature of player performance and its impact on fair market valuation.
The authors introduce a multi-disciplinary approach that integrates financial models with stochastic player performance models, borrowing from social network theory. Their main objectives are
The authors identify at three foundational works that have provided in-depth analyses of various pieces of this puzzle:
The first major work to address the link between performance and pay for athletes can be found in the paper of Scully (1974)
In this paper, the author investigated the connection between a baseball team’s revenue and the salary paid to the team’s players. By utilizing well-known tools in labor economics, such as the marginal revenue product (MRP) related to units of labour, the author derived a framework to link on-field play with salary.
Classical economic theory suggests that a profit maximizing firm will hire by setting wage equal to marginal revenue product of labor, which is the increase in revenue attributable to a unit increase in labor
A common sports analytics application is that a player should be compensated in relation to their contribution to team performance.
The second major work referenced by the authors is that of Tunaru, Clark, and Viney (TCV) (2005), which addressed the value \(V\) of a player to a team, in continuous-time, via stochastic modeling and a subsequent Black-Scholes partial differential equation (PDE).
The solution \(V\) of that PDE also includes the assumption that contracts can be sold at the end of the term
A Black-Scholes partial differential equation is used to model changes in the value of options (rights to buy and sell assets) over time, written as
\[ rS_t\frac{\partial V}{\partial S} + \frac{\partial V}{\partial t} + \frac{1}{2}\sigma^2 S^2 \frac{\partial^2 V}{\partial S^2} -rV = 0 \]
Here, \(V\) is the value of the option at time \(t\) with underlying asset price \(S\), \(r\) is the risk-free interest rate, and \(\sigma\) is the volatility of the underlying asset.
Let \(T\) be the turnover for the club, \(N\) be the number of Opta Index points for the individual player under evaluation, \(S\) be the sum of Opta Index points for all players playing for the club.
Financial modelling can be adapted to deal with the assumption that the jumps affect directly the evolution of the underlying.
Finally, the work of Rockerbie and Easton (2020) in their recent book offers a discrete, multi-period approach to contract pricing via expected present valuation and cohort analysis via a market beta.
In applying the principle of linking pay to a player’s MRP, the authors of this paper seek to connect the discrete landscape of Rockerbie and Easton with the continuous stochastic modeling of Tunaru, Clark, and Viney by defining a player peformance share as a stochastic process correlated to a team’s revenue.
The continuous-time TCV model proposes that a player’s performance is measured by a point (or performance) score that lives on the interval (0, \(\infty\))
The authors model this score for player \(j\) as a geometric Brownian motion (GBM) \(p_j\) that is correlated to another GBM in \(R\) that models the team’s revenue.
The result is that the revenue earned, per point, by the team is \(R/\sum_j p_j\) and the number of points \(n\) that multiplies this fraction returns the performance-linked value of player \(j\) to that club, \(Z_j := p_j R/\sum_j p_j\)
A standard Brownian Motion \(W(t)\) is defined on \(0 \leq t \leq T\) as
A stochastic process \(S(t)\) follows a geometric Brownian motion if it satisfies this stochastic differential equation (SDE)
\[ dS(t) = \mu S(t) + \sigma S(t)dW(t) \]
The authors propose linking this stochastic model to that of Rockerbie and Easton by rewriting the performance linked value \(Z_j\):
\[\begin{equation} \label{eq:perf} \tag{1} Z_j = p_j \cdot \frac{R}{\sum_j p_j} = \frac{p_j}{\sum_j p_j} \cdot R : = \pi_jR \end{equation}\]
For example, in soccer if one simplifies to think of 11 key players, a uniform distribution of performance load would suggest that each player would provide \(\pi_t = 1/11 = 0.0909\) per game
This obviously does not happen, as there are injuries, substitutions, matchups that enhance or detract from a specific player’s ability to contribute, and even worries about upcoming contract negotiations, to name a few (of a multitude of factors).
Finally, there are many in-game decisions that can lead to performance estimation, and this requires the special attention of analytics providers that may keep such metrics away from public view.
To address these and other issues, the author’s model can be calibrated using publicly available data
In doing so, they compute (augmented) passing matrices derived from in-game stats to apply network analysis methodologies similar to that of Pena and Touchette (2012) and Duch et.al. (2010; 2015)
Applying network theory allows us to define, according to managerial input, what drives performance share. Once this is established, we can calibrate the model in (Rockerbie and Easton 2020) below:
Consider the univariate approach of a single player \(j\).
Thus, we need to determine the evolution of \(Y_t = \pi_ta_tR_t\), which appears as the product of the two stochastic variables \(\pi\) and \(R\), with an update for a dynamic fraction \(a_t\) that represents the portion of revenue set aside for player salaries.
To value scoring attempts and devalue missed passes and turnovers, the authors augement the passing matrix wth two rows and columns: one for shots (\(S\)) and another for unsuccessful passes and/or turnovers (\(U\)).
This augmented passing matrix is denoted by \(P\). Here, “shots” can refer to successful shots, total shots, or a weighted combination of successful shots and missed shots.
\[ P = \left[ \begin{array}{ccccccc} 0 & 0.40 & 0.25 & 0.20 & 0.10 & 0 & 0.05 \\ 0.15 & 0 & 0.34 & 0.25 & 0.10 & 0.01 & 0.15 \\ 0.05 & 0.15 & 0 & 0.20 & 0.30 & 0.05 & 0.25 \\ 0.05 & 0.15 & 0.20 & 0 & 0.20 & 0.10 & 0.20 \\ 0 & 0 & 0.25 & 0.25 & 0 & 0.15 & 0.30 \\ 0 & 0 & 0 & 0 & 0 & 1 & 0 \\ 0 & 0 & 0 & 0 & 0 & 0 & 1 \\ \end{array} \right] \]
\[\begin{equation*} q_j = \mathbb{P}[\text{player } j \text{ is involved in possession} \mid \\\text{possession ends in }S \text{ not } U] \end{equation*}\]
Extreme cases provide some quick insight: if \(q_j = 1\), it means all shots were filtered through player \(j\) in some manner. Similarly, \(q_j=0\) means that player \(j\) is not involved in any possession that ended in a shot.
Even for strong defensive players, we would still expect \(q_j\) to be relatively large, as many shot-based possessions will involve them beginning with the ball.
In order to relate to the \(\pi_t\) process,
\[ \pi_{j,t} = \frac{q_{j_t}}{\sum_iq_{i,t}} \]
Note here we emphasize the subscript \(\pi_{j,t}\), instead of \(\pi_t\) as we are comparing across players on the team.
This ensures \(0\leq \pi_{j,t} \leq 1\) and transforms \(q_{j,t}\) into a measure of relative player importance.
\[\begin{equation} \begin{aligned} Y_t &= \pi_t a_t R_t \\ d\pi_t &= -\theta(\pi_t - \pi^*)dt + \sigma_\pi \sqrt{\pi_t (1 - \pi_t)}\, dW_t^{\pi} \\ dR_t &= \mu_R R_t\, dt + \sigma_R R_t\, dW_t^{R} \\ \rho\, dt &= \langle dW_t^{\pi}, dW_t^{R} \rangle \end{aligned} \tag{4} \end{equation}\]
Consider a scenario where a player can sign a contract with a team for multiple years, with
the negotiated salary payments, \(C_1, C_2, \ldots, C_T\) with the amount \(C_t\) computed as if paid at the end of the \(t^{\text{th}}\) year, and
the random dollar-value of the athlete’s performance-linked contributions on the field, \(Y_1, Y_2, \ldots, Y_T\).
There is a baseline expectation for performance, but management allows for the possibility that the player can under- or over-perform, in terms of expected statistical contribution
\[ \sum_{t=1}^Te^{-rt}C_t=\sum_{t=1}^Te^{-(r+\lambda)t}\mathbb{E}_0[Y_t] \]
To see how a contract negotiated at time 0 has aged by the middle of the term, we define the time \(t\) valuation of a player \(j\)’s performance-linked value as \(S_{t,t+h}\).
In words, this is the new (rolling) value of a fixed contract if the manager could re-negotiate with a single payment at time \(t\) and single stochastic valuation at time \(t+h\).
For such a 1-period swap, the agreed upon time \(t\) value at time \(t+h\) is \(S_{t,t+h}\), where
\[ e^{-rh}S_{t,t+h}=e^{-(r+\lambda)h}\mathbb{E}[Y_{t+h}\mid\mathcal{F}_t] \]
Consider a case where a player’s agent has asked for a three-year contract where
the salary is front-loaded with year 1 salary \(C_1\) triple that of years 2 and 3, i.e. \(C_1 = 3C_2 = 3C_3\), which could act as a signing bonus,
the risk-free rate \(r\) is set at the current 3-year UK government bond rate of \(r = 0.0375\),
an agreed-upon risk-premium of \(\lambda = 0.01\), and
a constant fraction of \(a_1 = a_2 = a_3 = 0.5\) of revenue directed towards player salaries over the next three years.
Additionally,
initial revenue \(R_0 = 100\) (in \(\text{ €}1,000,000\)) and performance share \(\pi_0 = 0.08\)
long-run performance expectation \(\pi^* = 0.10\)
correlation between revenue and performance fraction set at either (a) \(\rho = 0\) or (b) \(\rho = 0.30\),
corresponding (R,\(\pi\)) volatilities fixed at \((\sigma_{\text{R}}, \sigma_\pi)\) = (0.20, 0.25), mean reversion rate of \(\theta = 4\)), and a growth rate \(\mu = 0.10\) for the team’s revenue
Applying our swap model, the resulting balance equation for the full contract is now
\[\begin{equation} e^{-r}(3C) + e^{-2r}(C) + e^{-3r}(C) = e^{-r} S_{0,1} + e^{-2r} S_{0,2} + e^{-3r} S_{0,3} \tag{8} \end{equation}\]
Plugging in \(r = 0.0375\), the year 2 and 3 salaries \(C\) can be solved to be:
\[\begin{equation} C = \frac{e^{-0.0375} S_{0,1} + e^{-0.0375(2)} S_{0,2} + e^{-0.0375(3)} S_{0,3}}{3e^{-0.0375} + e^{-0.0375(2)} + e^{-0.0375(3)}} \tag{9} \end{equation}\]
Now, using the parameters described on the previous slide, we can use Taylor Series to approximate \(\{S_{0,1},S_{0,2},S_{0,3}\}\)
\[ (C_1, C_2, C_3) = (10.6046, 3.5349, 3.5349) \]
\[ (C_1, C_2, C_3) = (10.7236, 3.5745, 3.5745) \]
The authors emphasize here that numerical methods for pde’s could also be used to solve for the values \(\left\{e^{-\lambda j}\mathbb{E}[Y_j\mid\mathcal{F}_0]\right\}^3_{j=1}\) detailed in the swap (8).
The authors make the following observations:
The salaries \((C_1,C_2,C_3) \mid_{\rho=0.30} > (C_1, C_2,C_3) \mid_{\rho=0}\).
This would coincide with rewarding a player whose increased contributions have a positive effect on team revenue. A general formula for the corresponding “Greek” \(\frac{\partial \text{C}}{\partial \rho}\) would be interesting and valuable in player-team negotiations, and is the subject of further research.
A player agent would be expected to argue that their client’s performance over the proposed contract has at least some partial (positive) correlation (i.e. \(\rho > 0\)) with team revenue
Assume \(\rho = 0.30\) for the \(K= 5\) seasons considered
Revenue data, available annually, is from Deloitte
Salary data from Sportrac
Game data is from Whoscored
Game appearance data is from Transfermarkt, and is available for every game over the seasons of interest
Mohamed Salah (right winger), the team’s primary offensive threat;
Trent Alexander-Arnold (right-back), recognized for his creative playmaking; and
Virgil Van Dijk (center-back), known for his defensive stability and leadership.
All players were starters for the considered seasons.
Eddie Nketiah (striker), a young goal poacher;
Granit Xhaka (central midfielder), a key figure in controlling the game’s tempo; and
Rob Holding (center-back).
Granit Xhaka was the only consistent starter for all seasons.
Pascal Gross (attacking midfielder), the team’s creative engine;
Solly March (winger/full-back), valued for his versatility and work rate; and
Lewis Dunk (center-back), the cornerstone of Brighton’s defense.
All players were starters for the considered seasons.
With a variety of teams and consistency of positions, these players provide a practical basis for validating the estimates.
The contributions in different positions provide a comprehensive view of the dynamics of \(\pi_t\) over time, reflecting both the offensive and defensive aspects of the game.
Additionally, it considers players who are not starters as well as a team lacking superstars (Brighton).
Data are collected per game and player shares are statistically estimated
Estimates related to the performance share process for each player under analysis. The quantity in parentheses is the standard error of the estimate. Game-specific data is from https://www.whoscored.com
Table: Estimates of risk premium \(\hat{\lambda}\) over players. Trent refers to Trent Alexander-Arnold.
If the full revenue data were available, all parameters could be estimated more accurately, including a more precise estimation of \(\rho\).
Similarly, the player share \(a_k\) over season \(k\) should practically be known from the club’s financial statements.
As these data are unavailable, we estimate it as the historic fraction that went out to the club’s players over season \(k\) for each season.
Estimates of \(\pi\), \(\sigma_R\) and calibrated values for each team’s player share \((a_k)\) process.
For the figures on the next two slides:
1st slide: Calculated \(\pi_t\) processes using historical game data. The thick dark line is the average over that season, and the dashed line is reference of \(1/11 \approx 0.0909\), the share if a team consisted only of starters each with equal share
2nd slide: Time \(t\) on x-axis, \(\text{\$}1,000,000\) (annually) on y-axis. Plotted are \(C_t\) (thick dark line) and \(S_{t,T}\) (colored time series plots), where \(T = 1,2,\ldots,5\) represents seasons, and \(T-1 \leq t < T\) for each season.
See Figure 1 of (Cohen and Risk 2025)
See Figure 2 of (Cohen and Risk 2025)
The valuation model reveals significant patterns in how player compensation aligns with performance, highlighting both consistencies and discrepancies within current salary structures in European football.
Several players were consistently undervalued relative to their contributions before receiving pay increases.
For example, Mohamed Salah’s performance data from 2018-2021 suggested a need for a substantial pay increase.
Indeed, his value was higher in 2022-2023 but still below the salary provided, indicating a possible overvaluation, or most likely factors not considered in the model.
Trent Alexander-Arnold was under-compensated until the 2020-2021 season; his significant pay raise in 2021-2022 aligns with our model’s valuation. However, he continues to outperform his compensation, suggesting that he remains undervalued even after the increase
Granit Xhaka also consistently demonstrated performance exceeding his compensation. The minor salary increase he received in 2022-2023 does not fully reflect his contributions, indicating that he is still undervalued according to the model.
Proposed approach that expands mapping \(\text{performance} \rightarrow \text{salary}\) to \(\text{performance} \rightarrow \text{revenue} \rightarrow \text{salary}\)
Is there a dynamic Pythagorean exponent that could perhaps help with some of this analysis?
Could closeness centrality measures be used to imply the ratio of offense to defense in a team?

STAT 468 - Introductory Sports Performance Analysis