IN PROGRESS: Optimal Pension Design: A Naive Approach

In this post we look at why pension plans exist, how they work in Belgium, and how we can naively search for an optimal plan design using a model. The goal is to explain the reasoning behind the model: why you would want such a model in the first place, and how it works under the hood.

Why a Pension Plan?

The honest answer is that most of us are bad at saving for a future self we have never met. This is less a character flaw than a wiring problem: we discount the future steeply, and the person who will actually need the money — retired, decades older — reads to us almost like a stranger we are being asked to sacrifice for. Left to our own devices we under-save, retire into a consumption cliff, and fall back on the state.

And that stranger keeps living longer. Rising life expectancy sounds like unambiguous good news, but for anyone financing their own retirement it introduces a genuinely hard problem: longevity risk, the risk of outliving your money. You cannot simply divide your savings by your remaining years and spend the quotient, because you do not know the denominator. Plan to run dry at 85 and you are destitute at 90; hoard against reaching 100 and you deny your younger retired self a decent life along the way.

The institutional response to both problems — our present bias and our uncertain horizon — is the pension plan: a contractual commitment, made while you are earning, to a future version of you that you cannot be trusted to provide for spontaneously. Concretely, your employer sets aside part of your salary in a fund that is invested until you retire.

But there is a second, less obvious answer, and it is the one that motivates this whole project. A pension plan is not merely a savings vehicle; it is a risk-sharing contract between three parties — employee, employer, and, through guarantees and regulation, the state. Each wants something different. The employee wants a stable, predictable replacement income and as little investment risk as possible. The employer wants costs that are affordable and, above all, foreseeable — a liability it can budget for, not one that balloons the moment markets turn. The state wants neither party to fail: not the employee retiring into poverty and falling back on the first pillar, nor the employer defaulting on a promise it cannot keep.

These wants are not just different; they are in direct, structural tension. The very instrument that protects the employee — the guaranteed minimum return we are about to meet — is precisely the liability that threatens the employer, because a guarantee is a promise to cover the gap whenever the market underdelivers. Dial the guarantee up and you reassure the employee while loading risk onto the employer; dial it down and you do the reverse. No setting is costless for everyone. A pension plan is, at bottom, a particular choice about where on that spectrum to sit.

And notice that the three parties do not enter symmetrically. The employee's and the employer's objectives are the two actually traded off against each other — more for one is, at the margin, less for the other. The state does not sit on that spectrum at all: it draws the boundaries within which any trade is allowed to happen — minimum guarantees, contribution caps, non-discrimination — and rules out the corners where one party walks away holding an empty bag. That distinction — objectives to be balanced against each other, versus constraints that simply must be respected — turns out to be exactly how the model in this project is structured. Which raises the question the post is really about: given that a plan's design is a choice about how to balance these competing claims, can that choice be made optimally — and what would "optimal" even mean?

Belgium Legal Framework: the Three pillars

Belgian retirement income rests on three pillars. The distinction between them is not bureaucratic tidiness — each pillar places the risk of funding your old age on a different party, and the project lives entirely inside the one that shares that risk most explicitly.

The first pillar is the statutory state pension. It is a pay-as-you-go system, and that phrase is worth taking literally: there is no personal pot of money growing in your name. The contributions today's workers pay are handed almost straight back out as today's retirees' pensions — a transfer between generations, not a fund. That design is exactly what makes it demographically sensitive. Its solvency rests on the ratio of contributing workers to drawing retirees, and in Belgium, as across most of Europe, that ratio is deteriorating: people live longer and draw for more years, while relatively fewer workers pay in behind them. The same longevity that complicates individual saving squeezes the collective system too. And for most private-sector careers the first pillar replaces only a modest fraction of final salary — a replacement rate, the share of your working income your pension reproduces, that is low enough to leave a conspicuous gap between the salary you earned and the pension the state provides. That gap is the entire reason the other two pillars exist, and closing it is what a second-pillar plan is created for. If you want more on the first pillar, I refer you to Wikifin.

The second pillar is occupational: pension plans that an employer — or an entire sector — sets up for its employees, funded by contributions paid in across the working career. This is where the project lives. The plan is governed by the WAP (Wet op de Aanvullende Pensioenen), and its defining feature is a guaranteed minimum return on those contributions. That guarantee is what makes the second pillar more than a savings account: whatever the plan's assets actually earn, the employee is promised their contributions grow at least at a legally fixed rate. The risk of falling short doesn't sit with the employee — it sits with the employer, who underwrites the guarantee. If the assets backing the plan return less than the guaranteed rate, the employer makes up the difference out of pocket.

That guarantee used to be a fixed number — historically 3.75% on employee contributions and 3.25% on employer contributions. But a fixed guarantee is a fixed liability, and once safe market yields collapsed well below those levels in the 2010s, employers were underwriting a spread they had no safe way to earn. The 2016 reform resolved this by letting the guarantee float with the market. Since then it has been a single unified rate, recomputed from the yield on the 10-year OLO — the Belgian government bond whose curve we simulated in [previous section]. Concretely: 85% of the trailing 24-month average of the 10-year OLO yield, rounded to the nearest 25 bps, floored at 1.75% and capped at 3.75%, applied separately to the employer and employee legs:

\[g_t = \operatorname{clip}\!\Big(\operatorname{round}_{25\text{bps}}\big(0.85 \cdot \bar{y}^{\,10Y}_{[t-24m,\,t]}\big),\ 1.75\%,\ 3.75\%\Big).\]

Every term there is a risk-sharing knob. The 85% haircut and the 24-month averaging make the guarantee lag and dampen the market rather than track it tick-for-tick; the floor protects the employee when rates are near zero; the cap protects the employer when they spike.

And this is exactly why the OLO had to come first. Read the formula forward in time: contributions accrue over a whole career, and the rate applying to each future year is set off the OLO yield prevailing in that year. So the guarantee the employer is on the hook for isn't a single number — it's a trajectory, a deterministic function of the entire future path of the 10-year OLO. To know the distribution of shortfalls the employer might eventually have to cover, we first need the distribution of future OLO paths. That is precisely what the Hull-White simulator produces: each simulated short-rate path is mapped back to a future 10-year yield (reconstructFutureYield), the formula above is applied to its trailing average, and out comes the guarantee rate that year's reserve has to clear. The OLO model isn't scene-setting — it's the engine that turns the WAP guarantee from a legal definition into a number the environment can compute at every step.Note that the current implementation makes abstraction of this reasoning and just works with a fixed rate for simplicity reasons.

Sitting alongside the return guarantee are two further hard constraints that bound the plan rather than price it: the 80% rule, capping total pension build-up relative to salary, and the non-discrimination requirements. These enter the model as feasibility constraints the policy must respect, not as terms in the objective it trades off.

The third pillar is individual, tax-favoured saving — a personal pension product you take out and fund yourself, nudged along by a tax break. It completes the arc: where the first pillar pools risk across generations and the second shares it between employer and employee, the third places it squarely and voluntarily on the individual. Because no employer or collective guarantee sits behind it, it falls outside the risk-sharing problem this project is about, and I leave it aside here — relevant to an employee's full retirement picture, but not to the design question that follows.

Different types of Pension Plans

Within the second pillar the classic dichotomy is, at heart, about who carries the risk — which makes it the natural place to pin down exactly where this project sits.

Defined Benefit (DB): the plan promises an outcome — typically a formula on final or career-average salary — and the employer bears the full weight of delivering it. Both the investment risk (the backing assets may underperform) and the longevity risk (retirees may live longer than the plan funded for) land on the employer. That concentration is exactly why DB is increasingly rare: it is an open-ended liability sitting on the sponsor's balance sheet, and few employers still want to hold it.

Defined Contribution (DC): the plan promises an input — a contribution rate — and the retirement outcome is whatever those contributions happen to grow into. In most countries this is the mirror image of DB: fix the input, and every risk on the outcome side passes to the employee, who now carries the market and longevity exposure alone.

Belgium is the interesting exception — and it is the reason a risk-sharing model has anything to chew on here at all. The WAP guarantee bolts a minimum return onto an otherwise DC structure, producing a hybrid sometimes called "DC with a DB flavour": contributions are defined, but the employer still stands behind a floor. Risk here is neither fully concentrated (as in DB) nor fully offloaded (as in pure DC) — it is split. The employee keeps the upside above the guaranteed rate; the employer owns the downside below it. And that split is precisely what makes the central question of this project non-trivial. If all the risk sat with one party, there would be nothing to optimise jointly — you would simply minimise that party's pain. It is the shared middle ground, where a euro of reassurance to the employee is a euro of liability to the employer.

The insurance vehicle matters too, because it sets how the reserve behaves over time — and therefore how often the guarantee actually has to bite. Branch 21 products credit a contractually guaranteed return, plus rretionary profit sharing, directly to the reserve. The crucial modelling consequence is that the reserve is a credited-return process, not a marked-to-market bond portfolio: it grows by a smoothing crediting rule rather than by revaluing an underlying portfolio day to day, so short-term market swings are absorbed rather than passed straight through. That smoothing is itself a cushion between the market and the guarantee. Branch 23 products are unit-linked: the reserve simply tracks fund value, with more upside and no crediting cushion. When markets fall, the fund falls with them — and with nothing to smooth the drop, the WAP floor is reached far more readily, forcing the employer's guarantee to make up the difference. In that sense the guarantee behaves like a put option the employer has written to the employee: dormant while the fund clears the floor, triggered precisely when it does not — and Branch 23 leaves it triggered more often. Comparing how these two vehicles behave under one and the same contribution policy is one of the axes of the wider project.

This post works with the simplest member of the family: a single-employee Belgian DC plan with Branch 21-style deterministic crediting — the guarantee credited at a fixed rate, the put left dormant by construction. That is the same first-rung move as holding the WAP rate flat: strip out the stochastic crediting to get an environment whose behaviour you can check by hand, and switch it back on only once the mechanics are proven. The reasons for starting exactly here are what the next section makes precise.

What is Optimal?

"Optimal" is meaningless until you say for whom. The same contribution policy is a different thing seen from each side of the contract: one that thrills the CFO — low, predictable outlays — can quietly starve the employee's replacement rate, while one generous enough to delight the employee can make plan costs balloon exactly when the firm can least afford them, in a downturn. There is no policy that is optimal in the abstract; there is only optimal-for-a-chosen-balance. So rather than pick a side, the project's central object makes the balance itself explicit — a joint value function:

\[V \;=\; \lambda\, V_{\text{employee}} \;+\; (1-\lambda)\, V_{\text{employer}}.\]

Read it as a single dial with the two parties at its ends. Each term encodes what that party is actually trying to get out of the plan.

The employer side, \(V_{\text{employer}}\), is the negative of a cost sensitivity index — negative because from the sponsor's chair the plan is a cost, and value means having less of it to worry about. What goes into that index is not merely how large the cost is but how unpleasant its shape: its expected level, its tail risk (the rare, severe funding shortfalls that do real balance-sheet damage), and its procyclicality — whether the cost tends to spike in bad states of the world, when the firm is already under strain. A cost that is high but steady is easier to carry than one that is lower on average but lands its worst blows at the worst possible moment; the index is built to prefer the former.

The employee side, \(V_{\text{employee}}\), is a utility over retirement consumption, anchored to a replacement-rate target — the very quantity introduced back at the first pillar, now doing formal work. Framing it as utility rather than raw expected wealth matters: it builds in that employees are risk-averse and that shortfalls below the target hurt more than equivalent surpluses help, so the objective rewards reliably clearing a decent standard of living over gambling for a larger but less certain pension.

And then the weight \(\lambda\) itself — which, crucially, is not something to be optimised. There is no "correct" \(\lambda\) the model should rover; a higher one does not mean a better plan, only a plan tilted further toward the employee. It is the negotiation dial between the two parties: the numerical form of exactly the question the risk-sharing contract has been posing all along — how much of the risk does each side agree to carry? Fixing \(\lambda\) is choosing a point on the shared-risk spectrum; sweeping it from one end to the other and re-optimising at each setting traces out the whole design frontier — the menu of employee-optimal-given-employer-cost trade-offs, every point on it Pareto-sensible, the choice between them a matter of negotiation rather than mathematics. The model does not tell you where to stand on that frontier. It tells you what the frontier is — which is the genuinely useful thing, because it turns a vague tug-of-war into an explicit, priced menu of options.

That is the full framework. This post uses a radically stripped-down version of it, because the framework only means something if the machinery underneath it can be trusted — and before you optimise anything real, you have to prove that the environment computes what it claims to, on a problem whose answer you already know. That proof is what the rest of this post is about.

The Reward Function

At this first rung the reward is nothing more exotic than a present value: each cash flow discounted at an explicit rate and booked at the moment it occurs. Concretely, we take the two sides of the contract and reduce each to a single number at the retirement date \(T\).

Start with the employer, whose relationship to the plan is a stream of costs. Over the employee's career the employer pays in a contribution \(c_t\) each period; then, at retirement, it faces the guarantee. The plan has accrued a legal reserve \(L_T\) — what the contributions are worth grown at the WAP guaranteed rate — while the insurer holds a mathematical reserve \(R_T\), the actual Branch 21 reserve backing the contract. If the reserve falls short of what the guarantee legally owes, the employer tops up the difference; if it does not, the employer owes nothing. That is exactly the written-put payoff we described earlier — dormant when the reserve clears the floor, triggered only when it does not:

\[C_{\text{employer}} \;=\; \sum_{t=0}^{T}\frac{c_t}{(1+\text{r})^{t}} \;+\; \frac{\max\!\big(L_T - R_T,\,0\big)}{(1+\text{r})^{T}}.\]

Because from the sponsor's chair this is a cost, its value is the negative of it:

\[V_{\text{employer}} \;=\; -\,C_{\text{employer}}.\]

The employee's side is, at this rung, deliberately the simpler of the two: the value is just the present value of the lump-sum benefit \(B_T\) received at retirement —

\[V_{\text{employee}} \;=\; \frac{B_T}{(1+\text{r})^{T}}.\]

It is worth being honest that this is a risk-neutral stand-in for the employee objective the framework calls for. The full version, sketched earlier, is a concave utility over retirement consumption anchored to a replacement-rate target — curved precisely so that shortfalls hurt more than surpluses help. Here we strip that curvature out and reward raw discounted wealth. That is not the final employee value; it is the flattest thing that still points in the same direction, chosen so the first rung has as few moving parts as possible.

Combining the two through the negotiation dial \(\lambda\) gives the joint reward the agent actually optimises:

\[V \;=\; \lambda\,\frac{B_T}{(1+\text{r})^{T}} \;-\;(1-\lambda)\left(\sum_{t=0}^{T}\frac{c_t}{(1+\text{r})^{t}} \;+\; \frac{\max\!\big(L_T - R_T,\,0\big)}{(1+\text{r})^{T}}\right).\]

This is a radically simplified version of the reward the framework ultimately reaches for — linear where it should be curved, deterministic where it should be stochastic. But that is the point. Before trusting the agent on a reward that reflects reality, we need to know it behaves correctly on one whose optimum we can work out by hand. Once we understand how the agent responds to this stripped-down environment, each later rung swaps one simplification for its realistic counterpart — curvature into the employee term, stochastic crediting into the reserve — against an agent we already know to be sound.

Reinforcement Learning: A Basic Introduction

Most financial plans are fixed at the outset: a contribution schedule, a rebalancing rule, a glide path decided on day one and followed regardless of what happens next. Reinforcement learning starts from the opposite instinct. Strip the jargon and it is control — steering a system toward a goal by choosing actions as it unfolds — where the strategy is learned from simulated experience rather than fixed in advance. Instead of committing to a plan, you discover a rule by simulating across thousands of possible futures. The basic RL framework is shown below.

Agent Environment state St reward Rt action At Rt+1 St+1

A state \(s\) is what is known at the moment the agent makes a decision — here, for our minimal model, deliberately spare: the current year. An action \(a\) is the decision taken in that state — here about as brutal as a decision gets: contribute this year, or don't. A policy \(\pi\) is a rule mapping each state to an action, which is to say it is exactly a contribution strategy — but a dynamic one, reacting to the state the plan has actually reached rather than a fixed schedule laid down at inception. The framing has already bought us something: we are no longer searching for the best contribution number, we are searching for the best contribution rule.

The last piece is the one that should feel most familiar. A value function \(Q(s,a)\) is the expected sum of future rewards from taking action \(a\) in state \(s\) and following the policy thereafter. Given the reward built in the previous section, unfold that definition and it is nothing other than an expected present value, conditional on a state and a first decision. You have computed thousands of these by hand; RL changes only how they are obtained — estimated by simulating many careers rather than read off a formula — and what is done with them once they are.

How is the agent discovering this? The thing to hold onto is that the agent is never shown any of the machinery. It does not see the reward formula, nor the law of motion of the reserve; it only ever sees what the environment hands back at each step — a realised reward and a next state — exactly the loop of the diagram above. Everything it comes to know about value, it infers from those samples alone. Concretely, it keeps a table of estimates, one entry per state–action pair, all initially blank. In each episode it plays a whole career to \(T\), choosing actions by its current best guesses; and only once the career is complete — this is what makes it Monte Carlo rather than a bootstrapping method — does it look back over every state–action pair it actually visited and nudge each estimate a fraction of the way toward the return that followed it. That nudge is the update on the bottom line of the figure below: no single return is trusted outright, each merely tugs the running estimate toward itself, and holding the step size \(\alpha\) constant rather than letting it shrink keeps the junk returns from early training — when the policy was still bad — from anchoring the estimate once the policy has improved.

Estimating the value of a state Monte Carlo: simulate many careers from (s, a), then average their present values STATE s year  =  t reserve  =  Rt ACTION a ▸  contribute ▸  don’t follow π to T SIMULATE N CAREERS t T G1 G2 G3 each career → one return Gi Q (s, a)  ≈ (1/N) Σ Gi sample mean of returns Each career yields one return  G = discounted sum of rewards  (its present value). Q (s, a)  ←  Q (s, a)  +  α · [ G − Q (s, a) ] constant step size α keeps early-training noise from dominating the estimate

Learning, then, is the loop between estimating and acting. Better estimates make for better decisions: acting greedily with respect to the table means, in each state, taking whichever action the estimates now rank highest. But an agent that only ever exploits its current best guess will never find out that some neglected action was secretly better — so it explores, taking a random action a small fraction \(\varepsilon\) of the time, just often enough that no promising branch stays permanently untried. Evaluate, improve, repeat: each pass sharpens the estimates, each sharpening shifts the greedy policy, and under the right conditions the policy climbs toward the one that maximises the joint value \(V\). Nobody hands the agent the optimal contribution rule — it discovers it by playing thousands of careers and letting their realised present values vote.

evaluation Q ⤳ qπ π Q π ⤳ greedy(Q) improvement

At this first rung that loop is doing something you could, in principle, do without it: the environment is deterministic and the action is a single binary switch each year, so the optimal policy can be found by direct enumeration and checked by hand. That is not a weakness — it is the entire reason to start here. The point of a rung whose answer is independently knowable is precisely that the learner's answer can be checked against it. The reason to reach for RL at all only arrives later, when the reserve stops being deterministic and its stochastic path enters the state: at that point there is no formula to evaluate and no schedule to enumerate, and a policy learned from simulation is the only object that still scales. What we validate now is that the machinery finds the right answer where we can grade it — so that we can trust it where we cannot.

The Environement

𝜅 κ; contributions carry a burden term; the terminal reserve funds retirement consumption.

Why start this stupid? Because a bug in a complex RL environment is invisible: the agent trains, the reward curve rises, the policy looks plausible — and the objective it optimised is not the one you wrote down. The methodology is a ladder: each rung adds exactly one new element, is validated against an independent benchmark, then frozen as a regression test before the next rung is added.

Tuning the Learning Behaviour

The Results

Final Thoughts and Next Steps