The pilot

We’d Rather Be Measured Than Believed.

The value model makes claims across three tiers of confidence. Tier 1 rests on published research. Tier 3 rests on nothing yet, because the category is young enough that nobody — us or anyone selling against us — has credible outcome data.

This pilot is designed to convert the second thing into the first. It is a measurement exercise with pre-agreed thresholds, including the ones at which we tell you to walk away.

Start here

What 90 Days Cannot Tell You

The largest line in the value model is regrettable attrition avoided. You cannot measure that in a quarter. Attrition is an annual signal with seasonal structure, and any vendor who offers you a retention result in 90 days is showing you noise with a narrative attached.

So we don’t. The pilot measures the leading indicators that precede a retention decision — whether career conversations are happening, whether people told “not yet” received criteria, whether internal moves became visible — and it establishes the baseline against which the real answer arrives at month twelve. We will tell you at day 90 whether the mechanism is working. We will not tell you the retention number, because we won’t know it.

The instrument set

Six Measurements, Mapped to the Tier They Test.

Tests Tier 1

Career-Conversation Incidence

Share of employees who had a substantive conversation about their path in the period, and — the part that predicts departure — the share who received explicit criteria after a “not yet.” Baselined at week zero from your engagement instrument, remeasured at day 90.

Why it matters: Gallup’s finding that most exiting employees had no such conversation in their final three months makes this the closest available proxy for preventable attrition.

Tests Tier 1

Internal Application and Fill Rate

Internal applications per open role and share of roles filled internally, pulled from your ATS. Movement here is fast, unambiguous, and directly priced in the model through the external-hire premium.

Why it matters: most people who leave for growth would have stayed for the right internal move. They never saw one, or feared that asking would read as disloyal.

Tests Tier 2

Committed-Benefit Utilisation

Uptake of tuition assistance, learning budget, mentorship and EAP against the same period last year. Straight from your benefits administrator, no new instrumentation.

Why it matters: this is the fastest-moving line in the model and the one most likely to be visibly true by day 60. It is also the easiest for a sceptical CFO to verify independently.

Tests Tier 2

Self-Reported Friction

A four-item instrument at week zero and day 90: how long it takes to find out who decides something, how often work is redone because of misalignment, how often people wait on an answer that never comes, and how clear the standard for good work is.

Why it matters: it is self-reported and we will label it as such. It is the honest instrument available in 90 days — time-and-motion study is not.

Tests Tier 3

Observed Capability Movement

Competency levels at entry versus day 90 across the population, computed from goal-condition state rather than self-assessment or course completion. Reported with confidence intervals, which will be wide at this duration.

Why it matters: this is the artefact nobody else in the category can produce. Ninety days is enough to show the measurement works. It is not enough to show the movement is durable.

Tests Tier 3

Manager Attention Distribution

Where development attention landed across each team’s performance distribution, at week zero and day 90. Most managers concentrate on the bottom and believe they’re being even-handed.

Why it matters: it is the clearest observable behaviour change in the manager layer, and behaviour change in that layer is the mechanism the entire model depends on.

Ninety days

How It Runs

Weeks −2 to 0

Baseline and Consent

Nothing is deployed until the baseline exists, because a pilot without a week-zero measurement can only produce anecdotes.

  • Pull existing data: attrition by function and level, internal fill rate, benefit utilisation, engagement items on career support
  • Run the four-item friction instrument across the population
  • Works council, legal and IT review of the consent architecture and the audit trail
  • Agree the success and failure thresholds in writing, before anyone has a result to defend
Weeks 1–2

Org-Wide Activation

Everyone, not a cohort. The strategic layer cannot see where development stops if you deploy only to the people least likely to stall.

  • Every employee and every people manager provisioned
  • Manager layer briefed separately — their coaching is private from their own manager, and they need to know that before they use it
  • No mandatory usage targets: coerced usage produces coerced data and destroys the signal we are here to test
Weeks 3–6

First Strategic Artefact

The stall analysis lands at week six. This is the meeting where you find out whether the platform produces anything you didn’t already know.

  • Goal-condition state populated across the population
  • First stall analysis: what is blocking development, ranked, with the sponsorship-versus-capability split
  • First friction map by function and interface
  • Benefit utilisation trend visible against last year
Weeks 7–11

Manager Layer Under Load

This is the period that tells you whether behaviour changed or whether people just logged in.

  • Calibration and comp cycle support if one falls in the window — deliberately the hardest test
  • Manager attention distribution remeasured
  • Mid-pilot review: what is moving, what is not, and what we got wrong
Weeks 12–13

Readout, Honestly

The report states what moved, what didn’t, and what remains unproven. Written to survive your analytics team rather than to close a renewal.

  • All six instruments remeasured against baseline, with confidence stated
  • The value model re-run on your actual observed rates rather than our defaults
  • Written statement of every result that failed to move, given equal prominence
  • The 12-month attrition measurement plan, since that answer is not available yet
Thresholds

Agreed Before We Start, Including the Ones That End It

Set in writing during baseline, when neither side has a result to protect. Illustrative defaults — yours will be calibrated to your baseline:

What Counts as Working

Career-conversation incidence up meaningfully against baseline, with the criteria-after-a-no measure moving at least as much as the headline. Internal application rate up against the prior comparable period. Benefit utilisation up against the same period last year. Friction instrument improved on at least two of four items. Stall analysis produced findings that changed at least one funding or programme decision. Manager attention measurably redistributed toward the middle of the distribution.

What Counts as Failing

If activation stays below a third of the population by week six, the mechanism isn’t reaching people and the rest is irrelevant. If the friction instrument is flat at day 90, the largest Tier 2 line is unsupported in your environment and should be struck from your model. If the stall analysis tells you only things you already knew, the strategic layer is not earning its price here. If manager attention distribution has not moved at all, the delivery layer has not changed and the retention thesis does not hold.

If two or more of those are true at day 90, we will say so in the readout and recommend you don’t renew. We would rather lose a contract than be the reason a CHRO has to defend a number that didn’t hold twelve months later.

What Each Side Commits

From You

An executive sponsor at CHRO or Chief People Officer level. Access to attrition, ATS and benefits data for baseline. Your engagement instrument’s career-support items. Works council or employee representative review where applicable — we would rather be examined early. Ninety minutes of manager time in week one, and nothing mandatory after that.

From Us

Org-wide deployment, not a cohort. Baseline instrumentation before anything ships. A named person on the pilot for its full duration. The full audit trail, handed over rather than described. A written readout including every result that failed to move. And the recommendation not to renew if the thresholds above are missed.