Product Thinking

What preference modeling means in the context of employee benefits

Gianluca Enrietti 7 min read
Back to Blog

Preference modeling is a phrase that travels. It starts in recommendation systems: Amazon, Netflix, Spotify all run some version of a model that infers what you are likely to want next based on what you and similar users have chosen in the past. From there, the term migrates into adjacent domains. HR technology vendors have been using it for several years to describe everything from pulse surveys to engagement scoring to benefits configuration tools.

The result is a phrase that means almost everything and therefore risks meaning nothing. When we talk about preference modeling at Toduba, we mean something specific. This post is an attempt to say precisely what that is, where it is genuinely useful, and where it breaks down.

The core idea

In the benefits context, a preference model takes claim history as its primary input and tries to answer one question: given what employees in a particular workforce segment have actually chosen in the past, what category-level allocation would produce the highest expected utilization within a given budget?

That is a narrower question than it might initially sound. It is not asking what employees want in some abstract sense. It is not trying to predict individual behavior. It is estimating the expected aggregate claim rate across benefit categories for a proposed mix, using segment-level historical patterns as the underlying data. The goal is to surface which reallocation options are worth considering and which ones the data does not support.

How it differs from the Netflix model

The recommendation system analogy is useful up to a point, and then it breaks down in important ways.

Netflix makes millions of recommendations per day across a user base where individual preferences can be tracked at a granular level. The signal-to-noise ratio is very high because there is so much data. The feedback loop is tight: a user watches or skips something, and that behavior is incorporated into the model within a short cycle.

Benefits preference modeling operates in a completely different data environment. A mid-market employer with 300 employees might have two or three complete annual cycles of claim data. The segments you can meaningfully analyze are much larger than individual users. The categories are broader and fewer in number. And the feedback loop is annual, which means a mis-calibrated model can go a full year before its errors become visible.

These differences matter because they define what the model can and cannot do. The Netflix model can usefully recommend at an individual level. A benefits preference model cannot, and trying to operate at that granularity would be statistically unreliable and raise legitimate questions about how individual employee data is being used.

The right unit of analysis in benefits preference modeling is the workforce segment, not the individual. This is not a technical limitation we are working to overcome. It is the appropriate level of analysis for the question being asked.

What the model actually reads

The inputs that matter in a benefits preference model are:

  • Category-level claim rates by workforce segment over multiple annual cycles
  • Budget allocation per category in each cycle
  • Segment definitions: age band, tenure cohort, role type where available
  • The trajectory of each category over time (growing, stable, declining in claim rate)

What the model does not read, and should not read, is individual employee identifiers linked to specific claim amounts. The analysis operates on aggregated segment patterns, not on personal records. This is both a privacy design choice and a practical one: segment-level patterns are more stable and more useful for benefits design decisions than individual-level data would be.

What a preference model does not do

We want to be direct about the limits, because overstating what a model can do is the fastest way to undermine confidence when the model turns out to be less than advertised.

A preference model cannot tell you why claim rates are what they are. It can identify that a category has a 14% claim rate among employees in the 35-44 age band. It cannot tell you whether that is because the category is not well suited to that segment, because the category is poorly communicated, because the reimbursement ceiling is too low to be worth the paperwork, or because that specific cohort of employees is going through an unusual life-stage transition. Surfacing the pattern is the model's job. Understanding the cause requires human judgment and, often, direct conversation with employees.

A preference model also cannot optimize for outcomes that are not reflected in claim data. Employee wellbeing, stress levels, long-term retention, or the sense of being valued by an employer, these are real outcomes that benefits design can influence. But they do not appear directly in claim records. A model that optimizes for claim rates may produce a benefits mix that is highly utilized but misses the things that matter most to long-term employee satisfaction. HR teams need to keep that gap in view when using model outputs.

Finally, a preference model cannot help you with entirely novel categories that have no prior claim history. If you want to introduce eldercare support for the first time, or add a category tied to a benefit type that is new to your workforce, the model has nothing to learn from. For new categories, you need employee research, market benchmarking, and HR expertise, not a model.

The relationship between the model and the HR professional

The positioning question we spent the most time on in building Toduba is how the model should relate to the HR professional who is using it. The wrong answer is "the model knows best, trust the output." That framing positions the HR professional as a button-presser executing algorithmic recommendations, which is not how benefits decisions actually work, and not how they should work.

The right answer, in our view, is that the model makes the HR professional's analysis more efficient and their decisions more grounded. An experienced total-rewards lead at a northern Italian manufacturing company with 280 employees knows things about her workforce that no model will ever capture: which managers have been pushing for expanded parental leave for three years, which segment of employees is going through a wave of first-home purchases, which recent hire cohort is younger and more mobile than the company's historical profile. That knowledge is the context the model cannot produce.

What the model can do is take two years of claim data that currently lives in three separate spreadsheets across two platforms, surface the segment-level patterns in that data in a readable format, and show which reallocation options the data would support. The HR professional then applies their contextual knowledge to those options and makes a decision. The model narrows the decision space. It does not make the decision.

Where this is heading

Preference modeling in HR benefits is early. The theoretical approach is sound, the data requirements are achievable for most mid-market employers, and the problem it addresses, that most benefits allocations are sticky artifacts of past decisions rather than current evidence, is real and widespread. But the implementations are mostly at an early stage, and the empirical evidence on what drives utilization across different workforce types and geographies is still being built.

We started Toduba with a fairly simple hypothesis: that a benefits configuration built on what employees actually claim will produce meaningfully better fit than one built on surveys and habit. We believe that hypothesis is correct. We also believe that building the data infrastructure and modeling discipline to test it properly takes several years of consistent measurement. We are in the early stages of that process, and we are honest about what we do not yet know.