Product Thinking

From the annual survey to a predictive preference model

Sofia Barbieri 6 min read
Back to Blog

Every year, HR teams at mid-market employers send the same questionnaire. "Which of the following benefits do you value most?" The options are the same as last year. The responses cluster around the same answers. The results get summarized into a table that looks almost identical to the one from the previous cycle. And then the benefits mix stays largely the same.

This is not a complaint about HR professionals. The annual survey is a reasonable tool given the constraints most HR teams work under: limited time, limited analytical resources, a vendor contract that needs to be reviewed in a six-week window, and a workforce that has been mostly satisfied with the current arrangement. The survey answers the practical question of whether to make changes. It does not answer the harder question of which changes would actually improve fit between the workforce's needs and the benefits mix on offer.

The move toward preference modeling is an attempt to answer the harder question. Here is what that transition involves in practice, and where the limits are.

What a survey measures and what it misses

Surveys measure stated preferences, which is a reasonable proxy for genuine preferences under one condition: the person answering has thought carefully about the tradeoffs. In most annual benefits surveys, that condition does not hold. The survey lands in an inbox during a busy period, the options are framed in ways that prime certain answers, and most employees do not have strong opinions on the relative value of childcare versus pension contributions versus learning allowances because they have never had to choose explicitly between them.

This is not a criticism of the employees or the survey designers. It is a structural feature of stated-preference research in any domain. When you ask people to rank things they have not had to choose between under real conditions, the answers are weakly predictive of what they will actually do when it matters.

Claim history bypasses this problem. When an employee claims a benefit, they have made a real choice with real consequences. They filled out a form, they went through whatever friction the process involves, and they decided that the value was worth it. That decision contains more information than any survey response.

What a predictive model actually does

When we talk about a preference model in the context of benefits, we are describing something more specific than a machine learning system that reads minds. The model is an attempt to take a structured look at historical claim behavior across a workforce, segment that behavior by employee characteristics that are meaningfully correlated with different patterns, and use those patterns to produce a ranked estimate of which categories will see higher or lower uptake in the next cycle.

The inputs are claim history by category and time period, workforce segmentation by age band and tenure cohort, and the current budget allocation across categories. The output is not a prediction of what individual employees will claim. It is an estimate of what category-level utilization will look like across segments given the current or a proposed mix, and a recommendation for reallocation that would improve expected utilization within the existing budget envelope.

This is a narrow and specific task. It is not trying to predict everything. It is trying to answer one question more accurately than the annual survey does: given what this workforce has shown us about their actual choices, where is the current mix underallocated and where is it overallocated?

Where the model adds value, and where it does not

The model adds the most value when the gap between stated preferences and revealed preferences is large. In our experience building toward this approach, that gap is often largest among the non-claiming population, which is to say employees who are eligible for categories but have never claimed them. Survey responses from this group tend to overrepresent aspirational preferences. Their claim data is a clean zero, which is a different kind of signal entirely.

The model does not add value when the workforce is too small to produce statistically stable segment-level patterns. A company with 80 employees split across four departments has segments of 20 people, and 20-person claim histories are noisy enough that the model's output is not meaningfully better than an experienced HR professional's intuition about the same workforce. We are honest about this boundary. Preference modeling at the category-level starts to produce reliable signal in workforces of around 150 employees or more, and becomes substantially more useful above 300. Below that, the benefit of structured analysis is in organizing the data for better human judgment, not in generating predictions that should override that judgment.

A predictive model also cannot tell you about categories that have never been offered. If a workforce has never had access to eldercare support and no competing survey explicitly asked about it, there is no revealed-preference data and only weak stated-preference data to build on. Novel category introduction is a place where HR expertise and direct employee conversation will always matter more than any model.

The survey and the model together

This is not an argument for replacing the annual survey with a model. The two tools answer different questions and complement each other when used well.

The survey is better for surfacing novel categories, understanding the context behind utilization patterns, and giving employees a voice in the process. It is also better for picking up on morale and sentiment signals that claim data does not contain. An employee who claims nothing may be perfectly satisfied or deeply disengaged, and a well-designed survey can tell you which one.

The model is better for making the existing utilization data legible across segments, identifying where reallocating budget would improve fit, and stress-testing proposed redesigns before committing to them. It surfaces the questions the survey should be asking more specifically.

An HR team that uses the model to identify which segments are likely to see the largest improvement from a proposed reallocation, then asks targeted follow-up survey questions to validate the direction with employees in those segments, is getting the best of both approaches. That is a more disciplined process than either the survey alone or the model alone can produce.

The practical starting point

Moving from an annual survey to a preference model is not a one-year project. The prerequisite is two to three years of clean, segmented claim data. If that data does not exist yet, the work starts there: establishing a consistent data capture practice, making sure category definitions are stable year over year, and building a segmentation structure that will support meaningful pattern analysis.

The annual survey in the meantime is not wasted effort. It is the instrument that captures employee experience and identifies categories worth investigating further. Its value in the transition period is in directing the model's attention, not in replacing it.

Most HR teams at mid-market employers are a few years away from having the longitudinal data needed to make a preference model genuinely reliable. The teams that invest in data discipline now will be able to make substantially better decisions in three years' time. That investment costs very little. What it requires is consistency of measurement, which is more about process discipline than tooling.