Back to work
Case study·Gympass / Wellhub·AI class discovery

Helping members find classes that fit their lives.

Gympass ranked classes by popularity, so everyone nearby opened the app to the same feed. That one popular list could be too intense for a low-intensity yoga regular and too basic for an experienced runner. I redesigned class discovery around each member's goals, experience, and schedule, then added a clear reason and controls to every recommendation.

Role
Product design + UX research
Team
2 product managers · data analyst · iOS + Android engineers
Platform
iOS · Android · watchOS · Wear OS
Company
Gympass (now Wellhub)
Year
2023
Before and after of the class discovery feed: the before feed ranks by popularity with no match score or reason; the after feed shows match scores, a 'Why this class?' link, and personalized rows.
Fig. 01The old popularity feed and the redesigned class recommendations.
The same change, in motion
The old popularity-ranked feed. The personalized home, ranked by fit with match scores. The Why this class explanation sheet.
Before: a popularity-ranked feed, the same classes for everyone.
01
Problem

Popular classes were often the wrong classes.

Gympass gives employees access to nearby gyms, studios, trainers, and classes. Bookings were falling after New Year, and the team's first explanation was seasonality. Before changing the product, I wanted to check whether members were losing interest or simply failing to find something worth booking.

A quick product primer

01

What Gympass is

A corporate wellness benefit. Employees get access to thousands of gyms, studios, and fitness classes through their employer, all in one app.

02

Where people choose a class

The home feed is the main place members browse and decide what to book.

03

What was missing

The catalog was large enough. Members needed help finding classes suited to their level, goals, and schedule.

What members saw

The feed showed the same popular classes to everyone nearby. It did not account for experience, goals, or past bookings, and it did not explain why a class appeared. If CrossFit was popular in your area, it ranked first for a low-intensity yoga regular and for everyone else there too. Popularity was local and averaged across everyone, so a class could top the list on booking volume alone, even when it suited very few of the people seeing it.

The old discovery feed: nearby classes ranked by popularity, no match score, no explanation. The same popular list could put high-intensity CrossFit in front of a low-intensity yoga regular and beginner mobility in front of an experienced runner.

The card gave members no way to judge fit. There was no match score or explanation, and popularity determined the order for everyone.

That raised the question I took into research: were members losing interest, or were they failing to find something worth booking?

02
Research

The seasonal explanation did not match member behavior.

To answer that, I interviewed 50 members and reviewed three months of in-app activity, about half a million sessions. They were still browsing. They just were not booking what the feed showed them.

Research summary: members were browsing, but the classes did not fit
Research synthesis: the assumed seasonal dip versus a relevance problem. 50 interviews, ~500K sessions, members browsing far more than booking, and two mismatch anchors from the same popularity feed. High-intensity CrossFit was too much for a low-intensity yoga regular, and beginner mobility was too basic for an experienced runner.

Who I spoke with

I spoke with members across different goals, routines, and experience levels. I deliberately included active members whose booking had started to slow, because analytics flagged them as the most likely to cancel. The mismatch went both ways. The same popularity feed could feel too intense for a low-intensity yoga regular and too basic for an experienced runner.

Why fitness level was too blunt

Fitness level was a useful start, but too rough to explain what people actually booked. Someone who trains a lot might still be exploring, and a newer member might be pushing hard. So I looked at what members really weigh when they pick a class, how hard they want it, their goal, how experienced they are, when and where they go, and whether they want routine or variety. I synthesized those patterns into five behavioral personas that shaped the ranking's inputs and the experience requirements. They were not permanent labels on people, and they were not classes the algorithm assigns. I led the research and the decision framework, then designed the screens that capture preferences, explain each pick, and let members correct it. Analytics grounded the personas in real bookings and sized how often each one showed up, and the product manager decided which to focus on first. Because people shift between these behaviors during the week, the app pays attention to what you are doing right now instead of stamping you with one label for good.

Each persona below captures a recurring decision pattern. It shows what the app has to respect, what the ranking can use, and when it should adapt.

The Restorative SeekerRestore
~22% of members
IntensityLow
GoalStress relief
ExperienceAny
ContextEvening · home
NoveltyRoutine
Shows up in data as
Gentle evening classes on a steady weekly cadence; short-to-mid sessions.
What the model must respect
GateKeep inside the low-intensity band; never surface HIIT as 'similar'.
RankPrefer stress-relief goals; cross category only within low intensity (yoga → pilates → stretch → barre).
Override If a member books or browses higher intensity this session, the low-intensity gate lifts for that visit.
The Goal-Driven BuilderPush
~18% of members
IntensityHigh
GoalStrength
ExperienceSeasoned
ContextMorning
NoveltyStructured
Shows up in data as
Progressive booking toward strength and HIIT; consistent streaks; tracks difficulty.
What the model must respect
GateStay within their proven difficulty band, challenging but not unsafe.
RankPrioritize the goal (strength) over category; allow progressive difficulty; favor morning slots.
Override A recovery or rest-day search drops intensity for that session without changing the profile.
The Time-PressedEfficiency
~24% of members
IntensityModerate
GoalGeneral health
ExperienceAny
ContextLunch · near work
NoveltyLow
Shows up in data as
Short-format classes, lunch slots, near work or home; schedule-constrained.
What the model must respect
GateHard limits first: session within their available window, and near work or home.
RankThen fit by goal and intensity; offer short-format alternatives across categories.
Override A weekend with a longer free window relaxes the duration gate for that session.
The Social MoverCommunity
~16% of members
IntensityMod-high
GoalFun · accountability
ExperienceMixed
ContextWeekend · group
NoveltySome
Shows up in data as
Group and studio formats; weekend bookings; responds to social cues.
What the model must respect
GateRequire group or studio formats; weekend and evening slots.
RankPrioritize community signals and social formats over strict category match.
Override A solo weekday search is honored as an efficiency session, not forced into group classes.
The ExplorerDiscovery
~12% of members
IntensityVariable
GoalVariety
ExperienceConfident
ContextFlexible
NoveltyHigh
Shows up in data as
High category diversity; tries new studios and formats; low routine.
What the model must respect
GateStay inside their comfortable intensity band, their one safety rail.
RankWiden discovery: cross category broadly, weight novelty high.
Override If they repeat one category several times, ease off novelty and reinforce the emerging routine.
stated by membersinferred from behavior% = share of members whose dominant persona this is, sized by analytics

The problem was not a shrinking interest in fitness. Members were active. The popularity feed was putting the same classes in front of everyone, whatever each member actually needed.

03
Choosing an approach

We compared rules with machine learning.

Because the research pointed to fit, the team agreed that class discovery had to move past one popular list and respond to each member. The open question was whether hand-written rules could do enough, or whether the number of members, classes, and changing preferences called for a learning model.

Rule-basedMachine learning chosen
PersonalizationLimited to the rules we defineCan respond to more combinations of goals and behavior
Finding patternsOnly finds the patterns written into itCan find useful patterns we did not anticipate
EffortQuicker and less expensiveNeeds more data and engineering work
ExplanationEasier to explainNeeds the reason and controls designed into the experience

I recommended machine learning because a fixed set of rules would struggle with all the combinations we had already seen: class type, schedule, location, goals, experience, and changing habits. A prior internal prototype had also shown a learning model was feasible. But the extra complexity was only worth it if members could understand and correct the results.

What the experience needed to do

That condition became my brief for the experience. To make the model usable, I designed around three questions: Does the class fit this member? Can they see why it appeared? Can they change what the app uses?

Fit

Use goals and history

Classes should reflect the member rather than general popularity.

Reason

Explain each suggestion

The app should say why a class appeared in plain language.

Choice

Let members change it

Members should be able to adjust individual inputs or turn personalization off.

Of the three, Fit reached furthest into the system. So I mapped every member decision to something the app could capture and the ranking could read.

What a member wantsWhere the app learns itWhat the system matches on
How hard they want itOnboarding questionHow intense the class is
Their main goalOnboarding questionWhat the class is for (strength, stress, weight…)
How experienced they areOnboarding + classes they finishClass difficulty and beginner-friendly flag
When and where they goTheir booking times and locationClass schedule and location
Routine vs. varietyHow varied their past classes areHow much variety to mix in

I designed the first two columns, how members decide and the screens that capture it. The data and engineering teams turned those into what the ranking actually uses.

04
The experience

The recommendation starts in onboarding and ends with a booked class.

Onboarding asks about goals, workout types, and experience. The home feed uses those answers with recent activity, each card explains its reason, and settings let members change or stop personalization.

Member flow: from initial preferences to booking and feedback
The member flow from onboarding through the personalized home, why-this-class, class detail, booking, and post-class feedback.
Onboarding: multi-select chips for workout types and experience level, capturing the member's goals and preferred intensity up front.
Starting preferences

Onboarding asks only what the feed needs

New members do not have a workout history yet, so onboarding asks about goals, workout types, and experience. Adding questions can lose people. I kept it to quick multi-select chips and used the answers to improve the first set of classes.

A short list of strong matches

The home feed ranks classes by the member's goals and activity rather than making general popularity the main order. I used a carousel so people could compare a few options with one thumb without scanning the full catalog. I used unmoderated testing to set the card size, amount of detail, and scroll behavior. The rest of the page includes filters, familiar workout types, occasional variety, and classes popular with coworkers.

The full personalized home feed: greeting with a weekly stat, search, filter chips, a plain-language note on why today's picks were chosen, the ranked recommendation carousel with match scores and 'Why this class?' links, 'Because you do Strength' continuity cards, a 'Surprise me' variety prompt, a 'Popular at your company' row, and 'New near you' classes.
Personalized home · full scrollRanked classes with match scores, explanations, filters, and a prompt to try something different.

Anatomy of a recommendation

The card gives members enough information to judge a class without opening it, plus a way to ask why it appeared.

A recommendation card: a 92% match badge, a Strength type tag, the title Functional Body, Strength 45 min, instructor and distance, and a Why this class button. 1 2 3 4
  1. 1
    Match scoreA fit percentage, so a member can judge relevance at a glance instead of trusting a popularity rank.
  2. 2
    Class typeThe category up front, so the kind of workout is obvious before tapping in.
  3. 3
    Title & detailsFocus, length, instructor, and distance, enough context to decide without opening the class.
  4. 4
    Why this class?Opens the explanation behind the pick, and the controls to tune it or turn it off.
The Why this class sheet: a plain-language summary, then recent activity, stated goals, and members with similar habits, plus Show fewer classes like this and Manage AI recommendations.
Transparency

“Why this class?” explains the recommendation

The sheet names the information behind the suggestion: recent activity, stated goals, and patterns among members with similar habits. It avoids technical language. From the same place, a member can ask for fewer classes like it or open personalization settings.

AI settings: individual controls such as location, plus an option to turn personalization off.
Control

Members can change or stop personalization

A member can stop using one input, such as location, without turning off everything. They can also switch personalization off completely and return to manual browsing. Those choices remain available at any time.

The 'Surprise Me' card: deliberate variety that pulls classes outside the member's usual pattern, from categories the member chooses.
Variety

“Surprise Me” widens the feed

Personalization can keep showing more of the same. The "Surprise Me" card offers classes outside a member's normal pattern, but only from categories they choose. It reads as an optional invitation rather than a recommendation error.

From class details to feedback

The explanation follows the class into detail and booking. Afterward, a simple thumbs up or down helps improve future suggestions.

Class detail screen.
Class detailThe full pick, with the match reason carried through from the card.
Booking confirmed screen.
Booking confirmedClear confirmation, with the class added to the member’s plan.
Post-class feedback screen.
Post-class feedbackA thumbs up or down changes future suggestions.

The components behind it

I designed the repeated pieces as reusable components with the states shown below.

Recommendation card component with variants.
Recommendation cardMatch score, image, title, and the “Why this class?” affordance, across states.
Preference chips component with variants.
Preference chipsThe multi-select inputs used during onboarding.
Explanation module component with variants.
Explanation moduleThe plain-language “why” behind a pick, reused wherever transparency is needed.
Controls component with variants.
ControlsThe settings for changing or turning off personalization.

Directions I tried and rejected

I kept these explorations rough so I could compare the interaction before polishing the screens.

Early concept sketch: before the carousel
Early rough concept sketch of the discovery experience.
Rejected list-view exploration.
Rejected · list view, too dense
Rejected one-at-a-time exploration.
Rejected · one at a time, no comparison
Explored 'why' placement.
Explored · ‘why’ placement
Cut long-form onboarding exploration.
Cut · long-form onboarding, too costly
05
Limits and recovery

The design also covers weak or wrong recommendations.

Personalization needs clear limits. Members should know what information is used, new members should not see made-up confidence, and a poor suggestion should be easy to correct.

Information used

The app explains that suggestions use onboarding goals, workout history, and patterns from similar members. Individual inputs, including location, can be turned off.

New members

On day one, the app uses onboarding answers only. It does not show a confident match percentage when there is not enough history to support one.

A poor suggestion

Members can ask for fewer classes like it, give a thumbs down, or turn personalization off.

The ordinary failure states

I designed loading, no history, low confidence, unavailable recommendations, missing permissions, and no results. Each state tells the member what is happening and gives them another way to continue.

Adapting the recommendation for a watch

Later that year, I designed the watchOS and Wear OS versions. I did not shrink the phone feed. Each watch shows one recommendation with enough information for a quick glance, following the conventions of that watch.

watchOS recommendation as a notification.
watchOSA notification, using the platform’s own glance pattern.
Wear OS recommendation as a Tile.
Wear OSA Tile, using its own glance pattern.
06
Results

Bookings increased, but one number is only a projection.

After launch, I compared early results against the previous three months. I treat them as directional, not a clean experiment: other teams were changing sales messaging at the same time, so I cannot attribute all of the change to the feed, and the premium upgrade figure is a projection, not a measurement.

By the numbers

Class bookings per user+27%Compared with the previous three months.
Monthly active users+18%During the same early period.
Members leaving Gympass−20%During the same period.
Estimated premium upgrades+30%Projected, not measured.

What I looked at

I tracked how quickly members found a relevant class, what they opened, how far they moved through the carousel, completed bookings, repeat attendance, and how often they turned personalization off. Usability sessions, in-app surveys, and trainer feedback helped explain the activity data.

To measure the difference more cleanly, I would show the old and new feeds to comparable groups instead of comparing time periods.

Reflection

The explanation could not be added after the recommendation was designed. The score, reason, and controls had to work together on the card and in the detail sheet. I would keep that order next time: first confirm why discovery is failing, then choose the recommendation method, then design the explanation and correction states with it.