The score should act as a prioritization signal: it helps teams decide where to look, not what a customer will definitely do. The practical requirement is inspectability. Preserve every component, document the weights and peer baseline, account for lifecycle and product cadence, and show what caused the score to move.
What is a customer health score?
A customer health score is a numeric score, category, or status intended to summarize the condition of a customer relationship. Companies create one to help teams scan a portfolio, notice meaningful changes, prioritize accounts, and apply a more consistent review process.
In customer-success practice, health scores often combine product usage with support history, relationship activity, sentiment, commercial data, or manually recorded judgment.[1] That can make an overall score convenient, but it also makes the term imprecise. Two companies can both say “customer health score” while measuring very different things.
For that reason, define the scope before defining the formula. A clear scope statement might be:
This score summarizes observed product usage for mature mid-market accounts over the last 30 days. It does not measure customer sentiment, commercial intent, or renewal probability.
That one sentence prevents a product-usage signal from being mistaken for complete customer health.
Companies commonly create health scores to:
- rank accounts for review;
- identify changes before a scheduled customer conversation;
- apply a shared definition of product-usage strength;
- distinguish broad adoption from activity concentrated in one person;
- connect a signal to an appropriate next investigation;
- monitor whether an account is moving relative to its own baseline or relevant peers.
A score becomes less useful when its purpose is simply “predict churn” without a defined outcome, time horizon, validation method, or visible explanation.
Product-usage health versus complete customer health
Several related concepts are often treated as interchangeable. They are not.
| Concept | What it represents | Typical use | What it should not be confused with |
|---|---|---|---|
| Product-usage health score | A structured summary of observed account behavior inside the product | Prioritizing adoption, engagement, concentration, consistency, or friction reviews | Complete customer health or customer intent |
| Overall customer health score | A broader synthesis that may include product usage, relationship, support, commercial, outcome, and sentiment evidence | Portfolio reviews and cross-functional account discussions | An objective fact that explains every dimension of the relationship |
| Churn prediction model | A statistical or machine-learning model trained to estimate a defined future outcome over a defined horizon | Forecasting and risk ranking after out-of-sample validation | A manually weighted score that has never been tested on unseen periods |
| Prioritization signal | A rule, score, or flag used to decide what the team should inspect first | Directing limited product or customer-success attention | Proof that an account will churn or renew |
| Manually assigned customer-success status | A human assessment such as “on track,” “blocked,” or “executive risk” | Recording context that may not exist in behavioral data | An automatically measured product-usage state |
A product-usage score is one evidence layer. A complete account review may also need:
- Product-usage health: Adoption, workflow coverage, meaningful activity, cadence, user distribution, trend, and observed friction.
- Relationship health: Stakeholder engagement, champion strength, executive alignment, and the quality of recent conversations.
- Support health: Ticket severity, unresolved issues, recurring incidents, and escalation patterns.
- Commercial health: Contract status, payment or procurement issues, renewal timing, and changes in budget or scope.
- Customer outcomes and sentiment: Whether the customer is reaching its intended outcomes and how users or stakeholders describe the experience.
These layers can inform an overall view, but collapsing all of them into one unexplained number can reduce interpretability. A score of 68 does not tell the team whether the account has a product-adoption problem, a procurement delay, a missing champion, or a severe support escalation.
For a wider explanation of why B2B analytics needs account, user, and product context rather than only aggregate events, link to B2B product analytics.
Why one number is dangerous
Composite indicators are useful because they compress several dimensions into something easy to scan. They are also risky because a seemingly precise result can hide the assumptions, weights, missing data, and trade-offs underneath it. The OECD and European Commission’s methodological guidance on composite indicators warns that poorly constructed composites can produce misleading conclusions, obscure serious weaknesses in individual dimensions, and make corrective action harder to identify.[2]
A customer health score has the same problem.
Suppose a company has:
- strong total activity;
- broad feature coverage;
- one highly active champion;
- almost no participation from the rest of the account.
A weighted average may still produce a healthy-looking score because high activity compensates for weak user distribution. The arithmetic is valid, but the account remains fragile.
The opposite can also happen. A customer may use a product only once per month because that is the correct cadence for its workflow. A generic recency rule may classify it as inactive even though usage is completely normal.
A score therefore needs two forms of transparency:
- component transparency: show the underlying values and weights;
- interpretation transparency: show the lifecycle, peer baseline, data coverage, exceptions, and reason for movement.
What product events cannot tell you on their own
Behavioral data is valuable, but its scope is limited.
Customer intent is not directly observable in events. Events show what happened in the product. They do not directly reveal why a user acted, whether a stakeholder is satisfied, or what the buying committee intends to do. Large-scale behavioral metrics are more useful when they are connected to product goals and triangulated with other evidence rather than treated as a substitute for research or customer context.[3]
Budget and procurement changes are often absent. An account may use the product consistently while facing a spending freeze, legal review, procurement delay, or vendor-consolidation decision.
Leadership changes may not appear in usage immediately. A new executive sponsor, reorganization, acquisition, or strategy shift can change the renewal context before product behavior changes.
Seasonal or low-frequency products can look inactive. Tax software, quarterly reporting tools, annual planning products, and incident-response systems do not share the cadence of daily collaboration software.
Onboarding accounts require different expectations. Low breadth in week two may be normal. The same breadth six months later may indicate incomplete adoption.
One champion can produce healthy totals while the account remains fragile. Account-level volume can hide the fact that most activity belongs to one person.
Integrations and automated activity may create false engagement. Background jobs, API traffic, monitoring, scheduled exports, or integration users can inflate usage without representing active human adoption.
Usage can decline because a workflow was successfully completed. A migration, setup, audit, or planning cycle may end. Less activity can mean the work is done rather than the customer is disengaging.
These limitations do not make a product-usage health score useless. They define what the score can responsibly claim.
Seven product-usage dimensions
Breadth, depth, and frequency are common starting points in product-usage and customer-health frameworks.[4] A B2B account model usually needs more context because several users, roles, workflows, and lifecycle stages contribute to the same customer relationship.
The following seven dimensions provide a transparent framework. A product does not have to place every dimension in one formula. The important requirement is to define each dimension, expose its evidence, and document why it is included or excluded.
1. Adoption
What it measures
Adoption measures whether the account has crossed a meaningful threshold in the workflows that matter for its use case. It should focus on value-bearing actions rather than simple discovery.
A workflow should not count as adopted merely because someone opened its page. Define a threshold that represents meaningful use, such as completing a core task, producing an output, configuring a required capability, or returning to the workflow on more than one relevant occasion.
Possible metrics
- percentage of eligible essential workflows adopted;
- number of core product areas with meaningful use;
- completion of required setup or activation milestones;
- percentage of weighted adoption milestones completed;
- number of roles that have completed their expected first-value workflow.
An illustrative formula is:
Adoption
Adoption = weighted adopted essential workflows ÷ weighted eligible essential workflows
What a high or low value may mean
A high adoption score may indicate that the account has reached the workflows most closely connected to its intended use case. A low value may indicate incomplete onboarding, poor discovery, a setup dependency, a role mismatch, or a product area that has not become relevant.
Important exceptions
Not every workflow applies to every account. Plan access, customer use case, role permissions, purchased modules, and delegated processes all affect eligibility. A one-time setup workflow can be fully adopted even if it is not used again.
What to inspect next
Inspect which essential workflows remain unadopted, which role is expected to perform them, whether the account has access, whether initial use progressed to completion, and which users or visits provide evidence.
2. Breadth
What it measures
Breadth measures how much of the relevant product the account uses. Adoption asks whether important workflows crossed a meaningful threshold. Breadth asks how widely usage extends across all eligible product areas, grouped pages, modules, or workflows.
For a practical account-level measurement model, see product usage by company.
Possible metrics
- adopted eligible product areas divided by eligible product areas;
- grouped pages used by the account;
- number of modules with active usage;
- percentage of expected workflow categories represented;
- breadth by team, role, or business unit.
An illustrative formula is:
Breadth
Breadth = adopted eligible product areas ÷ eligible product areas
What a high or low value may mean
High breadth may indicate that the account depends on several parts of the product and has moved beyond one narrow use case. Low breadth may indicate incomplete adoption, a specialized use case, a missing workflow, or usage concentrated around one job.
Important exceptions
Broad usage is not automatically better. Some customers purchase a product for one specialized workflow. A viewer role may need only one product area. Plan restrictions may make apparently missing areas irrelevant.
What to inspect next
Compare the account with peers that share the same plan, use case, lifecycle, and expected roles. Inspect whether breadth is distributed across users or produced by one explorer, and whether previously adopted areas have disappeared.
3. Depth
What it measures
Depth measures how meaningfully or intensively an adopted workflow is used. It distinguishes a brief trial from recurring work that produces a relevant output.
Possible metrics
- completed core actions per active user or active day;
- percentage of started workflows that reach a defined completion;
- outputs created, reviewed, shared, exported, or used downstream;
- repeated use of a core workflow across relevant periods;
- volume normalized by account size or eligible users;
- task success, error rate, or time to complete a known workflow.
What a high or low value may mean
High depth may indicate that the workflow is embedded in the customer’s work. Low depth may indicate shallow discovery, incomplete onboarding, a workflow that lacks relevance, or a user who opened a feature without progressing.
Important exceptions
More activity is not always healthier. Excessive repetition can indicate retries, poor efficiency, or an automated process. Time spent is especially ambiguous: longer time can represent valuable work, but it can also mean difficulty, waiting, or confusion. The HEART framework separates engagement from task success for this reason and notes that common behavioral totals can have ambiguous interpretations.[3]
What to inspect next
Inspect task completion, outputs, workflow duration, repeat attempts, downstream use, errors, and relevant qualitative or session evidence. Do not raise the score simply because users spent longer in the product.
4. Consistency and recency
What it measures
Consistency and recency measure whether relevant activity repeats at an expected cadence and how recently the account performed a meaningful action.
Recency should be evaluated against the product’s natural workflow frequency rather than a universal number of days.
Possible metrics
- active expected intervals divided by expected intervals;
- number of active days or weeks within the evaluation period;
- days since the last meaningful workflow;
- longest unexplained gap;
- variation in usage between expected periods;
- percentage of scheduled cycles with at least one completed key task.
An illustrative consistency formula is:
Consistency
Consistency = expected intervals with meaningful use ÷ expected intervals observed
What a high or low value may mean
High consistency may indicate that the product has become part of a recurring process. Low consistency may indicate declining habit, a blocked workflow, a completed project, seasonal use, or a cadence that was defined incorrectly.
Important exceptions
A monthly reporting tool should not be evaluated like a daily collaboration product. Holidays, fiscal calendars, project completion, maintenance windows, and event-driven products can all produce legitimate gaps.
What to inspect next
Compare the account with its own historical cadence first. Then inspect equivalent peers, seasonality, lifecycle, last meaningful activity, and visits around the point where the pattern changed.
5. User distribution
What it measures
User distribution measures how many appropriate users participate and whether activity is concentrated in too few people.
This is one of the most important B2B distinctions. Healthy account totals can hide a single point of failure when one champion performs most meaningful actions.
Possible metrics
- active eligible users divided by known eligible users;
- percentage of licensed users who completed a meaningful workflow;
- number of active roles, teams, or departments;
- share of meaningful activity produced by the most active user;
- share produced by the top three users;
- number of users with recurring rather than one-off activity.
Two useful formulas are:
Eligible-user penetration
Eligible-user penetration = active eligible users ÷ known eligible users
Top-user concentration
Top-user concentration = meaningful activity from the most active user ÷ account meaningful activity
What a high or low value may mean
Broad, role-appropriate distribution may indicate that usage is resilient and not dependent on one person. Low penetration or high concentration may indicate incomplete rollout, permissions problems, role mismatch, weak onboarding, or champion dependency.
Important exceptions
Some products are legitimately admin-led or specialist-led. Licensed seats may not equal eligible users. An account with one expected operator should not be penalized for not having ten active users.
What to inspect next
Inspect roles, permissions, team structure, invite and onboarding progress, inactive eligible users, backup champions, and whether the denominator is accurate. If the eligible-user count is unknown, do not silently convert missing information into a score of zero.
6. Trend
What it measures
Trend measures whether comparable usage is stable, increasing, or declining. It is usually more informative than a single current-period value because it shows movement.
Possible metrics
- current period versus the immediately preceding equal-length period;
- rolling slope over several comparable periods;
- change from the account’s own historical baseline;
- change in adoption breadth, active users, consistency, or depth;
- movement relative to matched peers;
- percentage-point change for percentage metrics.
What a high or low value may mean
An improving trend may indicate onboarding progress, increased adoption, a new team rollout, or a temporary spike. A declining trend may indicate lost users, narrowing workflows, seasonality, completed work, instrumentation changes, or genuine loss of momentum.
Important exceptions
A release, data migration, automated integration, company headcount change, holiday, or tracking change can move the trend without changing customer health. A high-growth onboarding account and a stable mature account should not share the same expectation.
What to inspect next
Decompose the trend into users, workflows, product areas, and periods. Check whether the change is broad or concentrated, verify tracking quality, and inspect relevant visits or account context before assigning a cause.
7. Friction or failed progress
What it measures
This dimension captures evidence that users repeatedly fail to progress, encounter errors, abandon a known workflow, or need unusual effort to complete it.
Friction should be scored only where the product can measure it reliably. In many cases it is safer to preserve friction as a separate diagnostic flag rather than burying it inside a positive engagement score.
Possible metrics
- error rate within a defined workflow;
- abandonment after a known start event;
- repeated retries or corrections;
- backtracking between workflow steps;
- repeated incomplete setup sequences;
- unusually long task completion relative to comparable successful sessions;
- support-linked sessions or confirmed usability problems.
What a high or low value may mean
More confirmed friction may indicate a product problem, configuration issue, missing permission, unclear requirement, or reliability problem. A low observed-friction value does not prove that the experience is easy; the instrumentation may simply miss the problem.
Important exceptions
Repeated actions can be normal. A session can look incomplete because the user changed goals. One unusual visit is anecdotal. Automated traffic can create retries that do not represent a human experience.
What to inspect next
Inspect multiple affected users and accounts, the exact workflow, errors, session evidence, and support or research context. Confirm the pattern before treating it as a cause.
An illustrative customer health score formula
The following customer health score formula is an example, not a Hymetry standard and not an industry best practice.
Assume a team has selected five components that it can define and measure consistently:
- adoption;
- breadth;
- consistency;
- user distribution;
- trend.
Normalize each component to a 0–100 scale, then calculate:
Illustrative product-usage health score
Product usage health = Adoption × 25% + Breadth × 20% + Consistency × 20% + User distribution × 15% + Trend × 20%
The weights express a product-specific judgment. Another product may give more weight to consistency, treat breadth as diagnostic only, or include depth in the score. A product with reliable task-completion instrumentation may include friction as a negative component. A specialist product may use a different distribution model.
A transparent scoring process
- Choose the account cohort and evaluation window: Define lifecycle, plan, use case, account size, expected cadence, and minimum data requirements.
- Define every component before normalizing it: Document the numerator, denominator, eligibility rules, direction, time window, missing-data behavior, and source events.
- Normalize the components: Use fixed target ranges or clearly defined peer baselines. Version the bounds so historical values remain reproducible.
- Apply documented weights: Make the weights visible and explain why each component is included.
- Calculate the composite: Keep the arithmetic simple enough for a product or customer-success team to reproduce.
- Preserve the evidence: Store and display the raw value, normalized component, weight, data coverage, previous value, and reason for movement.
For a higher-is-better metric with a documented lower and upper target, one illustrative normalization method is:
Illustrative component normalization
Normalized component = clamp(0, 100, 100 × (observed value − lower target) ÷ (upper target − lower target))
Invert the direction for a lower-is-better metric. Do not use changing global minimums and maximums without understanding how they will alter historical comparability.
Depth and friction are omitted from this illustrative formula because they are often highly product-specific or unevenly measured. They should still remain visible as diagnostic evidence or guardrails.
Worked example: four fictional B2B SaaS accounts
All company names and values in this example are fictional. The decimals show the arithmetic; a production interface may round the final score because the input definitions are not precise to tenths.
| Fictional account | Lifecycle | Adoption | Breadth | Consistency | User distribution | Trend | Illustrative score |
|---|---|---|---|---|---|---|---|
| ArborGrid | Mature | 88 | 90 | 86 | 82 | 78 | 85.1 |
| BeaconDesk | Mature | 82 | 78 | 86 | 25 | 68 | 70.7 |
| CedarOps | Mature | 88 | 82 | 75 | 84 | 25 | 71.0 |
| DeltaPilot | Onboarding | 42 | 35 | 50 | 55 | 88 | 53.4 |
ArborGrid: broad, consistent, distributed use
ArborGrid has adopted most relevant workflows, uses a broad portion of the eligible product, returns consistently, and distributes activity across several appropriate users. Its score is high because the component pattern is balanced.
The next review should not stop at “healthy.” The team should still inspect whether depth represents successful work, whether any important role is absent, and whether friction is hidden behind high activity.
BeaconDesk: strong activity concentrated in one champion
BeaconDesk has good adoption, breadth, and consistency. Its user-distribution score is low because one person performs most meaningful work.
The account’s total activity can look healthy while its adoption remains fragile. The appropriate action is not a generic engagement email. The team should identify backup champions, inspect which roles are missing, and understand whether permissions or onboarding are limiting broader use.
CedarOps: broad adoption with a declining trend
CedarOps has broad, distributed usage but a low trend score. Its current state remains substantial, yet comparable activity has declined.
The team should inspect which workflows and users account for the change, whether the decline is seasonal, and whether recent visits show a blocked or completed workflow. The account does not need the same action as BeaconDesk.
DeltaPilot: onboarding momentum that should not be compared with mature accounts
DeltaPilot has low adoption and breadth because it is newly onboarded. Its positive trend shows that activity is increasing.
A mature-account threshold would classify the account too harshly. DeltaPilot should be assessed against onboarding milestones, expected time to first value, setup completion, and peers at the same stage. It may be progressing normally even though the arithmetic score is 53.4.
Similar scores can require different actions
BeaconDesk and CedarOps have almost identical final scores: 70.7 and 71.0.
Their component patterns are very different:
- BeaconDesk has a concentration problem.
- CedarOps has a trend problem.
A single number would place both accounts in the same queue. A transparent score tells the team what to inspect next.
Thresholds and status labels
Status labels such as healthy, needs attention, and high risk can make a score easier to scan. They should be treated as operational categories, not objective truths.
An illustrative threshold set might be:
| Illustrative range | Status | Intended meaning |
|---|---|---|
| 75–100 | Healthy | No immediate product-usage review is indicated, but components remain visible |
| 55–74 | Needs attention | At least one component or movement deserves review |
| 0–54 | High risk | High product-usage review priority; not a deterministic churn prediction |
| Not scored | Onboarding or insufficient data | Use lifecycle milestones or resolve data coverage before classifying |
Do not copy these ranges into production without calibration. They exist only to show how labels may map to a score.
Thresholds should be:
- calibrated using the team’s own historical data;
- compared with relevant peers;
- adjusted for lifecycle and expected cadence;
- periodically reviewed;
- versioned when definitions change;
- accompanied by component-level explanations.
A label should answer “why?” For example:
That is more actionable than a red badge with no explanation.
Use guardrails when averaging would hide a serious weakness
A weighted average allows compensation: a high component can offset a low one. That may not always be acceptable.
Examples of guardrails include:
- an account cannot be labeled healthy when required identity data is missing;
- extreme champion concentration can force a review status even when the overall score is high;
- confirmed severe friction can create a separate alert rather than subtracting a few points;
- an onboarding account can remain in an onboarding state until minimum milestones or observation periods are complete.
Document every override. Hidden overrides create a different kind of opacity.
Peer baselines and lifecycle context
A B2B SaaS health score should compare an account with a relevant expectation, not one global average.
Useful peer dimensions include:
- account size;
- plan or purchased modules;
- lifecycle stage;
- product use case;
- onboarding status;
- expected workflow frequency;
- industry, where it materially changes product behavior;
- known role or team structure.
A practical baseline hierarchy is:
- compare the account with its own prior behavior;
- compare it with accounts in the same lifecycle and use case;
- add plan, size, or cadence where those dimensions materially change expectations;
- use industry only when there is a defensible behavioral difference and enough data.
Lifecycle-specific score design is a common recommendation in customer-success practice because onboarding, adoption, maturity, and renewal stages produce different expected behavior.[5]
Avoid over-segmenting. A cohort with only a few accounts may produce unstable percentiles and thresholds. When the peer sample is too small, show that limitation and rely more heavily on the account’s own historical baseline.
How to validate whether the score is useful
A customer health score is useful when it consistently directs attention toward accounts and questions that matter. It is not useful merely because the formula produces a number for every account.
Test the score on the following dimensions.
| Test | Question to ask | Practical check |
|---|---|---|
| Stability | Does the score remain reasonably stable when behavior is unchanged? | Recalculate across adjacent equivalent windows and inspect unexplained movement |
| Sensitivity | Does it move when a meaningful workflow, user group, or cadence changes? | Create known test cases and confirm that the relevant component responds |
| Noise resistance | Can one import, integration user, retry loop, or brief spike distort it? | Replay the score with noisy events removed or capped |
| Interpretability | Can a product manager or CSM explain the score without reading implementation code? | Ask reviewers to identify the top contributors and next investigation |
| Review yield | Do flagged accounts produce a meaningful issue or question when reviewed? | Sample the review queue and record whether the signal was useful |
| False-positive rate | How often are stable accounts incorrectly flagged under the defined outcome? | Use a confusion matrix when the score is tied to a binary prediction |
| Actionability | Does each material component map to a sensible investigation or action? | Verify that every low component has an evidence path and owner |
| Time-window consistency | Does the interpretation remain coherent across appropriate 7-, 30-, or 90-day views? | Compare windows that match the product’s cadence |
If the score is described as predictive
A manually weighted product-usage score is not automatically a churn prediction model.
To describe a model as predictive, define:
- the exact outcome, such as non-renewal;
- the prediction horizon;
- the population being scored;
- the observation period;
- the features available at prediction time;
- the validation period;
- the decision threshold;
- the costs of false positives and false negatives.
Evaluate it on periods the model did not see during development. For time-ordered customer data, random splitting can leak future information into training. Time-based validation preserves ordering and avoids training on future periods while evaluating past ones.[6]
Report metrics that reflect the actual decision. Precision measures how many flagged cases match the defined positive outcome, while a confusion matrix exposes true positives, false positives, false negatives, and true negatives.[7] For a review queue, also report how many of the top-ranked accounts produce a useful review.
Compare the model with simple baselines. A complex model that does not outperform last meaningful activity, active-user change, or a basic rule may not justify its opacity.
Finally, correlation with churn does not establish causation. A pattern can help identify accounts worth reviewing without proving that changing the metric will change the outcome. Product usage may be a symptom, a consequence, a shared effect of another cause, or a genuinely useful leading signal. The score alone cannot decide which.
From signal to action
A score should lead to a specific investigation, not a generic playbook triggered by the final number.
| Visible signal | What to inspect | Possible next action |
|---|---|---|
| Lost adoption breadth | Which workflows disappeared, which roles used them, and whether plan or use case changed | Confirm the dropped workflow and address discovery, fit, access, or product problems |
| Rising user concentration | Top-user share, backup champions, inactive roles, and team distribution | Develop backup champions and broaden role-appropriate adoption |
| Declining consistency | Recency, expected cadence, seasonality, project completion, and recent visits | Determine whether the gap is normal before contacting the account |
| Low eligible-user penetration | Denominator quality, permissions, invitations, onboarding, and role fit | Resolve access or onboarding gaps for the appropriate users |
| High activity with repeated friction | Task completion, errors, retries, support evidence, and relevant sessions | Investigate a product or configuration problem rather than celebrating volume |
| Missing or inconsistent data | Identity mapping, instrumentation, integration users, and event coverage | Fix data quality before assigning poor health |
| Sudden positive trend | Which users and workflows drove it and whether activity is human or automated | Confirm durable adoption before labeling expansion readiness |
Common customer health score mistakes
Using logins as the dominant signal
A login proves access, not meaningful adoption. Include it as context when useful, but do not let it outweigh workflow-level evidence.
Combining unrelated data into an opaque score
Product usage, support severity, payment status, sentiment, and CSM judgment represent different dimensions. If they are combined, preserve the sub-scores and the reason each one moved.
Treating missing data as poor health
Unknown eligible-user counts, incomplete identity, absent sentiment, or broken instrumentation should produce a coverage warning—not an automatic zero.
Comparing onboarding and mature customers
New accounts have different goals, time windows, and expected breadth. Use lifecycle-specific milestones or score profiles.
Weighting what is easy to collect instead of what represents value
A reliable page-view event can still be a weak health signal. Start with the customer workflow and product goal, then choose measurable evidence.
Letting one power user hide account fragility
Always inspect user distribution and concentration alongside account totals.
Changing score definitions without preserving historical comparability
Version the formula, weights, component definitions, normalization bounds, and effective date. Backfill history only when it can be done consistently, and label discontinuities.
Presenting the score as guaranteed churn prediction
A prioritization score, a descriptive composite, and a validated predictive model are different tools. Name the tool accurately.
Automating outreach solely from the final score
Use the score to open the right company, user, product-area, or visit evidence. Let a person choose the response.
How Hymetry approaches company health signals
Hymetry is account-centric product intelligence for B2B SaaS. It connects product behavior across companies, users, product areas, grouped pages, and Visits so teams can inspect the evidence behind an account-level signal.
In Companies, a team can begin with company-level context such as adoption breadth, active users, engaged time, recent movement, user-health distribution, and product-area usage.[8] The account signal remains connected to the source views that can explain it.
The investigation can then move through four levels:
- Company-level change: Identify the account whose usage, adoption, or user pattern changed.
- Product-area adoption: Inspect which relevant product areas or grouped workflows were adopted, missed, or dropped.
- User distribution: Use Users to see which people are active, passive, dropped, at risk, or gaining momentum, and whether one champion carries the account.[9]
- Relevant visits: Open Visits when session-level evidence is necessary to understand what actually happened in the product.[10]
This approach keeps the score or status separate from its components until the team needs to connect them. Aggregate analytics identify where attention may be useful; account and user context show who is affected; visits provide evidence for deeper investigation.
Hymetry does not claim to observe every cause of renewal risk, infer customer intent from events, or replace customer-success judgment. Its customer-success workflow is designed to connect account behavior, a directional signal, source evidence, and the next review step.[11]
Frequently asked questions
What is a customer health score?
A customer health score is a numeric score, category, or status that summarizes selected evidence about a customer account. In B2B SaaS, it may include product usage, relationship, support, commercial, outcome, or sentiment data. The score is only meaningful when its scope and components are defined.
How do you calculate a customer health score?
First define the purpose, account cohort, lifecycle, expected cadence, and source metrics. Normalize the selected components, apply documented weights, calculate the composite, and preserve every component and its movement.
An illustrative product-usage formula is:
Illustrative product-usage formula
Adoption × 25% + Breadth × 20% + Consistency × 20% + User distribution × 15% + Trend × 20%
These weights are an example, not a universal standard.
Which customer health score metrics should a B2B SaaS company include?
For product usage, consider adoption, breadth, depth, consistency and recency, user distribution, trend, and reliably observed friction. Keep relationship, support, commercial, outcome, and sentiment evidence in separate layers or visible sub-scores.
Choose metrics that represent the product’s value and expected workflows rather than metrics that are merely easy to capture.
Is a customer health score the same as churn prediction?
No. A health score may be a descriptive composite or prioritization rule. A churn prediction model has a defined outcome and horizon, is trained on historical data, and is evaluated on unseen periods. A score should not be described as predictive unless it has been validated as such.
What is a good customer health score?
There is no universal good score. A value depends on the formula, normalization, lifecycle, plan, account size, use case, expected cadence, and peer baseline. Thresholds should be calibrated with the company’s own data and accompanied by component-level explanations.
How often should a customer health score update?
Update it often enough to detect relevant changes without moving faster than the product’s natural cadence. A daily-use product may benefit from frequent updates with a multi-week comparison window. A monthly or quarterly workflow needs longer observation periods. The calculation frequency and interpretation window do not have to be the same.
Should onboarding accounts use the same score as mature accounts?
Usually not. Onboarding accounts should be evaluated against setup, activation, role participation, and time-to-value milestones appropriate to their stage. Applying mature-account breadth or consistency thresholds can create false risk signals.
Can product usage alone reveal customer intent?
No. Product events show observed behavior, not budget decisions, procurement changes, leadership priorities, satisfaction, or renewal intent. Product usage can provide valuable health signals, but it should remain one inspectable layer of the account review.
Sources
Vendor sources below document common customer-success and product-usage practices. They are not treated as independent proof that a particular score predicts churn.
- Gainsight — Customer Health Score Explained: Metrics, Models & Tools — https://www.gainsight.com/blog/customer-health-scores/
- OECD and European Commission Joint Research Centre — Handbook on Constructing Composite Indicators: Methodology and User Guide — https://www.oecd.org/content/dam/oecd/en/publications/reports/2008/08/handbook-on-constructing-composite-indicators-methodology-and-user-guide_g1gh9301/9789264043466-en.pdf
- Google Research — Measuring the User Experience on a Large Scale: User-Centered Metrics for Web Applications — https://research.google.com/pubs/archive/36299.pdf
- Pendo Help Center — Improve customer health and retention — https://support.pendo.io/hc/en-us/articles/360042454072-Improve-customer-health-and-retention
- Vitally — How To Build a Customer Health Score by Customer Lifecycle Stage — https://www.vitally.io/post/why-you-should-build-customer-health-scores-by-lifecycle-stage
- scikit-learn — TimeSeriesSplit — https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.TimeSeriesSplit.html
- scikit-learn — precision_score — https://scikit-learn.org/stable/modules/generated/sklearn.metrics.precision_score.html
- Hymetry — Companies: Account Intelligence — https://www.hymetry.com/product/companies/
- Hymetry — Users: User Intelligence for B2B SaaS — https://www.hymetry.com/product/users/
- Hymetry — Visits: Session Visits & Replay — https://www.hymetry.com/product/visits/
- Hymetry — Customer Success — https://www.hymetry.com/use-cases/customer-success/