Published on : Sep 16, 2026

North Star Metric vs KPI: How to Choose the Right Metrics for Your Business

Screening a candidate metric, decomposing it into owned inputs, and pairing it with guardrails

5 Minutes Read
Rutvik Acharya, Principal Data Scientist at Atlassian

Rutvik Acharya

Principal Data Scientist Atlassian

North Star Metric vs KPI: How to Choose the Right Metrics for Your Business thumbnail

North Star Metric vs KPI: How to Choose the Right Metrics for Your Business

The argument usually starts as a definitional one and stays there. Someone says the north star should be revenue. Someone else says revenue is a lagging indicator. A third person points out that the company already tracks forty KPIs and nobody looks at them. Nothing gets decided, and the dashboard grows by two tiles.

The distinction is not academic, and it is not about which number is most important. A north star metric is a single leading measure of the value customers receive, chosen so that moving it reliably moves long-term revenue. A KPI is any indicator a team is accountable for. One is a coordination device for a whole company. The other is a management tool for a function. Confusing them produces either a north star nobody can influence or a set of KPIs pulling in different directions.

This article covers how the two differ in practice, how to screen a candidate north star before committing to it, how to decompose it into inputs that individual teams actually own, what guardrails are for, and how to write a metric definition that survives contact with three teams.

The running example is a B2B collaboration SaaS product. Figures are illustrative and constructed so the arithmetic is checkable.


Four Kinds of Metric, and What Each Is For

Most metric confusion dissolves once you separate four categories that get lumped together as "KPIs".

Type

Purpose

Time horizon

How many

North star

Aligns the whole company on delivered value

Leading, weekly or monthly

Exactly one

Input metric

Decomposes the north star into movable parts

Leading, weekly

One per major driver

KPI

Holds a team accountable for its area

Mixed

A handful per team

Guardrail

Detects damage caused by pursuing the north star

Lagging or leading

A small fixed set

Two clarifications that resolve most disputes.

Revenue is usually a poor north star and an excellent KPI. It is the outcome you want, but it is lagging, it moves for reasons no single team controls, and a product team cannot act on it this week. That does not demote it. It stays on the executive dashboard as the outcome the north star is supposed to predict.

Having one north star does not mean having one metric. The north star is the alignment device. Teams still need their own KPIs, and the company still needs guardrails. What "one metric that matters" is arguing against is having five competing top-level goals, not against measuring anything else.


Screening a Candidate North Star

A candidate north star has to clear five tests. Most popular candidates fail at least two, and the failures are predictable.

1. Does it reflect value the customer receives? Not value you extracted, value they got. Registered users measures your success at collecting signups; it says nothing about whether anyone got anything.

2. Is it leading rather than lagging? It should move before revenue does, so it functions as a steering wheel rather than a rear-view mirror.

3. Can teams actually move it within a quarter? A metric influenced mainly by macroeconomic conditions or a two-year sales cycle cannot coordinate weekly work.

4. Is it hard to game, or paired with guardrails that catch gaming? Any metric under pressure will be optimised, including in ways nobody intended.

5. Is it one number a new joiner can understand? If explaining it takes a paragraph, it will not survive as a shared reference.

Screenshot 2026-09-02 184628.png

Applying that to four candidates for the collaboration product:

Registered users fails on customer value and on gaming resistance. It rises when marketing runs a giveaway and never falls, which makes it a ratchet rather than a signal.

Monthly recurring revenue fails on leading and on team influence. It is the right outcome and the wrong steering instrument.

Weekly active users is closer, but "active" usually means "opened the app", which is not value received, and it is trivially inflated by notification volume.

Weekly active teams completing at least three collaborative actions passes all five, at the cost of being longer to say. It captures the product's actual value (teams working together, not individuals logging in), it leads revenue because collaborating teams renew, teams can move it, the collaborative-action threshold makes notification spam ineffective, and a new joiner understands it immediately.

The threshold in that definition is not arbitrary and should not be invented. Derive it from your own data by finding the usage level above which retention meaningfully separates, then state that you derived it and revisit it annually.


Decomposing the North Star Into Owned Inputs

A north star that no team owns is a poster. The decomposition into input metrics is what makes it operational.

The north star here is the count of weekly active collaborating teams. That count is a stock, and stocks are governed by flows:

Plain text
Active teams (steady state) = weekly inflow ÷ (1 - weekly retention rate)

With the running example's numbers:

Input

Value

Owner

New accounts per week

1,400

Growth

Activation rate (reach 3 collaborative actions in week 1)

35%

Onboarding

Newly activated teams per week

490

Onboarding

Resurrected dormant teams per week

110

Lifecycle

Total weekly inflow

600

Weekly retention rate

88%

Core product

Active teams at steady state

5,000

Check it: 600 ÷ (1 − 0.88) = 5,000.

Screenshot 2026-09-02 184551.png

Now the question that decides next quarter's roadmap. Which input is worth improving?

Change

New steady state

Gain

Baseline

5,000

Activation 35% to 40% (a 5 point gain)

5,583

+583

Retention 88% to 90% (a 2 point gain)

6,000

+1,000

Two points of retention beat five points of activation, because retention applies to the entire base while activation applies only to this week's new accounts. That is not a general law, it is arithmetic that depends on the current base size relative to inflow, and it will change as the company grows. The point is that the tree lets you compute it rather than argue about it.

The activation term is a funnel, and improving it means finding which step of onboarding loses teams before the third collaborative action, which is the analysis covered in the guide to finding where customers drop off in a funnel.


Choosing KPIs From the Tree

Once the tree exists, KPI selection stops being a brainstorm and becomes an assignment. Each team takes the input it owns as its primary KPI, plus a small number of secondary measures that explain movement in it.

Team

Primary KPI (from the tree)

Secondary KPIs

Growth

New accounts per week

Cost per account, share of accounts from target segments

Onboarding

Week-1 activation rate

Time to first collaborative action, step-level onboarding completion

Core product

Weekly retention rate

Feature adoption among retained teams, session depth

Lifecycle

Resurrected teams per week

Re-engagement campaign response, dormancy duration before return

Two properties make this work. Every KPI traces upward to the north star, so no team can succeed on its own number while damaging the company's. And every KPI is influenceable by the team that holds it, which is the difference between accountability and blame.

The count matters too. A team with a dozen KPIs has none, because attention is the scarce resource and a dashboard with everything on it gets read as decoration. Keep the primary KPI visually dominant and the secondaries subordinate to it, which is a layout decision as much as a selection one; the relevant encoding principles are in the guide to data visualisation for analysts, and the mechanics of building the tracking views themselves are covered in Power BI for beginners alongside the Power BI documentation on KPI visuals and refresh.


Guardrails: What Stops the North Star Being Gamed

Any metric that determines what teams are rewarded for will be optimised, including through routes nobody intended. Guardrails exist to detect that, and they should be chosen at the same time as the north star, not after the first incident.

For the collaboration product, the obvious gaming routes and their guardrails:

Gaming route

What it looks like in the north star

Guardrail

Notification volume drives activity

Active teams rise

Notification opt-out rate, unsubscribe rate

Counting trivial actions as collaborative

Active teams rise

Eight-week retention of newly activated teams

Buying low-quality signups

Inflow rises

Contribution margin per account, activation rate by channel

Onboarding pushes teams through the threshold artificially

Activation rate rises

Week-4 activity among week-1 activated teams

The pattern is consistent: every guardrail is a metric that would deteriorate if the north star were being moved the wrong way. That is what makes it a guardrail rather than just another KPI.

Guardrails need thresholds, and the honest way to set one is from the metric's own historical variation rather than from a round number. Compute the measure across many past periods, look at the spread, and set the alert outside the range that normal operation produces. The NIST/SEMATECH e-Handbook of Statistical Methods is a reliable reference for how that variability is characterised.

When a guardrail fires and the north star looks good, the resulting conversation is uncomfortable, because someone's quarter is about to be reinterpreted. Leading with the guardrail number and the specific mechanism, rather than with an accusation of gaming, is the same skill as presenting difficult findings to senior leaders.


Writing a Definition That Survives

A metric that three teams compute differently is worse than no metric, because it produces confident disagreement rather than uncertainty. Every metric on the tree needs a written specification with six parts.

Numerator and denominator. For the activation rate: teams reaching three collaborative actions, over teams that created an account. Not sessions, not users, teams.

Population and exclusions. Internal test accounts out, trial accounts in or out (state which), accounts from a specific region excluded if a data feed is unreliable there.

Window. Week 1 means the seven days from account creation, not the calendar week the account was created in. This single ambiguity produces more metric disagreements than any other.

Grain and unit. One row per team per week, and the metric is a count of teams rather than of actions.

Owner and cadence. Who is accountable for the number and who is accountable for the definition, which are frequently different people, and how often it refreshes.

Aggregation rule. Rate metrics must be computed as a ratio of sums rather than as an average of ratios, or weekly and quarterly roll-ups will silently disagree with the daily figures.

Plain text
-- Correct: sum the components, then divide.
SELECT
    date_trunc('week', signup_ts)                      AS cohort_week,   -- PostgreSQL
    count(*)                                           AS accounts,
    count(*) FILTER (WHERE activated_wk1)              AS activated,
    round(100.0 * count(*) FILTER (WHERE activated_wk1)
          / NULLIF(count(*), 0), 1)                    AS activation_pct
FROM team_accounts
WHERE is_internal = false
GROUP BY 1
ORDER BY 1;
-- FILTER is PostgreSQL; use count(CASE WHEN ... THEN 1 END) in other engines.
-- Wrong, and it will not error: avg(activation_rate_per_day)

That aggregation behaviour follows from how the counting aggregates handle conditions and nulls, documented in the PostgreSQL aggregate function reference. Where the organisation runs a modelling layer, these definitions belong there rather than in individual dashboards, which is the case the dbt semantic layer documentation makes for centralising metric logic.


When to Change the North Star

Changing it is expensive, because every target, dashboard, and historical comparison is anchored to it. Three situations justify the cost.

The business model changed. A company that moves from self-serve to enterprise sales has a different value delivery mechanism, and a metric built for the old one will steer the new business wrong.

The metric stopped predicting the outcome. This is the empirical test and it is worth running annually: does the north star still correlate with revenue and retention over the following two or three quarters? If the relationship has decayed, the metric has stopped doing its job regardless of how well understood it is.

It has been thoroughly gamed and cannot be repaired. Sometimes a guardrail reveals that the metric's definition is the problem rather than the behaviour, and patching the definition repeatedly is worse than replacing it.

What does not justify a change: a new executive preferring a different number, or the current metric looking bad. Changing the metric because it is unflattering is the clearest possible signal that measurement is not being used to make decisions, which is one of the recurring patterns behind why analytics projects fail.

When you do change it, run both metrics in parallel for at least a quarter, and publish the mapping between them so historical comparisons remain interpretable.


Where to Go From Here

For improving the activation term in the tree, which is fundamentally a drop-off problem, see finding where customers drop off in a funnel.

For building the tracking views that make a metric tree visible to teams rather than living in a spreadsheet, the walkthrough in Power BI for beginners covers the measures and layouts involved.

For presenting a metric layout where the primary KPI stays dominant and the secondaries stay subordinate, the data visualisation guide covers the encoding choices.

And for the organisational conditions that determine whether a metric framework changes behaviour or becomes wall decoration, why analytics projects fail is worth reading alongside this.


Quiz

TEST WHAT YOU LEARNED

Question 1 of 15

Q1: Which characteristic most clearly disqualifies monthly recurring revenue as a north star metric for a product organisation?

FAQ

FREQUENTLY ASKED QUESTIONS

One per business, but a large company with genuinely separate business lines can have one per line, provided each is decomposed from a shared company-level outcome. What does not work is two competing north stars for the same product, because the entire function of the metric is to resolve trade-offs, and two of them just relocate the argument. If you find yourself needing a second, you usually need a guardrail instead.
For a business where revenue is close to the value delivered and teams can move it within a quarter, such as a transactional marketplace, it can work. For most subscription and product businesses it lags too far behind the work and moves for reasons no single team controls. The practical test is whether a product team could look at it on Monday and know what to do differently, and for revenue the answer is usually no.
The north star is a metric that stays constant across quarters; an OKR is a time-boxed goal that usually targets a movement in that metric or one of its inputs. The north star answers 'what are we optimising', and the objective answers 'how far, by when'. Replacing the north star every quarter because the OKRs changed defeats the purpose, since its value comes from being stable enough for teams to build intuition around it.
Watch the guardrails rather than the metric. Gaming almost always shows up as the north star improving while a quality or cost measure deteriorates, so choose guardrails specifically for the routes you can already imagine. Also check whether an improvement is concentrated in a single team or surface, since a genuine product improvement usually shows up broadly and a gamed one usually does not.
Either the relationship between them has broken down, the lag is longer than you assumed, or the metric is being moved in a way that does not create value. Check the lag first by comparing historical movements in the metric against revenue two and three quarters later. If the relationship held historically and has now decayed, that is grounds to revisit the metric, and it is exactly the annual test worth building into the review cycle.
Few enough that every member can recite them, which in practice means one primary and a small number of secondaries. The failure mode of a long list is not that the extra numbers are wrong, it is that attention gets distributed evenly across things of unequal importance. If a metric would not change a decision this quarter, it is context rather than a KPI, and context belongs in a reference view rather than on the team's board.
Some have a legitimate diagnostic use even though they are poor goals. Total registered users is useless as a target because it only ever rises, but it is a reasonable denominator for rates and a reasonable input to capacity planning. The problem is not the number, it is setting it as a goal, since a metric that cannot fall provides no feedback about whether anything is working.
Counts are easier to communicate and align teams around; rates are better at detecting quality problems. A common resolution is a count as the north star with a rate as a guardrail, so growth is the goal and quality is the constraint. What you should avoid is a compound score that combines several metrics with weights, because when it moves nobody can say why.
Start from the historical distribution rather than from ambition. Compute what the metric has done over the past several quarters, establish its normal rate of change, then set a target that represents a real stretch against that baseline rather than a number chosen for how it sounds. Targets set without reference to historical variation get either missed by a mile or hit in week two, and both outcomes teach the organisation to ignore targets.
Start with the definition, not the data. Check the numerator, the denominator, the population exclusions, and above all the window, since 'week 1' meaning days since signup versus the calendar week is a common cause. Most disputes of this kind are definitional, and fixing the definition centrally in a modelling layer prevents them from recurring every time someone builds a new dashboard.
It needs the north star and the tree; it probably does not need per-team KPI hierarchies, because at small scale everyone works on everything. The value of writing the tree down early is that it makes the growth constraint explicit, and the arithmetic of inflow versus retention is often surprising in a way that changes what a small team builds next. The formal KPI layer becomes necessary when teams specialise.
They are outcome and efficiency KPIs rather than north star candidates, since they are lagging and depend on inputs from several functions. LTV to CAC in particular needs a modelled lifetime value, a decision about horizon and discounting, and stable cohort data, all of which make it unsuitable as a weekly steering metric. It belongs on the executive view alongside revenue rather than in the operational decomposition.
A movement should be judged against the metric's own historical variation before it is treated as a meaningful change, and weekly metrics on small bases move a great deal by chance. Doing it properly means addressing seasonality, autocorrelation between consecutive weeks, and the multiple comparisons created by watching many KPIs at once. The practical interim step is to plot several quarters of history behind every KPI so the reader can see the noise band.
Someone accountable for the outcome across functions, typically a general manager or product leader, with an analyst owning the definition. Separating these two roles matters: the business owner decides what the company optimises, and the definition owner decides what the number means and defends it against convenient reinterpretation. Where the same person holds both, the definition tends to drift toward whatever makes the current quarter look better.
Writing the metric tree with an owning team next to every input. It converts an abstract company goal into a set of specific assignments, it exposes immediately whether any input has no owner, and the arithmetic of the tree answers the roadmap prioritisation question with a calculation rather than a debate.