The objective usually reads fine. Ambitious, clear, easy to rally around. It's the key results underneath it that quietly turn the whole thing into a task list, while the dashboard goes green and nothing about the business actually moves.
That's fixable, and it's simpler than it looks. The key results just need to be written differently, nobody needs to try harder. Here's how to spot the difference, and how to fix what's broken.
Free download: Rewriting key results is easier with a template in front of you. Grab the OKR Quick Start Guide and the OKR vs KPI Comparison from our free resource library.
The difference in one line
A strong key result measures an outcome you can drive, not an activity you can guarantee. A weak key result measures effort instead of impact.
That is the whole distinction, and almost every other rule follows from it. It also matches how the discipline is defined at the source. The OKR Institute defines a Key Result as a quantitative metric that measures whether the Objective has been achieved, and is explicit on the point that trips up most teams: "Key Results measure outcomes, not activities or outputs."
- Output: "Ship the new onboarding flow." This is something a team can promise on day one, which is exactly why it tells you nothing about whether onboarding got better.
- Outcome: "Increase the share of new users who reach first value in week one from 20% to 45%." The team can influence this metric, but they cannot command it, which is precisely why it is worth measuring.
Weak key results are safe. That is the problem with them.
Four tests for a key result
Run any key result through these four questions. A strong one passes all four. A weak one fails at least one, usually the first. Each of these gets a full treatment elsewhere on the blog, so the tests here are the fast version, not the deep dive.
- Outcome or output? An output is work you produced (a feature shipped, a migration completed). An outcome is what changed because of it (a metric moved, a behavior shifted). If you hit 100% on the key result without moving the business forward, you are measuring the wrong thing. Full breakdown: why so many goals feel like busy work.
- Could you hit 100 percent by working the number instead of the goal? Any measure with stakes attached becomes a target to satisfy, not a signal to learn from. That is how a dashboard ends up green on the outside and red on the inside, and it is one of the most common mistakes that quietly kill OKR programs.
- Is it a KPI wearing an OKR costume? A key result describes a change from one state to a new one, then it is done. "Keep uptime above 99.9 percent" never ends, which makes it a KPI, not an OKR. Full distinction: OKRs vs KPIs.
- Are you aiming for the whole objective, not a safe fraction of it? Set the target at the full, ambitious outcome you actually want to achieve. The danger arises when teams confuse the target with the evaluation threshold. You should always aim for 100% of your stretch goal; if you push hard and land around 0.7, that is considered a strong result for a true stretch goal. But aiming for 0.7 from day one isn't setting a stretch goal, it's just sandbagging with a lower bar. We make the full case, including why the 0.7 is where you land and never what you aim for, in what score should you aim for on your OKRs.
Weak to strong, three examples
The fastest way to internalize this is to watch a weak key result get rewritten.
- Weak: Launch the customer referral program.
- Strong: Grow referred signups from 3% to 15% of new accounts by the end of Q3.
- Why: The launch is an output you control. The referral rate is the outcome you actually wanted, forcing the team to care whether the program works, not just whether it shipped.
- Weak: Hold weekly cross-team syncs.
- Strong: Cut average time from feature request to first deployment from 40 days to 20.
- Why: Meetings are just activity. Delivery speed is the change the meetings were supposed to produce. If the syncs do not move the metric, you can stop holding them.
- Weak: Migrate 100% of services to the new platform.
- Strong: Reduce median deployment failures from 12 per month to fewer than 4.
- Why: A full migration can succeed while nothing gets more reliable. Naming the reliability outcome keeps the migration pointed at the reason you started it.

In every pair, the strong version is less certain and more honest. That discomfort is the feature, not a flaw.
A well-written key result still has to be seen
Writing a key result well is necessary, but it is not sufficient. A perfectly designed key result that lives in a document nobody opens will change nothing. Making the measurement visible is what turns it into behavior.
We watched this directly on one engagement. A team was using AI tools to modernize a set of slow, manual, error-prone processes. They tracked three metrics simultaneously: process duration, error rates, and cost.
The moment those numbers were made visible to everyone and reviewed regularly, the team's behavior shifted. They stopped asking only "how fast can we make this" and started weighing speed against quality and cost to choose the right problem first. One process that used to take eight or nine hours dropped down to fifteen minutes.
The key results were well-written, but the real impact came because they were visible and used by the people doing the work. A good key result hidden in a spreadsheet is just paperwork with better grammar.
That is the same discipline most organizations are missing right now with AI spend. Most finance and technology leaders can tell you the total. Very few can tell you which teams, tools, or use cases are actually driving that cost, or whether any of it is paying off, because the numbers exist somewhere without ever becoming visible enough to act on. A well-written key result and a well-instrumented AI budget fail for the identical reason: the measurement was never put where anyone would see it.
Why weak key results are a design problem, not a people problem
When teams are judged on completed tasks, they deliver completed tasks. When they are evaluated on safe, predictable numbers, they play it safe. In both cases, people are responding rationally to the incentives in front of them. The issue is rarely effort or attitude, it is how the goals were designed..
This has less to do with industry than people assume. What really predicts it is a team's history. Teams that have spent years as "order takers", handed to-do lists and measured solely on completion, default to output metrics. A command-and-control environment that hands out checkboxes and calls it management will keep getting output-based goals back, because output is the only language it has ever reinforced.
Fixing wording one goal at a time only gets you so far. The key is learning to recognize the pattern so the next set of key results comes out right the first time.
When a leader is attached to the wrong key result
Sometimes the person holding onto an output-based key result is the one who set it, and they are certain it is the right goal. Telling them directly that the output is not the outcome they actually want rarely lands well, and it can turn into a standoff nobody wins.
A better opening is a question:
- “Why is it important to do this activity? What do we hope to accomplish?”
- "Once this initiative is finished, what does success look like? What do you hope actually changes?"
- "What happens if some part of this doesn't go the way we expect after release?"
Let them describe the target state in their own words. Once they lay out the intended benefits, most leaders will talk their own way to an outcome-based goal when given room to think out loud.
A question worth sitting with
Take one key result your team is carrying this quarter and ask: If we hit it completely, are we certain the business is better, or only certain the work got done?
If it is the second one, you have found a design problem. And design problems can be fixed.
Two quick answers
What is the difference between a strong and a weak key result? A strong key result measures an outcome you can see but do not fully control, a metric that could move or fail to move regardless of how hard the team worked. A weak key result measures an activity you can guarantee, like shipping a feature or holding a meeting, which tells you the work happened but not whether anything changed.
What is a good OKR score? Most practitioners, including the OKR Institute, put a good score at 0.6 to 0.7. Consistently landing at 1.0 usually means the goal was not ambitious enough to begin with.
Learn to write key results that stretch
Most OKR problems are not motivation problems. They are design problems, and design problems can be fixed once you can see them. That is exactly what we teach.
Our next OKR Practitioner course runs online over two half-days, starting September 9, 2026, live from 9:00 AM to 1:00 PM. It is accredited by the OKR Institute, the certification exam fee is included, and it counts for 7 PDUs or 7 SEUs toward your renewals.
You will leave able to look at any key result and tell, in seconds, whether it will drive change or just generate paperwork, and how to rewrite the ones that will not.

