Did That Feature Work? How to Actually Know
Last updated:
Features are rarely evaluated
Most products ship features and move on. Nobody checks whether adoption happened, whether the intended behaviour changed, or whether it was worth the cost.
The result is products that accumulate features indefinitely, each carrying maintenance and complexity, none of which can be removed because nobody knows what they do.
Define success before building
- Which behaviour should change? Named, specific, measurable.
- By how much, and among whom?
- By when should the effect be visible?
- What would tell us it failed, and what would we do then?
The fourth question is the one that gets skipped and the one that matters. Deciding in advance what failure looks like is the only reliable protection against post-hoc justification.
Measure adoption and effect separately
A feature can be widely adopted and change nothing, or barely adopted and be essential to the few who use it. Track both: what proportion of relevant users engaged, and what changed for those who did.
Segment by customer type. A feature that works for one segment and not another is useful information, not a failure.
Watch for the effects you did not intend
- Did support volume rise because the feature confused people?
- Did another feature's usage fall in a way that matters?
- Did the added complexity slow onboarding for new users?
- Did it change who signs up, in a direction you wanted?
Give it a fair window
Adoption of anything non-trivial takes time: people must discover it, try it, and change a habit. Judging at two weeks measures novelty; judging at three months measures value.
Set the review date when you set the success criteria, and put it in a calendar so it actually happens.
Be willing to remove
A feature that failed its criteria and shows no path to improvement should be removed, with proper notice to the few users it has. That is discipline, not waste — the waste was building it, and keeping it compounds the cost.
Products that never remove anything become products that are hard to use and expensive to maintain.
Frequently asked questions
What if we cannot measure the intended effect?
How long should we wait before judging?
Who should decide whether it worked?
Does this slow down shipping?
Shipping features and never checking?
The success definition takes ten minutes per feature and changes what gets built. Happy to share the template we use.
Related services
What we build for problems like this one