Comparison of four methods for evaluating productivity software impact
Image: Work Stack Lab

Measurement

Part of Measure productivity software with a metric contract and a baseline you can revisit

Four ways to tell whether productivity software actually changed anything

Compare four ways to assess whether productivity software contributed to change, using consistent criteria for evidence strength, data, assumptions and cost.

Business productivity software attribution methods differ in the question they can answer. A simple before-and-after view records change. A credible causal claim needs a defensible account of what would have happened without the software.

This comparison is for an England-based organisation planning an evaluation. It uses four consistent criteria: evidence question, data needed, main assumption and practical burden. No method guarantees a valid result. A qualified analyst should design the study, with privacy, employment and financial review where relevant.

What to take away

  • Before-and-after views show change but cannot prove the software caused it.
  • Causal claims need a defensible counterfactual and comparable data for intervention and comparison groups.
  • Interrupted time series can detect level or trend changes but remains vulnerable to concurrent events.
  • Phased rollout and randomised designs require careful planning to avoid selection and spill-over problems.
  • Match the strength of your claim to the decision at stake and the design you can support.

Four methods compared

Question it can address

Descriptive before and after
What changed after launch?
Interrupted time series
Did the level or trend change at launch?
Phased rollout comparison
Did treated groups change differently from later groups?
Randomised rollout
What average difference did allocation produce?

Minimum evidence

Descriptive before and after
Stable metric definition and comparable periods
Interrupted time series
Enough observations before and after, with intervention timing
Phased rollout comparison
Comparable groups, shared measures and controlled rollout dates
Randomised rollout
Eligible units, random assignment, pre-set outcomes and sufficient sample

Main limitation

Descriptive before and after
Other changes can explain the movement
Interrupted time series
Seasonality and concurrent events can distort the estimate
Phased rollout comparison
Selection and spill-over can break the comparison
Randomised rollout
May be impractical, disruptive or ethically unsuitable

HM Treasury's Magenta Book explains that experimental and quasi-experimental approaches use a counterfactual and require comparable data for intervention and comparison groups. Its audience is central government. The design logic transfers, but the scale and assurance should match the business decision.

Table comparing four evaluation methods by question, evidence and limitation (Four ways to tell whether productivity software actually changed anything)
The four methods differ in what they can claim and what evidence they require. Image: Work Stack Lab

Descriptive before and after

Freeze the workflow, measure and periods before reading the result. Show demand, work mix and known operational changes beside it. Use language such as "cycle time was lower in the later period", not "the platform saved time".

Descriptive before and after

Before launch

Cycle time
Higher
Demand
Recorded
Work mix
Documented
Operational changes
Noted

After launch

Cycle time
Lower
Demand
Recorded
Work mix
Documented
Operational changes
Noted

This is suitable for service monitoring and hypothesis generation. It becomes especially weak when the launch coincides with staffing, policy, seasonal or customer changes. A short baseline can mistake ordinary variation for improvement.

Interrupted time series

Plot a consistent outcome across multiple time points on both sides of a clear intervention date. Model a change in level, trend or both, chosen in advance. Document seasonality, autocorrelation and events that could alter the series.

Interrupted time series

  1. Before
    Multiple observations
  2. Intervention
    Launch date
  3. After
    Multiple observations
  4. Model
    Level, trend or both

The Magenta Book's analytical-method annex warns that series length, slow trends and external influences matter. The method can be stronger than a two-period snapshot, but it is not immune to another event occurring at launch.

Phased rollout comparison

Use the same outcome, unit, extraction rule and dates for groups that adopt now and later. Compare pre-launch trends before treating the later group as a counterfactual. Record why each group entered its phase.

This design can fit a planned operational rollout, but manager choice may put the most capable or troubled team first. Work shared across groups can also contaminate the contrast. Adjusting for observed differences cannot guarantee that hidden differences disappeared.

Randomised rollout

Random assignment can balance known and unknown influences on average when carried out properly. Pre-register the eligible population, allocation, primary outcome, analysis and stopping rule. Keep the service available where withholding it would cause unacceptable harm.

Randomisation does not repair a vague metric, low adherence or underpowered sample. It estimates an effect for the studied population and conditions, not every team or future product version.

Choose claim strength before cost

If the decision only needs operational learning, careful descriptive monitoring may be proportionate. A large renewal, workforce change or public claim of savings deserves stronger design and independent review. The 2026 Green Book separates monitoring from evaluation and asks public bodies to plan evaluation early; it is not a private-company mandate.

Write the intended sentence before selecting the method. If the available design cannot support that sentence, weaken the claim or improve the study. Never upgrade association to causation in the final report.

Before you act

  • Write the intended sentence before selecting a method.
  • Freeze the workflow, measure and periods before reading results.
  • Document seasonality, autocorrelation and concurrent events.
  • Pre-register the eligible population, allocation and primary outcome.
  • Weaken the claim if the design cannot support it.
  • Never upgrade association to causation in the final report.

Common questions

What is the main limitation of a descriptive before-and-after comparison?

Other changes can explain the movement. The method records what changed after launch but cannot rule out staffing, policy, seasonal or customer changes that coincide with the launch. It is suitable for service monitoring and hypothesis generation, not for making causal claims about the software's effect.

How does an interrupted time series improve on a simple before-and-after view?

It uses multiple observations before and after a clear intervention date to model changes in level or trend. This can be stronger than a two-period snapshot, but it is not immune to another event occurring at launch. Seasonality and concurrent events can still distort the estimate.

When might a randomised rollout be unsuitable?

It may be impractical, disruptive or ethically unsuitable. Randomisation does not repair a vague metric, low adherence or underpowered sample. It estimates an effect for the studied population and conditions, not every team or future product version. Withholding the service where it would cause unacceptable harm is not acceptable.

More in Measurement