
Measurement
Part of Measure productivity software with a metric contract and a baseline you can revisit
A productivity benchmark is a reference class, not a target
Research productivity software benchmarks without false precision by testing definitions, populations, periods, geography, methods, uncertainty and sponsorship.
Business productivity software benchmark research is a comparability exercise. A published number is useful only when its measure, population, period and method fit the decision. There is no defensible universal target for task completion, adoption or time saved.
This guide is for an organisation in England seeking context for its own results. It does not provide a market benchmark or financial recommendation. A statistician or analyst should approve a material comparison.
What to take away
- A published productivity number is only useful when its measure, population, period and method fit the decision.
- Freeze your internal metric specification before searching for external figures to avoid reshaping your definition.
- Use a comparability screen to reject direct comparisons when a material field is unknown.
- A reference class provides context, not a target, and faster completion can conceal easier work.
- Report what the evidence permits, labelling close matches as comparable only after an analyst checks.
Freeze the internal measure first
Write a metric specification before searching. Include the formula, unit, eligible work, exclusions, start and stop events, period, time zone, software version and treatment of missing or reopened items. Otherwise, a tempting external figure can quietly reshape the internal definition.
Metric specification fields
- Formula and unit
- Eligible work and exclusions
- Start and stop events
- Period and time zone
- Software version
- Missing or reopened items
Create a baseline from comparable internal history where the records allow it. Preserve demand, work mix, staffing and process changes. Internal data are close to the operating context, but they are not automatically accurate or stable.
Search by construct and unit
Search for the actual quantity, such as median calendar time from ready to accepted closure, rather than the broad word productivity. Look first for sources that publish a methodology, underlying table and revision date. Record the precise page, not an organisation's homepage.
ONS measures vs task benchmark
ONS productivity measures
- Unit
- Output per hour
- Population
- UK economy, industries
- Use
- Economic context
Task closure benchmark
- Unit
- Median calendar time
- Population
- Platform tasks
- Use
- Direct comparison
The ONS productivity measures use economic concepts including output per hour, output per job and output per worker for the UK, industries and some regional analysis. They can provide economic context. They cannot be relabelled as a benchmark for tasks closed in a collaboration platform.
Apply a comparability screen
For every candidate, capture:
- metric formula and unit;
- target population and sample frame;
- geography, sector and organisation size;
- collection and reference dates;
- work or product mix;
- exclusions, missing-data treatment and weighting;
- uncertainty, revisions and methodological breaks;
- publisher, funder and commercial interest.
Reject direct comparison when a material field is unknown. The ONS account of quality in statistics covers relevance, accuracy, timeliness, accessibility, coherence and comparability. Although written for statistical products, those dimensions form a sensible screening record for business evidence.
Distinguish a reference class from a target
A reference class groups sufficiently similar observations to provide context. It does not prescribe good performance. Faster completion can conceal easier work or premature closure; higher adoption can reflect mandatory use without better outcomes.
Prefer a distribution, such as median and quartiles, when the source provides enough valid observations. Keep the sample count and period attached. Never manufacture a percentile from a chart or combine supplier studies whose units differ.
HM Treasury's 2026 Green Book asks public bodies to use historical forecast errors and relevant prior evidence when considering optimism bias. That appraisal rule is not binding on private buyers. It supports a cautious practice: the organisation's own comparable outturns may be more informative than a generic optimistic claim.
Report what the evidence permits
Label a close match as comparable only after an analyst checks the screen. Call weaker material contextual and explain the mismatch beside the number. If a vendor publishes a benchmark, state the sponsor and separate its methodology from any commercial conclusion.
Government Analysis Function guidance on communicating uncertainty advises authors to explain sources, definitions, adjustments and uncertainty so readers can judge intended use. Apply that discipline to the benchmark note.
End with the decision, not a league table. The organisation might investigate an internal tail, collect a longer baseline or commission a better comparison. When no external source passes the screen, report that boundary and use well-defined internal evidence instead.
Before you act
- Write a metric specification before searching for external figures.
- Search for the actual quantity, not the broad word productivity.
- Capture the metric formula, unit, population and dates for each candidate.
- Reject direct comparison when a material field is unknown.
- Prefer a distribution with median and quartiles when available.
- Label weaker material contextual and explain the mismatch beside the number.
Common questions
Why should I freeze my internal measure before looking for external benchmarks?
Writing a metric specification first prevents a tempting external figure from quietly reshaping your internal definition. It forces you to record the formula, unit, eligible work, exclusions, start and stop events, period, time zone, software version and treatment of missing or reopened items before you search.
Can I use ONS productivity measures as a benchmark for my collaboration platform tasks?
No. The ONS productivity measures use economic concepts such as output per hour, output per job and output per worker for the UK, industries and some regional analysis. They can provide economic context but cannot be relabelled as a benchmark for tasks closed in a collaboration platform.
What should I do if no external source passes the comparability screen?
Report that boundary and use well-defined internal evidence instead. The article advises ending with the decision, not a league table. You might investigate an internal tail, collect a longer baseline or commission a better comparison rather than forcing an unsuitable external figure.



