
Reviews
Part of How to compare eight productivity software tools before reaching a verdict
A desk review protocol for productivity software, from evidence labels to scoring anchors
How to run a desk review protocol for productivity software, labelling evidence, separating gates from scores and recording unresolved questions.
This business productivity software review methodology is designed to stop a documentation search being mistaken for a product test. It creates a traceable route from an organisation's work to the evidence behind a conclusion.
Reviewer: Codex editorial research desk. Research date: 5 September 2026. This was desk research only. No account, free trial, paid plan, integration, support request or accessibility journey was tested, and no customer interview took place. No supplier contributed, provided access or paid for inclusion. There are no affiliate links or known commercial conflicts.
What to take away
- A desk review of productivity software cannot replace a controlled trial with real users and data.
- Label every evidence item as provider claim, authoritative guidance, observed test or customer evidence.
- Run one protocol across the shortlist with neutral samples and consistent scoring anchors.
- Score only what the record supports and keep mandatory gates outside the arithmetic.
- Publish raw observations with the total and perform a sensitivity check on non-critical weights.
Define the case before choosing candidates
Write a compact scenario with the users, workflow, data sensitivity, volume and required outcome. Convert it into testable statements. "Easy to use" is too vague; "a new coordinator can accept, assign and find an overdue request using the approved instructions" can be observed.
From vague need to testable statement
- Define users, workflow, data sensitivity, volume
- State required outcome
- Convert to testable statements
- Separate mandatory gates from preferences
- Document exclusions and triggering evidence
Separate mandatory gates from weighted preferences. Data handling, an essential permission or a usable exit export may be gates. Colour choices or an extra view may be scored. Document exclusions, including the evidence that triggered them, before calculating a total.
Label every evidence item
Use four labels: provider claim, authoritative risk guidance, observed test and customer evidence. Provider documentation is necessary for availability. For example, Microsoft's current Planner service description records task fields, board use, several views and plan-dependent feature availability. That page supports a shortlist note; it does not prove Planner suits a particular team.
Four evidence labels and what they prove
Label
- Provider claim
- Availability only
- Risk guidance
- Review questions
- Observed test
- Actual behaviour
- Customer evidence
- Real-world use
What it supports
- Provider claim
- Risk guidance
- Observed test
- Customer evidence
Authoritative guidance supplies review questions rather than a product verdict. The NCSC's SaaS security guidance covers user lifecycle, authentication, privileged administration, data protection, sharing, recovery and monitoring. A reviewer should turn relevant controls into requests and trial steps, then obtain specialist assessment.
Run one protocol across the shortlist
Prepare neutral sample records. Use the same roles, browser conditions, task wording and scoring anchors for every candidate. Test a normal request, an incomplete request, reassignment, approval, restricted access, reporting, export and user removal. Add an unusual case that could expose a costly weakness, such as an integration failure or a departing administrator.
For each step capture the edition, configuration, expected result, actual result, time only where measured consistently, screenshot or log reference, limitation and tester. Government Digital Service advice on quality-assurance testing applies to public services, but its emphasis on repeated, multidisciplinary testing under normal and less common conditions is a sound review discipline.
Score only what the record supports
A practical four-point scale is enough: 0 not evidenced, 1 materially constrained, 2 meets the scenario and 3 exceeds it in a relevant way. State what each anchor means for every criterion. Keep gates outside the arithmetic so an unresolved security or exit requirement cannot disappear behind a strong presentation score.
Publish raw observations with the total. Perform a sensitivity check by changing non-critical weights within a declared range. If the preferred candidate changes easily, describe the outcome as uncertain and commission another test instead of forcing a winner.
Handle reviews and conflicts openly
The CMA's current review and endorsement guidance says review material should be genuine and accurate and incentives made clear. Legal review is needed for the final publication. Operationally, retain the research date, reviewer identity, evidence type, corrections, free access, gifts, sponsorship, referral arrangements and relevant supplier relationships.
Finish with a decision record listing unresolved questions, risk owners and expiry dates for time-sensitive evidence. A desk review can narrow a field. Only a controlled trial, reviewed contracts and competent risk checks can support the later buying decision.
Before you act
- Define a compact scenario with users, workflow, data sensitivity and volume.
- Separate mandatory gates from weighted preferences before scoring.
- Label each evidence item by its type and source.
- Run the same test steps for every candidate using neutral samples.
- Keep gates outside the arithmetic and publish raw observations.
- Record unresolved questions, risk owners and evidence expiry dates.
Common questions
What does the article say about the limits of desk research?
The article states this was desk research only. No account, free trial, paid plan, integration, support request or accessibility journey was tested. No customer interview took place. A desk review can narrow a field, but only a controlled trial, reviewed contracts and competent risk checks can support the later buying decision.
How should a reviewer handle mandatory gates versus weighted preferences?
Separate mandatory gates from weighted preferences. Data handling, an essential permission or a usable exit export may be gates. Colour choices or an extra view may be scored. Document exclusions, including the evidence that triggered them, before calculating a total. Keep gates outside the arithmetic so an unresolved security or exit requirement cannot disappear behind a strong presentation score.
What should be published alongside a total score?
Publish raw observations with the total. Perform a sensitivity check by changing non-critical weights within a declared range. If the preferred candidate changes easily, describe the outcome as uncertain and commission another test instead of forcing a winner. Also retain the research date, reviewer identity, evidence type, corrections, free access, gifts, sponsorship and referral arrangements.



