Writing a Use-Case Brief with a Measurable Outcome
A brief turns an idea into an assessable proposal
“Help purchasing staff process invoices faster” is a useful direction, but it leaves several decisions open. Which invoices? Faster at finding records, preparing an approval packet, or completing payment? Compared with what, and with which errors still acceptable? A use-case brief is a concise agreement about the user, problem, intended change, evidence of success, and boundaries. It lets people decide what to investigate or test without pretending that the solution is already proven.
Consider a candidate proposal: help authorized purchasing clerks retrieve matching order and goods-receipt evidence while preserving the existing approval process. We will make that proposal measurable. All organizations, figures, thresholds, and review periods below are fictional teaching choices. They illustrate how to write a brief, not measured business results or a recommended financial-control policy.
Write the user’s outcome before the deliverable
A searchable screen is an output: something the team delivers. Less effort to prepare a correct evidence packet is an outcome: a change in the user’s work. Reduced overall operating cost would be a further business effect requiring additional evidence. Keep these separate. Shipping the screen does not demonstrate saved effort, and saved minutes do not automatically become lower payroll expense.
Name the actor, trigger, and terminal result: an authorized clerk receives an eligible invoice and finishes either a checked evidence packet or a documented escalation. The approver and receiving team are affected users even if they never open the new screen. State what will not change: this proposal does not authorize payment or alter approval rules. That boundary prevents a retrieval aid from quietly becoming an automated payment project.
Define eligibility before seeing results. For example, the trial may include invoices from selected suppliers with usable order identifiers, in the supported document formats. Missing identifiers follow the current process and their volume remains visible. A case that meets entry conditions but later defeats the tool remains an eligible case; it must not be reclassified as out of scope to improve the reported result.
Separate the problem from the explanation you want to test
A business task is the question the analysis must answer to support a decision. “Build an AI search screen” names a deliverable; “Which evidence-preparation steps consume avoidable effort, and would shared references reduce it without shifting work?” names an investigation. The problem domain includes the records, people, rules, and handoffs that affect this work. The trial scope selects the part being tested; a receiving team can remain part of the problem domain even when its software is outside the delivery scope.
Ask why a delay occurs, but record each proposed cause as a hypothesis until checked. Repeated “why” questions can reveal missing identifiers, inconsistent naming, access delays, or several contributing causes. Five questions are not a proof or a stopping rule. Compare staff accounts with permitted case records. AHRQ’s Five Whys example shows how answers can branch into multiple paths; the same questioning discipline is useful here without importing a clinical procedure.
A gap is the difference between the current and desired state; it does not identify the cause or the remedy. Sketch the workflow with staff, locate where effort accumulates, and list evidence that would change the proposal. If most delay is waiting for an unrecorded receipt, a faster search interface may leave the real constraint untouched. A manual reference list or a recording-process change belongs among the alternatives.
A baseline needs a definition and a source
A baseline describes performance of the comparison process before or alongside the proposed change. Record the workflow version, eligible population, measurement period, and source of observations. A manager’s estimate of eight minutes is a planning assumption until measured with the agreed method. The GOV.UK guide to setting performance metrics connects service purpose, meaningful measures, and comparison with existing performance.
For this example, active handling begins when the clerk starts evidence preparation and ends when a checked packet or documented escalation is ready. Sum all active sessions, including verification, corrections, tool failures, and manual fallback. Exclude inactive waiting from active time but track it separately as elapsed time. Time spent reviewing a suggestion is work, even if the model itself responds instantly.
This primary measure counts purchasing-clerk effort up to the stated terminal result, not the full cost of resolving an invoice. Track work transferred to approvers and receiving staff separately, using the same case identifiers. Also record review effort, maintenance, and elapsed time through final resolution. The brief’s workload guardrail requires no observed increase in receiving-team active minutes per eligible entry under matched follow-up; missing records block a success decision. A documented escalation ends the clerk’s preparation stage, but it does not mean the underlying case is resolved.
Use comparable cases and conditions. Easier suppliers, better-trained staff, or a quieter week can change the result independently of the tool. The brief should identify these concerns and the comparison plan; a detailed pilot design can resolve assignment and sample size later. A simple before-and-after difference is a descriptive result unless the evaluation design supports attributing the change to the intervention.
Define a primary measure and the conditions around it
The primary measure is the main quantity used to judge the intended improvement. For the clerk, use mean active handling minutes per completed eligible case, under the completion and follow-up rules below. The numerator is total active time for those cases; the denominator is their count. Report both, plus eligible entries, unfinished cases, escalations, and manual fallbacks. Counting only successful tool responses would answer a different question.
A cohort here means cases grouped by the agreed entry conditions and enrollment period. Suppose a completed comparison cohort has 100 eligible cases and 800 active minutes: 800 / 100 = 8 minutes per case. A proposed target of 6 minutes means a reduction of 2 minutes, or (8 − 6) / 8 = 25%. The six-minute value is a target, not a result. A tool’s one-second response time cannot be substituted for the clerk’s end-to-end active work.
Guardrails are conditions that must hold while pursuing the primary target. Faster preparation is not success if packets contain wrong orders, private records cross access boundaries, or routine cases are dumped into an escalation queue. Define how correctness is checked and by whom, and give adverse outcomes an explicit consequence. A single blended score would conceal which condition failed.
Put the agreement in a compact brief
Here is a filled example. It is a proposal for review, not authorization to run a trial. Its observation plan and thresholds must be agreed by the named roles before results are collected. The brief points to supporting records rather than embedding every technical detail; “one page” means a compact decision aid, not a guarantee of a particular printed layout.
| Brief field | Illustrative entry |
|---|---|
| User and problem | Authorized purchasing clerks spend active time locating order and receipt evidence across systems. |
| Proposed change | Retrieve candidate matching records with source references; the clerk verifies the packet. Compare with current lookup and a simpler shared-reference screen. |
| Scope | Selected suppliers and supported formats with usable order identifiers. Preserve original entry eligibility after tool failure. No payment execution or approval-rule changes. |
| Primary outcome | Reduce mean active preparation time, including verification, correction, and fallback, from the illustrative 8-minute baseline to at most 6 minutes. |
| Baseline evidence | Example only: 100 completed eligible cases, 800 active minutes. Actual baseline must be measured using the same event definitions and comparable case mix. |
| Observation plan | Proposed two-week enrollment; follow each case for five business days. Record unique case ID, eligibility, actor role, active-session start/end, workflow version, tool use, fallback, and terminal status; review follow-up corrections. If any eligible case remains unfinished at cutoff, report provisional results and do not declare success. |
| Quality and workload conditions | A designated reviewer checks every trial packet against source records and agreed rules. No observed wrong-order packet; escalation share (escalated cases / all eligible entries) must not exceed the measured comparison baseline. Receiving-team active minutes per eligible entry must not increase under matched follow-up. Missing workload evidence blocks success. A breach blocks a success decision. |
| Access and operational limits | Only permitted records; no write access to payments. Any observed unauthorized disclosure or action stops the trial and triggers incident review. Maintain manual fallback. |
| Data and dependencies | Permitted invoice, order, and receipt records; reliable linking; source owners; access approval; event logging that excludes unnecessary sensitive content. |
| Ownership and next decision | Purchasing service owner owns the outcome; analyst owns measurement; security owner reviews access evidence. Confirm baseline, comparison design, sample adequacy, staffing, and approval before deciding whether to test. |
The proposed two-week enrollment and five-day follow-up are scheduling choices, not proof of an adequate sample or a complete error window. If important corrections arrive later, extend or redesign follow-up and document the reason. The illustrative quality threshold is deliberately strict; observing no errors in a finite trial still does not prove a zero production error rate. Agreement on a threshold and evidence that it was met are separate from the uncertainty that remains.
Make failure and missing data visible
Suppose a candidate appears to average five minutes, but the report includes only 80 successful cases out of 100 eligible entries. The other 20 required manual work or remain unfinished. The report cannot establish the brief’s target for the eligible cohort. Recover their handling and status records, include completed fallbacks, and report unresolved cases explicitly. The example’s completion rule prevents a quick average from hiding slow cases left in the queue.
Missing timing records are not zero-minute cases. Decide how to flag them and whether the remaining evidence can support a decision; do not silently impute a favorable value. Report adoption separately too: eligible cases offered the tool, cases where staff used it, and reasons for non-use. A system may be accurate when used but fail to improve the service because it is difficult to access or adds duplicate work.
Record open questions with an owner and a next action. “Can the receipt records be linked reliably?” might require a permitted sample review by the data owner. “Is six minutes a useful target?” requires the service owner to weigh verification and operational effort. A brief can be ready for a discovery decision while remaining unready for a live trial. Label that status instead of filling unknowns with invented certainty.
Check that a familiar metric answers the intended question
The same brief structure works outside invoice handling, but the measure changes. Consider an illustrative inventory cohort of 50 orders requesting 100 units of one product variant. At the agreed first-shipment cutoff, 97 units ship and three orders are each missing one unit. Unit fill rate is 97 / 100 = 97%; complete-order fill rate is 47 / 50 = 94%. These are different results from the same records. State the product identifier (SKU), location, entry period, and fulfillment window. Oracle’s retail reporting guide provides an example of defining fill performance at the first shipment.
A stockout means inventory cannot meet demand at the specified place and time. An unfilled unit may be backordered for later delivery, canceled, or replaced by another purchase; it is not automatically a permanently lost sale. Track these statuses separately. Estimating lost sales needs assumptions about unobserved demand and substitution. A promotion may shift purchases from another product in the same portfolio, so apparent growth of one product is not necessarily additional business.
An inventory brief should also constrain excess stock and the share of slow-moving product variants under stated thresholds. Specify whether “overstock” compares stock with a target level or a demand horizon; those denominators differ. Safety stock is a buffer against uncertainty, and a reorder point is a replenishment trigger, not a guarantee against shortages. Supplier lead time, capacity, minimum order quantities, and delivery reliability can limit an intervention. Supplier coordination and timely inventory records are possible enablers; neither guarantees the desired service outcome.
For lead-generation software, separate contacts captured from qualified prospects and later customer outcomes; define duplicates, qualification, permission, and follow-up. For software supplied as a hosted service (SaaS), a subscription or installed integration is a delivery arrangement, not proof of user benefit. Adoption, supported-task completion, retention, service effort, and cost require their own populations and time windows. Select measures for the decision rather than copying a vendor’s metric names.
Review the brief as a decision document
Ask someone outside the authoring team to explain who benefits, what changes, how success is counted, what must not deteriorate, and who decides the next step. If they interpret “completed” or “accurate” differently, revise the definitions before collecting results. Keep a version and approval record, and make later changes to scope or thresholds visible. Quietly relaxing a target after seeing results changes the question the trial answered.
1. The baseline is 800 active minutes for 100 completed eligible cases. A candidate uses 650 minutes for 100 comparable completed cases with no other reported issue. Does it meet the illustrative six-minute target?
Solution
No. Its mean is 650 / 100 = 6.5 minutes. The descriptive reduction from 8 minutes is 1.5 / 8 = 18.75%, below the target reduction of 25%. Other conditions must still be checked; an improvement alone is not attainment of the stated target, and the difference alone does not establish causality.
2. A brief says “95% accuracy” but does not define the unit or how errors are reviewed. What must be clarified?
Solution
Define whether the unit is a field, a matched record, or an entire packet; identify the reference and reviewer; specify numerator, denominator, eligible scope, missing cases, and observation window. State how consequential errors affect the decision. A high field-level score can coexist with many packets containing at least one wrong field.
Discover more from Insightful Data Lab
Subscribe to get the latest posts sent to your email.
