Evaluating Data Sources with ROCCC
Good for which question?
A retailer wants the return rate for orders delivered in July across its online shop and physical stores. A glossy dashboard says 4%, while an internal export suggests 10%. Choosing the more impressive publisher or the more recent download would not settle the difference. The two files may describe different orders, periods, or return rules.
ROCCC stands for Reliable, Original, Comprehensive, Current, and Cited. Use these five prompts to investigate whether a source supports a particular analysis. They are not a certification or five equally weighted points. A file can have an excellent citation and still omit the population you need. Data quality depends on intended use: evidence suitable for exploring a trend may be insufficient for paying an individual refund.
Define the unit, population, and observation window
Before comparing sources, write the question precisely. In this fictional exercise, count each delivered order once, include both sales channels, and exclude orders cancelled before delivery. A returned order is an eligible order with at least one item accepted for return within 30 days of delivery. The rate is returned eligible orders divided by all eligible delivered orders. It is an order rate, not an item rate or a share of sales value.
Use the agreed business time zone. Wait until the return window has closed for every July delivery and the agreed processing delay has elapsed. The example assumes a frozen September 5 extract contains all such accepted returns after a completed reconciliation. A later download date alone would not establish that. Every count below is invented for the exercise.
Reliable: inspect how the records were made
Ask how orders and returns were captured, which checks ran, and what failures remain. Compare eligible identifiers with an independent control where available; test duplicates, missing records, invalid dates, and return-to-order links. A valid date or unique ID does not prove that the event happened as recorded. Check a justified sample against transaction evidence and record what that sample cannot establish.
Selection matters too. A voluntary satisfaction survey may overrepresent customers who are unusually pleased or upset. A dataset containing every survey response is complete for respondents, not necessarily representative of every buyer. “No obvious bias” is not evidence that selection or measurement bias is absent.
Distinguish a flawed presentation from flawed source values. A truncated baseline in a bar chart can exaggerate differences because bar length encodes magnitude; inspect the table and labels. That does not by itself prove that the recorded measurements are false. Assess the collection, transformation, and presentation separately.
Original: trace the route back to evidence
Find who collected the data and how the file reached you. A direct export can still contain collection errors. A secondary publisher can be useful when it documents its source, selection rules, transformations, and corrections. Originality supports traceability; it does not automatically rank every primary source above every secondary one.
For a dashboard percentage, ask for the underlying counts, definitions, dataset version, and transformations. If you cannot inspect personal records, an authorized aggregate with documented methods may be the appropriate evidence. Do not acquire unnecessary identifying data merely to claim that your source is original.
Comprehensive: check coverage for the defined question
Comprehensive means having the critical coverage and variables for the question, not collecting every available field. For the return-rate question, both channels and their eligible order denominators matter. A file with no blank cells can still omit all store orders. Confirm the expected population, observation window, required fields, and exclusions.
Suppose the reconciled counts are 600 online orders with 24 returned orders, and 400 store orders with 76 returned orders. Online is 24 / 600 = 4%; stores are 76 / 400 = 19%. Across both channels, (24 + 76) / (600 + 400) = 100 / 1,000 = 10%. Averaging 4% and 19% gives 11.5%, which incorrectly gives equal weight to channels with different order counts. The appropriate weights here are 600/1,000 and 400/1,000.
Current: use the right time, not just the newest file
Separate the period described by the data, when it was collected, when late events became available, and the publication or extraction date. A newly exported file can contain old observations. Conversely, historical data is necessary for a historical question; age alone does not make it bad.
The July return cohort needs its full return window. An August 1 extract may miss later returns for orders delivered at the end of July. For a real-time stock decision, even yesterday’s accurate inventory could be too old. Choose timeliness requirements from the decision, and preserve the version used when later corrections restate history.
Cited: make the source identifiable and inspectable
Record the producer, dataset title, stable identifier or location, version, observation period, extraction date, and the method or query used. Include definitions and known limitations. A citation tells another analyst which evidence to inspect; it does not certify the evidence or mean that many people have cited it.
Government statistics, academic publications, and commercial datasets can all be useful starting points. Check their methods and fitness for this task instead of granting a pass based on the organization’s reputation. Separately confirm permitted access and reuse: a complete citation does not grant a licence or permission to disclose personal information.
Compare two sources and record a bounded decision
Source A is a dashboard for online orders only. Its method, version, and underlying 24/600 counts are documented, and the complete observation window has been verified. Source B is the reconciled combined extract with the 100/1,000 counts. Its store return codes required a documented mapping to the common accepted-return definition; the reviewer checked that mapping against the agreed rules.
A can support the online-only question, but not the requested all-channel rate. B supports this exercise’s combined descriptive rate under the stated checks. Neither source alone explains why customers returned products or proves that a new policy caused returns. Missing evidence should be recorded as unverified, not converted to a passing check or an arbitrary zero score.
Write a decision record: “Use combined extract version v2 for the July all-channel order return rate; 10%; both channels reconciled; observation window closed; store-code mapping reviewed; no causal claim. Operations owns the source checks. Reopen the result if accepted returns are corrected or the eligibility definition changes.” The decision is tied to a version, purpose, evidence, limitation, and owner. If a critical population or definition cannot be verified, withhold the combined claim or narrow the question explicitly.
Try the decision yourself
1. An analyst calls Source A bad because it gives 4% rather than 10%. What should change in that judgment?
Solution
The figures answer different population questions. A is supported for online orders under its documented checks. It does not cover all channels. Compare definitions and denominators before calling a disagreement an error.
2. A revised combined file has 1,000 rows, no empty cells, and a citation. Is it ready to replace version v2?
Solution
Not from those facts. Confirm which orders the rows represent, uniqueness and expected coverage, return eligibility and mappings, observation window, and the reason for the revision. The same row count can hide missing orders replaced by unrelated ones. Retain the old version and record verified differences before approving the new result.
For a broader approach to fitness for purpose and ongoing quality management, see The Government Data Quality Framework. Its quality dimensions complement these source-evaluation questions; they are not the ROCCC acronym itself.
Discover more from Insightful Data Lab
Subscribe to get the latest posts sent to your email.
