Abstract
The Purchase Mechanics Standard defines how Measured Choice separates product identity from retailer commerce, preserves missing and conflicting information, distinguishes source classes, and applies decision methods only when the evidence can support them. Its purpose is not to make consumer research look mathematical. Its purpose is to make unsupported precision difficult to introduce.
The standard also separates the consumer layer from the technical layer. A shopper should be able to understand the decision quickly. An auditor should be able to trace a material claim to a dated observation, source class, field definition, calculation and release.
1. Core operating rules
- Separate identity from commerce. A product configuration and a retailer offer are different records.
- Preserve provenance. Decision-critical facts retain source class, timestamp and material qualifications.
- Preserve missingness. Unknown price, lifespan, failure rate or repair cost stays unknown unless a modeled value is explicitly authorized and labeled.
- Preserve conflicts. Credible disagreements are carried into the decision rather than silently averaged away.
- Separate evidence types. Manufacturer claims, independent tests, owner signals, direct observations and modeled assumptions are not interchangeable.
- Use hard gates before preference trade-offs. A favorable score cannot rescue a product that fails a basic evidence, identity, safety or ownership-quality requirement.
- Allow zero winners. A category may finish with no overall Choice.
- Stop advanced modeling when inputs are not ready. More mathematics does not repair weak evidence.
2. Data architecture
Configuration Master
The canonical configuration record holds the relatively stable identity and specification layer. Internal configuration IDs remain stable even when retailer naming changes. GTIN, UPC, MPN and model aliases are identity evidence where available; they are not required to exist for every record.
Offers
Retailer offers are time-stamped commerce observations linked to a configuration. They may contain retailer, seller, price, regular price, promotion, shipping, availability, condition, source class and qualification status.
Raw Facts
Raw Facts preserve source-level observations before they are compressed into decision fields. They support traceability for warranty terms, cleaning instructions, hands-on tests, owner-duration signals, parts findings and similar evidence.
Decision Dataset
The Decision Dataset is deliberately narrower. It contains only fields that materially affect buyer value or recommendation confidence.
Release Manifest
The Release Manifest records population counts, observation counts, collection dates, method version, QA status, limitations and supersession history. Source: release summary It is the current reproducibility layer. Measured Choice does not currently claim a fully event-sourced or bitemporal production database.
3. Market population and freeze discipline
A category starts with written inclusion and exclusion rules covering geography, product function, form factor, configuration treatment and release timing. A release may then freeze the configuration set for analysis.
A frozen set is a versioned research frame, not a claim that no other product exists. For the nugget-ice pilot, the frozen U.S. frame contains 84 configurations. MCDD-ICE-001 release Post-freeze discoveries are recorded separately so the denominator is not silently changed after analysis starts.
4. Entity resolution and shared-platform evidence
Entity resolution decides whether multiple listings represent the same configuration, related variants, or different products. Strong evidence includes exact model numbers, GTIN/UPC/MPN, manufacturer documentation and exact retailer/manufacturer crosswalks. Supporting evidence can include dimensions, wattage, capacities, manuals and distinctive chassis details.
| Public state | Meaning |
|---|---|
| Confirmed shared platform | Strong identifiers or a direct crosswalk support the match. |
| Likely shared platform | Several specific technical matches exist, but no decisive identifier is available. |
| Possible relationship | Partial or visual similarities require more evidence. |
| No crosswalk established | The available evidence is insufficient. |
Similarity by appearance or dimensions alone is not enough to declare two brands identical.
5. Offer qualification and price statistics
A qualified current price is an observed, actionable offer—not a transaction price and not a permanent market norm. Displayed base price and promotional codes are stored separately when possible.
Snapshot statistics may be calculated over the qualified-price population. Worked release table In MCDD-ICE-001 Screening V0.1:
| Statistic | Qualified model price |
|---|---|
| Q1 | $199.00 |
| Median | $239.99 |
| Q3 | $299.99 |
These values describe the dated snapshot. They do not prove that $239.99 is the long-run “normal” transaction price.
Longitudinal pricing
MODAL_14D_PRICE, PRICE_30D_RANGE and PRICE_VOLATILITY_30D remain in development. A single sale is not labeled “normal” until sufficient repeated observations exist and collection coverage is documented.
6. Missing data, conflicts and imputation
Missing decision facts remain missing. Conflicting credible values remain conflicts until resolved. The current standard does not use carry-forward, peer-class mean or hedonic imputation to manufacture current product facts.
If modeled values are introduced later, they must remain visibly different from observed values and must carry their own method version and assumptions.
7. Evidence provenance
| Evidence class | Typical use |
|---|---|
| Official / manufacturer | Specifications, manuals, warranty and support policy. |
| Major retailer | Commerce observations and retailer-listed configuration facts. |
| Independent test | Hands-on performance and usability evidence. |
| Owner signal | Failure modes, real-use issues and duration-qualified ownership context. |
| Public safety / regulatory | Recalls, incidents and relevant certifications where appropriate. |
| Derived | Transparent arithmetic from observed inputs. |
| Modeled assumption | Scenario or estimate not directly observed. |
Evidence strength is claim-specific. A manual can be stronger than a review for cleaning instructions; an independent test can be stronger than a specification sheet for observed first-hour output.
8. Public claim labels
Measured data is directly observed or computed from a defined release. Evidence signal is real and decision-relevant but not mature enough for a population statistic. Our read is a transparent Measured Choice interpretation. In development is not allowed to influence the published verdict.
Examples currently held in development include the proposed Biofilm Vulnerability Tier and normalized CPSC Incident Ratio. Current status
9. Decision derivations and Choice logic
Diagnostic ratios
Ratios such as price per claimed daily output can illuminate one dimension of value, but they cannot stand in for reliability, cleaning burden, support, footprint or buyer fit.
Pareto screening
Pareto analysis can identify non-dominated products under explicitly chosen dimensions. Non-dominated does not mean “indisputably best.” The result depends on the dimensions and evidence quality.
Hard gates and role frontiers
Products with materially different functions may remain on different buyer-specific frontiers. Identity, current-price qualification, utility floor, evidence quality, safety handling and lifecycle evidence can all operate as gates before a product is eligible for an overall Choice designation.
Stop rule
For MCDD-ICE-001 V0.1, 0 of 9 finalists cleared the full upstream evidence plus lifecycle/cost-to-own requirements. Decision release Downstream Monte Carlo uncertainty propagation, Kneedle and DEA were therefore stopped rather than used to force a winner.
Category result: ZERO CATEGORY-LEVEL CHOICE PRODUCTS — V0.1.
10. Reliability, repairability and cost to own
A long warranty can reduce buyer downside; it is not proof of population lifespan. Manufacturer durability testing can be useful evidence about the test performed; it is not automatically a field failure rate.
Repairability should state the path actually verified: open parts, support-only parts, authorized service, no open catalog found, or unknown. Evidence-status example “No catalog found” must not become “no parts exist.”
Cost-to-own models can include acquisition, energy, consumables, repair and replacement only when the inputs are supportable. If lifespan or failure probability is unknown, the output should remain partial or scenario-based.
11. Versioning and change control
Every decision release should identify the dataset ID, dataset version, schema version, method version, frozen population, observation counts, collection dates, known limitations, QA status and superseded release. Example release
Published decisions should also state their reopen triggers—for example material price movement, new recall information, stronger duration-qualified reliability evidence, repair-path changes, warranty changes or a new product that materially changes the frontier.
12. Publication and citation design
The technical layer should publish as stable canonical HTML first, with a PDF mirror and downloadable CSV/JSON when the data rights and quality support it. Finished papers should include title, paper ID, version, date, abstract, definitions, methods, results or worked examples, limitations, source mapping, change log and a short citation block.
Structured data should match visible content. No markup format can guarantee citation by search systems or AI tools; the strategy is to become the original source for facts Measured Choice actually measured or derived.
Measured Choice. “The Purchase Mechanics Standard for Consumer Market Datasets.” WP-METHOD-001, Version 0.1, 18 Aug. 2026.Measured Choice. “MCDD-ICE-001: Countertop Nugget Ice Maker Public Data Release 0.7.” Version 0.7-PUB-QA-V0.1, 18 Aug. 2026.Open data release →
13. Source map
Method rules are normative statements defined by this paper. Worked examples and category facts must point to a versioned release or evidence class. The table below shows the current mapping used in this publication candidate.
| Paper section | Claim type | Underlying source / class | Public locator |
|---|---|---|---|
| Data architecture | Implemented method | MCDD Config Master, Offers, Raw Facts, Decision Dataset, Release Manifest | Release summary |
| Market freeze example | Measured category fact | Release Manifest | 84-config release frame |
| Qualified-price quartiles | Measured category fact | MCDD Screening V0.1 | Price summary |
| Evidence labels | Publication method | MCDD Publication QA V0.1 | Public claim map |
| Zero-Choice stop-rule example | Measured decision output | Choice Analysis V0.1 / Stop-Rule Audit | Decision release |
| Repairability / reliability rules | Method + evidence signal | Decision Dataset fields and source-class notes | Evidence status |
| Unfinished metrics | In development | MCDD Data Insights V0.1 | Held metrics |
The public source map intentionally exposes release facts and locators without publishing private notes, personal data, or copyrighted source material.
14. Current limitations
- The market frame is broad and versioned but is not claimed to be an exhaustive global census.
- Current prices are offer observations, not transaction data.
- 14/30-day longitudinal price metrics are not mature for the nugget-ice release.
- Shared-OEM resolution remains partial.
- Long-horizon reliability and population failure denominators remain weak across the finalist set.
- Cost-to-own modeling is incomplete for most finalists.
- Biofilm vulnerability scoring and normalized incident ratios are not published.
- The current system uses release snapshots, not a fully implemented bitemporal datastore.
- No measured operational error-rate thresholds are yet published for the pipeline.
Change log
V0.1 — Aug. 18, 2026. First publication candidate. Audited against the earlier research draft. Removed or downgraded unimplemented claims concerning exhaustive census coverage, continuous scraping, modal 14-day current price, automatic imputation, bitemporal storage, fixed accuracy thresholds and a rigid deterministic Choice formula.