NOWMethodology in development

Methodology and limitations

This page is written to be used against NOW. Everything the instrument cannot currently do is stated at the same weight as everything it can, because a research instrument that is easier to trust than to check is not a research instrument.

Dataset now-dataset-2 · corpora dated 2026-09-06 · baseline run 2026-09-06T18-15-baseline-01

Detector observed — independent human review pending.

One pass, one instrument

Every retailer is evaluated against the same pinned snapshot, the same frozen ontology, the same frozen detector and the same evidence rules. That is the whole design. The previous version joined three separate labelling passes, and a retailer’s figures depended on which pass it happened to fall into; there is now only one pass, so there is nothing for that confound to attach to.

241 pages across 96 retailers were scanned, producing 474 signals. The other pages of the 114-retailer universe were hard-blocked, errored or 404, and are recorded as coverage gaps rather than as absence of pressure.

The instrument, and its hashes

The split was drawn first, before the ontology existed and before any page was read: by retailer rather than by page, because a site-wide banner appears on all three pages of a store and would otherwise sit on both sides of the split. Then the ontology was written. Then the detector was developed against the development half only. Then all five artefacts were hashed. Only then did the detector meet the holdout.

ontology075e6bfb27f515f1fe47d52221d4c5d8
detector3ca19f7b1664b8b4ca290716e37b13cc
wide_netd83e88836030e984b867bea6c10a2c9c
splite91d5c5ae87471bf7140c6b0a9dbae48
sampler653b03bc04ffce3a0bcec5d4e7d965ce
acceptance_criteria16df988c3ab3fb074d33a05eccf0b866

The detector was written and revised against the 125 development pages only. Four revisions were made after inspecting development evidence: a spelled-out countdown format, a negated-mention veto, a flash-deal trigger, and page-level reconciliation of free-shipping threshold against unconditional. No holdout page was read during any of them.

What is measured, and how well

The figures below are author-adjudicated. They were made blind to detector output, against a frozen ontology, from a worksheet in seeded random order that carried no verdict. They are still not independent, and they do not satisfy the acceptance criteria. They are what an independent reviewer will be measured against.

DimensionTPFPFNTNPrecisionRecallF1Ambiguous
Urgency100297100.0%83.3%90.9%7
ScarcityExperimental — evidence base currently limited.41111080.0%80.0%80.0%0
Promotion850425100.0%95.5%97.7%2
Shipping Pressure5501441100.0%79.7%88.7%6
Payment Softening1201103100.0%92.3%96.0%0

Ambiguous ground truth is excluded from the matrix and counted separately. Neither present nor absent is true for those items, and folding them into either cell would be the adjudicator quietly resolving what the ontology says is unresolved.

Scarcity is experimental

It is the only dimension not at perfect precision, and its holdout rests on five positive examples. At that size the interval is so wide that 80/80 is barely a measurement. It is held, not killed: recorded like the others, shown like the others, and excluded from any future composite unless it independently clears every acceptance criterion on its own evidence.

Its single false positive exposed a real gap that is recorded rather than patched: s_exact_count explicitly excludes size pickers and s_low_stockdoes not, so a per-size “low stock” label is under-specified. Resolving it either way changes whether the dimension measures banner scarcity or inventory display, which makes it material.

The shipping limitation, measured and left in place

h_free_threshold requires 'free' immediately followed by 'shipping' or 'delivery'. An adjective between them defeats it: 'Free standard shipping on orders over $75', 'Free carbon off-set shipping over $100+'. Because a threshold IS present in the window, h_free_unconditional vetoes itself, so the page falls through both rules and records nothing.

Measured cost: 12 of the 14 shipping_pressure false negatives on the holdout. Recall 0.797 instead of an estimated ~0.90.

Ruling: NOT PATCHED. Fixing it against the holdout that measured it would convert a real 79.7% into a number with no meaning. First fix in v2, against a freshly drawn split.

The error runs one way. Under-reporting. The dimension misses real shipping pressure; it does not invent any. A dimension that under-reports is a dimension whose positives can still be trusted.

Ontology boundaries

order_by_deadline_placement · CONFIRMED, frozen

h_order_by_deadline remains EXCLUSIVELY under shipping_pressure.

A page whose only time claim is a dispatch cut-off fires shipping_pressure and does NOT fire urgency.

An actual offer deadline may separately trigger urgency on the same page when the evidence supports both. The two are different facts and both are recorded; this is not double counting and does not merge the mechanisms.

Why no composite is issued

No dimension has been through independent human review. The acceptance criteria were recorded before the 181-item blind package was sent, and until it comes back no dimension can be shown to clear them.

The bar was recorded before the review package was built, and the file is hashed. A dimension qualifies only on independent review: precision ≥ 90.0%, recall ≥ 75.0%, F1 ≥ 80.0%, at least 10 independently judged positive examples, and no unresolved ontology defect that materially changes the construct. 4 of 5 dimensions must clear all five. Four of the five must clear every criterion above. Fewer means the composite is an average over a construct with a hole in it.

Weighting is not decided by that file and is not decided anywhere else yet. Passing the gate makes a composite arguable; it does not make one exist.

Unknown is not absence

A page that was blocked, errored or 404 carries validation_status: unknown and a null presence. A dimension nobody has observed carries not_measured. Neither is ever written as false or as 0. The four statuses stay distinct: detector_observed, human_verified, human_rejected, unknown. Nothing in this codebase writes the second or third.

The distinction matters because the two failure modes point in opposite directions. Reading a blocked page as “no pressure” would flatter the retailer that blocks hardest; reading an unobserved dimension as zero would flatter every retailer equally and make the corpus look more complete than it is.

What is not measured

Transaction speed — not measured

The three-page model observes home, category and product. Checkout and cart are not observed at all, and the earlier cart pass reached 9 of 63 observable retailers against a 60% threshold predeclared before that run. There is no honest way to score it.

Decision friction — not measured

No observation method exists. Reserved, not measured.

The wide-net ceiling

Presence is only ever judged inside windows the wide net surfaced. 344 of 580 holdout page-dimensions carried no candidate and were recorded absent on the wide net's authority. The 14-item recall probe in the blind package is the first measurement of that ceiling. It is the one limitation that cannot be bounded from inside the instrument, which is why the independent package carries a recall probe of full page texts where nothing was flagged at all.

Known ambiguities in the source data

abandoned_partial_run

observations/ holds a second run, 2026-09-06T17-41-weekly: 9 retailers, 18 pages, 34 minutes before the pinned baseline.

Excluded. Half an hour is intra-session churn, not history. The baseline run is pinned by id rather than discovered.

What NOW does not claim

Presence is not deception. A countdown can be honest, a stock count can be accurate, a deadline can be real. NOW records what a page said and where it said it. Whether a claim is truthful needs longitudinal verification, and that is a different instrument.

Higher pressure therefore means stronger conditions for immediate action, and nothing else. It is not a rating, not a grade, and not a judgement about a retailer’s conduct.

History

None exists. Every retailer has one dated observation, on 2026-09-06. The data model appends dated observations and the weekly run begins 2026-09-13; until a second comparable capture exists there are no arrows, no movers and no trend.


Sources

The dataset is generated read-only from frozen corpora and records a hash for every input, so any row can be traced to the bytes that produced it. The corpora themselves are never modified.

InputRolesha256
freeze_manifestthe five hashes that define the instrumenta07d9aa0d58fc362
acceptance_criteriathe bar a dimension must clear to enter a composite, recorded before review16df988c3ab3fb07
holdout_resultauthor-adjudicated holdout metrics, blind to detector outputdd3a218c2a724bd2
ontologyfrozen ontology v1075e6bfb27f515f1
v1_page_dimensionone row per page per dimension, full universe pass4ab27a8c84c51e9c
v1_signalsone row per detected signal, with evidence window41915b544f8ab44e
universethe 114-retailer panel49a2507fe186236c
baseline_runcoverage and evidence hashes265 ledgers, hashed individually inside each.per-file