Methodology

Every number here comes out of the same procedure, whatever the result. This page describes it in full, including the cases where the procedure tells me to publish no number at all.

Plain language is part of the method

A result that only a statistician can read is not much use to the person who has to set the price. So every report is written to be understood without a background in maths or statistics, and every term that cannot be avoided is explained where it appears.

Three of them come up in nearly every report:

The estimate is my best single answer. For example, a 10% discount improved rank by about 7%.

The range around it says how precisely I know that answer. A narrow range means the events agreed with each other. A wide one means they scattered, and the true figure could sit anywhere inside it.

Not demonstrable means the range runs into the area where nothing happens. I measured, and the data cannot tell me whether there was an effect at all. It is a result, and I publish it.

Writing plainly does not soften anything. The thresholds, the exclusions and the publication gates below are applied exactly as written, whatever the result. → Full glossary

Where the estimates come from

Elasticity is often quoted as a fixed property of a category. On a marketplace it isn't fixed. Rank comes out of a ranking system reacting to demand, while competitors react to each other and the season moves underneath both.

Sellers change prices constantly, usually for reasons that have nothing to do with the rest of the category. Each change is a small experiment. A few hundred of them add up to an answer.

Control groups

A price change only tells you something next to a product that didn't change. For each price event I assemble a control group from the same category: products whose price stayed within 2% through the same window. At least five of them are required, otherwise the event is dropped.

If the control group moves with the product that changed price, the movement came from the season or the platform. I don't credit it to the price.

better rank ↑price change−14 d0+14 dcontrol groupwithout the changethe effect
The gap between what happened and what would have happened is the part I attribute to the price.

How the estimate is calculated

Take the rank change for the products that moved their price. Subtract the rank change for the products that didn't. What's left is the part I attribute to the price move. Repeat across hundreds of events and you get an estimate with an interval around it.

I report the interval next to every estimate. It is usually the more useful number.

Where the design comes from

The comparison used here is a standard tool in applied economics, not something invented for this publication. The version I follow is set out in Chapter 18 of Nick Huntington-Klein's "The Effect: An Introduction to Research Design and Causality", which is freely readable online: Chapter 18 — Difference-in-Differences

Reading the chapter is the fastest way to check whether the design here is sound, including the parts where it can go wrong.

The tool behind it

The data collection and the estimates run through a pipeline I built for this purpose. It pulls the histories, finds the price events, applies the eight checks, assembles the control groups and computes the estimates and intervals.

All of it is deterministic code: the same set and the same settings produce the same numbers, and no language model is involved in producing any figure. Each published report points to one specific run of that pipeline, recorded with its data state and code version, so a number can be traced back to exactly the state it came from.

Coverage and provenance

Review history is not available for every product, so the review-burst check cannot be applied to every event. Each report states how many of its events went unchecked. Every published figure carries a reference id pointing to one recorded run; republishing a category creates a new reference, and the old one stays valid for the report it produced.

What gets excluded, and why

Eight checks run before anything is estimated. An event has to pass all of them.

  1. Gaps in the price history around the change.
  2. Too little ranking data in the window, less than 70% of days covered.
  3. Another price change on the same product overlapping the window.
  4. Deal or coupon markers, where Keepa recorded them.
  5. Signs the product was out of stock.
  6. A burst of new reviews: more than 10% and at least 15 additional reviews, which moves rank on its own.
  7. Listings with less than 60 days of history.
  8. Fewer than five comparable control products.

Each of these moves rank through something other than the shelf price, and leaving them in would credit price with work it did not do.

Price changes found1,312Passed every check653Used for the figure328
About half of all detected price changes are dropped before anything is estimated.Figures shown are illustrative examples taken from published reports.

How results are graded

Each direction is graded from its 90% interval. Reliable means the interval excludes zero and I am willing to state a direction. Not demonstrable means the interval includes zero. I looked and could not show an effect. I publish that in the same typeface as everything else.

Insufficient data means the category did not clear the publication gates below. I ran the estimate and I withhold the number.

nothing happensreliable0.74reliable1.82not demonstrable0.670.01.02.0
When the range crosses zero, I cannot tell an effect from noise, and I say so.Figures shown are illustrative examples taken from published reports.

Study protocol v1.1

· frozen October 2026
  1. 01Data: Keepa price and sales-rank histories for a curated set of ASINs per category, one marketplace at a time.
  2. 02Price event: a change of at least 5% and at least €0.30, sustained for at least 7 days.
  3. 03Windows: baseline days −14 to −1 before the change; effect window days +3 to +14 after it. The first two days are skipped because rank reacts with a lag.
  4. 04Control group: products in the same category whose price stayed within 2% through the same window; at least 5 controls required, otherwise the event is dropped.
  5. 05Outcome: change in category sales rank, log-transformed, so that halving the rank counts the same at every level.
  6. 06Estimator: difference-in-differences against the control group. Intervals from a cluster bootstrap over products (1,000 iterations, 90%), because several events of the same product are not independent.
  7. 07Pre-trend guard: events whose product was already moving unusually before the price change are excluded from the fit.
  8. 08Category guard: Amazon often files the same product type in several parallel category trees. The guard accepts every category branch that holds at least 20% of the set's products, matched by name across trees. Products outside these branches are excluded from the fit and from control groups. A generic parent category is never accepted as the only branch.

Publication gates

  • At least 50 qualifying price events in the fit set.
  • At least 20 distinct ASINs contributing events.
  • A directional claim is only made when the 90% interval excludes zero.

Changes to the protocol

Version 1.1, October 2026. The category guard was changed. Version 1.0 kept only products in the single category branch shared by most products of a set. Where Amazon files a product type in parallel category trees, no specific branch reached a majority and the guard fell back to a generic parent category. Genuine products were excluded as a result, which made fit sets and control groups smaller than they should have been. Version 1.1 accepts every branch held by at least 20% of the set, matched by name across trees. All categories from Q4 2026 onward are measured under version 1.1.

Q3 2026 reports were measured under version 1.0 and stay published as they are, each with its original reference id. Re-running them under version 1.1 changes the estimates in five of fifteen categories. In three of them the estimate for price cuts moves by between 0.02 and 0.11, in the other two only the ranges shift slightly. No reliability grade changes.

The review-burst rule was changed once, in August 2026: it now requires both a 10% jump and at least 15 additional reviews. A purely relative threshold was flagging ordinary growth on small review bases, where 13 to 15 reviews counts as a 15% shock. It applies to every category measured since.

The control-stability threshold stands at 2% and has never been changed for a published category. A 3% variant is documented for sets whose control groups are too thin to measure, and is in use nowhere on this site.

Nothing else has been adjusted. Any future change is listed here with its date and reason, together with its effect on published figures.

What I never claim

  • That a measured effect is causal beyond the observation window.
  • That an elasticity holds outside the price range actually observed.
  • That a category result applies to any individual listing.
  • That an absent effect proves an effect does not exist. It only means I could not demonstrate it.
  • That rank movement equals revenue.

Scrutiny welcome

If you think the design is wrong, the controls are badly chosen or an exclusion is doing too much work, write to me. Corrections appear alongside the report they correct. hello@marketplace-economics.com