Price & Competitive Intelligence: How It Actually Gets Built
Introduction
Almost every price intelligence programme that disappoints does so for the same reason, and it is not collection difficulty. It is that prices were compared before products were matched. This guide covers the five layers of a working programme — matching, collection, normalisation, metrics, delivery — what each actually requires, the failure modes that quietly produce wrong numbers, and how to scope a first phase that proves something.
What Price Intelligence Is, and What It Isn't
Price intelligence is the continuous measurement of your price position against competitors, at the level of individual products, in each market where you compete.
It is not a price feed. A file listing competitor prices is an input, not an output. The intelligence is in knowing that these two products are the same product, that this comparison holds in this city, and that this gap changed last Tuesday.
The distinction matters commercially because the market sells the input and the buyer needs the output. A brand that buys a price feed and discovers six months later that a third of its comparisons were matching different pack sizes has not bought price intelligence. It has bought a spreadsheet that produced confident, wrong decisions.
The Five Layers
Layer 1 — Product matching
This is where programmes succeed or fail, and it is almost always underestimated.
To compare your price against a competitor's, you must establish that the two listings describe the same thing. Where a shared identifier exists — EAN, UPC, ASIN, GTIN — this is comparatively easy. On many platforms, and in most B2B and cross-border contexts, it does not exist.
What matching actually requires:
Attribute extraction from unstructured listing titles: brand, product descriptor, pack size, unit, variant specification
Normalisation of units and pack expressions to a common convention, so "500 g", "0.5 kg" and "500gm" resolve to one value
Candidate scoring on attribute agreement, producing a confidence level rather than a binary decision
A three-state model — accepted, review queue, rejected — so uncertainty is visible and reviewable
An explicit unmatched state, because products with no competitor equivalent are a market fact, not a data defect
A persistent map, maintained across runs, so coverage compounds instead of being rebuilt
The test of whether a programme has this layer: ask what percentage of the category is matched, and at what confidence. If the answer is unavailable, comparisons are running on name similarity.
Layer 2 — Collection
Collection is the layer everyone thinks is the hard part. It is largely a solved problem, with three specific complications that are not.
Location resolution. On any platform that personalises by delivery location — marketplaces, quick commerce, most grocery — price and availability are functions of the shopper's location. Collection must specify a location per request and verify the resolved location on the returned page. Silent fallback to a platform default produces well-formed rows describing the wrong market, and no downstream check catches it.
Timing consistency. Comparing prices across competitors requires holding time roughly constant. Collect competitor A in the morning and competitor B in the evening and you have measured time as well as competition.
Frequency matched to volatility. Base grocery prices move weekly. Quick-commerce prices move within the day. Marketplace prices can move hourly during promotional periods. Applying one frequency to everything either wastes budget or misses the movement.
Layer 3 — Normalisation
Comparable prices are not the prices on the page.
Per-unit price — required wherever pack sizes differ, and mandatory for cross-border comparison where pack conventions vary by market.
Landed price — item price plus shipping, and where relevant taxes. A competitor at your price with inflated shipping is undercutting you while appearing matched. This is the single most common measurement error in MAP and price-position work.
Promotional versus base price, held separately. Collapsing them loses the ability to distinguish a structural price move from a two-day promotion — which is usually the actual question.
Currency handling for multi-market programmes, with the exchange basis recorded rather than applied invisibly.
Layer 4 — Metrics
Raw comparisons are not decisions. The metrics that get used:
Price index — your price divided by a competitor's, or by the market median, per product per location per period. The core output. Compute it per location; a national index averages your best and worst markets into a number that describes neither.
Price gap distribution — not just the average gap but its spread. An average index of 1.00 can mean "matched everywhere" or "20% expensive on half the range and 20% cheap on the other half." Only the distribution distinguishes them.
Availability rate — percentage of tracked product-location combinations in stock. Price position is meaningless on an out-of-stock product, and availability often explains sales movement that price cannot.
Share of shelf — your products as a proportion of visible listings in a category or search result. The digital equivalent of facings, and frequently a better predictor of sell-through than price index.
Promotional overlap — how often your promotional windows coincide with competitors'. Regularly reveals discounting into a competitor's discount, which buys nothing.
Assortment gap — products a competitor lists where you are absent, per market. Often the highest-value output, and the one that most surprises commercial teams.
Movement flags — changes beyond threshold since the last period. Nobody diffs two large files; the programme has to surface what moved.
Layer 5 — Delivery
Delivery format determines whether the data gets used, and this is more consequential than it sounds.
Analysts want files — CSV, Excel or JSON in a schema that loads into an existing BI layer without transformation.
Commercial and marketing teams want a view. A weekly file that requires building a pivot table gets opened twice and then ignored. A dashboard showing index by brand by market, week over week, with movement highlighted, gets used in meetings.
Engineering teams want an API. Where price data feeds a pricing engine or an application, scheduled files are friction.
Providing more than one costs little and substantially widens the internal audience. Most programmes that quietly die do so because they were delivered in one format to one team.
Four Failure Modes That Produce Confident Wrong Numbers
Matching on names. Produces comparisons between different pack sizes, variants and configurations. The output looks precise and is commercially misleading. Symptom: nobody can state the match rate.
Location fallback. A session loses its location context and reverts to a default; rows keep arriving, prices are plausible, and part of the dataset describes a market nobody asked about. No downstream validation catches this, because the prices are real. Only collection-time verification prevents it.
List price instead of landed price. Systematically understates competitor aggression wherever shipping is used as a lever.
Overwriting history. The most common first-generation mistake. Without accumulated history you cannot measure whether a price change achieved anything, cannot see seasonality, and cannot demonstrate a pattern of competitor behaviour. Append, never overwrite.
Build or Buy
Build if price data is core to your product, you can staff it permanently, and you need collection logic no vendor will customise. The honest cost is maintenance, not construction — platforms change, and an unattended pipeline produces quiet errors rather than obvious failures.
Buy if you need data to make commercial decisions rather than to build a product, you want multi-platform and multi-market coverage without a proportional engineering effort, and you would rather own the analysis than the plumbing.
The middle path most organisations land on: buy collection and matching, own the metrics. Take normalised feeds into your own BI layer and build the indices that fit your commercial process. It keeps the compounding asset — the product map and the price history — in a maintained state while leaving interpretation with the people who own the pricing decision.
How to Scope a First Phase
A first phase that proves something has five parameters, and none of them is "comprehensive":
One category where you compete directly.
Three to five named competitors. Naming them converts an open-ended crawl into a bounded matching problem.
Two markets or two to three locations. Enough to reveal whether geographic variance exists.
Fifty to two hundred products. Enough for a distribution, small enough to review matches by hand.
Four weeks at intended frequency. Long enough to capture at least one promotional cycle.
The deliverable that matters from phase one is not the price data. It is the match rate and the measured gap distribution. Those two numbers tell you whether a wider programme is worth building and roughly what it will find. A phase one that hands over prices without them has skipped the layer that makes the rest valid.
If the gap distribution turns out to be tight and geographically flat, that is a genuinely useful and cheap finding — your category does not need continuous monitoring. If it is wide, you now have a sized business case rather than an argument about methodology.
FAQ
What is price intelligence?
The continuous measurement of your price position against competitors at individual product level, per market. It requires product matching, location-aware collection, price normalisation and derived metrics — a competitor price feed alone is an input, not price intelligence.
Why do price intelligence projects fail?
Most commonly because products were not properly matched before prices were compared. Comparisons run on name similarity, silently comparing different pack sizes or variants, producing precise-looking figures that lead to wrong pricing decisions.
How do you match products without a shared identifier?
By extracting structured attributes from listing text, normalising units and pack expressions, then scoring candidate matches on attribute agreement with a confidence level. Matches are classified as accepted, needing review, or rejected, and products with no counterpart are recorded as unmatched rather than force-matched.
Should prices be compared using list price or landed price?
Landed price — item price plus shipping and, where relevant, taxes. Competitors commonly use shipping as a pricing lever, so list-price comparison systematically understates their aggression.
How often should competitor prices be collected?
Match frequency to the platform's volatility: weekly for grocery base prices, daily for marketplaces, up to several times daily for quick commerce. Applying a single frequency across all sources either wastes budget or misses movement.
Does price position need to be measured per location?
Yes, on any platform that personalises by delivery location. A national price index averages your strongest and weakest markets into one number, which can be simultaneously accurate and useless — the market where you are badly mispriced is absorbed by the ones where you are fine.
What should a first price intelligence phase deliver?
A match rate, a measured price-gap distribution, and a validated schema — not just a price file. Those three outputs determine whether a wider programme is justified and what it is likely to find.
Is it better to build or buy price intelligence?
Build if price data is your product and you can staff maintenance permanently. Buy if you need it for commercial decisions across multiple platforms and markets. The common middle path is buying collection and matching while owning the metrics in your own BI layer.
Conclusion
You can also reach us for all your mobile app scraping, data collection, web scraping , and instant data scraper service requirements!
Comments
Post a Comment