E-commerce data pipeline · 2025
Reliable daily competitor pricing, without manual copying
ShippedIndependent engineer
Competitor prices lived in a morning ritual. Open several stores, copy figures, paste a sheet. A layout change on one site broke the day. There was no retry, no isolation per store, and no record of what failed.
The constraint
Multiple storefronts, different pagination, rate limits that are not documented. One broken parser must not take down the rest of the run.
Built with
PythonBeautifulSoupSelenium
How it works
A scheduler starts the job. Each storefront is fetched and parsed on its own. Records are normalized into one table. A failed source is retried or logged, not allowed to abort the others.
- Scheduler
- Storefronts
- Parsers
- Normalize
- Morning table
- Source log
The calls that shaped it
Each decision with the pressure that forced it and the price it keeps costing.
One job, many sources
A scheduled process pulls each storefront independently, so a selector change fails one source without taking down the rest of the run.
BeautifulSoup where the HTML is stable, Selenium where it is not
Not every page needs a browser. The ones that do are isolated so the rest stay cheap.
Normalize, then write
The morning artefact is a table the client can open. Raw markup stays in the job's output when a source needs a replay.
A job on a schedule, each storefront parsed on its own, a table in the morning, a log when a page changes. That is the shape of a pricing or product-data pipeline.
If you price against competitors, tell me how many storefronts you watch and how often the table needs to refresh.
Where it stands
Shipped. Competitor prices arrive as a clean table each morning, with each store isolated and failures logged.
What was handed over
- How the job is scheduled
- Per-source notes and what a selector change looks like
- The output table and who opens it