Ingestion · 2026
Collecting web data without losing the failures
LaboratoryIndependent engineer
A data collection process that restarts on every crash, shares one rate limit across every source, and mixes broken records into the same table as valid ones quickly becomes unreliable and hard to trust.
The constraint
This is a laboratory pattern, not a client metric page. Four demo sources ship with it, including a JavaScript-rendered page. Tests run offline against frozen HTML so a live site cannot flake the suite.
Built with
PythonCamoufoxSQLiteDocker
How it works
Sources are YAML. Fetch writes a raw cache. Parse and validate, then deduplicate. Bad records go to a dead-letter queue. Tests use frozen HTML.
- YAML sources
- Fetch + cache
- Parse
- Dedup
- Clean table
- Dead letter
The calls that shaped it
Each decision with the pressure that forced it and the price it keeps costing.
YAML per source
Rate limits, selectors, and retries live next to the source, not in a shared bag of globals.
Cache the raw HTML
A crash does not refetch the corpus. Replay is a local problem.
Two-layer deduplication and a dead letter
Duplicates and invalid records do not enter the clean table. Retries follow the server's rate-limit signal rather than a fixed timer.
Reliability here is earned by structure: configuration lives with each source, raw HTML is cached so a crash does not refetch the corpus, and bad records go to a dead-letter queue instead of the clean table. See the e-commerce pipeline case for that pattern running as a daily job.
Where it stands
Laboratory. A reliable pattern for collecting and validating web data, ready to apply to a client's data source.
What was handed over
- How to add a source in YAML
- How to run the tests offline
- What the dead-letter queue looks like