Schema test
Every row is validated against the schema from the approved sample: required fields present, types correct, enumerations in range. Rows that fail are quarantined and reported, never silently dropped.
How we work
The same process whether you hire an engineer or order a dataset: a written scope, a sample before the full run, quality gates on every delivery, and a thirty-day fix window.
Process
You send URLs and fields. We probe the site, name the protection stack, and reply in writing within 24 hours: doable or not, approach, timeline, price.
A real sample from the live site, in your target format, usually within 3 days. You check fields and quality before committing to the full run.
Scrapers are written in Python (Scrapy, Playwright, Crawlee) in a repo you can read, with the schema from the approved sample enforced in code.
The full run goes through the quality gates below. Anything that trips a threshold is investigated before data leaves our side.
CSV, JSON, Sheets, a bucket, a warehouse table or a webhook. Each delivery comes with a short run report: rows, duplicates removed, failures, run time.
Thirty-day fix window on every data project. For recurring runs and hired engineers, monitoring and fixes are part of the monthly price.
Quality gates
Row counts alone prove nothing. Every delivery runs through the same checks, and the run report says which ones fired.
Every row is validated against the schema from the approved sample: required fields present, types correct, enumerations in range. Rows that fail are quarantined and reported, never silently dropped.
Natural keys (URL, SKU, listing ID) are defined per site and enforced across pages, sort orders and runs.
A run that returns 20% more or fewer rows than the previous run or the estimate is held and checked before delivery. It is usually a site change, a block, or a real change in the data; we tell you which.
Per-column empty rates are compared against the sample. A field that was 2% empty and is now 40% empty means a selector broke or the site is serving placeholders.
403s, captchas and challenge pages are counted per run. A rising block rate is fixed before it becomes missing data.
A handful of rows from every delivery are compared by hand against the live page, field by field.
Reporting
What shipped, what is blocked, what is next. Posted at the end of each working day by every engineer on your project. No meeting needed.
// example, not a real runtue · pyle_product_urls: pagination fixed, 4,120 rows · blocked: none · next: detail pages Sites worked on, scrapers shipped or fixed, rows delivered, blocks handled, plan for next week. Two to four paragraphs every Friday, read by a lead before it reaches you.
// example, not a real runweek 41 · 3 sites live · 2 fixes · 86k rows delivered · 1 new protection (DataDome) assessed Legal stance
We would rather decline a job than deliver something you cannot use. These rules apply to every quote.
Questions
One line per engineer per day in your tracker or chat: what shipped, what is blocked, what is next. No meetings required; if you run a weekly call we join yours.
Sites worked on, scrapers shipped or fixed, rows delivered, blocks encountered and how they were handled, and the plan for next week. Written, two to four paragraphs, sent every Friday.
Within thirty days of a data project delivery we fix and re-run at no charge. After that, or for recurring runs, maintenance is included in the monthly price.
Python for nearly everything: Scrapy and httpx for HTTP, Playwright for browser work, Crawlee where it fits. Apify for hosted Actors. Delivery through boto3, the Google APIs or plain SQL. When you hire an engineer, we use your stack.
On a hired-engineer engagement the code lives in your repo from day one. On a data project you receive data by default; the scraper source can be included for a fee quoted up front.
Send a URL and the fields. The written scope arrives within 24 hours.
Hi! Send me the site you need data from and I'll get an engineer to look at it.
Chat on WhatsApp Or get a quote