Data scraping & pipelines: the tech stack of 38 real products
Scraping, ETL, data integration and enrichment.
What sets them apart
Picks at least twice as common here as among products overall, in decisions at least 10 of them show.
The typical stack
The leading pick where at least 8 of them show the decision and the leader has at least a quarter of it.
| Decision | Most common pick | Share | At default usage |
|---|---|---|---|
| Database | 10 of 27 · 37% | — | |
| Frontend Framework | 13 of 22 · 59% | — | |
| Backend Framework | 9 of 17 · 53% | — | |
| LLM API | 9 of 15 · 60% | $10/mo · GPT-6 Luna | |
| AI SDK & Agent Framework | 5 of 12 · 42% | — | |
| Database Access & ORMs | 7 of 10 · 70% | — | |
| Hosting | 4 of 10 · 40% | $12/mo · Workers Paid | |
| Transactional Email | 3 of 9 · 33% | $20/mo · Pro 50k |
Monthly bills add up to about $42 at the calculators' default usage, list prices. Set your own usage →
What they chose, decision by decision
Among the data scraping & pipelines that show each choice, from makers' products and open-source code alike.
Small samples, fewer than 8 products: Authentication (Clerk 3, Auth.js 2) · Background Jobs & Cron (Celery 3, RabbitMQ 2) · Payments (Stripe 5, Paddle 1) · Vector Database (LanceDB 4, Qdrant 3) · Image & Video Hosting (sharp 3, Cloudinary 1) · Product Analytics (PostHog 4, Mixpanel 1)
Data scraping & pipelines we track
21 makers' products and 17 open-source projects. Makers' products first.