Love turning messy web data into clean pipelines, with AI in production rather than in a demo? This one's for you.
🏢 About the company
An AI-native startup helping companies find, analyse and win public tenders, a market worth around 10% of EU GDP.
- Traction: 120+ customers within a year of the first paying client
- Funding: €6M Seed just closed, backed by top European VCs
- Growth: expanding to 6 countries, with Barcelona as the first hub abroad
- Team: ~30 people, 14 in engineering
💼 About the role
- 📍 Location: Barcelona hub (remote for now, hybrid once the office opens)
You'll own the pipelines that carry a tender from the source portal to the customer's screen. No tickets handed down: the sources, jobs and tables are yours.
Day-to-day:
- Scrape & ingest: build and maintain scrapers for tender portals across Spain and Italy
- Keep it running: catch broken portals, endpoints and rate limits before the customer does
- Clean & merge: turn messy HTML, XML and duplicates into one reliable record
- Orchestrate: flows with retries, backfills, alerting and data quality checks
- AI enrichment: batch LLM extraction, embeddings and OCR
What you'll need:
- Production Python: typed, tested, readable
- Junior (1-2 yrs): real scraping experience and solid SQL
- Mid (2-4 yrs): pipelines run with an orchestrator and hands-on PostgreSQL tuning
- Curiosity about LLM-based extraction
- Stack: Python · PostgreSQL · Prefect · K3s / ArgoCD · AWS
Who should apply?
Early-career engineers who want ownership, enjoy the detective work of scraping and thrive in a fast, high-trust team.