Data Engineer – Demographic & Geospatial Data
About the Role
We are seeking a hands-on Data Engineer to build Databricks pipelines that combine demographic data with geographic reference data, delivering reliable, reconciled datasets for population analysis and service planning.
Key Responsibilities
- Design and deploy pipelines using Databricks, Python, PySpark, SQL and Delta Lake
- Standardise postcodes and enrich records using NHS and ONS postcode products and ODS reference data
- Implement spatial joins and point-in-polygon matching, handling edge cases and multiple matches
- Version lookups and boundaries so historical outputs stay reproducible
- Automate data quality checks and reconcile counts before and after joins
- Optimise Spark and spatial processing as data volumes grow
- Apply access, pseudonymisation and disclosure controls to sensitive data
Essential Skills
- Strong Databricks, Python, PySpark and SQL in production
- Geospatial engineering: CRS (BNG, WGS84), geometry validation, GeoJSON, WKT/WKB, GeoParquet
- Apache Sedona or native Databricks spatial functions
- Git, CI/CD, automated testing and clear communication with analysts
Desirable
- NHS demographic or registration data; LSOAs, MSOAs, ICBs; UPRNs
- GeoPandas, Shapely or QGIS; regulated data environments