Requirements
- 4+ years of experience in Big Data or Data Engineering
- Strong proficiency in Python (PySpark) and SQL
- Hands-on experience with AWS cloud services for data engineering solutions
- Practical experience with Databricks on AWS, including building and managing data pipelines
- Strong knowledge of Apache Spark and large-scale data processing
- Experience with batch and streaming data processing
- Familiarity with Delta Lake and Databricks components such as Workflows and Jobs
- Experience with orchestration tools such as Apache Airflow or MWAA
- Knowledge of streaming technologies such as Apache Kafka, Amazon MSK, or Kinesis
- Upper-intermediate or higher level of English
Offer description
ABOUT THE ROLE
IIn this role, you will contribute to developing scalable cloud-based data solutions on AWS using Databricks and modern big data technologies. You will work on batch and streaming data pipelines, support optimisation and modernisation initiatives, and collaborate with cross-functional teams to deliver reliable and efficient data platforms across different stages of the project lifecycle.
Your responsibilities
- Design, develop, and maintain scalable batch and streaming data pipelines
- Build and optimize data processing solutions using Python (PySpark) and SQL
- Develop and manage data workflows using Databricks on AWS
- Work with Apache Spark and Databricks for large-scale data processing
- Implement and maintain Delta Lake-based data architectures
- Utilize orchestration tools such as Apache Airflow or MWAA to manage workflows
- Support implementation of streaming solutions using Apache Kafka, Amazon MSK, or Kinesis
- Collaborate with technical and business stakeholders to deliver effective data solutions
- Participate in architecture discussions and contribute to continuous improvement initiatives within the team
- Participate in the full project lifecycle, from PoC and MVP stages to production implementation