Stellenausschreibung
Freelancer Projekt -
Data Engineer (m/f/d) with Python and PySpark on a Databricks Tech-Stack ID24957-0
Project duration: 08/23 - 12/23
Project volume: 528 hrs remote
Project location: Remote, Hamburg
Project description:
- Advice and designing business-critical data engineering use cases, from the business problem to delivery and operation.
- Programming of data ingestion pipelines from various sources, e.g Graph Database or Data Lakehouse.
- Writing of production feature engineering code in Python and PySpark on a Databricks Tech-Stack.
- Build and maintain data pipelines for static, mixed, and time-series data.
- Design, implementation and maintenance of data infrastructure for ML algorithms in Azure Databricks.
- Responsible for the design and implementation of the CI/CD pipelines
- Data modeling and architecture (schemes, sources, optimizations)
Project skills: ?
Good communicator, but great at independent work. Solution oriented. Analytical thinking
Experience in machine learning engineering, exploratory data analysis, and software development and writing of ETL-Pipelines. Optimally, candidate should have a degree in mathematics, physics, computer sciences or in a related field.
Experience in Python programming is mandatory, especially with PySpark, whereas XGBoost, Seaborn, Matplotlib and dbutils (Databricks) are nice-to-have.
Expertise in Git, Gitlab, and CI/CD are beneficial (including Azure CLI and Azure-Cloud specific APIs).
Experience in working with Azure, Databricks, Kubernetes and Docker.
Candidate should be familiar or inclined to working in an agile environment. Prior experience with predictive maintenance tasks is a plus.
Freelancer