Develop and maintain reliable data pipelines, ELT processes, workflow orchestration, custom data connectors, and CI/CD pipelines using Airflow, Python, PySpark, Hive/Trino, SQL, and Snowflake. Collaborate with cross-functional teams, monitor and troubleshoot data workflows, validate and test data, document technical processes, and support data governance.
Must Needed Skills: Apache Airflow, Python, PySpark ,Hive(Trino), SQL and Snowflake
- Develop and maintain data pipelines, ELT processes, and workflow orchestration using Apache Airflow, Python, PySpark ,Hive(Trino) and Snowflake to ensure the efficient and reliable delivery of data.
- Design and implement custom connectors to facilitate the ingestion of diverse data sources into our platform, including structured and unstructured data from various document formats .
- Collaborate closely with cross-functional teams to gather requirements, understand data needs, and translate them into technical solutions.
- Design and implement data CI/CD pipelines to enable automated and efficient data integration, transformation, and deployment processes.
- Monitor and troubleshoot data pipelines, proactively identifying and resolving issues related to data ingestion, transformation, and loading.
- Conduct data validation and testing to ensure the accuracy, consistency, and compliance of data.
- Stay up-to-date with emerging technologies and best practices in data engineering.
- Document data workflows, processes, and technical specifications to facilitate knowledge sharing and ensure data governance.
Similar Jobs
Agency • Information Technology
Design, build, and maintain ELT/data pipelines and workflow orchestration using Airflow, Python, and PySpark. Develop custom connectors to ingest diverse data, implement DataOps and data CI/CD pipelines, monitor and troubleshoot pipelines, validate data quality, collaborate with cross-functional teams, document workflows, and work with large-scale distributed datasets and visualization tools.
Top Skills:
Apache AirflowApache KafkaSparkApache SupersetGitHiveKubernetesLinuxOpenshiftPysparkPythonSnowflakeSQLSsisTrino
Agency • Information Technology
Design, build, and maintain ELT/data pipelines and workflow orchestration using Airflow, Python, and PySpark. Develop custom connectors, implement DataOps and CI/CD for data, monitor and validate data flows, collaborate with cross-functional teams, and document data workflows and governance.
Top Skills:
Apache AirflowApache KafkaSparkApache SupersetGitHiveKubernetesLinuxOpenshiftPysparkPythonSnowflakeSQLSsisTrino
Agency • Information Technology
Design, build, and maintain scalable data pipelines and ETL processes using Python and PySpark. Orchestrate workflows with Airflow, query data with Trino and Hive, and write efficient SQL. Collaborate in Agile/Scrum teams to deliver data solutions and optimize data platform performance.
Top Skills:
Agile ScrumAirflowHivePysparkPythonSQLTrino
What you need to know about the Melbourne Tech Scene
Home to 650 biotech companies, 10 major research institutes and nine universities, Melbourne is among one of the top cities for biotech. In fact, some of the greatest medical advancements were conceptualized and developed here, including Symex Lab's "lab-on-a-chip" solution that monitors hormones to predict ovulation for conception, and Denteric's vaccine for periodontal gum disease. Yet, the thousands of people working in the city's healthtech sector are just getting started, to say nothing of the tech advancements across all other sectors.
