Design, build, and maintain ETL pipelines and data platforms using Python, PySpark, SQL, and Airflow. Support cloud data platforms (AWS/Azure/GCP), maintain relational databases, convert unstructured data into vectors, collaborate with data/ML teams, troubleshoot pipeline issues, and document processes.
Key Responsibilities:
- Design, develop, and maintain ETL (Extract, Transform, Load) processes to ensure the seamless integration of raw data from various sources into our data lakes or warehouses.
- Utilize Python, PySpark, SQL and AirFlow etc., to process, analyze, and store large-scale datasets efficiently.
- Write and maintain SQL queries for data retrieval, transformation, and storage in relational databases like Redshift or PostgreSQL.
- Support cloud-based data platforms such as AWS, Azure, or GCP, with a focus on orchestrating AI retraining cycles, versioning, and automated pipeline monitoring.
- Familiarity in converting unstructured data into vectors using frameworks like LangChain or LlamaIndex and storing them.
- Collaborate with cross-functional teams, including data scientists, ML engineers, and domain experts to design and implement scalable solutions.
- Troubleshoot and resolve performance issues, data quality problems, and errors in data pipelines.
- Document processes, code, and best practices for future reference and team training.
Requirements
Additional Information:
- Experience level 3+ years.
- Strong understanding of data governance, security, and compliance principles is preferred.
- Ability to work independently and as part of a team in a fast-paced environment.
- Excellent problem-solving skills with the ability to identify inefficiencies and propose solutions.
- Experience with version control systems (e.g., Git) and scripting languages for automation tasks.
Similar Jobs
Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Design, build, and maintain cloud-native data pipelines and ETL/ELT workflows for high-volume trade lifecycle data. Integrate DTCC sources (RTTM, Settlement Web) into a Unified Trade Record, process ISO 15022 messages, manage trade exception databases, optimize SQL Server performance, and ensure data quality, lineage, governance, and observability while collaborating with business stakeholders.
Top Skills:
AWSDockerDtcc RttmDtcc Settlement WebEksEltETLFixFxmlIbm MqIndexingIso 15022KafkaKubernetesRdsRest ApiS3SQLSQL ServerStored ProceduresUnified Trade Record (Utr)
HR Tech • Legal Tech • Software • Consulting
Lead technical direction for data platforms, design and maintain data pipelines for reporting and AI (RAG, embeddings, vector stores), reduce architectural debt, mentor engineers, partner with product leadership, and represent Mitratech externally.
Top Skills:
AirbyteAnalytics PlatformsAWSBi ToolsCi/CdClaude CodeCloud StorageCopilotCursorDbtDbt Semantic LayerEmbeddingsEtl/EltFivetranGitInfrastructure-As-CodeMonitoringPostgresRag IngestionReactRuby On RailsSnowflakeSQLTerraformVector Stores
Information Technology • Insurance • Professional Services • Consulting
Design, develop, and maintain ELT/ETL Snowflake data warehouse solutions; implement data standards, CI/CD, monitoring, and production support; collaborate with product, business, and vendors; mentor junior developers and improve data engineering practices.
Top Skills:
AzureCi/CdData LakeData WarehouseEltETLGitMdmPower BISnowflakeSQLSsisSsrs
What you need to know about the Melbourne Tech Scene
Home to 650 biotech companies, 10 major research institutes and nine universities, Melbourne is among one of the top cities for biotech. In fact, some of the greatest medical advancements were conceptualized and developed here, including Symex Lab's "lab-on-a-chip" solution that monitors hormones to predict ovulation for conception, and Denteric's vaccine for periodontal gum disease. Yet, the thousands of people working in the city's healthtech sector are just getting started, to say nothing of the tech advancements across all other sectors.



