Hyphen Connect Limited
LLM Pre-training & Distributed Engineer (AI Infrastructure)
Be an Early Applicant
Lead orchestration and optimization of large-scale LLM pretraining across 1,000+ GPUs. Manage distributed training with PyTorch/DeepSpeed/Megatron-LM, tune networking and memory (InfiniBand/RDMA), and implement checkpointing and robust failure recovery for long-running jobs.
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes.
Responsibilities:
- Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
- Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
- Automate checkpointing and failure recovery during month-long training runs.
Required Skills:
- Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
- Experience managing SLURM or Kubernetes-based GPU clusters.
- Strong systems engineering background (C++, CUDA, Python).
Similar Jobs
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Design, build, and maintain scalable, high-performing frontend systems and user interfaces. Collaborate with engineers, designers, and managers to solve user problems; lead technical design, implementation, launches, code reviews, documentation, and complex fixes. Mentor engineers, contribute to frontend testing and release practices, and help deliver reliable software across Atlassian’s cloud product suite.
Top Skills:
AngularjsChaiCSSCypressHTML5Javascript (Es6)JestMochaReactVue
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Manage retention and growth for complex enterprise customers across the Greater China Region. Own the customer lifecycle, account strategy, forecasting, renewals, expansions, upsells, and cross-sells. Partner with sales, services, channel, and customer success teams to execute strategic plans, navigate enterprise procurement and security workflows, and maintain customer health. The role requires fluent English and Mandarin, enterprise SaaS sales experience, strong CRM and forecasting skills, and familiarity with GCR business dynamics.
Top Skills:
Analytics ToolsCRMSaaSSalesforce
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
Research Staff will develop foundational voice AI technologies, including neural audio codecs, steerable generative models, latent-space embeddings, synthetic audio data generation, multimodal speech-to-speech systems, and hardware-efficient architectures. The role combines mathematical modeling, algorithmic innovation, large-scale data engineering, real-time deployment optimization, and rigorous experimental design to advance scalable, low-latency voice interaction.
Top Skills:
Data PipelinesDeep LearningEmbedding SystemsFoundation ModelsGenerative ModelsMamba-2Multimodal LearningNeural Audio CodecsSelf-Supervised LearningSoundstreamSpeech-To-TextTensor Processing Units (Tpus)Text-To-SpeechTransformersTriton KernelsVoice AiVq-VaeWhisper
What you need to know about the Melbourne Tech Scene
Home to 650 biotech companies, 10 major research institutes and nine universities, Melbourne is among one of the top cities for biotech. In fact, some of the greatest medical advancements were conceptualized and developed here, including Symex Lab's "lab-on-a-chip" solution that monitors hormones to predict ovulation for conception, and Denteric's vaccine for periodontal gum disease. Yet, the thousands of people working in the city's healthtech sector are just getting started, to say nothing of the tech advancements across all other sectors.


