Design and optimize multimodal AI systems integrating vision and audio models. Improve voice-to-voice streaming latency, embed vision encoders and audio-native models into agent reasoning, and architect multimodal RAG systems for retrieving insights from videos and PDFs.
We are seeking a talented Multimodal AI Systems Architect to develop and optimize AI systems that seamlessly integrate vision and audio models. This role focuses on enhancing our voice-to-voice interactions and multimodal retrieval capabilities, ensuring our systems are efficient and innovative.
Responsibilities:
- Integrate vision encoders and audio-native models into core agent reasoning loops.
- Optimize streaming latency for voice-to-voice AI interactions.
- Architect multimodal RAG systems capable of retrieving insights from videos and PDFs.
Qualifications:
- Experience with Whisper, CLIP, and multimodal LLM integration.
- Knowledge of streaming architectures and WebRTC.
- Expertise in cross-modal alignment.
Similar Jobs
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Design, build, and maintain scalable, high-performing frontend systems and user interfaces. Collaborate with engineers, designers, and managers to solve user problems; lead technical design, implementation, launches, code reviews, documentation, and complex fixes. Mentor engineers, contribute to frontend testing and release practices, and help deliver reliable software across Atlassian’s cloud product suite.
Top Skills:
AngularjsChaiCSSCypressHTML5Javascript (Es6)JestMochaReactVue
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Manage retention and growth for complex enterprise customers across the Greater China Region. Own the customer lifecycle, account strategy, forecasting, renewals, expansions, upsells, and cross-sells. Partner with sales, services, channel, and customer success teams to execute strategic plans, navigate enterprise procurement and security workflows, and maintain customer health. The role requires fluent English and Mandarin, enterprise SaaS sales experience, strong CRM and forecasting skills, and familiarity with GCR business dynamics.
Top Skills:
Analytics ToolsCRMSaaSSalesforce
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
Research Staff will develop foundational voice AI technologies, including neural audio codecs, steerable generative models, latent-space embeddings, synthetic audio data generation, multimodal speech-to-speech systems, and hardware-efficient architectures. The role combines mathematical modeling, algorithmic innovation, large-scale data engineering, real-time deployment optimization, and rigorous experimental design to advance scalable, low-latency voice interaction.
Top Skills:
Data PipelinesDeep LearningEmbedding SystemsFoundation ModelsGenerative ModelsMamba-2Multimodal LearningNeural Audio CodecsSelf-Supervised LearningSoundstreamSpeech-To-TextTensor Processing Units (Tpus)Text-To-SpeechTransformersTriton KernelsVoice AiVq-VaeWhisper
What you need to know about the Melbourne Tech Scene
Home to 650 biotech companies, 10 major research institutes and nine universities, Melbourne is among one of the top cities for biotech. In fact, some of the greatest medical advancements were conceptualized and developed here, including Symex Lab's "lab-on-a-chip" solution that monitors hormones to predict ovulation for conception, and Denteric's vaccine for periodontal gum disease. Yet, the thousands of people working in the city's healthtech sector are just getting started, to say nothing of the tech advancements across all other sectors.


