Machine Learning Engineer
On-siteSydney, New South Wales, Australia
Job Summary
Develop large-scale distributed training pipelines to manage datasets and complex models while building and optimizing low-latency inference pipelines for real-time production predictions. Design scalable model frameworks to handle high-volume trading data, develop libraries to improve machine learning framework performance, and maximize GPU hardware acceleration using CUDA, TensorRT, and distributed training tools. Collaborate with quantitative researchers to automate ML experiments, hyperparameter tuning, and model retraining, and partner with HPC specialists to optimize workflows and reduce costs. Evaluate and roll out third-party tools to enhance model development capabilities and dig into open-source internals to extend their functionality.
Required Qualifications
- 3+ years of experience in machine learning with a focus on training or inference systems
- Strong engineering skills, including Python, CUDA, or C++
- Knowledge of machine learning frameworks such as PyTorch, TensorFlow, or JAX
- Proficiency in GPU programming for training and inference acceleration (e.g., CuDNN, TensorRT)
- Experience with distributed training for scaling ML workloads (e.g., Horovod, NCCL)
- Exposure to cloud platforms and orchestration tools
Desired Qualifications
- Hands-on experience with real-time, low-latency ML pipelines in high-performance environments is a strong plus
- A track record of contributing to open-source projects in machine learning, data science, or distributed systems is a plus
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.