Inferact logo
InferactPosted 2 months ago

Member of Technical Staff, Inference

On-siteSingapore, Singapore

Full TimeMid LevelBachelors DegreeStartup

Job Summary

Inferact seeks an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger; architectures shift toward mixture-of-experts and multimodal deployments. You will work at the core of vLLM, optimizing how models execute across diverse hardware and architectures, and your work will directly impact how the world runs AI inference. Responsibilities center on implementing inference techniques from research papers, contributing performant and maintainable code, and enhancing inference engine capabilities across hardware. Preferred qualifications cover advanced KV-cache memory management, prefix caching, hybrid model serving, RL frameworks for LLMs, multimodal inference, and open-source contributions. The role is based in Singapore with compensation in SGD and visa sponsorship considered case-by-case. Benefits include medical, dental, and vision coverage.

Required Qualifications

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce