Senior/Principal Visual ML Engineer
RemoteUnited States
Job Summary
Architect, develop, and deploy production-grade AI systems for image, video, and creative generation using Vision Language Models, diffusion architectures, and multimodal frameworks. Build scalable pipelines for automated creative workflows, fine-tune foundation models with LoRA and PEFT, and optimize performance across quality, latency, and cost. Lead decisions on RAG pipelines, embeddings, and multi-agent systems while establishing long-term Visual AI strategy and communicating complex concepts to executive leadership. Own end-to-end ML system design from data ingestion to monitoring and collaborate with Product, Data, and Platform teams to power creative automation and dynamic optimization.
Required Qualifications
- MS or PhD in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, Robotics, or a related field
- 8–12+ years of experience developing production ML systems
- 5+ years of experience in Deep Learning and Computer Vision
- 3+ years of hands-on experience with Generative AI for images and video
- Expert-level proficiency in Python and PyTorch
- Strong software engineering fundamentals with production-quality code
- Solid understanding of distributed systems, GPU optimization, batching, and cost-aware inference
- Excellent software engineering fundamentals (Python, APIs, microservices, Docker, Kubernetes)
- Must have mobility to attend meetings remotely and in person
- Must be able to write, type and use a telephone system 100% of the time
Desired Qualifications
- Deep expertise in Generative AI, multimodal foundation models, Vision Language Models (VLMs), Large Vision Models (LVMs), diffusion models, transformers, and autoregressive architectures
- Hands-on experience building image and video generation systems using leading models such as FLUX, Stable Diffusion, Imagen, Veo, Runway, Kling, and open-source video diffusion models
- Strong experience with computer vision and multimodal AI frameworks (CLIP, Florence, Qwen-VL, LLaVA, SAM, YOLO, Grounding DINO)
- Proven ability to productionize large-scale AI models using modern ML infrastructure including Hugging Face, Diffusers, DeepSpeed, FSDP, TensorRT, ONNX, CUDA/GPU optimization, and cloud-native MLOps platforms (AWS/GCP/Azure, Kubernetes, Kubeflow, MLflow, distributed inference)
- Architect and oversee scalable LLM/GenAI systems for MarTech/AdTech use cases
- Lead development and building of AI systems for: Text-to-video generation, Image-to-video generation, Video editing, AI avatars, Motion transfer, Storyboarding, Creative sequencing, Marketing video generation, Dynamic creative optimization
- Design and deploy multi-agent systems using frameworks such as LangGraph, AutoGen, CrewAI, MCP, or equivalent
- Own end-to-end ML system design: data ingestion, feature pipelines, training, inference, evaluation, and monitoring
- Build robust model evaluation frameworks to measure creative quality, visual fidelity, consistency, brand alignment, safety, hallucination risk, and human preference alignment
- Improve model performance across quality, latency, scalability, and cost through continuous experimentation, benchmarking, and production optimization
- Communicate complex ML concepts clearly to executive leadership, stakeholders, and the Board
- Contribute to technical narratives used for fundraising, company valuation, and strategic planning
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.