Clarvos LLC logo
Clarvos LLCPosted 2 weeks ago

Senior/Principal Visual ML Engineer

RemoteUnited States

Full TimeSenior LevelSmall

Job Summary

Architect, develop, and deploy production-grade AI systems for image, video, and creative generation using Vision Language Models, diffusion architectures, and multimodal frameworks. Build scalable pipelines for automated creative workflows, fine-tune foundation models with LoRA and PEFT, and optimize performance across quality, latency, and cost. Lead decisions on RAG pipelines, embeddings, and multi-agent systems while establishing long-term Visual AI strategy and communicating complex concepts to executive leadership. Own end-to-end ML system design from data ingestion to monitoring and collaborate with Product, Data, and Platform teams to power creative automation and dynamic optimization.

Required Qualifications

  • MS or PhD in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, Robotics, or a related field
  • 8–12+ years of experience developing production ML systems
  • 5+ years of experience in Deep Learning and Computer Vision
  • 3+ years of hands-on experience with Generative AI for images and video
  • Expert-level proficiency in Python and PyTorch
  • Strong software engineering fundamentals with production-quality code
  • Solid understanding of distributed systems, GPU optimization, batching, and cost-aware inference
  • Excellent software engineering fundamentals (Python, APIs, microservices, Docker, Kubernetes)
  • Must have mobility to attend meetings remotely and in person
  • Must be able to write, type and use a telephone system 100% of the time

Desired Qualifications

  • Deep expertise in Generative AI, multimodal foundation models, Vision Language Models (VLMs), Large Vision Models (LVMs), diffusion models, transformers, and autoregressive architectures
  • Hands-on experience building image and video generation systems using leading models such as FLUX, Stable Diffusion, Imagen, Veo, Runway, Kling, and open-source video diffusion models
  • Strong experience with computer vision and multimodal AI frameworks (CLIP, Florence, Qwen-VL, LLaVA, SAM, YOLO, Grounding DINO)
  • Proven ability to productionize large-scale AI models using modern ML infrastructure including Hugging Face, Diffusers, DeepSpeed, FSDP, TensorRT, ONNX, CUDA/GPU optimization, and cloud-native MLOps platforms (AWS/GCP/Azure, Kubernetes, Kubeflow, MLflow, distributed inference)
  • Architect and oversee scalable LLM/GenAI systems for MarTech/AdTech use cases
  • Lead development and building of AI systems for: Text-to-video generation, Image-to-video generation, Video editing, AI avatars, Motion transfer, Storyboarding, Creative sequencing, Marketing video generation, Dynamic creative optimization
  • Design and deploy multi-agent systems using frameworks such as LangGraph, AutoGen, CrewAI, MCP, or equivalent
  • Own end-to-end ML system design: data ingestion, feature pipelines, training, inference, evaluation, and monitoring
  • Build robust model evaluation frameworks to measure creative quality, visual fidelity, consistency, brand alignment, safety, hallucination risk, and human preference alignment
  • Improve model performance across quality, latency, scalability, and cost through continuous experimentation, benchmarking, and production optimization
  • Communicate complex ML concepts clearly to executive leadership, stakeholders, and the Board
  • Contribute to technical narratives used for fundraising, company valuation, and strategic planning

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce