Cerence AI logo
Cerence AIPosted 1 month ago

Sr. Principal Software Scientist

$185,000–$280,000 year

RemoteBurlington, Massachusetts, United States

Full TimeSenior LevelDoctorate Or Professional DegreeMediumAI Services

Job Summary

Design and train large-scale transformer and hybrid foundation models, owning architecture choices across text, multimodal, and emerging paradigms. Diagnose and resolve training instabilities at scale by navigating scaling tradeoffs across data, compute, and architecture. Define technical direction for next-generation models while optimizing dynamics, including scheduler choices and learning-rate warmup. Build models from first principles rather than adapting pre-existing codebases, implementing novel designs like MoE routing and SSM/hybrid architectures. Execute large-scale distributed training using FSDP, ZeRO-3, and tensor parallelism with mixed precision. Ensure stable convergence and principled architectural decisions as models scale in size and complexity.

Required Qualifications

  • Strong theoretical and practical understanding of modern deep learning
  • Hands-on experience training large models from scratch
  • Ability to reason about optimization, not just tune hyperparameters
  • Comfort operating in ambiguous, research-driven environments
  • Transformer internals and attention mechanisms
  • Optimization algorithms and training dynamics
  • Scaling laws and compute/data tradeoffs
  • Distributed training strategies and mixed precision
  • Architecture innovation for large, real-world models
  • Basic knowledge of information security and data privacy requirements
  • Demonstrative knowledge of information security through internal training programs

Desired Qualifications

  • Design and train large-scale transformer and hybrid foundation models
  • Own model architecture choices across text, multimodal, and emerging paradigms
  • Diagnose and resolve training instabilities at scale
  • Navigate scaling tradeoffs across data, compute, and architecture
  • Define the technical direction for next-generation models
  • Apply strong fundamentals in deep learning and representation learning
  • Design and modify transformer architectures, including: Attention variants RoPE, ALiBi Grouped Query Attention (GQA) Mixture-of-Experts (MoE)
  • Build models from first principles, not just adapt pre-existing codebases
  • Own optimizer and scheduler choices, including: AdamW Lion Adafactor
  • Learning-rate and warmup schedulers
  • Understand and debug: Optimizer instability Gradient pathologies Divergence at large scale
  • Apply and validate scaling laws
  • Navigate Chinchilla-style compute vs data tradeoffs
  • Make informed decisions about model size, dataset size, and training duration
  • Design and experiment with loss functions including: Next-token prediction Contrastive objectives RLHF, DPO, GRPO
  • Understand how loss design impacts convergence, generalization, and alignment
  • Design and execute large-scale training using: FSDP ZeRO-3 Tensor parallelism Pipeline parallelism
  • Apply Mixed precision (bf16, fp8)
  • Gradient checkpointing
  • Partner closely with ML systems teams while retaining architectural ownership
  • Explore and implement novel model designs, including: MoE routing strategies Multimodal fusion architectures SSM / hybrid architectures
  • Design architectures with KV cache efficiency and inference implications in mind
  • Training remains stable as models scale in size and complexity
  • Architectural decisions are principled and defensible
  • Models converge faster and generalize better due to architecture and optimisation choices
  • Failure modes are understood, not mysterious
  • The organization develops true in-house foundation model expertise

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce