Cloud Platform Engineer
$132,700–$206,800 year
HybridSan Francisco, California, United States
Job Summary
Write, review, and ship production-quality Python and Go code to build tools, automation, and AI services that improve platform accessibility and developer productivity. Assist in designing shared infrastructure components for scalability and help integrate LLM-powered workflows and agentic features into the cloud foundation. Develop and maintain Infrastructure as Code using Terraform or Ansible to manage resources across AWS, Azure, and GCP environments. Participate in on-call rotations for incident response, root-cause analysis, and self-healing solutions while monitoring cloud environments with Datadog and Grafana to optimize performance and cost-efficiency. Implement Zero Trust security controls and partner with InfoSec teams to maintain compliance across multi-cloud setups.
Required Qualifications
- BS/BA degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent experience
- 5+ years of experience in Cloud Engineering, DevOps, Platform, or Infrastructure Engineering roles
- Experience writing, reviewing, and testing production-quality code (Python and/or Go) - not just configuring tools
- Experience with LLMs, prompt engineering, and/or AI agent frameworks with the ability to apply AI to automate infrastructure and developer workflows and comfort using AI coding assistants to accelerate delivery
- Experience with scripting and automation (Python, Bash, and/or PowerShell) for infrastructure management
- Experience managing resources in a cloud environment (AWS, Azure, or GCP) and building immutable infrastructure with IaC tools such as Terraform or Ansible
- Experience with CI/CD pipelines and DevOps practices
- Hybrid: Employee divides their time between in-office and remote work. Access to an office location is required. (Frequency: Minimum 2 days per week; may vary by team but will be weekly in-office expectation)
Desired Qualifications
- Hands-on experience building or operating LLM-powered or agentic services in production (e.g., RAG, model routing/gateways, orchestration frameworks)
- Experience with containerization and orchestration tools (Docker, Kubernetes) and microservices architecture
- Knowledge of Linux and/or Windows administration and optimization
- Experience with monitoring, telemetry, or AIOps tooling
- Understanding of network architecture and security best practices, including VPNs, firewalls, and load balancing in cloud environments
- Strong problem-solving skills and excellent communication and collaboration skills, with the ability to articulate complex technical concepts to non-technical stakeholders
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.