Infrastructure Engineer
HybridAustin, Texas, United States
Job Summary
Own the chain from ambiguous business problems to running controls by building production TypeScript in a shared monorepo, designing Temporal workflows, and managing SQL migrations. Run the cloud platform using EKS, AWS defined in Pulumi, and CI/CD pipelines while instrumenting observability to detect issues before they impact brokers. Evaluate vendor permission models and integrate identity systems, ensuring least-privilege IAM, BAA compliance, and audit trails that hold up in carrier reviews. Treat cost as a design input by reconciling estimated versus billed LLM spend and alerting on runaway jobs. Simplify complex, multi-tenant environments for regulated healthcare operations while acting as the technical interface to security, data, and finance.
Required Qualifications
- 5–10 years in infrastructure, platform, SRE, or DevOps engineering, with meaningful production ownership
- Production Kubernetes experience — EKS preferred: cluster upgrades, autoscaling, networking, and hands-on troubleshooting
- Strong AWS fundamentals: IAM, VPC and networking, compute, storage, KMS, and a working grasp of the cost model
- Infrastructure as code in a real codebase — Pulumi preferred; Terraform or CDK acceptable with a willingness to work in Pulumi
- Production TypeScript or Node — you will ship application code in a shared monorepo, not only infrastructure definitions
- Relational data modeling and SQL, Postgres preferred; comfortable writing and reviewing migrations
- Built and owned CI/CD pipelines (GitHub Actions or comparable), including deployment strategy and rollback
- Observability in practice — metrics, logs, and distributed tracing, plus alerting that engineers trust
- Experience operating in a HIPAA, SOC 2, or otherwise regulated and audited environment
- A writing habit — design docs, RFCs, or postmortems you can point to
- Fluency with AI coding tools in day-to-day work
Desired Qualifications
- Durable workflow engines — Temporal especially — or comparable orchestration and job systems
- Hands-on administration of an IdP or SaaS admin surface (Okta, Entra) including SCIM, RBAC, and API integration
- Multi-tenant SaaS, or M&A integration work folding acquired companies onto a single platform
- Data warehouse pipelines (Redshift, BigQuery, Snowflake) and event-ingestion plumbing
- LLM infrastructure: gateways and virtual keys, provider rate limits, token and spend telemetry
- Secrets management at scale, policy-as-code, and supply chain hygiene (image signing, SBOMs)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.