Platform Engineer (Agent Runtime)
On-siteTaguig, Metro Manila, Philippines
Job Summary
Own the health of the agent execution environment day to day by managing session isolation, resource limits, and failure modes that emerge under real concurrency. Investigate and resolve cases where agents behave differently in production than in testing, working directly with Agent Engineers to determine if the issue stems from design or the runtime environment. Tune cold-start and concurrency settings for critical-path functions on a regular cadence as usage patterns shift. Build and maintain observability for tracing, behavioral drift detection, and quality signals alongside standard metrics. Track per-agent cost and efficiency, flagging agents that burn excessive tokens or run longer than required. Deploy and manage runtime-layer infrastructure resources through CI/CD pipelines, including IAM roles and observability configurations. Run and improve tests that validate isolation holds between agents.
Required Qualifications
- At least 5 years in a platform, SRE, or infrastructure engineering role, with real production experience running serverless or containerized workloads on AWS (Lambda, Fargate, Docker, Kubernetes or equivalent) at meaningful scale.
- Hands-on experience with AWS observability tooling (CloudWatch, X-Ray, or a comparable distributed tracing stack) — able to go from 'something's wrong' to a root cause using traces and logs, not just dashboards.
- Real experience deploying and managing infrastructure through CI/CD independently — comfortable owning IaC changes (Terraform, CDK, or similar) end to end rather than handing them off to someone else.
- Working understanding of compute isolation concepts (containers, microVMs, or similar sandboxing models) and where their guarantees actually stop, since a lot of this role is reasoning about the edges of what a platform promises versus what it might not fully cover on its own.
- Comfort investigating cost and performance problems at a granular level — able to trace an unexpectedly expensive or slow workload back to a specific cause, not just flag that costs went up.
- Strong incident response instincts: staying calm and methodical while root-causing a live production issue, and following through with an actual fix rather than a workaround.
Desired Qualifications
- Direct experience with AWS Bedrock, SageMaker, or any managed AI/agent runtime platform.
- Experience with cold-start or concurrency tuning specifically (Lambda provisioned concurrency, container warm pools, or similar).
- Exposure to a security-conscious or regulated environment where infrastructure changes go through a formal review or approval process.
- Familiarity with LLM-specific cost drivers (token pricing, tool-call volume, model tiering) even if it wasn't the primary focus of a past role.
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.