Mothership logo
MothershipPosted 1 month ago

Devops Engineer

On-siteSan Antonio, Texas, United States

Full TimeSmall

Job Summary

Design and deploy scalable, fault-tolerant real-time data processing pipelines using Apache Flink and Kafka on AWS. Manage and optimize Kafka clusters (MSK or self-managed) while configuring, monitoring, and troubleshooting Flink jobs to ensure low-latency, high-throughput performance. Automate infrastructure provisioning and deployment using Terraform, CloudFormation, and CI/CD tools like Jenkins or GitHub Actions. Collaborate with Data Engineers and DevOps teams to integrate services and maintain security compliance for data platforms. Optimize deployments for scalability and fault tolerance, addressing issues related to message delivery or service outages.

Required Qualifications

  • Proficiency in AWS services such as Amazon MSK (Managed Streaming for Kafka), Amazon Kinesis, AWS Lambda, Amazon S3, Amazon EC2, Amazon RDS, Amazon VPC, and AWS IAM
  • Ability to manage infrastructure as code with AWS CloudFormation or Terraform
  • Understanding of Apache Flink for real-time stream processing and batch data processing
  • Familiarity with Flinks integration with Kafka, or other messaging services
  • Experience in managing Flink clusters on AWS (using EC2, EKS, or managed services)
  • Deep knowledge of Kafka architecture, including brokers, topics, partitions, producers, consumers, and zookeeper
  • Proficiency with Kafka management, monitoring, scaling, and optimization
  • Hands-on experience with Amazon MSK (Managed Streaming for Kafka) or self-managed Kafka clusters on EC2
  • Strong experience in automating deployments and infrastructure provisioning
  • Familiarity with CI/CD pipelines using tools like Jenkins, GitLab, GitHub Actions, CircleCI, etc
  • Experience with Docker and Kubernetes, especially for containerizing and orchestrating applications in cloud environments
  • Strong scripting skills in Python, Bash, or Go for automation tasks
  • Ability to write and maintain code for integrating data pipelines with Kafka, Flink, and other data sources
  • Knowledge of CloudWatch, Prometheus, Grafana, or similar monitoring tools to observe Kafka, Flink, and AWS service health
  • Expertise in optimizing real-time data pipelines for scalability, fault tolerance, and performance

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce