Site Reliability Engineer (Data warehouse administration)
On-siteOntario, Canada or Clearwater, Florida, United States
Job Summary
Operate and support data warehouse platforms including Hive, Hadoop, ClickHouse, and Vertica in production environments. Monitor system health, query performance, and resource utilization while troubleshooting incidents across data pipelines, ETL workflows, and Spark jobs. Perform root cause analysis for system failures and design observability solutions to drive operational excellence. Manage deployment activities, configuration changes, and patching for large-scale distributed systems. Participate in on-call rotation to ensure timely resolution of critical issues. Maintain operational documentation and collaborate with data engineering and analytics teams to improve system robustness.
Required Qualifications
- Bachelor's degree in Computer Science, Information Systems, or equivalent practical experience
- 3+ years of experience in data warehouse administration, data platform SRE, or production support roles
- Proven experience supporting large-scale data platforms in production environments
- Experience managing incidents, troubleshooting system failures, and driving resolution
- Familiarity with data warehouse architecture and large-scale data processing concepts
- Experience working with cross-functional teams including data engineering, platform, analytics teams, and external vendors
- Experience supporting business-critical systems with high availability and performance requirements
- Strong SQL skills with experience working on large-scale data systems
- Hands-on experience with data warehouse platforms (Hive, ClickHouse, Vertica, or similar)
- Experience supporting production data platforms with SLA/SLO awareness (required)
- Strong understanding of distributed systems, preferably within the Hadoop ecosystem
- Experience supporting ETL pipelines, data ingestion, and workflow orchestration tools (e.g., Azkaban, Airflow)
- Solid troubleshooting skills across data pipelines, query performance, and system/infrastructure layers
- Experience with monitoring, alerting, and observability tools
- Working knowledge of Linux systems in production environments
- Experience with scripting (Shell, Python, or similar) for automation
- Familiarity with deployment processes, environment management, patching, and upgrade activities
- Experience with Spark, PySpark, or SparkSQL troubleshooting
Desired Qualifications
- Familiarity with cloud storage (e.g., Azure Blob) and BI integrations (e.g., Power BI Gateway) is a plus
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.