Principal Data Engineer – (Hadoop/Big Data, AWS, Python, Kinesis) - Irvine, CA - Onsite - Contract to Hire - AS
On-siteIrvine, California, United States
Job Summary
Lead creation of data environments and sets to serve data scientists, analysts, and business users. Perform offline analysis of large datasets using big data software ecosystems and validate root cause analysis for escalated technical issues. Own product datasets from definition through production deployment, designing and implementing Hadoop EMR clusters and deploying Hadoop, Spark, and database storage infrastructures in AWS. Lead design of distributed, scalable data pipelines for real-time ingestion and process big data technologies to improve architecture. Monitor HDFS and Spark releases to optimize system performance while developing tools to automate server and cluster tasks. Collaborate with developers and architects to deploy data tools supporting operations and product use cases.
Required Qualifications
- Bachelor's degree in computer science, computer engineering, or a related technical field
- 12+ years of professional experience as a data software engineer
- 16+ years of related experience as a data software engineer in lieu of 4-year degree
- 2+ years of experience with AWS cloud or other cloud Big Data computing design, provisioning, and tuning
- Previous experience as a Data Engineer / Database Administrator and/or Business Intelligence Analyst
- Expertise in database concepts, object and data modeling techniques and design principles
- Expertise in database architectures, software, and facilities
- Expertise with programming languages - Python (required)
- Expertise with programming languages - Scala
- Expertise with programming languages - Ruby
- Expertise with programming languages - R
- Expertise with database technologies - SQL
- Expertise with database technologies - performance tuning concepts
- Expertise with database technologies - AWS RDS
- Expertise with database technologies - RedShift
- Expertise with database technologies - MySQL
- Expertise with big data batch processing tools: Hadoop MapReduce
- Expertise with big data batch processing tools: ElasticSearch
- Expertise with big data batch processing tools: PIG
- Expertise with big data batch processing tools: Hive
- Expertise with big data batch processing tools: Cascading/Scalding
- Expertise with big data batch processing tools: Apache Spark
- Expertise with big data batch processing tools: AWS EMR
- Expertise with stream-processing systems: Kinesis
- Expertise with stream-processing systems: Kafka
- Expertise with stream-processing systems: MQTT
- Expertise with relational NoSQL databases including DyanamoDB
- Expertise in writing JSON, XML, YAML and other data definition schemas
- Ability to work on advanced complex technical projects or business issues requiring state of the art technical knowledge or industry
- Ability to work on significant and unique issues where analysis of situations or data requires an evaluation of intangibles
- Exercises independent judgment in methods, techniques and evaluation criteria for obtaining results
- Ability to lead and mentor junior engineers and colleagues
Desired Qualifications
- Related AWS certification, preferred
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.