Data Engineer, Gen AI
RemoteUnited States
Job Summary
Design and build scalable data pipelines to ingest, process, and store large volumes of training data for generative AI models. Implement data preprocessing and feature engineering workflows using Spark and Delta Lake to prepare data for model training and inference. Develop and maintain data quality checks and monitoring systems to ensure data integrity, while optimizing infrastructure for high-performance AI model serving in production environments. Collaborate with ML engineers and data scientists to refine data flows, and implement data governance and security best practices for sensitive training data. Troubleshoot data-related issues in AI pipelines and deploy solutions. Requires a Bachelor's or Master's degree, 3+ years of experience, and strong Python/SQL skills with Databricks, Apache Spark, and Delta Lake. Familiarity with AWS, GCP, or Azure is essential.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Data Science, or related field
- 3+ years of experience in data engineering roles
- Strong programming skills in Python and SQL
- Experience with Databricks, Apache Spark, and Delta Lake
- Familiarity with cloud platforms (AWS, GCP, or Azure) and their data services
- Knowledge of data modeling, ETL processes, and data pipeline architectures
- Experience with version control systems (e.g. Git) and CI/CD practices
- Understanding of data privacy and security considerations
- Experience with Delta Live Tables for building reliable, maintainable data pipelines
- Familiarity with Databricks SQL for querying and analyzing large datasets
- Knowledge of Unity Catalog for data governance and access control
- Strong problem-solving and analytical skills
- Excellent communication and collaboration abilities
Desired Qualifications
- Experience supporting machine learning or AI projects in production environments
- Familiarity with containerization and orchestration tools like Docker and Kubernetes
- Knowledge of streaming data technologies like Kafka or Kinesis
- Experience with MLOps practices and tools
- Understanding of large language models and generative AI architectures
- Ability to work in a fast-paced, dynamic environment
- Passion for staying up to date with the latest developments in AI and data technologies
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.