Lifelancer logo
LifelancerPosted 2 weeks ago

Senior AI Data Engineer

On-siteBudapest, Budapest, Hungary

Full TimeSenior LevelSmall

Job Summary

Design and maintain scalable data pipelines and ETL processes supporting AI research and development, including autonomous AI agents. Architect, implement, and maintain robust data models for analytics, ML training, and agent workflows while collaborating with ML engineers and AI scientists. Establish observability, lineage tracking, and monitoring frameworks to detect anomalies and ensure data quality. Drive data platform reliability, scalability, and cost optimization across cloud-based infrastructure using modern orchestration frameworks and distributed computing tools. Monitor and troubleshoot pipeline issues to ensure continuity, implementing automated validation and governance measures. Stay current with emerging data engineering and AI technologies to shape data strategy for next-generation systems.

Required Qualifications

  • Bachelor's or Master's degree in Computer Science, Data Engineering, or a related field
  • 5+ years of professional experience in data engineering
  • at least 2 years focused on ML/AI data infrastructure
  • Advanced proficiency in Python
  • Advanced proficiency in Scala
  • experience with Rust, Go, Java, or Julia
  • Expert-level knowledge of SQL
  • Expert-level knowledge of NoSQL databases
  • Hands-on experience with vector databases (e.g., Pinecone, Weaviate, Milvus)
  • Proficiency with modern data orchestration platforms (e.g., Airflow 2.x)
  • Extensive experience with at least one major cloud platform (AWS, Azure, or GCP)
  • Expertise in containerization and orchestration (Docker, Kubernetes)
  • Experience with Infrastructure as Code tooling (e.g., Terraform)
  • Experience with distributed computing frameworks (Spark, Dask, Ray)
  • Proficiency with streaming technologies (Kafka, Flink)
  • Knowledge of modern data lakehouse architectures

Desired Qualifications

  • advanced degree
  • Certifications in cloud platforms, big data technologies, engineering, or ML operations
  • Experience collaborating with ML engineers on CI/CD pipelines for data processing and model deployment
  • Working knowledge of ML frameworks (PyTorch, TensorFlow)
  • Experience with feature stores and experiment‐tracking platforms
  • Understanding of LLM fine‐tuning data requirements and processing
  • Experience developing data systems for autonomous AI agents or agentic AI applications
  • Background in prompt engineering or retrieval‐augmented generation systems
  • Experience with semantic caching and efficient storage/retrieval of AI‐generated artifacts
  • Familiarity with LLM evaluation metrics and benchmarking frameworks
  • Expertise in hybrid architectures combining traditional databases with vector stores
  • Experience with RAG systems and related data pipelines
  • Knowledge of RLHF data workflows
  • Experience mentoring junior engineers, establishing best practices, and contributing to architectural decisions

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce