U.S. Pharmacopeia logo
U.S. PharmacopeiaPosted 1 month ago

Lead Data Engineer

On-siteHyderabad, Telangana, India

Full TimeSenior LevelBachelors DegreeLarge

Job Summary

Design and implement scalable data collection, storage, and processing pipelines to support enterprise-wide data needs while maintaining data governance frameworks and quality checks for compliance. Build and optimize data models and marts to enable self-service analytics with Tableau, Looker, and Power BI, then partner with data scientists to operationalize models into production-grade pipelines. Collaborate with cross-functional stakeholders to translate business requirements into scalable solutions, lead design and code reviews, and influence the data roadmap through automation and cost optimization. Provide technical leadership for data platform design, ensuring alignment with enterprise standards and long-term strategy.

Required Qualifications

  • Bachelor's degree in relevant field (e.g. Engineering, Analytics or Data Science, Computer Science, Statistics) or equivalent experience
  • 7+ years of experience in big data technologies such as Python, PySpark, and SQL for processing structured, semi-structured, and unstructured data
  • Strong experience with AWS data services including Redshift, S3, Glue, Lambda, EventBridge, Postgres, Neo4j (Azure/GCP equivalents such as ADLS, Synapse, ADF acceptable)
  • Experience in building batch, micro-batch, and streaming pipelines (real-time / near real-time) using Lambda/Kappa architectures
  • Hands-on expertise in designing and delivering enterprise-scale data platforms, including data lakehouse, data warehouse, data lake, and data marts
  • Strong understanding and hands-on implementation of data modeling techniques including - Data Vault 2.0, Dimensional Modeling, Knowledge Graphs, One Big Table (OBT) approaches
  • Experience with medallion architecture and metadata-driven data pipeline frameworks
  • Strong expertise in data governance frameworks, including- Data discovery, Data quality, Data security, Hands-on experience with DQ tools such as Great Expectations, Pydantic, etc
  • Strong SQL and programming skills for data transformation, modeling, and analysis
  • Hands-on experience in building and maintaining complex ETL/ELT pipelines and managing day-to-day data operations
  • Experience with workflow orchestration tools such as Airflow (2+ years or equivalent)
  • Knowledge of streaming/event-driven architectures and modern data processing patterns
  • Good understanding of dashboarding and visualization techniques with tools like Tableau, Power BI, or equivalent
  • Experience with Agile (SAFe) methodologies, CI/CD pipelines, and modern deployment practices for data platforms
  • Ability to explain complex technical issues to a non-technical audience

Desired Qualifications

  • Certification in at least one area of data modeling techniques (Data Vault 2.0, Dimensional Modeling, Knowledge Graphs, One Big Table)
  • Experience with scientific chemistry nomenclature or prior work experience in life sciences, chemistry, or hard sciences or degree in sciences
  • Experience with pharmaceutical datasets and nomenclature
  • Experience in developing Machine learning & Deep learning models; Familiarity with MLOps and deploying ML models in production environments
  • Exposure to AI/ML concepts, with familiarity in Generative AI patterns (e.g., RAG, chunking techniques) as an added advantage

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce