Abstra logo
AbstraPosted 1 week ago

Text Analytics Engineer (Conversational Data Pipelines & Analytics)

RemoteMexico

Full TimeBachelors DegreeSmall

Job Summary

Design, build, and maintain scalable data pipelines for processing, sampling, and analyzing customer conversations and unstructured text data. Develop ingestion, transformation, enrichment, and processing workflows within Athena Studio to support conversational insights, prompt training data preparation, and downstream AI applications. Implement sampling methodologies for high-volume data analysis and establish validation, monitoring, and quality-control checks for production reliability. Collaborate with global Data Science, Engineering, and Product teams to deliver robust data flows while documenting pipeline logic for technical stakeholders. Work with structured and unstructured datasets to ensure performance and maintainability.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, Data Science or related technical discipline, or equivalent practical experience
  • 5+ years of relevant work experience in Data Engineering, Analytics Engineering, Text Analytics, Data Science or related technical areas
  • Hands-on experience building and supporting data pipelines and ETL/ELT workflows
  • Demonstrated proficiency in Python and SQL
  • Experience working with large-scale structured and unstructured datasets
  • Experience processing text data, conversational data, customer interaction data or similar unstructured data sources
  • Working knowledge of data quality, validation, monitoring and troubleshooting practices for data pipelines
  • Ability to collaborate effectively with cross-functional and geographically distributed teams
  • Hardware and software setup (mandatory)

Desired Qualifications

  • Experience with Athena Studio and AWS Athena for analytics or data processing workflows
  • Experience with AWS data services such as S3, Glue, Lambda, EMR, SageMaker or similar cloud-based platforms
  • Experience with Spark, PySpark, Databricks or distributed data processing technologies
  • Experience developing sampling methodologies for analytics, NLP, prompt training or machine learning data preparation
  • Familiarity with Natural Language Processing, Text Analytics, Conversational Analytics or customer experience analytics domains
  • Familiarity with Generative AI, Large Language Models, prompt engineering concepts or LLM data preparation workflows
  • Strong debugging, performance optimisation and data pipeline reliability skills
  • Strong communication skills, especially in describing data workflows, pipeline behaviour and technical trade-offs to technical and non-technical audiences
  • Experience working with teams across India, Argentina or other global delivery locations

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce