Senior Data Analyst (Graph focused), Fixed Term - Toronto
$145,600–$145,600 year
RemoteToronto, Ohio, United States or Ontario, California, United States
Job Summary
Validate datasets in Databricks and CDS to confirm completeness, accuracy, and structural integrity relative to source data. Verify benchmark query outputs across technologies to ensure consistent results regardless of the execution system. Identify, document, and trace the root cause of data discrepancies, distinguishing between ETL issues, schema errors, and upstream quality problems. Develop validation test cases and expected outputs for compliance use cases while assessing whether they require graph database traversal or conventional approaches. Apply statistical methods to evaluate dataset representativeness, sampling quality, and measurement reliability across benchmark runs. Produce summary statistics and data quality reports that inform the team's architecture assessment and document findings for engineering action.
Required Qualifications
- Demonstrated experience validating large, complex datasets — identifying discrepancies, tracing root causes, and documenting findings clearly
- Strong SQL skills; ability to write analytical queries against relational databases
- Experience working with data at significant scale — hundreds of millions of records — where manual spot-checking is insufficient and systematic validation approaches are required
- Familiarity with ETL pipelines and the types of data quality issues that arise in data loading and transformation
- Practical experience applying statistical methods to data quality assessment: distribution analysis, outlier detection, variance analysis, sampling validation
- Ability to interpret benchmark result data and distinguish meaningful performance differences from noise
- Comfort working across multiple database technologies and query languages — this role will need to query data in PostgreSQL, graph databases, and Databricks as part of normal validation work
- Experience with Databricks or similar distributed data platforms (Spark, Delta Lake)
- Strong written communication — validation findings, defect reports, and root cause analyses must be clear enough for both engineers and product stakeholders
- Ability to work independently under minimal supervision, taking direction from peers rather than requiring structured management oversight
- Experience embedded in a cross-functional engineering team
- Validated datasets loaded into each technology to confirm completeness, accuracy, and structural integrity relative to source data in Databricks and CDS
- Verify benchmark query outputs across technologies — confirm that the same logical query against the same underlying data produces consistent, correct results regardless of which system executes it
- Identify, document, and trace the root cause of data discrepancies and defects discovered during validation; distinguish between ETL issues, schema translation errors, technology-specific behavior, and upstream data quality problems
- Develop and maintain validation test cases and expected outputs for benchmark queries and compliance use cases
- Support the collection and documentation of compliance use cases from the business unit
- Help assess each use case: can it be fulfilled by a conventional database approach, or does it require the traversal and pattern-matching capabilities of a purpose-built graph database?
- Contribute analytical rigor to use case triage — this is a cost-and-complexity decision as much as a technical one
- Apply statistical methods to evaluate dataset representativeness, sampling quality, and measurement reliability across benchmark runs
- Analyze benchmark result distributions — identify outliers, assess variance across cold/warm/concurrent runs, and flag results that require deeper investigation before scoring
- Produce summary statistics and data quality reports that inform the team's architecture assessment
- Document validation findings, defect reports, and root cause analyses in a format the engineering team can act on
- Maintain a running record of known data issues and their resolution status across each technology under evaluation
- Must be able to lift 50 lbs
Desired Qualifications
- PostgreSQL experience preferred
- Experience embedded in a cross-functional engineering team
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.