Skip To Main Content
backBack to Search

Lead Data Engineer

Hybrid in Argentina, Brazil
Data Integration& 5 others
Looking for something else?

Find a vacancy that works for you. Send us your CV to receive a personalized offer.

Find me a job

We're seeking a Lead Data Engineer to come aboard and strengthen our team. Among all engineering positions tied to this engagement, this one carries the heaviest data responsibility. Should any pipeline lose data, positions shift out of alignment, or signals miscalculate, the entire calibration and UAT process collapses. Nothing else can function properly unless accuracy and operational resilience are locked in here first.

Responsibilities
  • Carry out every database migration set across the Unified Data Store so that schema alignment remains intact throughout the platform
  • Construct and sustain the complete collection of Source Adapters that link outside systems into the broader data platform
  • Build out the Field Mapper, setting up field bindings specific to each project and franchise that standardize source system data into the unified schema upon intake
  • Develop the Link Resolver, tasked with connecting CL-to-ticket, ticket-to-ticket, and case-to-defect/requirement relationships spanning every source family, supplying input directly to the Attribution Resolver
  • Construct the Attribution Resolver to map the change-ticket-area-test pathway, together with the Counted Signals Aggregator, which generates file-to-area and area-to-area association tables complete with path standardization and count validation
  • Assemble the Area Vocabulary along with its related Translation Tables to maintain uniform area categorization throughout the system
  • Create the complete set of Signal Catalogue computation jobs addressing all eight primary signals — area fragility, recency, recent failures, change-touch, coupling, windowed area change volume, testing alignment, validation recency, and defect impact/volume/age — with provenance tracking built into each
  • Take charge of writing the data dictionary for every storage element, updating it progressively as new functionality rolls out
Requirements
  • Five-plus years of direct, practical experience in Python software development
  • A minimum of one year spent leading or managing development teams
  • Prior exposure to AI Data Engineering, applying relevant practices to support systems driven by AI/ML
  • Real-world experience leveraging Apache Airflow to coordinate and schedule data-related workflows
  • Functional understanding of Machine Learning principles and how they integrate into data systems
  • Demonstrated capability in building data pipelines from initial design through full deployment
  • Direct experience working with PostgreSQL for both storing and querying data
  • Background creating idempotent ingestion mechanisms, covering natural keys, upsert auditing, duplicate identification, and safe replay handling
  • Experience bringing together various heterogeneous data sources into one cohesive system
  • Competency in data quality and telemetry approaches, including fill-rate tracking, volume measurement, and reconciliation reporting
  • Exposure to the Spec Driven Development approach
  • Reasonable communication skills combined with working English fluency (B2 level or above) to interpret business needs and convert them into agentic architectures
Nice to have
  • Background building API clients for platforms like Perforce or Code Hub
  • Experience connecting with the JaaS (Jira) REST API
  • Exposure to test management APIs such as QMetry, Zephyr, or comparable solutions
  • Background developing Snowflake connectors, covering key-pair authentication and warehouse extract querying
  • Understanding of temporal data systems along with as-of read patterns