Lead Data Software Engineer
Remote in Argentina, & 4 others
Data Software Engineering& 11 others
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobChoose an option
We are seeking a Lead Data Software Engineer to design reusable data-sharing adapters across a cloud lakehouse and external analytics platforms with governed access. You will partner with engineers to deliver modular integrations and dependable pipelines, while showcasing AI-assisted development in day-to-day work.
Responsibilities
- Architect a UniForm lakehouse write layer with dual-format metadata (Delta and Iceberg) to serve multiple consumers
- Create and validate GCS-to-BigQuery ingestion pipeline patterns for structured operational datasets
- Deliver CDC pipelines with Kafka to enable real-time and near-real-time lakehouse updates
- Establish dependency-aware bookkeeping and data lineage tracking patterns across data pipelines
- Develop modular, version-controlled adapter code that can be reused for new data source integrations
- Set up Snowflake external table definitions and enable governed access using Horizon catalog metadata
- Build and certify Delta Sharing adapters to support zero-copy data sharing for Databricks consumers
- Configure Delta Sharing endpoints, registrations, and sharing agreement management
- Test end-to-end freshness, sharing latency, and SLA compliance for external data sharing flows
- Create connector registry entries, RBAC, and tenant-scoped authorization for all data-out paths
- Add metering hooks aligned with billing requirements for governed external data flows
- Document integration patterns and operational procedures for reuse across additional data products
Requirements
- Proven data software engineering experience (5+ years) using Python to build and maintain data pipelines
- Hands-on experience with Google Cloud BigQuery, including advanced usage and performance optimization
- Solid background with data lakehouse table formats, including Apache Iceberg and Delta Lake
- Demonstrated experience integrating Databricks and applying governed data access patterns
- Practical knowledge of Snowflake, including external table access to lakehouse data
- Strong understanding of Kafka and CDC patterns for real-time and near-real-time ingestion
- Deep architecture skills in data lake design, modular adapter development, and version control practices
- Working proficiency with AI-assisted development tools such as Claude Code, GitHub Copilot, or Cursor
- Excellent documentation skills for reusable integration patterns and operational runbooks
- Upper-Intermediate English proficiency (B2) for technical collaboration and written communication
Nice to have
- Apache Spark experience to validate shared reads and pipeline patterns
- Databricks Unity Catalog experience for governed metadata and access control
- Delta Lake expertise in sharing and interoperability patterns
- Gen AI Assisted Development experience with measurable workflow improvements
- Snowflake Horizon Catalog experience for metadata governance and access controls
