Senior Data Software Engineer, Spark, Java, Scala
Remote in Georgia, & 4 others
Data Software Engineering
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobChoose an option
We are seeking a Senior Software Data Engineer to join our team in a software engineering capacity. This is not a data-science or analytics position centered on ad-hoc data exploration; instead, the role focuses on building software, data processing jobs, and data pipelines consumed by internal and external partners. Data is our main product and first-class citizen, and we value correct, high-quality data as much as clean and maintainable code.
Responsibilities
- Write new data pipelines and jobs to produce new outputs (datasets) in scope of new features development
- Adopt existing data pipelines to integrate with new org-wide platforms, tools, services, and languages
- Fix bugs in code and correct data caused by incorrect logic or implementation
- Perform ad-hoc data exploration, validation, and investigation to help select the right tech design and support Product Management team decisions
- Monitor and troubleshoot production issues with pipelines owned by the team
- Develop and adopt data quality checks to monitor data issues in the systems
- Scope and plan new development, including assessing level of effort and providing timelines
- Maintain tickets hygiene in Radar (ticketing system)
- Evolve jobs, apps, and systems to a better state across all aspects: code quality, complexity, maintainability, and documentation
- Communicate with other data engineers in the team, peer teams (QA, UAT, Platform, etc), project managers, and engineering managers on status, blockers, estimates, and timelines
Requirements
- 3+ years of hands-on experience in the big-data field, including Hadoop (HDFS, YARN or Mesos) and Spark
- Excellent knowledge and hands-on experience of SQL in context of Big Data: Spark SQL, HiveQL
- Excellent knowledge of Spark, including ability to understand and optimize Spark execution plans via Spark UI, with upcoming migration to Spark 3
- Excellent knowledge of Scala or Java
- Understanding of batch processing and ETL principles in Data Warehouses
- Familiarity with data completeness signals and orchestration
- Knowledge of approaches for historical reprocessing and data correction
- Skills in handling bad data and late data in inputs and outputs
- Understanding of schema migrations and datasets evolution
- Strong speaking English, with ability to rely on information heard verbally in meetings and to explain own ideas clearly to native speakers
- Capability to learn fast new set of tools and technology used internally at the company: platform services, telemetry providers, Spark-as-a-Service, build system, and more
Nice to have
- Understanding of functional programming ideas and principles
- Experience in building and using web services
- Familiarity with any of Teradata, Vertica, Oracle, Tableau
- Skills in Spark Streaming and Kafka
- Knowledge of Apache Iceberg, Trino (Presto), Druid, Cassandra, or Blob storage like AWS
- Experience with Splunk
- Experience with Snowflake
