Skip To Main Content
backBack to Search

Senior Data Software Engineer, Spark, Java, Scala

Remote in Georgia, & 4 others
Data Software Engineering
hot
Looking for something else?

Find a vacancy that works for you. Send us your CV to receive a personalized offer.

Find me a job

We are seeking a Senior Software Data Engineer to join our team in a software engineering capacity. This is not a data-science or analytics position centered on ad-hoc data exploration; instead, the role focuses on building software, data processing jobs, and data pipelines consumed by internal and external partners. Data is our main product and first-class citizen, and we value correct, high-quality data as much as clean and maintainable code.

Responsibilities
  • Write new data pipelines and jobs to produce new outputs (datasets) in scope of new features development
  • Adopt existing data pipelines to integrate with new org-wide platforms, tools, services, and languages
  • Fix bugs in code and correct data caused by incorrect logic or implementation
  • Perform ad-hoc data exploration, validation, and investigation to help select the right tech design and support Product Management team decisions
  • Monitor and troubleshoot production issues with pipelines owned by the team
  • Develop and adopt data quality checks to monitor data issues in the systems
  • Scope and plan new development, including assessing level of effort and providing timelines
  • Maintain tickets hygiene in Radar (ticketing system)
  • Evolve jobs, apps, and systems to a better state across all aspects: code quality, complexity, maintainability, and documentation
  • Communicate with other data engineers in the team, peer teams (QA, UAT, Platform, etc), project managers, and engineering managers on status, blockers, estimates, and timelines
Requirements
  • 3+ years of hands-on experience in the big-data field, including Hadoop (HDFS, YARN or Mesos) and Spark
  • Excellent knowledge and hands-on experience of SQL in context of Big Data: Spark SQL, HiveQL
  • Excellent knowledge of Spark, including ability to understand and optimize Spark execution plans via Spark UI, with upcoming migration to Spark 3
  • Excellent knowledge of Scala or Java
  • Understanding of batch processing and ETL principles in Data Warehouses
  • Familiarity with data completeness signals and orchestration
  • Knowledge of approaches for historical reprocessing and data correction
  • Skills in handling bad data and late data in inputs and outputs
  • Understanding of schema migrations and datasets evolution
  • Strong speaking English, with ability to rely on information heard verbally in meetings and to explain own ideas clearly to native speakers
  • Capability to learn fast new set of tools and technology used internally at the company: platform services, telemetry providers, Spark-as-a-Service, build system, and more
Nice to have
  • Understanding of functional programming ideas and principles
  • Experience in building and using web services
  • Familiarity with any of Teradata, Vertica, Oracle, Tableau
  • Skills in Spark Streaming and Kafka
  • Knowledge of Apache Iceberg, Trino (Presto), Druid, Cassandra, or Blob storage like AWS
  • Experience with Splunk
  • Experience with Snowflake