Data Engineer - Tech Lead (Databricks, Pyspark)
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobWe're looking for a Senior Data Engineer – Tech Lead (Databricks, PySpark) to join our team in London, UK, in a hybrid working mode.
In this role, you will lead the design, development and optimization of scalable cloud-native data architectures, focusing on Azure Databricks, PySpark and Lakehouse principles. You will work hands-on to deliver performant data solutions for high-volume workloads, ensuring governance, reliability and best practices for enterprise-grade platforms.
As a technical leader, you will define data strategies, drive modernization initiatives and mentor engineers, fostering excellence and innovation throughout the team. This position offers the opportunity to shape large-scale data ecosystems, implement modern engineering practices and enable next-generation analytics and AI-driven solutions.
- Lead the architecture, design and build of large-scale data platforms using Azure Databricks and modern cloud technologies
- Implement and optimize ETL workflows and streaming pipelines with PySpark and Delta Live Tables following Lakehouse principles
- Enhance performance, manage cloud costs and ensure platform reliability for structured streaming workloads
- Define data governance, security and quality standards to maintain consistency across the platform
- Collaborate with stakeholders to translate complex business requirements into actionable technical solutions
- Develop integration approaches using Azure-native services such as Data Factory, Synapse and Blob Storage
- Mentor data engineers, promote modern engineering practices and perform technical reviews
- Drive adoption of CI/CD, Infrastructure as Code and automated testing in data engineering environments
- Implement observability and monitoring using tools like Databricks Workflows and related frameworks
- Contribute to AI-driven initiatives by leveraging Databricks ML/MosaicML to integrate Generative AI and LLM-based solutions
- Bachelor’s or Master’s degree in Computer Science, Software Engineering or related field
- Extensive experience designing and implementing production-grade platforms using Azure Databricks
- Expertise in PySpark, including advanced optimization, data skew mitigation and query tuning
- Strong programming skills in Python with knowledge of modern software design principles
- Practical experience with structured streaming, Delta Lake and Delta Live Tables
- Proven experience in Lakehouse migration and modernization using open table formats such as Delta Lake or Apache Iceberg
- Proficiency with cloud-native services on Azure and knowledge of multi-cloud environments (AWS or GCP)
- Hands-on experience with CI/CD and Infrastructure as Code tools (Terraform, GitHub Actions, Jenkins)
- Strong leadership ability to guide teams, define epics/user stories and ensure delivery in agile environments
- Excellent communication and stakeholder management skills for both technical and non-technical audiences
- Experience operationalizing LLM or Generative AI workflows in Databricks pipelines
- Familiarity with frameworks like LangChain, LlamaIndex or Databricks ML/MosaicML
- Knowledge of AI governance, security practices and enterprise integration controls
- Background in financial trading data or related domains
- Official Databricks certifications such as Certified Data Engineer Professional or Apache Spark Developer
