Spark / Data Lakehouse Developer

Więcej niż jeden projekt. Więcej niż jeden rynek.

Jesteśmy częścią Hiberus, firmy technologicznej wywodzącej się z Hiszpanii z silną pozycją na rynku polskim: 4700+ specjalistów, obecnej w 14+ krajach i realizującej projekty dla klientów na całym świecie.

Pracujemy z technologiami AI, Data, Cloud, Software Development, BI i Cybersecurity. Ale przede wszystkim stawiamy na ludzi – gdy jeden projekt się kończy, szukamy kolejnych możliwości w ramach Hiberus, również w innych krajach.

Naszą wartość dodatkową stanowi Hiberus University - międzynarodowe środowisko i realna możliwość rozwijania swoich kompetencji. Ponad tysiąc osób przeszkolonych w zeszłym roku z wybranych technologii.

Pierwszy projekt to początek. Z Hiberus możesz budować swoją karierę długoterminowo. 

About the company

More Than One Project. More Than One Market.

We are part of Hiberus, a technology company founded in Spain with a strong presence in the Polish market: 4,700+ specialists, operating in 14+ countries and delivering projects for clients worldwide.

We work with technologies including AI, Data, Cloud, Software Development, BI, and Cybersecurity. But above all, we put people first. When one project comes to an end, we look for new opportunities within Hiberus — including projects in other countries.

One of our key advantages is Hiberus University — an international learning environment offering real opportunities to develop and expand your skills. More than 1,000 people were trained last year in selected technologies.

Your first project is just the beginning. With Hiberus, you can build a long-term career.

Opis stanowiska

Dla naszego Klienta z sektora finansowego poszukujemy doświadczonego Spark / Data Lakehouse Developera, który dołączy do zespołu Data Platform.

Osoba na tym stanowisku będzie odpowiedzialna za projektowanie i rozwój rozwiązań do rekoncyliacji i synchronizacji danych w warstwie ODS (Operational Data Store), opartych na MongoDB, a także za budowę i utrzymanie warstw danych w architekturze Data Lakehouse – od surowych danych po warstwy Curated, Silver i Gold.

 

Główne obowiązki : 

  • Projektowanie i implementacja procesów Spark SQL/PySpark do walidacji i porównywania danych (MongoDB/ODS).
  • Implementacja reguł jakości danych, detekcji duplikatów, anomalii i braków oraz raportowania/alertowania.
  • Integracja i optymalizacja Spark–MongoDB (MongoDB Spark Connector, partitioning, push-down, schema-aware reads).
  • Projektowanie warstwy danych Raw → Bronze → Silver → Gold z wykorzystaniem Delta Lake/Iceberg/Hudi oraz CI/CD dla schematów i migracji.
  • Monitoring i observability procesów Spark/batch/stream (Prometheus, Grafana, OpenTelemetry, Dynatrace) oraz definiowanie SLA/SLO.
  • Automatyzacja deploymentów Spark na Docker/Kubernetes (Spark Operator).
  • Diagnozowanie problemów z danymi, wydajnością i opóźnieniami procesów.
  • Tworzenie dokumentacji technicznej, runbooków i diagramów przepływu danych.
  • Współpraca w Scrum/Kanban z Data Engineerami, BA, testerami i DevOps; prowadzenie warsztatów i code review.

Job description

For our client in the financial sector, we are looking for an experienced Spark / Data Lakehouse Developer to join the Data Platform team.

The person in this role will be responsible for designing and developing data reconciliation and synchronization solutions within the ODS (Operational Data Store) layer, based on MongoDB, as well as building and maintaining data layers within the Data Lakehouse architecture — from raw data through Curated, Silver, and Gold layers.

Key Responsibilities

  • Design and implementation of Spark SQL/PySpark processes for data validation and comparison (MongoDB/ODS).
  • Implementation of data quality rules, duplicate, anomaly, and missing data detection, as well as reporting and alerting.
  • Integration and optimization of Spark–MongoDB workflows (MongoDB Spark Connector, partitioning, push-down, schema-aware reads).
  • Design of Raw → Bronze → Silver → Gold data layers using Delta Lake/Iceberg/Hudi, including CI/CD for schemas and migrations.
  • Monitoring and observability of Spark/batch/streaming processes (Prometheus, Grafana, OpenTelemetry, Dynatrace) and definition of SLAs/SLOs.
  • Automation of Spark application deployments using Docker/Kubernetes (Spark Operator).
  • Troubleshooting data quality, performance, and processing latency issues.
  • Creation and maintenance of technical documentation, runbooks, and data flow diagrams.
  • Collaboration in Scrum/Kanban teams with Data Engineers, Business Analysts, QA, and DevOps; conducting workshops and code reviews.

Benefity

  • Karta Multisport
  • Prywatna opieka medyczna

Benefits

  • Multisport card
  • Private medical care

Wymagane doświadczenie

  • 4+ lata doświadczenia w pracy z Apache Spark w środowisku produkcyjnym.
  • Bardzo dobra znajomość PySpark, Spark SQL, DataFrames i Structured Streaming.
  • Doświadczenie z MongoDB i Spark Connector oraz znajomość modelowania i indeksowania danych.
  • Praktyczna znajomość Apache Iceberg i architektury Data Lakehouse (upsert/merge, optymalizacja zapisów).
  • Doświadczenie w CI/CD z Jenkins, w tym automatyzacji testów i deploymentów aplikacji Spark.
  • Bardzo dobra znajomość Python oraz Spark SQL.
  • Znajomość Prometheus i Grafana w obszarze monitoringu i observability.
  • Doświadczenie w pracy w Scrum/Kanban oraz znajomość Jira i Confluence.

Experience required

  • 4+ years of experience working with Apache Spark in production environments.
  • Very good knowledge of PySpark, Spark SQL, DataFrames, and Structured Streaming.
  • Experience with MongoDB and the Spark Connector, with a good understanding of data modeling and indexing.
  • Hands-on experience with Apache Iceberg and Data Lakehouse architecture, including upsert/merge operations and write optimization.
  • Experience with CI/CD using Jenkins, including automation of testing and deployment of Spark applications.
  • Very good knowledge of Python and Spark SQL.
  • Knowledge of Prometheus and Grafana for monitoring and observability.
  • Experience working in Scrum/Kanban environments, with practical knowledge of Jira and Confluence.
ID: 279 job_post.published_on: 01/09/2026
announcement.apply