Tuyển dụng Middle Data Engineer Apache Spark, Trino tại Hà Nội
Chưa có CV? Tạo CV miễn phí ở đây →
Mô tả công việc Middle Data Engineer Apache Spark, Trino
Role Overview
We are seeking a Middle Data Engineer to build and scale our Production Datahouse on-premise environment. This is a purely technical, hands-on role focused on implementing a high-performance, on-premise data architecture.
You will own the end-to-end technical execution, managing the entire lifecycle of data as it moves from landing ingestion to the final serving layer. This role is not about high-level theory; it is about the hard engineering required to optimize on-premise hardware, tune distributed processing engines, and ensure a seamless data flow across our internal ecosystem.
Key Responsibilities
- Data Pipeline Engineering: Design, build, and maintain automated data ingestion pipelines from several source systems using Airbyte and Airflow, ensuring reliability and scalability.
- Lakehouse Architecture Implementation: Develop and manage the Medallion Architecture (Bronze/Silver/Gold) using dbt-spark to enable structured, high-quality analytical datasets.
- Platform Performance Optimization: Optimize Spark and Trino workloads on Kubernetes and manage the full lifecycle of Apache Iceberg tables (compaction, snapshotting, and maintenance) to improve query performance, storage and reduce latency.
- Infrastructure & Reliability Management: Proactively monitor and maintain Data Platform resources using Grafana and Prometheus; analyze query patterns to enhance system stability and performance for production users.
- Data Security & Governance: Enforce row- and column-level security using Open Policy Agent (OPA) and develop foundational frameworks for Data Governance and Data Quality, embedding lineage, automated testing, and compliance directly into pipelines.
- Enterprise Data Integration: Architect and implement high-performance data bridges to deliver on-premise Lakehouse data into Microsoft Fabric, enabling seamless and secure hybrid-cloud connectivity.
- 3+ years of hands-on experience in Data Engineering or Platform Engineering, building and operating production-grade data platforms.
- Strong expertise in distributed processing with Apache Spark and Trino, including performance tuning and optimization.
- Proven experience designing and maintaining reliable data pipelines using Airbyte and Apache Airflow.
- Solid understanding of modern Lakehouse architectures, dbt-based transformations, and Iceberg table management.
- Hands-on experience with Kubernetes, containerized deployments, and CI/CD for data workloads.
- Knowledge of data security, governance, and access control practices (row/column-level security, policy enforcement, IAM).
- Experience supporting BI/analytics use cases and integrating with enterprise reporting tools (e.g., Microsoft Fabric, Power BI).
- Strong troubleshooting, system reliability, and problem-solving skills in complex distributed environments.
- Ability to work independently in a highly technical, ownership-driven role.
**easy going and friendly environment
**Remuneration Package:
• Working hours: 5 days per week (in a professional and yet young, dynamic environment);
• Salary: Competitive remuneration package (based on skills and experience);
• 12 leave days per year
Nhà tuyển dụng IMIP Technology And Solution Consultancy
IMIP Technology And Solution Consultancy · 📍 Hà Nội
Việc khác tại IMIP Technology And Solution Consultancy
- I
Middle Data Engineer Apache Spark, Trinohôm nay · Còn 35 ngày