Data Engineer with experience in building and operating large-scale data platforms, developing batch and real-time data pipelines, and automating data processing workflows.
Experienced in building and operating data platforms based on Spark, Hadoop, Kafka, and Airflow. More recently, I have worked on improving Kubernetes-based Airflow environments, building and validating CDC-based real-time data pipelines, and processing streaming data with Flink.
My experience covers not only data processing logic but also the broader operation of data platforms, including Kubernetes, Docker, GitLab CI/CD, monitoring, and incident response. I focus on identifying performance bottlenecks and diagnosing system issues based on data and operational metrics to improve platform stability and efficiency.
- Design and operation of large-scale distributed data platforms
- Development of batch and real-time streaming data pipelines
- Real-time data processing with Kafka and Flink
- Airflow-based workflow automation and operational optimization
- Building and operating Kubernetes-based data platforms
- Data pipeline performance analysis and bottleneck optimization
- Monitoring, incident detection, and response systems
- GitLab CI/CD and operational automation
- Contributions to Airflow Provider testing and permission policy improvements


