Data Engineer focused on scalable data platforms, reliable pipelines, and data products for analytics and operational use cases. Experienced with Python, PySpark, SQL, AWS, Terraform, CI/CD, data quality, and and event-driven data workflows.
Real-time data pipeline with Kafka, Flink, Iceberg, Trino, MinIO, and Superset.
Medallion batch data pipeline with Airflow, DuckDB, Delta Lake, Trino, MinIO, and Metabase. Full observability and data quality.
Docker containerized and configurable Airflow data pipeline for collecting and storing stock and cryptocurrency market data.
ETL pipeline using Pulumi for infrastructure as code, integrating AWS services and Snowflake for automated data flow.
Async Python microservices with FastAPI, CDC, real-time WebSocket, and gRPC for high-performance document management.
Improved the library’s data manipulation and reporting functionalities, supporting the development of efficient data pipelines and enabling scalable data solutions for analytics and modeling.