Profile

Data Engineer focused on building scalable data platforms, reliable pipelines, and automation tools across cloud-native environments. Experienced with Python, SQL, AWS, Terraform, Kubernetes, CI/CD, data quality, and software engineering practices for maintainable data systems.

Professional Experience
Dec 2025 – Present
Daimler Truck⁠, Data Engineer
  • •Built and maintained data pipelines and data products on Palantir Foundry, integrating vehicle, warranty, production, and quality data into reliable, analytics-ready datasets.
  • •Developed data transformations using Python and PySpark, supporting scalable ingestion and processing of multiple enterprise data sources.
  • •Designed and maintained data models (Foundry Ontology), mapping business entities and relationships to support consistent, reusable analytics across teams.
  • •Implemented data validation and quality checks within pipelines to ensure reliability, consistency, and trust in downstream datasets.
  • •Supported access controls, documentation, and governance, enabling self-service analytics for business and quality teams across Europe and North America.
  • •Collaborated with product owners, IT, and business stakeholders to translate requirements into scalable data solutions.
  • Oct 2024 – Dec 2025
    Continental⁠, Data Platform Engineer
  • •Working on continental.datalake, a cloud-based self-service analytics platform for scalable data processing, storage, and BI use cases.
  • •Managed AWS infrastructure using Terraform and Terragrunt, supporting S3 storage, networking, secure Linux appliances, and platform configurations.
  • •Provisioned and maintained EMR clusters for PySpark processing, notebook-based exploration, and Spark job execution.
  • •Operated AWS Redshift environments, including maintenance, access management, external schemas, and query troubleshooting.
  • •Delivered containerized platform services and automation tools using Docker, ECR, and Kubernetes.
  • •Improved developer experience with standardized local environments using Devbox, direnv, and automated dependency updates with Renovate.
  • Nov 2023 – Oct 2024Porto, Portugal
  • •Worked on the Cloud Data & Analytics Framework, a scalable AWS-based lakehouse platform using S3, Apache Iceberg, Glue, EMR, and Redshift.
  • •Built and maintained event-driven data pipelines using EventBridge, SQS, Lambda, and Glue PySpark jobs.
  • •Implemented automated data validation checks with PyDeequ, improving data quality, compliance, and pipeline reliability.
  • •Optimized Spark and Iceberg workloads through partitioning, bucketing, repartitioning, and file-size tuning.
  • •Expanded integration testing for incremental loads, late-arriving data, and SLA validations.
  • •Maintained a Python wrapper for Terraform and supported CI/CD pipelines with GitHub Actions.
  • Nov 2021 – Nov 2023Porto, Portugal
    Infineon Technologies⁠, Data Analytics
  • •Automated reporting workflows using Python, pandas and SQL, reducing manual effort and improving accuracy.
  • •Built Tableau dashboards and KPI reports to support financial and operational decision making.
  • •Designed and implemented key business metrics to monitor business operations and drive insights for optimization.
  • •Managed delivery of custom reporting solutions, ensuring data governance and alignment with stakeholder requirements to improve accessibility and reliability for decision-making.
  • Mar 2021 – Nov 2021Porto, Portugal
    Hotel Black Tulip, Business Intelligence Analyst
  • •Implemented a Power BI reporting solution to track financial and operational performance.
  • •Automated data collection and reporting workflows using Python and SQL.
  • Projects

    Real-time data pipeline with Kafka, Flink, Iceberg, Trino, MinIO, and Superset.

    Medallion batch data pipeline with Airflow, DuckDB, Delta Lake, Trino, MinIO, and Metabase. Full observability and data quality.

    ETL pipeline using Pulumi for infrastructure as code, integrating AWS services and Snowflake for automated data flow.

    Docker containerized and configurable Airflow data pipeline for collecting and storing stock and cryptocurrency market data.

    Streamlit Python-based web application to analyze historical stock data.

    Certificates
    Open Source and Community
    pandas⁠, Contributor

    Improved the library’s data manipulation and reporting functionalities, supporting the development of efficient data pipelines and enabling scalable data solutions for analytics and modeling.

    Education
    Postgraduate Degree in Big Data Engineering, ISEP - Porto School of Engineering⁠
    Master's Specialization in Corporate Finance, ISCAP - Porto Accounting and Business School⁠
    Bachelor's degree in Business Management, UPT - Portucalense University⁠
    Skills
    — Languages & Engineering: Python, SQL, PySpark, pandas, FastAPI, OOP | Cloud & Data Platforms: AWS, S3, EMR, Glue, Lambda, Redshift, Spark, Iceberg, Delta Lake, Snowflake | Pipelines & Orchestration: Airflow, dbt, EventBridge, SQS, Kafka, Flink, DuckDB, Trino | Infrastructure & DevOps: Terraform, Terragrunt, Pulumi, Docker, Kubernetes, GitHub Actions, CI/CD | Data Quality & Testing: PyDeequ, pytest, data validation, integration testing, SLA checks | BI & Observability: Tableau, Power BI, Superset, Metabase, Grafana, CloudWatch
    Languages
    Portuguese
    English