Data Engineer

Diego PardoMontero

I build and operate data platforms in production: databases and pipelines that move information reliably, models that make it usable and the monitoring that keeps everything running. Systems Engineer with a solid backend and cloud foundation.

Data Engineer with 3+ years building and operating data platforms in production: databases that hold up, pipelines that run every day, data that lands complete and on time and systems that keep working when something breaks.

What I do

Design relational and analytical databases (Oracle, PostgreSQL, MySQL, Snowflake); build ETL/ELT pipelines in Python, SQL and PySpark; model data so it answers the questions a team actually asks; run it in production with monitoring, root cause analysis and permanent fixes.

Background

Systems Engineer from Pontificia Universidad Javeriana (GPA 4.4/5.0, honours thesis). Two years designing schemas and backend services in Java and Spring before specialising in data, so I speak both sides of the stack.

Looking for

Data Engineer, Analytics Engineer or Data Platform roles, hybrid in Bogotá or fully remote, on teams where data has to be reliable at scale.

Apr 2026 – Present
Bogotá · hybrid

Data Engineer · Diametral

  • Designed and built a configuration-driven Python framework for large-scale Oracle data migrations, using declarative YAML and SQL definitions with multiple loading strategies (catalogue, SCD Type 2, fact, aggregate), auditing and concurrency control
  • Implemented a control-table-driven scheduler and a reporting module for pipeline observability, orchestrating daily and weekly batch cycles on a Medallion (Bronze/Silver/Gold) architecture with PySpark and Starburst (Trino)
  • Monitor production data pipelines and resolve incidents end to end in an enterprise financial-services environment: root cause analysis, permanent fixes and documentation, working directly with DBA and DevOps teams

Python · SQL · PySpark · Oracle · Starburst/Trino · Airflow · GCP · Databricks · IBM Cloud

Aug 2025 – Apr 2026
Bogotá · hybrid

Data Engineer Jr. Advanced · Globant

  • Optimised large-scale ETL pipelines and backend workflows in Python and SQL, reducing execution times by ~20% through query tuning and process automation
  • Developed automated data validation and quality checks, cutting manual verification by 65% and improving data reliability across production datasets
  • Resolved ~7 production data incidents per day within SLA targets, with root cause analysis and cross-team fixes

Python · SQL · Java · Spring · AWS · Docker · Angular · TypeScript · Linux · Bash

Jan 2025 – Jul 2025
Bogotá

Information Technology Trainee · Globant

  • Maintained and debugged Snowflake and Airflow ETL pipelines, improving stability and cutting job failures by 25%
  • Built Python validation scripts to detect data discrepancies and inconsistencies across datasets with thousands of records
  • Automated recurring backend processes with Python and Bash, saving ~6–8 hours of manual work per week

Python · SQL · Snowflake · Airflow · Spark · Java · Docker

Jan 2023 – Dec 2024
Remote

Data Engineer · Pacífico Emprende

  • Designed the relational model in MySQL for a platform with over 200 users: entities, keys, normalisation and integrity constraints that kept the data consistent as the product grew over two years
  • Optimised the SQL behind the heaviest views, cutting page load times by ~20%
  • Built the backend data services from scratch in Java and Spring, so the schema and the code reading it were designed together
  • Two years on the same data layer, from the first schema to the queries the product ran on every day

MySQL · SQL · Java · Spring Boot · Django · PHP · WordPress

2022 – 2024

Teaching assistant, Pontificia Universidad Javeriana: Distributed Systems, Data Structures and Introduction to Programming.

Measured impact
−20%

execution time on the ETL pipelines I optimised

Globant · 2025–2026
−65%

manual verification, replaced by automated quality checks

Globant · 2025–2026
−25%

job failures in the Snowflake and Airflow pipelines I maintained

Globant · 2025
6–8 h

of manual work saved every week by automating recurring processes

Globant · 2025

Pontificia Universidad Javeriana

Systems Engineering · 2021–2025

GPA 4.4/5.0
Double emphasis: Advanced Software Development & Data Management Systems
Honours thesis 5.0/5.0
8 academic excellence awards
ACM Javeriana (Senior, 2020–2022) · IEEE RAS Javeriana

Languages

Spanish (native) · English C1, Professional Working Proficiency

Certification path
Certifications
  • AWS Certified Solutions Architect – AssociateAWS
  • AWS Certified Cloud PractitionerAWS
  • Google Cloud Digital LeaderGoogle
  • Architecting with GKE: FoundationsGoogle
  • Application Development using Microservices and ServerlessIBM
  • Introduction to DatabasesMeta
  • SQL (Intermediate) · JavaHackerRank
  • Project managementTecnológico de Monterrey

In progressDatabricks Data Engineer Associate · Snowflake SnowPro Core

+40 additional certifications, see LinkedIn

The path data takes through the systems I build. Tools set in full ink are in sustained production use; the lighter ones, in a narrower professional scope.

The pipeline

Sources

Oracle · Snowflake · PostgreSQL · MySQL

PL/SQL · Azure Data Factory · IBM Cloud

Processing

Python · PySpark · ETL / ELT

Pandas · NumPy · Databricks · Azure Databricks

Modelling

Medallion · SCD Type 2 · Dimensional

Starburst / Trino · marts · query tuning

Orchestration

Airflow · Linux · Bash

Docker · CI/CD · Git · GitHub

Quality & ops

Data validation · Incidents & RCA

Observability · SLA · Agile / Scrum

Around the pipeline

Databases

Oracle · PostgreSQL · MySQL · Snowflake

PL/SQL · relational modelling · SCD Type 2 · query tuning

Backend

Java · Spring Boot · REST APIs

TypeScript · Angular · Django · Node.js

Cloud

AWS · GCP · Azure Databricks · IBM Cloud

Certified: AWS SAA, AWS CCP, GCP Digital Leader

Ways of working

Git · GitHub · CI/CD · Docker

Agile / Scrum · Postman / Swagger · Linux

Studied, not yet in production

Kubernetes / GKE · Kafka · Microservices

Serverless · React · Machine Learning · GraphQL · Distributed systems

Personal project · 2026

Orión, booking platform for a language academy

Students book classes against each teacher’s availability rules. A modular Spring Boot backend with versioned PostgreSQL migrations, server side sessions and role based access; a Next.js front end; and an end to end test suite, all running on Docker.

JavaSpring BootPostgreSQLNext.jsDocker
Honours thesis · 5.0/5.0 · 2025

Enseñarte

A working prototype for teaching, learning and assessing Colombian Sign Language (LSC), with the data model and pipeline that capture, store and evaluate each learner’s progress.

KotlinTensorFlowPythonMachine Learning
Personal project · 2024

E-commerce platform on microservices

Microservices architecture with a database per service pattern, event driven communication between services, Docker containers and REST APIs.

JavaSpring BootDockerPostgreSQLKafka

Hiring a data engineer for your team? Let’s talk.

Open to full time Data Engineer, Analytics Engineer and Data Platform roles, hybrid in Bogotá or fully remote. Email, a call or WhatsApp works best.