
Juan Tamariz
Verified Expert in Engineering
Senior DevOps Engineer and Developer
Guadalajara, Mexico
Toptal member since March 19, 2021
Juan is a DevOps tech lead with 15+ years of experience in fintech, eCommerce, and data engineering. He leads small senior DevOps teams as a hands-on coder, shipping platforms that let engineers self-serve infrastructure. Juan's recent wins include a Pulumi Python library powering 100% of microservices, a Jenkins-to-Argo CD GitOps migration that cut deploys from 4 hours to 15 minutes with zero incidents, and $11,000+ in monthly cloud savings via a Karpenter and Kafka rearchitecture.
Portfolio
Experience
- Amazon Web Services (AWS) - 8 years
- Kubernetes - 7 years
- Terraform - 7 years
- Python 3 - 6 years
- Pulumi - 5 years
- Prometheus - 5 years
- Karpenter - 3 years
- Argo CD - 2 years
Preferred Environment
Python 3, Kubernetes, Amazon Web Services (AWS), Google Cloud, Terraform, Azure, Karpenter, Pulumi, Argo CD, Datadog
The most amazing...
...thing I've built is a Pulumi Python library that turned weeks of DevOps tickets into a single import—now powering 100% of microservices across 28 environments.
Work Experience
DevOps Tech Lead
A Series B Fintech
- Built a Pulumi Python platform library that wraps networking, IAM, observability, and Datadog automation, which now powers 100% of microservices (with 14 services and 28 environments) for an engineer self-service infrastructure.
- Led the migration of 14 microservices across 28 environments from standalone Jenkins to GitOps with Argo CD, cutting deploy time from 4 hours to 15 minutes and eliminating deployment-related incidents.
- Led a zero-disruption Karpenter migration across 3 EKS clusters (with around 50 nodes) with PDBs and topology spread. Saved around $4,000 monthly and made EKS upgrades non-disruptive (a 6-week project validated on the FinOps platform).
- Replaced AWS DMS, Kinesis, and Glue with an in-house Kafka cluster for the data-movement layer, saving around $7,000 monthly and reducing the operational surface across the data platform.
- Drove ongoing FinOps reviews and right-sizing. Reduced OpenSearch by $400 monthly at zero risk, rolled out S3 lifecycle tiering across the estate, and aligned Savings Plans and RIs with AWS quarterly.
- Owned the data platform end-to-end, including Airbyte, Airflow, DMS, Redshift, multi-region RDS replicas configured for under 5-minute DR with automated backups and scheduled patching across EKS, RDS, and EC2.
- Led the compliance initiative for SOC 2 and PCI infrastructure controls using Drata, Wiz, and Vanta, including a sub-4-hour remediation SLA for high-severity findings.
- Built an on-call playbook system that keeps response to high-severity events at single-digit minutes. Personally restored RDS cross-region replication during a NYE P0.
AWS EKS Expert
Entos, Inc.
- Coded different Terraform modules with their implementation through Makefile wrappers to guarantee consistency and easy inclusion on a CI/CD tool.
- Understood the existing code so I could contribute according to the internal guidelines and code style.
- Fixed different situations detected in previous implementations daily.
AWS | DevOps
Yields NV
- Worked on a migration from Jenkins-X to Concourse CI, which involved mastering the Concourse CI technology. The migration was completed by moving 29 projects from the old platform to the new one, resulting in more than 120 pipelines.
- Coded a generator of pipelines for Concourse CI, which takes a YTT template and variables as input to then output a pipeline definition in a YAML format. This helps the development team to be self-sufficient in maintaining the concourse pipelines.
- Maintained more than 120 different pipelines on Concourse CI. Managed the automation of pull requests and merged them into the release branch. Tested code, built artifacts, and published them in collaboration with the development team.
- Collaborated on the IBM Cloud project to set up the security recommendations for a Yields NV project. Worked together with the Yields NV team and other external contractors. Used relevant technologies like Kubernetes.
- Automated a way to perform smoke tests on ephemeral clusters. Used tools such as Kubernetes, Terraform, Concourse CI, Bash, Python, and Google Cloud.
- Maintained the CI setup "in-house" by using technologies like Kubernetes, Google Cloud, Bash, and Python.
- Automated the updates on the Concourse CI pipelines, which resulted in a product where developers need to push the pipeline changes to a specific repo for Concourse to update its own pipelines automatically.
Senior DevOps Engineer
Tacit Knowledge
- Improved monitoring for a Google Cloud project with the setup of Prometheus Operator on Kubernetes.
- Implemented CI/CD automation for “1-click” deployments with no downtime. Building custom AMIs as well as Docker images with AWS Code Build.
- Defined an internal workflow to continuously test Helm charts for Kubernetes with an internal repository.
- Defined and configured monitoring and alerting policies for site reliability engineering (SRE).
- Upgraded Jenkins and Ansible to guarantee service availability and maintainability of deployment scripts.
- Developed Python code to create lambda functions to automate firewall whitelisting and storage cleanup.
Senior DevOps Consultant
Levi Strauss & Co
- Supported and improved an AWS serverless architecture.
- Defined a model for support and escalations of user access requests.
- Established a CloudFormation library to be used for infrastructure deployments.
DevOps Engineer Consultant
Tacit Knowledge
- Deployed a private Chef Supermarket to promote common practices with wrapper and community cookbooks.
- Created custom Chef resources with Ruby scripts to automate backups with duplicity.
- Designed and developed environments in Kubernetes to production with Helm in Google Cloud.
- Designed and developed environments in AWS using Jenkins, Ansible, OpenVPN, OpenLDAP, and CloudFormation.
- Established CI/CD workflows for clients with virtual machines and containers in Google Cloud.
- Migrated a Kubernetes cluster from Google Cloud to Azure which provided service portability.
- Performed log parsing tuning for Stackdriver in Google and CloudWatch in AWS.
- Autoscaled a cluster of Java applications with CloudFormation in AWS, which provided highly available infrastructure.
DevOps and SysAdmin Manager
PriceTravel
- Managed projects with budgets of $2.5 million for a colocation setup expansion.
- Scripted policies and procedures to establish configurations in compliance with the PCI for credit card management.
- Developed an HA cluster with the SQL Server to provide an RTO of one minute in case of hardware failure.
- Composed shell scripting for the management of 350 network routers.
- Installed and built the configuration remotely, which resulted in a new record for the company, mounting 75 servers in one day.
- Managed the infrastructure by monitoring more than 300 servers with Nagios, Cacti, MRTG, and Datadog.
- Provided tier-three support in networking, VoIP, the email server, databases, and 3rd-party applications (server-side).
- Deployed SQL Monitor, Nagios, and New Relic for monitoring and proactive planning.
- Automated deployments of Java applications and implemented virtualization for production servers with Windows and Linux.
Experience
Zero-downtime Deployments
• Set up the infrastructure for 4 different environments, including production. CI/CD, monitoring, backups, and security tools.
• Established the automation to continuously introduce security patches from lower environments to production.
• Set up AWS Inspector to validate possible new vulnerabilities in the code.
• Set up a CI/CD pipeline including code testing, security assessment, and a no-downtime deployment strategy that covers database upgrades. This reduced the application's downtime in production and increased the production release frequency.
• Implemented core component upgrades to reduce costs and maximize performance for the client.
A CI/CD Framework to Speed-up Project Setups
The impact of my work was a significant reduction of implementation time for new pipelines from 30 days to 7 days.
JupyterHub Notebooks in Kubernetes
At a glance, for every user logged in, a new Kubernetes pod is created on-demand. When more resources are needed, the Kubernetes cluster will also auto-scale.
Terraform Modules to Speed-up Infrastructure Creation
I created an Agile project to track the creation of every Terraform module. We ended up on a set of authorized scripts that were instanced on several projects on Google Cloud.
As a result, project setup speed improved from a month to every 3 days with the scripts. The framework considered the usage of the latest available Terraform version, along with a shared back end/state to make collaboration easier.
Pulumi Python Platform Library – Self-service Infrastructure
OUTCOME
• 100% adoption across 14 microservices and 28 environments.
• Removed DevOps from the critical path of new service onboarding.
• IAM scoping, tagging, and Datadog checks are now applied uniformly across the estate with no human review needed for standard cases.
• Production-grade today; actively extended as new services come online.
Jenkins to Argo CD GitOps Migration – 14 Services & 28 Environments
OUTCOME
• Deploy time dropped from 4 hours to 15 minutes.
• Deployment-related incidents reduced to zero (100% reduction).
• CI cost reduced by retiring legacy Jenkins workers.
• Every release is now auditable through Git history rather than tickets, and rollbacks are a single revert.
Kafka Cluster Replacing AWS DMS, Kinesis, & Glue
OUTCOME
• Around $7,000 per month in cloud cost savings.
• Lower operational surface for the data platform team.
• Foundation is now in place for downstream stream-processing work—including change data capture, real-time analytics, and event-driven services.
Karpenter Migration – Zero-disruption Multi-cluster
OUTCOME
• Around $4,000 per month in cluster cost savings, validated by the FinOps platform.
• EKS version upgrades are now non-disruptive—a recurring source of toil was eliminated.
• The same runbook was reused on a second client engagement, replacing the AWS Cluster Autoscaler and avoiding a Cast AI license.
Education
Bachelor's Degree in Computer Systems
Universidad del Caribe - Cancun, Mexico
Skills
Libraries/APIs
OpenLDAP, Node.js
Tools
Helm, Ansible, Terraform, SAP Hybris, Google Kubernetes Engine (GKE), Azure Kubernetes Service (AKS), Amazon EKS, AWS CloudFormation, EFK Stack, Jenkins, AWS Key Management Service (KMS), Nagios, Bitbucket, Git, HashiCorp, Docker Hub, AWS Command Line Interface (CLI), AWS IAM, GitHub, Chef, Amazon ElastiCache, Loki, AWS Glue, OpenVPN, Google Stackdriver, Amazon CloudWatch, Hyper-V, Apache JMeter, SonarQube, HashiCorp Vault, Concourse CI, CloudOps, Packer, Kafka Connect, NGINX, Amazon OpenSearch, Claude, n8n, Istio, Kustomize, Amazon Elastic Container Registry (ECR), Grafana
Languages
Python, Python 3, Java, Go, Ruby, Bash Script, YAML, Bash
Frameworks
Flux
Paradigms
DevOps, Microservices, Continuous Deployment, Object-oriented Programming (OOP), Serverless Architecture, REST, Continuous Integration (CI)
Platforms
Kubernetes, Linux, Amazon Web Services (AWS), Docker, Google Cloud Platform (GCP), Amazon EC2, Vanta, Azure, AWS Lambda, Percona, Windows Server, AWS ALB, Jupyter Notebook, Apache Kafka
Storage
Google Cloud, Datadog, Amazon Aurora, Databases, Amazon DynamoDB, Redshift, Google Cloud Storage, PostgreSQL, Amazon S3 (AWS S3)
Other
Networking, Back-end Admin Systems, VoIP, Web Servers, Prometheus, DNS, Ubiquiti Wireless Gear, Groovy Scripting, Pulumi, Content Delivery Networks (CDN), Amazon RDS, Monitoring, CI/CD Pipelines, Site Reliability Engineering (SRE), Leadership, SSL Certificates, Infrastructure as Code (IaC), Configuration Management, Infrastructure as a Service (IaaS), GitHub Actions, Cloud Migration, Argo CD, GitOps, FinOps, Drata, Wiz Cloud Security Platform, Karpenter, Serverless, Amazon Inspector, Solution Architecture, Platform Engineering, APM, DHCP, SQL Server 2015, Fortinet Firewall Configuration, Active Directory Federation, Multiprotocol Label Switching (MPLS), Mail Servers, IIS 7, Agile DevOps, IBM Cloud, AWS Bedrock AgentCore, Document Management Systems (DMS), Amazon Kinesis, Cloudflare, AWS WAF, Cursor AI, AWS Database Migration Service (DMS), Data Engineering, spot instances
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring