
Sergio Francisco
Verified Expert in Engineering
Infrastructure and Platform Engineer and Developer
Rio de Janeiro - State of Rio de Janeiro, Brazil
Toptal member since July 17, 2019
Sergio is an infrastructure and platform engineer with experience, working where infrastructure meets the application. He built a PCI-DSS-certified infrastructure for a payments company in 3 months; cut a crypto exchange's costs by 87.5% and reduced scaling time from 45 minutes to 5; cut CI pipeline time by around 50%; and helped a company consistently meet a 99.99% SLO. Deep AWS, plus GCP and Azure, on Kubernetes, Terraform, and Datadog.
Portfolio
Experience
- Amazon Web Services (AWS) - 11 years
- Terraform - 10 years
- Cloud Architecture - 8 years
- Docker - 7 years
- DevOps - 7 years
- CI/CD Pipelines - 6 years
- Kubernetes - 6 years
- Google Cloud - 2 years
Preferred Environment
Terraform, Amazon Web Services (AWS), Kubernetes, Amazon Elastic Container Service (ECS), GitHub Actions, Amazon EKS, Datadog, PostgreSQL, Google Cloud Platform (GCP), Azure
The most amazing...
...project I worked on cut a crypto exchange's infrastructure cost by 87.5% and its scaling time from 45 minutes to 5, by moving EC2 workloads to ECS Fargate.
Work Experience
Senior Infrastructure Engineer
Kojo Technologies
- Migrated the main monolith's CI from CircleCI to GitHub Actions, cutting pipeline duration roughly 50% at near-neutral cost (around $2 per full run) by moving jobs to Graviton r8g.4xlarge instances and tuning Node heap allocation per step.
- Authored a 50-item performance program against an SLO of 99% of GraphQL API requests under 2 seconds, and shipped roughly half: Postgres index rework and Node.js heap tuning directly; Redis cache-aside and query optimization through the owning teams.
- Delivered key workstreams on the ECS and Elastic Beanstalk to EKS migration: the nginx sidecar required for parity, pod sizing from 30-day peak data, and a dedicated Karpenter node pool provisioned through Argo CD GitOps.
- Led the org-wide OpenTelemetry rollout on a hybrid Operator and Collector architecture, auto-instrumenting services across staging and production EKS and exporting to Datadog with near-zero application code changes.
- Hardened database access platform-wide, moving services off shared master credentials to least-privilege role-based access with Credstash and AWS Secrets Manager, with zero downtime and a repeatable cutover runbook.
- Owned incident response and root cause analysis in production, correlating Datadog APM, Sentry, and CloudWatch to resolve 504 bursts from V8 heap exhaustion, Aurora writer saturation, and recurring table deadlocks.
- Eliminated the CircleCI license after the migration, removing a recurring vendor cost and measurably improving the team's developer experience scores.
Senior DevOps Engineer
Toptal
- Served as the sole infrastructure engineer on a three-person team delivering a WCAG 2.2 AA document remediation platform, now running in production.
- Designed and built the full AWS footprint as reusable Terraform modules across staging and production: multi-service ECS Fargate for API, front end, workers, and SSO, plus RDS, S3, SQS, SES, and ACM.
- Ran the platform on a private VPC with interface endpoints for ECR, Secrets Manager, CloudWatch Logs, ECS, and SQS, keeping service traffic off the public internet.
- Built the delivery pipeline on GitHub Actions with AWS OIDC federation, removing long-lived cloud credentials from the deployment path entirely.
Senior Site Reliability Engineer
Sight Machine
- Operated a global manufacturing SaaS platform across roughly 30 Azure Kubernetes Service clusters spanning the Americas, APAC, and Europe, as part of a five-person SRE team.
- Took over the NVIDIA Omniverse digital twin project after its previous owner left, reworking the inherited Terraform modules for flexibility and deploying the solution to production inside the Azure accounts of large enterprise manufacturers.
- Deployed NVIDIA's dcgm-exporter on AKS to expose GPU node metrics to Prometheus, driving Grafana dashboards and AlertManager alerts routed to Slack for assisted operation of customer-hosted deployments.
- Onboarded and trained new teammates to full ownership of deployed workloads, distributing operational load across the SRE team.
- Automated infrastructure with Terraform and Atlantis, integrating infrastructure as code into the CI/CD pipeline for a mission-critical stack running nginx, PostgreSQL, Kafka, Prometheus, Grafana, AlertManager, and PagerDuty.
AWS/GCP Cloud Engineer
Colgate-Palmolive
- Reviewed a codebase containing over 40 Python functions deployed on Cloud Functions that performed operations in Colgate-Palmolive's GCP Infrastructure, which were not documented.
- Created comprehensive documents with diagrams designed using Mermaid and Lucidchart to assist the Colgate-Palmolive engineering team in understanding and enhancing their infrastructure operations.
- Worked on this review and documentation, which led to significant improvements in the reliability of the infrastructure.
Senior DevOps Engineer
4 Elements Music
- Implemented a complete CI/CD pipeline using GitHub Actions and AWS CodeDeploy to improve operational efficiency by adopting CI/CD pipelines to release software faster and securely.
- Architected and implemented a new solid platform to run a containerized Python/Django application on AWS using ECS and Fargate.
- Implemented all infrastructure resources on AWS and CloudFlare using Terraform and Terraform Cloud.
- Deployed a GitHub Actions CI/CD pipeline to automate software delivery (test, build, and deploy applications using Blue/Green as release model) on top of AWS ECS and Fargate.
- Created Terraform modules from scratch and published them in HCP Terraform (formerly known as Terraform Cloud) to manage infrastructure.
- Deployed Cloudflare to protect their API from bots and other basic attacks. It included migrating the 4elementsmusic.com zone from Amazon Route 53 to Cloudflare.
Senior Infrastructure Engineer
CoinList Services
- Modernized CoinList's infrastructure operations by migrating their workloads from virtual servers (EC2) to containers (ECS). This project reduced infrastructure costs by 87.5% (from $4/hour to $0.5/hour).
- Reduced main application scaling process time by 88% (from 45 minutes to approximately 5 minutes), a significant efficiency improvement that increased the confidence level in the deployment process.
- Migrated CI/CD pipelines from Jenkins to GitHub Actions, which standardized and sped up CoinList's software delivery process, improved the engineering team's efficiency, and simplified workflow operations.
Cloud Architect
Caylent
- Delivered projects for six clients: EVgo, Art of Problem Solving (AoPS), Whatnot, Web 3 Pro, TeleTracking and PlanetDDS.
- Deployed Sagemaker training pipelines and model-serving infrastructure to assist the client's migration of their Recommendation Engine from SpellML to support their MLOps practice.
- Designed a hub-and-spoke networking architecture that centralized ingress and egress networking access across three regions (the US and Europe) using services such as Transit Gateway AWS WAF and Load Balancers.
- Migrated workloads from Rackspace to AWS EKS using the lift-and-reshape migration method.
- Deployed AWS Control Tower for Terraform to create and customize new accounts complying with the client's organization's security guidelines.
- Delivered a PoC to showcase how the client's application could be modernized using ECS Fargate, CircleCI, and GitHub.
- Adopted Datadog for centralized logging, providing a unified view of products (Denticon, Apteryx, Legwork, Cloud 9) across AWS, GCP, Azure, and on-premises datacenter.
Lead Site Reliability Engineer
ETUS Media Holding
- Discovered, planned, and migrated all company services from Digital Ocean and an on-premises data center in the USA to Google Cloud, enhancing our reliability, including uptime, security, capacity, and performance.
- Re-architected and optimized the infrastructure of the company's main application and improved its reliability to handle tens of thousands of simultaneous clients using Google Cloud-managed services.
- Modernized applications by implementing containerization, deploying them on GCP Cloud Run, and setting up streamlined CI/CD pipelines.
- Implemented an observability solution using Google Cloud Operations Suite.
Infrastructure Architect
Dock
- Architected and implemented a multi-account infrastructure across two regions with multiple VPCs that used a broad range of AWS services such as EC2, S3, Route 53, RDS, ElastiCache, SQS, IAM, CloudTrail, Config, etc.
- Deployed and architected the infrastructure for a PCI-certified system that processed thousands of financial transactions daily and a microservices infrastructure for tens of RESTFul APIs developed in Java.
- Participated in recruiting and selecting new senior engineers for the team that I technically led and that migrated several systems and terabytes of data from a traditional data center to the AWS cloud.
- Developed a CD pipeline to deploy static websites (built using Angular) on AWS using S3 in conjunction with CloudFront. This solution allowed the company to perform more deployments without downtime, at any time, and without manual intervention.
- Deployed a GitLab autoscaling solution to automatically spin up and down Amazon EC2 Spot instances to process builds immediately and have a cost-effective, flexible/scalable solution.
Linux Support Analyst
Huawei Technologies Co.
- Collaborated during the planning and execution phases of the project that added the 9th digit to the phones of the "Gestor Online" platform with a 9x prefix.
- Supported, as an app and software engineer, a value-added services platform called "Gestor Online" for the carrier Claro Brazil; it had hundreds of thousands of corporative lines and used to process up to 100 call attempts per second.
- Installed a rack for the SDU project with two switches, one chassis with 12-blade servers, KVM Raritan, and single storage with four expansions totaling 36 terabytes of storage.
Linux Analyst
SONDA
- Managed Unimed Rio Hospital's virtual infrastructure comprising more than five Dell physical servers, Fibre Channel EMC storage, Cisco switches, and 50+ virtual machines.
- Administered 25+ GNU/Linux servers in six locations, running applications like database clusters, applications servers, and web servers.
- Handled highly complex requests and incidents requiring in-depth research and scaled for local support teams (Level 1).
Support Analyst and IT Intern
RNP
- Supported the computers and network equipment of the Rio de Janeiro unit of ESR, the training school of RNP, Brazil's national research and education network serving 800+ institutions and 3.5 million users.
- Built and documented the lab configurations used to run ICT training courses, preparing each environment to the local coordinator's requirements ahead of every class.
- Managed and tuned the operating system images installed across the unit's desktops, standardizing how lab machines were provisioned and reset between courses.
- Implemented and administered a VMware vSphere 4 platform on two ESXi hosts with HP iSCSI storage, consolidating physical servers into roughly 10 Windows and Linux virtual machines with high availability and fast recovery.
- Supported the IT infrastructure of ESR training units across several Brazilian states, covering roughly 60 network devices per unit, including Dell servers and desktops, 3COM switches, and Polycom video conferencing.
- Led the Online Reviews ESR project end-to-end, replacing paper post-course feedback forms with digital ones and cutting turnaround in post-sales information management.
- Authored and quality-assured course content for ESR's ICT training portfolio, including Server Virtualization and Introduction to Voice over IP with Asterisk.
Experience
Kojo | Org-wide OpenTelemetry Rollout on EKS
https://sergiofrancisco.com/case-kojoI designed a hybrid collector architecture, with an agent running as a DaemonSet on every node feeding a gateway collector running as a Deployment, then presented it to the team, collected feedback, and secured approval before building anything.
From there, I ran the project end-to-end: validated auto-instrumentation in staging, configured the collection, processing, and export pipeline to Datadog, verified that resource tagging was applied correctly, and confirmed the Datadog API integration. Then, reproduced the same approach in the production environment.
The outcome was distributed tracing across staging and production with near-zero application code changes, so service teams got traces without touching their own code.
I self-managed the project in Jira on an Agile cadence, defining my own sprints and publishing a weekly plan, progress, and blockers to the wider team.
CoinList | Infrastructure and DevOps Modernization
https://sergiofrancisco.com/case-coinlistCheck all details of this project by clicking on the case study link.
Sight Machine | NVIDIA Omniverse Digital Twin on Azure AKS
I continued and hardened that work, reworking the modules to be more flexible and reusable, then taking the solution to production in the Azure accounts of large enterprise manufacturers in the automotive and beverage sectors.
I also deployed NVIDIA's GPU telemetry exporter on AKS so Prometheus could scrape metrics from the GPU nodes, centralize them, and drive Grafana dashboards, as well as AlertManager alerts routed to a dedicated Slack channel. That gave us assisted operations over deployments that lived in customer accounts rather than in ours.
Before leaving, I trained the engineers who took over both module development and deployment.
Kojo | CircleCI to GitHub Actions Migration
Instead of throwing larger machines at the problem, I moved jobs to memory-optimized Graviton instances (r8g.4xlarge) and tuned Node.js heap allocation per step. Pipeline duration dropped roughly 50% at near-neutral cost, around $2 per full CI run: a more expensive instance running for far less time. I defended the instance choice in the developer experience guild once the numbers confirmed it.
The CircleCI license was retired, and developer experience scores improved.
I converted the entire workflow and drove most of the test iterations using Claude as an AI-assisted development tool, which brought the migration in on deadline.
Web3 Pro | AWS Control Tower Account Factory for Terraform
https://sergiofrancisco.com/case-web3-proI ran a security assessment of the existing environment, designed the AWS Landing Zone, and implemented AWS Control Tower with an Account Factory for Terraform, enabling new accounts to be created and customized from code within the organization's security guardrails rather than by hand. I enrolled the existing accounts into the new structure and delivered the training and documentation the team needed to operate it without me.
The result is a multi-account foundation where account creation is repeatable, auditable, and governed by default.
Toptal | WCAG 2.2 AA Document Remediation Platform
I was the sole infrastructure engineer on a three-person team. I designed and built the full AWS footprint as reusable Terraform modules across staging and production, covering multi-service ECS Fargate for the API, front end, workers, and SSO, as well as RDS, S3, SQS, SES, and ACM.
The platform runs on a private VPC with interface endpoints for ECR, Secrets Manager, CloudWatch Logs, ECS, and SQS, keeping service traffic off the public internet.
Delivery runs on GitHub Actions with AWS OIDC federation, which removed long-lived cloud credentials from the deployment path entirely.
EVgo | Migration from Rackspace to AWS EKS and S3 + CloudFront
https://sergiofrancisco.com/case-evgoI ran the discovery to map dependencies, then chose re-platforming over a straight lift-and-shift. I improved the Dockerfiles, built the CI/CD pipelines with Bitbucket Pipelines, configured Kubernetes objects, and deployed workloads to Amazon EKS, with the static front end served from S3 via CloudFront.
Traffic was migrated during a scheduled maintenance window. The applications ended up in a single consolidated AWS environment with automated delivery, replacing the manual deploys they had on the previous provider.
TeleTracking | Multi-region Hub-and-Spoke Architecture and TCO
https://sergiofrancisco.com/case-teletrackingI assessed the existing environment, conducted a total cost of ownership analysis to support the migration decision, and designed the target network architecture: a hub-and-spoke topology spanning three regions across the United States and Europe, centralizing ingress and egress through AWS Transit Gateway, with WAF and load balancers at the edge.
My part of the project was delivered in one month.
Muxi | Infrastructure Migration to AWS
https://br.claranet.com/case-studies/muxi-otimiza-infraestrutura-de-ti-com-cloud-e-managed-services-da-claranetAs the architect, I first evaluated the technical and financial aspects of several cloud vendors and ultimately chose AWS as the platform. Next, I reviewed their entire legacy infrastructure, designed a multi-account/region/VPC architecture, and collaborated with the engineering team to migrate a set of systems that processed millions of financial transactions daily.
This migration brought order to the client's infrastructure architecture and operations, which had been in a state of chaos. As a result of this project, I received an invitation from AWS and Claranet, an AWS partner, to present the migration case at AWS Summit São Paulo 2017.
4 Elements Music | Infrastructure and DevOps Modernization
https://sergiofrancisco.com/case-4-elements-musicThe old system lacked scalability and performance. The solution involved containerizing the app with Docker, building a new platform with Terraform, deploying secure networking with VPC, and automating software delivery with CI/CD.
This boosted efficiency and security and laid a solid foundation for the new platform's launch.
Whatnot | MLOps Migration from SpellML to Amazon Sagemaker
https://sergiofrancisco.com/case-whatnotI designed the SageMaker infrastructure, built the training pipelines, deployed the models as real-time endpoints, and integrated them with the client's existing MLOps tooling.
Moving off SpellML put both training and serving on managed AWS services, so the team could iterate on models without operating the underlying platform.
Art of Problem Solving | Infrastructure Modernization
https://sergiofrancisco.com/case-art-of-problem-solvingI delivered a proof-of-concept for a container-based platform on ECS Fargate, with CircleCI and GitHub handling build and delivery, so the team could evaluate the target architecture against their existing stack before committing to a full migration.
The proof of concept covered container orchestration, the delivery pipeline, and the operational model, giving AoPS a concrete basis to decide on modernization rather than an estimate.
Education
Bachelor's Degree in Information Systems
Faculdade de Informática Lemos de Castro - Rio de Janeiro, Brazil
Certifications
KCNA: Kubernetes and Cloud Native Associate
The Linux Foundation
HashiCorp Certified: Terraform Associate (002)
Hashicorp
AWS Solutions Architect Associate
Amazon Web Services
Google Cloud Certified Associate Cloud Engineer
Google Cloud
Certified Scrum Master (CSM) I
Scrum Alliance
Red Hat Certified Engineer (RHCE)
Red Hat
Red Hat Certified Systems Administrator (RHCSA)
Red Hat
CompTIA Network+ (N10-005)
CompTIA
Skills
Libraries/APIs
Node.js
Tools
VMware, Terraform, GitLab CI/CD, Amazon Virtual Private Cloud (VPC), Docker Compose, Apache Tomcat, Packer, Apache, NGINX, Grafana, Amazon EKS, Jira, GitLab, Ansible, Git, Iptables, RabbitMQ, Sentry, Google Compute Engine (GCE), Google Kubernetes Engine (GKE), Logging, GitHub, Bitbucket, CircleCI, Amazon Elastic Container Registry (ECR), Amazon Elastic Container Service (ECS), AWS IAM, Amazon SageMaker, Amazon CloudFront CDN, AWS Fargate, Amazon ElastiCache, Artillery, Amazon Firewall, Lucidchart, Helm, Amazon CloudWatch, Azure Kubernetes Service (AKS), Amazon Simple Queue Service (SQS), Amazon Simple Email Service (SES), Claude Code, Claude, VPN
Languages
Bash, Bash Script, SQL, Python, GraphQL
Paradigms
DevOps, Continuous Delivery (CD), Automation, Continuous Integration (CI), DevSecOps
Platforms
Docker, Amazon Web Services (AWS), Linux, Google Cloud Platform (GCP), Amazon EC2, Kubernetes, DigitalOcean, AWS Lambda, Azure, PagerDuty, Apache Kafka
Storage
Google Cloud, Amazon S3 (AWS S3), MySQL, Redis, Google Cloud Storage, Google Cloud SQL, PostgreSQL, Google Cloud Datastore, Datadog, Amazon Aurora
Frameworks
Laravel, Ruby on Rails (RoR)
Other
Certified ScrumMaster (CSM), Documentation, Data Center Migration, CI/CD Pipelines, AWS Cloud Architecture, Containers, Cloud Architecture, System Administration, Shell Scripting, Monitoring, Amazon RDS, GitHub Actions, GitOps, Network Security, Security Engineering, Security Monitoring, PCI DSS, NFS, Content Delivery Networks (CDN), Gunicorn, Google BigQuery, Information Systems, Architecture, Amazon API Gateway, Networking, DNS, Terraform Cloud, Elastic Load Balancers, AWS Transit Gateway, Web Application Firewall (WAF), AWS Control Tower, AWS Organizations, Flow Diagrams, OpenTelemetry, Argo CD, Karpenter, Site Reliability Engineering (SRE), Observability, Incident Response, Incident Management, AWS Secrets Manager, Atlantis, Prometheus, AWS ECS Fargate, Virtual Private Cloud (VPC), AWS Certificate Manager, Single Sign-on (SSO), Relational Database Services (RDS), GPU Computing, Graphics Processing Unit (GPU), Monitoring & Alerting, Alertmanager, NVIDIA DCGM Exporter, Endpoint Protection, IDS/IPS, Vulnerability Management, Proxies
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring