
Sergio Francisco
Verified Expert in Engineering
Infrastructure and Platform Engineer and Developer
Rio de Janeiro - State of Rio de Janeiro, Brazil
Toptal member since July 17, 2019
Sergio is an infrastructure and platform engineer with experience working at the intersection of infrastructure and applications. He built a PCI-DSS-certified infrastructure for a payments company in 3 months; cut a crypto exchange's web fleet cost by 87.5% and reduced scaling time from 45 minutes to 5 minutes; cut CI pipeline time by around 50%; and helped a company consistently meet a 99.99% SLO. He has deep AWS expertise, including GCP and Azure, as well as Kubernetes, Terraform, and Datadog.
Portfolio
Experience
- Amazon Web Services (AWS) - 11 years
- Terraform - 10 years
- Cloud Architecture - 8 years
- Docker - 7 years
- DevOps - 7 years
- CI/CD Pipelines - 6 years
- Kubernetes - 6 years
- Google Cloud - 2 years
Preferred Environment
Terraform, Amazon Web Services (AWS), Kubernetes, Amazon Elastic Container Service (ECS), GitHub Actions, Amazon EKS, Datadog, PostgreSQL, Google Cloud Platform (GCP), Azure
The most amazing...
...project took a payments company off its data center and onto AWS across two regions, earning an invitation from AWS to present it at Summit São Paulo 2017.
Work Experience
Senior Infrastructure Engineer
Kojo Technologies
- Migrated the main monolith's CI from CircleCI to GitHub Actions, cutting pipeline duration roughly 50% at near-neutral cost (around $2 per full run) by moving jobs to Graviton r8g.4xlarge instances and tuning Node.js heap allocation per step.
- Authored a 50-item performance program against an SLO of 99% of GraphQL API requests under 2 seconds, and shipped roughly half: Postgres index rework and Node.js heap tuning directly; Redis cache-aside and query optimization through the owning teams.
- Delivered key workstreams on the ECS and Elastic Beanstalk to EKS migration: the nginx sidecar required for parity, pod sizing from 30-day peak data, and a dedicated Karpenter node pool provisioned through Argo CD GitOps.
- Led the org-wide OpenTelemetry rollout on a hybrid Operator and Collector architecture, auto-instrumenting services across staging and production EKS and exporting to Datadog with near-zero application code changes.
- Hardened database access platform-wide, moving services off shared master credentials to least-privilege role-based access with Credstash and AWS Secrets Manager, with zero downtime and a repeatable cutover runbook.
- Owned incident response and root cause analysis in production, correlating Datadog APM, Sentry, and CloudWatch to resolve 504 bursts from V8 heap exhaustion, Aurora writer saturation, and recurring table deadlocks.
- Eliminated the CircleCI license after the migration, removing a recurring vendor cost and measurably improving the team's developer experience scores.
- Built and shipped a flaky-test detector on GitHub Actions using Anthropic's claude-code-action; it classifies a failing run as flaky or a genuine break and opens a GitHub issue automatically.
- Wrote a log-parsing engine that strips oversized test output before the model call, cutting both latency and false positives. Adopted by the DevEx guild and still in production at rolloff.
Senior DevOps Engineer
Toptal
- Served as the sole infrastructure engineer on a three-person team delivering a WCAG 2.2 AA document remediation platform, now running in production.
- Designed and built the full AWS footprint as reusable Terraform modules across staging and production: multi-service ECS Fargate for API, front end, workers, and SSO, plus RDS, S3, SQS, SES, and ACM.
- Ran the platform on a private VPC with interface endpoints for ECR, Secrets Manager, CloudWatch Logs, ECS, and SQS, keeping service traffic off the public internet.
- Built the delivery pipeline on GitHub Actions with AWS OIDC federation, removing long-lived cloud credentials from the deployment path entirely.
Senior Site Reliability Engineer
Sight Machine
- Operated a global manufacturing analytics SaaS platform across roughly 30 AKS clusters in the Americas, APAC, and Europe, as one of five SREs.
- Ran cost visibility on the fleet's shared multi-tenant cluster: roughly a dozen customers in dedicated namespaces across hundreds of nodes, the largest single cost driver in a six-figure annual Azure spend.
- Led Kubecost and Mavvrik proofs-of-concept end-to-end, configuring per-namespace cost attribution, thresholds, alerts, and the dashboards that engineering management used. Moved the evaluation to Mavvrik on tooling cost.
- Took over the NVIDIA Omniverse digital twin project after the engineer leading it left and scaled it from one customer in production to roughly ten enterprise deployments in the automotive and beverage sectors.
- Each deployment ran inside the customer's own Azure account. Evolved the inherited Terraform modules, fixing defects and reducing dependencies for VNet, AKS, APIM, and the Omniverse workloads, with Flux driven by the same Terraform.
- Added the GPU observability the project lacked: NVIDIA's dcgm-exporter into in-cluster Prometheus, Grafana dashboards, and alert thresholds I wrote for utilization, memory, errors, and temperature.
- Routed GPU alerts through AlertManager to a central Slack channel, escalating to PagerDuty. With deployments living in customer accounts rather than ours, that telemetry is what made remote-assisted operations possible.
- Trained the two engineers who took over the Omniverse project on rolloff, covering module internals, known failure modes, and the improvement roadmap.
- Automated infrastructure with Terraform and Atlantis on GitHub, and operated and debugged GitOps reconciliation with Flux across the cluster fleet.
AWS/GCP Cloud Engineer
Colgate-Palmolive
- Reviewed a codebase containing over 40 Python functions deployed on Cloud Functions that performed operations in Colgate-Palmolive's GCP Infrastructure, which were not documented.
- Created comprehensive documents with diagrams designed using Mermaid and Lucidchart to assist the Colgate-Palmolive engineering team in understanding and enhancing their infrastructure operations.
- Worked on this review and documentation, which led to significant improvements in the reliability of the infrastructure.
Senior DevOps Engineer
4 Elements Music
- Implemented a complete CI/CD pipeline using GitHub Actions and AWS CodeDeploy to improve operational efficiency by adopting CI/CD pipelines to release software faster and securely.
- Architected and implemented a new solid platform to run a containerized Python/Django application on AWS using ECS and Fargate.
- Implemented all infrastructure resources on AWS and CloudFlare using Terraform and Terraform Cloud.
- Deployed a GitHub Actions CI/CD pipeline to automate software delivery (test, build, and deploy applications using Blue/Green as release model) on top of AWS ECS and Fargate.
- Created Terraform modules from scratch and published them in HCP Terraform (formerly known as Terraform Cloud) to manage infrastructure.
- Deployed Cloudflare to protect their API from bots and other basic attacks. It included migrating the 4elementsmusic.com zone from Amazon Route 53 to Cloudflare.
Senior Infrastructure Engineer
CoinList Services
- Migrated all of CoinList's services from EC2 to containers on ECS, right-sizing on actual usage rather than peak and tuning autoscaling for token-sale traffic spikes.
- Contributed to the platform's primary API service—the entry point for site traffic and caller of the other services—cutting web fleet cost by 87.5% ($4/hour to $0.50/hour) and application scaling time by 88% (45 to around 5 minutes).
- Migrated CI/CD pipelines from Jenkins to GitHub Actions, which standardized and sped up CoinList's software delivery process, improved the engineering team's efficiency, and simplified workflow operations.
Cloud Architect
Caylent
- Delivered projects for six clients: EVgo, Art of Problem Solving (AoPS), Whatnot, Web 3 Pro, TeleTracking and PlanetDDS.
- Deployed Sagemaker training pipelines and model-serving infrastructure to assist the client's migration of their Recommendation Engine from SpellML to support their MLOps practice.
- Designed a hub-and-spoke networking architecture that centralized ingress and egress networking access across three regions (the US and Europe) using services such as Transit Gateway AWS WAF and Load Balancers.
- Migrated workloads from Rackspace to AWS EKS using the lift-and-reshape migration method.
- Deployed AWS Control Tower for Terraform to create and customize new accounts complying with the client's organization's security guidelines.
- Delivered a PoC to showcase how the client's application could be modernized using ECS Fargate, CircleCI, and GitHub.
- Adopted Datadog for centralized logging, providing a unified view of products (Denticon, Apteryx, Legwork, Cloud 9) across AWS, GCP, Azure, and on-premises datacenter.
Lead Site Reliability Engineer
ETUS Media Holding
- Discovered, planned, and migrated all company services from Digital Ocean and an on-premises data center in the USA to Google Cloud, enhancing our reliability, including uptime, security, capacity, and performance.
- Re-architected and optimized the infrastructure of the company's main application and improved its reliability to handle tens of thousands of simultaneous clients using Google Cloud-managed services.
- Modernized applications by implementing containerization, deploying them on GCP Cloud Run, and setting up streamlined CI/CD pipelines.
- Implemented an observability solution using Google Cloud Operations Suite.
Infrastructure Architect
Dock
- Architected and implemented a multi-account infrastructure across two regions with multiple VPCs that used a broad range of AWS services such as EC2, S3, Route 53, RDS, ElastiCache, SQS, IAM, CloudTrail, Config, etc.
- Deployed and architected the infrastructure for a PCI-certified system that processed thousands of financial transactions daily and a microservices infrastructure for tens of RESTFul APIs developed in Java.
- Participated in recruiting and selecting new senior engineers for the team that I technically led and that migrated several systems and terabytes of data from a traditional data center to the AWS cloud.
- Developed a CD pipeline to deploy static websites (built using Angular) on AWS using S3 in conjunction with CloudFront. This solution allowed the company to perform more deployments without downtime, at any time, and without manual intervention.
- Deployed a GitLab autoscaling solution to automatically spin up and down Amazon EC2 Spot instances to process builds immediately and have a cost-effective, flexible/scalable solution.
Linux Support Analyst
Huawei Technologies Co.
- Collaborated during the planning and execution phases of the project that added the 9th digit to the phones of the "Gestor Online" platform with a 9x prefix.
- Supported, as an app and software engineer, a value-added services platform called "Gestor Online" for the carrier Claro Brazil; it had hundreds of thousands of corporative lines and used to process up to 100 call attempts per second.
- Installed a rack for the SDU project with two switches, one chassis with 12-blade servers, KVM Raritan, and single storage with four expansions totaling 36 terabytes of storage.
Linux Analyst
SONDA
- Managed Unimed Rio Hospital's virtual infrastructure comprising more than five Dell physical servers, Fibre Channel EMC storage, Cisco switches, and 50+ virtual machines.
- Administered 25+ GNU/Linux servers in six locations, running applications like database clusters, applications servers, and web servers.
- Handled highly complex requests and incidents requiring in-depth research and scaled for local support teams (Level 1).
Support Analyst and IT Intern
RNP
- Supported the computers and network equipment of the Rio de Janeiro unit of ESR, the training school of RNP, Brazil's national research and education network serving 800+ institutions and 3.5 million users.
- Built and documented the lab configurations used to run ICT training courses, preparing each environment to the local coordinator's requirements ahead of every class.
- Managed and tuned the operating system images installed across the unit's desktops, standardizing how lab machines were provisioned and reset between courses.
- Implemented and administered a VMware vSphere 4 platform on two ESXi hosts with HP iSCSI storage, consolidating physical servers into roughly 10 Windows and Linux virtual machines with high availability and fast recovery.
- Supported the IT infrastructure of ESR training units across several Brazilian states, covering roughly 60 network devices per unit, including Dell servers and desktops, 3COM switches, and Polycom video conferencing.
- Led the Online Reviews ESR project end-to-end, replacing paper post-course feedback forms with digital ones and cutting turnaround in post-sales information management.
- Authored and quality-assured course content for ESR's ICT training portfolio, including Server Virtualization and Introduction to Voice over IP with Asterisk.
Experience
Kojo | Org-wide OpenTelemetry Rollout on EKS
https://sergiofrancisco.com/case-kojo-opentelemetryI designed a hybrid collector architecture, with an agent running as a DaemonSet on every node feeding a gateway collector running as a Deployment, then presented it to the team, collected feedback, and secured approval before building anything.
From there, I ran the project end-to-end: validated auto-instrumentation in staging, configured the collection, processing, and export pipeline to Datadog, verified that resource tagging was applied correctly, and confirmed the Datadog API integration. Then, reproduced the same approach in the production environment.
The outcome was distributed tracing across staging and production with near-zero application code changes, so service teams got traces without touching their own code.
I self-managed the project in Jira on an Agile cadence, defining my own sprints and publishing a weekly plan, progress, and blockers to the wider team.
CoinList | Infrastructure and DevOps Modernization
https://sergiofrancisco.com/case-coinlistCheck all details of this project by clicking on the case study link.
Sight Machine | NVIDIA Omniverse Digital Twin on Azure AKS
https://sergiofrancisco.com/case-sight-machineI received it with one customer in production and left it with roughly ten enterprise deployments across the automotive and beverage sectors. I evolved the inherited modules, fixing defects and reducing dependencies: they provisioned the full Azure footprint (VNet, AKS, APIM) and the Omniverse workloads themselves, with Flux installed and driven from the same Terraform, from a per-customer monorepo with remote state in Azure Blob Storage.
The project had no GPU observability, so I added it. NVIDIA's dcgm-exporter fed the GPU nodes into in-cluster Prometheus, driving Grafana dashboards and AlertManager alerts on thresholds I wrote for utilization, memory, errors, and temperature, routed to Slack and escalated to PagerDuty. With deployments in customer accounts rather than ours, that telemetry made remote-assisted operations possible.
Before rolling off, I trained the two engineers who took it over.
Kojo | CircleCI to GitHub Actions Migration
https://sergiofrancisco.com/case-kojo-ci-migration/Instead of throwing larger machines at the problem, I moved jobs to memory-optimized Graviton instances (r8g.4xlarge) and tuned Node.js heap allocation per step. Pipeline duration dropped roughly 50% at near-neutral cost, around $2 per full CI run: a more expensive instance running for far less time. I defended the instance choice in the developer experience guild once the numbers confirmed it.
The CircleCI license was retired, and developer experience scores improved.
I converted the entire workflow and drove most of the test iterations using Claude as an AI-assisted development tool, which brought the migration in on deadline.
Web3 Pro | AWS Control Tower Account Factory for Terraform
https://sergiofrancisco.com/case-web3-proI ran a security assessment of the existing environment, designed the AWS Landing Zone, and implemented AWS Control Tower with an Account Factory for Terraform, enabling new accounts to be created and customized from code within the organization's security guardrails rather than by hand. I enrolled the existing accounts into the new structure and delivered the training and documentation the team needed to operate it without me.
The result is a multi-account foundation where account creation is repeatable, auditable, and governed by default.
Toptal | WCAG 2.2 AA Document Remediation Platform
I was the sole infrastructure engineer on a three-person team. I designed and built the full AWS footprint as reusable Terraform modules across staging and production, covering multi-service ECS Fargate for the API, front end, workers, and SSO, as well as RDS, S3, SQS, SES, and ACM.
The platform runs on a private VPC with interface endpoints for ECR, Secrets Manager, CloudWatch Logs, ECS, and SQS, keeping service traffic off the public internet.
Delivery runs on GitHub Actions with AWS OIDC federation, which removed long-lived cloud credentials from the deployment path entirely.
EVgo | Migration from Rackspace to AWS EKS and S3 + CloudFront
https://sergiofrancisco.com/case-evgoI ran the discovery to map dependencies, then chose re-platforming over a straight lift-and-shift. I improved the Dockerfiles, built the CI/CD pipelines with Bitbucket Pipelines, configured Kubernetes objects, and deployed workloads to Amazon EKS, with the static front end served from S3 via CloudFront.
Traffic was migrated during a scheduled maintenance window. The applications ended up in a single consolidated AWS environment with automated delivery, replacing the manual deploys they had on the previous provider.
TeleTracking | Multi-region Hub-and-Spoke Architecture and TCO
https://sergiofrancisco.com/case-teletrackingI assessed the existing environment, conducted a total cost of ownership analysis to support the migration decision, and designed the target network architecture: a hub-and-spoke topology spanning three regions across the United States and Europe, centralizing ingress and egress through AWS Transit Gateway, with WAF and load balancers at the edge.
My part of the project was delivered in one month.
Muxi | Infrastructure Migration to AWS
https://sergiofrancisco.com/case-muxi/As the architect, I first evaluated the technical and financial aspects of several cloud vendors and ultimately chose AWS as the platform. Next, I reviewed their entire legacy infrastructure, designed a multi-account/region/VPC architecture, and collaborated with the engineering team to migrate a set of systems that processed millions of financial transactions daily.
This migration brought order to the client's infrastructure architecture and operations, which had been in a state of chaos. As a result of this project, I received an invitation from AWS and Claranet, an AWS partner, to present the migration case at AWS Summit São Paulo 2017.
4 Elements Music | Infrastructure and DevOps Modernization
https://sergiofrancisco.com/case-4-elements-musicThe old system lacked scalability and performance. The solution involved containerizing the app with Docker, building a new platform with Terraform, deploying secure networking with VPC, and automating software delivery with CI/CD.
This boosted efficiency and security and laid a solid foundation for the new platform's launch.
Whatnot | MLOps Migration from SpellML to Amazon Sagemaker
https://sergiofrancisco.com/case-whatnotI designed the SageMaker infrastructure, built the training pipelines, deployed the models as real-time endpoints, and integrated them with the client's existing MLOps tooling.
Moving off SpellML put both training and serving on managed AWS services, so the team could iterate on models without operating the underlying platform.
Art of Problem Solving | Infrastructure Modernization
https://sergiofrancisco.com/case-art-of-problem-solvingI delivered a proof-of-concept for a container-based platform on ECS Fargate, with CircleCI and GitHub handling build and delivery, so the team could evaluate the target architecture against their existing stack before committing to a full migration.
The proof of concept covered container orchestration, the delivery pipeline, and the operational model, giving AoPS a concrete basis to decide on modernization rather than an estimate.
Education
Bachelor's Degree in Information Systems
Faculdade de Informática Lemos de Castro - Rio de Janeiro, Brazil
Certifications
KCNA: Kubernetes and Cloud Native Associate
The Linux Foundation
HashiCorp Certified: Terraform Associate (002)
Hashicorp
AWS Solutions Architect Associate
Amazon Web Services
Google Cloud Certified Associate Cloud Engineer
Google Cloud
Certified Scrum Master (CSM) I
Scrum Alliance
Red Hat Certified Engineer (RHCE)
Red Hat
Red Hat Certified Systems Administrator (RHCSA)
Red Hat
CompTIA Network+ (N10-005)
CompTIA
Skills
Libraries/APIs
Node.js
Tools
VMware, Terraform, GitLab CI/CD, Amazon Virtual Private Cloud (VPC), Docker Compose, Apache Tomcat, Packer, Apache, NGINX, Grafana, Amazon EKS, Jira, GitLab, Ansible, Git, Iptables, RabbitMQ, Sentry, Google Compute Engine (GCE), Google Kubernetes Engine (GKE), Logging, GitHub, Bitbucket, CircleCI, Amazon Elastic Container Registry (ECR), Amazon Elastic Container Service (ECS), AWS IAM, Amazon SageMaker, Amazon CloudFront CDN, AWS Fargate, Amazon ElastiCache, Artillery, Amazon Firewall, Lucidchart, Helm, Amazon CloudWatch, Azure Kubernetes Service (AKS), Amazon Simple Queue Service (SQS), Amazon Simple Email Service (SES), Claude Code, Claude, VPN
Languages
Bash, Bash Script, SQL, Python, GraphQL
Paradigms
DevOps, Continuous Delivery (CD), Automation, Continuous Integration (CI), DevSecOps
Platforms
Docker, Amazon Web Services (AWS), Linux, Google Cloud Platform (GCP), Amazon EC2, Kubernetes, DigitalOcean, AWS Lambda, Azure, PagerDuty, Apache Kafka
Storage
Google Cloud, Amazon S3 (AWS S3), MySQL, Redis, Google Cloud Storage, Google Cloud SQL, PostgreSQL, Google Cloud Datastore, Datadog, Amazon Aurora
Frameworks
Laravel, Ruby on Rails (RoR)
Other
Certified ScrumMaster (CSM), Documentation, Data Center Migration, CI/CD Pipelines, AWS Cloud Architecture, Containers, Cloud Architecture, System Administration, Infrastructure as Code (IaC), Identity & Access Management (IAM), Load Balancers, Shell Scripting, Monitoring, Amazon RDS, GitHub Actions, GitOps, Network Security, Security Engineering, Security Monitoring, FinOps, AI Tools, Disaster Recovery (DR), PCI DSS, NFS, Content Delivery Networks (CDN), Gunicorn, Google BigQuery, Information Systems, Architecture, Amazon API Gateway, Networking, DNS, Terraform Cloud, Elastic Load Balancers, AWS Transit Gateway, Web Application Firewall (WAF), AWS Control Tower, AWS Organizations, Flow Diagrams, OpenTelemetry, Argo CD, Karpenter, Site Reliability Engineering (SRE), Observability, Incident Response, Incident Management, AWS Secrets Manager, Atlantis, Prometheus, AWS ECS Fargate, Virtual Private Cloud (VPC), AWS Certificate Manager, Single Sign-on (SSO), Relational Database Services (RDS), GPU Computing, Graphics Processing Unit (GPU), Monitoring & Alerting, Alertmanager, NVIDIA DCGM Exporter, Endpoint Protection, IDS/IPS, Vulnerability Management, Proxies, Kubecost, Multi-tenant SaaS
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring