
Felipe Rodrigues Maia
Verified Expert in Engineering
DevOps Engineer and Developer
Belo Horizonte - State of Minas Gerais, Brazil
Toptal member since May 2, 2018
Felipe is a staff-level platform, reliability, and cloud engineer with 20+ years of experience designing and operating large-scale production systems. He specializes in platform engineering, SRE, cloud architecture, data engineering, and secure infrastructure. Felipe has led large-scale cloud initiatives, standardized engineering platforms across 300+ codebases, and reduced cloud costs by 46%.
Portfolio
Experience
- Python - 17 years
- Site Reliability Engineering (SRE) - 15 years
- Back-end Development - 15 years
- Software Architecture - 14 years
- Infrastructure as Code (IaC) - 12 years
- Amazon Web Services (AWS) - 12 years
- Data Pipelines - 10 years
- Kubernetes - 6 years
Preferred Environment
Python, Go, Site Reliability Engineering (SRE), Platform Engineering, Cloud Architecture, Machine Learning Operations (MLOps), Distributed Systems, Databricks, Amazon Web Services (AWS)
The most amazing...
...thing I've done is architect and own large-scale platforms for big data and media streaming, balancing reliability, security, performance, and cost at scale.
Work Experience
Cloud Architect and MLOps Platform Engineer
Anova Sports
- Architected and owned end-to-end, production-grade computer vision systems, spanning data ingestion, large-scale video inference, and results persistence at scale.
- Delivered a scalable, cost-efficient CV platform with strong MLOps guarantees, including model versioning, inference traceability, and reproducible deployments for safe iteration in production.
- Worked cross-functionally with product and AI teams to translate ambiguous requirements into robust back-end systems, powering real-world AI products used by high-impact end users.
Senior DevOps Engineer
Whip Media
- Owned security, reliability, performance, and cloud operations for a large-scale, production platform serving millions of users, acting across SRE, platform engineering, and cloud architecture responsibilities.
- Led cloud cost optimization and architecture strategy, implementing best practices that resulted in a 46% reduction in total cloud spend while maintaining performance and reliability.
- Drove site reliability engineering practices, including pre-mortems and post-mortems, hands-on debugging from networking through application code, and operational readiness for systems at massive scale.
- Built and led platform engineering and DevOps initiatives, automating infrastructure and workflows to reduce operational overhead and significantly improve engineering team efficiency.
- Designed and standardized CI/CD pipelines, Kubernetes platforms, and infrastructure-as-code across multiple AWS accounts, supporting and unifying delivery for 300+ codebases.
- Developed data and MLOps pipelines for large-scale data ingestion, processing, training, and migration across databases, data warehouses, and data lakes.
DevOps | Site Reliability Engineer (SRE) | Software Architect
Toptal Clients
- Developed standards on infrastructure as code (IaC), continuous integration, alerting, and monitoring, helping data and application engineers to build, monitor, improve and debug eventual issues in production quicker and easier.
- Helped to prepare AWS accounts, using the best standards in terms of security and management simplicity, making it possible to continuously grow the company, and smoothing the absorption and integration with newly acquired companies and applications.
- Architected, developed, and scaled microservices for high loads of requests (API) and processing (media scraping, transcode, and metadata extraction), using Python or Go languages, running on Amazon EKS (Kubernetes), Lambda, and Fargate (serverless).
Data Architect Consultant
Cinnecta
- Helped the company to simplify and scale big data (data lake, data warehouse, and real-time) systems by rearchitecting the projects responsible for collecting, processing, and storing events-based data.
- Built POCs to present evaluations of cloud providers and develop data migration plans.
- Fostered culture building, provided training, and offered guidance to tech teams.
Head of Engineering
Samba Tech
- Headed the engineering area of the company, translating business strategy into technological solutions.
- Managed high-performance teams (30 people distributed in five teams) and helped them make better architectural decisions.
- Created and implemented new effective engineering strategies—encouraging "bottom-up" innovations over processes and tools and knowledge sharing, which engaged people in continuous improvements over operational and tactical scopes.
- Achieved a remarkable 56% average increase in productivity for all product development squads within the first six months.
- Managed the key performance indicators (KPI), objectives and key results (OKR), and budgets.
Software Architect | Site Reliability Engineer (SRE)
Samba Tech
- Designed the architecture and developed the services related to the multi-tenant SaaS (OVP and OTT) solutions.
- Designed, analyzed, and troubleshot large-scale distributed systems.
- Developed internal tools and explored emerging technologies—continuously improving performance, scalability, availability, and the direct and indirect costs.
- Led the technical aspects for the engineering and operations teams.
Full-stack Software Engineer | DevOps | Site Reliability Engineer (SRE)
Incode Software
- Contributed to various projects with roles in analysis, modeling, development, troubleshooting, and fine-tuning of systems engaging with Brazilian and international customers on the sundry focus of the business.
- Worked on various projects for banking automation, email marketing, a GUI for a Fibre Channel, a Gigabit Ethernet (GbE), and SAS/SATA protocol analyzer appliance, and applications for the management of F5 significant IPs.
- Modeled and developed architectural patterns in highly complex systems using open sources or Microsoft technologies.
Software Engineer | Full-stack Developer
International Syst
- Performed research and development of software solutions related to the Linux operating system, providing tools for automating Infrastructure.
- Developed and packed Linux distros in partnership with Intel, ASUS, Positivo, and the Brazilian Government.
- Utilized programming languages such as Python, Perl, shell scripting, PHP, and C++.
Systems Administrator
Hospital das Clinicas da UFMG
- Provided 3rd-level support and technical team leadership.
- Managed mission-critical servers based on Linux, Unix, and Windows; DBMS (Microsoft SQL Server), MySQL, and InterSystems Caché.
- Developed applications for monitoring and automation of mission-critical servers, aiming to comply with high availability, performance, and security requirements.
- Standardized and documented the processes and development of applications intending to simplify and automate procedures performed by the 1st-level support team.
- Analyzed, designed, developed, and operated various IT services and security projects.
IT Technician
SPM Infor
- Managed the IT systems and networks for a small business.
- Developed and maintained the IT infrastructure and security projects.
- Implemented technical support from the 1st to the 3rd level.
Experience
Anova Sports
https://www.anovasports.com/As a Computer Vision and AI Infrastructure Engineer, I designed GPU-accelerated processing pipelines for machine learning inference and image analysis. My work focused on optimizing inference performance, distributed processing, experimentation workflows, and cloud-native GPU infrastructure.
I developed high-performance computer vision pipelines that combine classical image processing techniques with deep learning models, while optimizing GPU utilization, reproducibility, and deployment in production environments.
Technologies: Python, PyTorch, OpenCV, Triton Inference Server, CUDA, Docker, Kubernetes, AWS, GitHub Actions.
TV Time
https://tvtime.com/I was responsible for the reliability and scalability of the cloud platform powering TV Time. My work included Kubernetes operations, AWS infrastructure management, CI/CD automation, production incident response, observability improvements, cost optimization, and performance tuning.
I led multiple infrastructure and data modernization initiatives, improved deployment reliability, optimized database performance, built data pipelines, reduced AWS operational costs, and enhanced platform resilience while supporting high-traffic production environments.
Technologies: AWS, Kubernetes, Terraform, GitHub Actions, MySQL, Redis, CloudFront, ECS, EKS, Docker, Prometheus, Elastic, Python, Go.
The TVDB
https://www.thetvdb.com/As a Senior Site Reliability Engineer and Platform Engineer, I modernized the cloud infrastructure supporting TheTVDB, improving platform reliability, deployment automation, observability, security, and operational efficiency.
I designed and implemented infrastructure-as-code using Terraform, migrated services to Kubernetes, optimized AWS resource utilization, strengthened CI/CD pipelines, and reduced operational overhead through automation. I also worked on database optimization, monitoring, incident response, and production reliability initiatives.
Technologies: Go, PHP, AWS, Kubernetes (EKS), Docker, MariaDB, PostgreSQL, Rest APIs, Databricks.
SambaVideos
https://www.sambatech.com.br/samba-videos-proAs a software engineer, I worked on the core back-end platform, developing highly scalable services responsible for video processing, content management, APIs, and distributed infrastructure. My work involved improving platform performance, reliability, and scalability while supporting millions of video deliveries.
I also participated in the evolution of the platform architecture, integrating cloud services, optimizing storage and media workflows, and contributing to continuous delivery practices across engineering teams.
Technologies: Java, Spring, Python, AWS, Azure, Google Cloud, Akamai, SQL, NoSQL, Caching, Distributed Search and Analytics engines, DRM, Cyber Security, HLS, MPEG-DASH, Media Processing, FFmpeg.
GitHub | Simple Scripts to Aid Site Reliability Engineers
https://github.com/frmaia/uscriptsEducation
Master of Business Administration (MBA) in Computer Vision
ICA PUC Rio - Rio de Janeiro, Brazil
Specialist's Degree in Computer Software Engineering
UFMG | Universidade Federal de Minas Gerais - Belo Horizonte, MG, Brazil
Bachelor's Degree in Systems Analysis and Development (Computer Science)
Centro Universitario Anhanguera | Fabrai - Belo Horizonte, MG, Brazil
Certifications
AWS Certified Solutions Architect - Associate
Amazon Web Services
Skills
Libraries/APIs
REST APIs, Amazon EC2 API, FFmpeg, PyTorch
Tools
Git, Apache Airflow, AWS CloudFormation, Terraform, Amazon Virtual Private Cloud (VPC), Amazon CloudWatch, Amazon EKS, GitHub, AWS Command Line Interface (CLI), NGINX, Vim Text Editor, CircleCI, Jenkins, Ansible, Makefile, Istio, Helm, Fastlane, VPN, Chef, Wowza, AWS Fargate, Amazon Simple Queue Service (SQS), AWS Batch, Observability Tools
Languages
SQL, Go, Python, Java, Bash, PHP, C#, Tcl, Perl, C++, Ruby, JavaScript, Bash Script, Snowflake
Paradigms
Continuous Delivery (CD), Continuous Integration (CI), DevSecOps, DevOps, Automation, Microservices, Asynchronous Programming
Platforms
Linux, Kubernetes, Docker, Databricks, Amazon Web Services (AWS), AWS Lambda, Amazon EC2, Windows, Visual Studio Code (VS Code), Google Cloud Platform (GCP), New Relic, Azure, AWS Cloud Computing Services, Amazon, Apache Kafka, NVIDIA CUDA
Storage
PostgreSQL, MySQL, Data Pipelines, Memcached, Elasticsearch, Amazon DynamoDB, NoSQL, Datadog, Database Architecture, On-premise, Databases, Database Administration (DBA), Google Cloud, Redis, AWS Elemental, MariaDB
Frameworks
Spark, Serverless Framework, Flask, Spring, Apache Spark, Media Players
Other
Cloud Architecture, Software Architecture, Site Reliability Engineering (SRE), Distributed Systems, Networks, Site Reliability, Cloud Patterns, CI/CD Pipelines, Digital Platform Development, Identity & Access Management (IAM), Network Architecture, Infrastructure as Code (IaC), Networking, TCP/IP, OSI Model, Video Transcoding, Video & Audio Processing, Back-end Development, Video Processing, Video Streaming, GitHub Actions, Cloud Monitoring, SaaS, Cloud, Architecture, Serverless, AWS Security Hub, IT Networking, AWS Certified Solution Architect, Multitenancy, Solution Architecture, Cloud Migration, Troubleshooting, AWS Cloud Architecture, Cloud FinOps, Machine Learning Operations (MLOps), IT Systems Architecture, Automation Scripting, Over-the-top Content (OTT), Amazon RDS, Cloud Cost Management, Data Engineering, Data Processing, Virtual Private Cloud (VPC), APIs, AWS DevOps, Serverless GPUs, Cloud Infrastructure, MPEG-DASH, HTTP Live Streaming (HLS), Video Codecs, Data Architecture, Content Delivery Networks (CDN), Shell Scripting, Azure Databricks, Network Design, Network Engineering, Team Leadership, IT Security, Scripting, Data Migration, GitOps, Prometheus, Cloudflare, Security, Mail Servers, DNS, Firewalls, Database Applications, Email Delivery, Network Security, GPU Computing, Digital Rights Management (DRM), Amazon Bedrock AgentCore, Lambda Functions, Engineering Software, Computer Science, Algorithms, Infor, Platform Engineering, Computer Vision, Artificial Intelligence (AI), Deep Learning, Machine Learning, Convolutional Neural Networks (CNNs), Cloud Computing, Scalability, Performance, ECS, AWS ECS Fargate, FastAPI, Async Batch Processes, Cloud Security, Observability, FinOps, Sandbox to Production
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring