Wenlong Dong, Developer in Sydney, New South Wales, Australia
Wenlong is available for hire
Hire Wenlong

Wenlong Dong

Database Developer

Sydney, New South Wales, Australia

Toptal member since January 21, 2022

Bio

Wenlong is a senior data engineer with over five years of experience building data and ETL solutions, primarily in SQL and Python. He has vast experience building data pipelines and is familiar with various tools like dbt, Snowflake, Redshift, Python, Airflow, Power BI, Excel VBA, and PowerShell. Wenlong has led projects including fuzzy mapping in Python, end-to-end data pipeline with Dataiku and Anaplan, Salesforce data migration, and omnichannel models in dbt, Redshift, and Airflow.

Portfolio

AstraZeneca
Microsoft Power BI, Snowflake, Apache Airflow, Python 3, SQL, DBeaver, Dataiku...
Toptal Client
SQL, Python, Salesforce, Data Engineering, Snowflake, Creatio, RESTFul APIs
IBM
Python, Salesforce, SQL, IBM Cloud, GitHub, Data Analysis, Data Engineering...

Experience

  • SQL Server 2016 - 5 years
  • SQL Server Integration Services (SSIS) - 5 years
  • Python 3 - 2 years
  • GitHub - 2 years
  • Visual Studio Code (VS Code) - 2 years
  • STATA - 2 years
  • Excel VBA - 2 years
  • R - 2 years

Preferred Environment

PyCharm, Windows, SQL Server 2016, Visual Studio Code (VS Code), SQL Server Integration Services (SSIS), Snowflake, Redshift, Python 3

The most amazing...

...project I've independently designed and completed is a complex medical data validation platform with built-in validation rules using Excel VBA.

Work Experience

Data Engineer

2022 - PRESENT
AstraZeneca
  • Supported the analytics team for Microsoft Power BI reporting.
  • Created a Power BI data flow and built report templates.
  • Developed and maintained a Snowflake-based data warehouse via DBT.
  • Administrated the Snowflake data warehouse and supported data users with troubleshooting issues.
  • Built and maintained Apache Airflow schedules. Completed BAU and troubleshooting tasks.
Technologies: Microsoft Power BI, Snowflake, Apache Airflow, Python 3, SQL, DBeaver, Dataiku, Data Visualization, Data Build Tool (dbt), Data Analytics, Data Analysis, Redshift, Analytics, Data Pipelines, Transact-SQL (T-SQL), SQL DML, Data Queries, SQL Performance, Performance Tuning, Automated Data Flows, Amazon Web Services (AWS), CI/CD Pipelines, ETL Tools, Business Intelligence (BI) Platforms, SQL Stored Procedures, Stored Procedure, JSON, PostgreSQL, Amazon S3 (AWS S3), Excel 2010, Excel 365, Excel 2016, MySQL, ELT, BI Reporting, Databases, Data Transformation, Data Profiling, Dashboard Development, Data Cleansing, Information Gathering, Relational Databases, Data Manipulation, Query Optimization, Data Warehouse Design, Microsoft Word, Windows

Data Engineer

2024 - 2025
Toptal Client
  • Designed and implemented an end-to-end data migration framework to support the transition from Salesforce to Creatio, enabling seamless data transfer from Snowflake with minimal disruption to business operations.
  • Developed scalable Python-based ETL pipelines leveraging Creatio APIs and OData protocols to migrate high-volume CRM data from Snowflake to Creatio, ensuring data integrity and consistency across systems.
  • Engineered bidirectional data pipelines between Snowflake and Creatio using API-driven architecture, supporting real-time synchronization and reducing data latency during the migration phase.
  • Implemented parallel processing strategies to optimize API throughput under rate-limiting constraints, significantly improving data transfer efficiency and reducing end-to-end migration time.
  • Built comprehensive data cleansing and transformation workflows, including address validation and standardization, improving overall data quality and downstream usability for business stakeholders.
  • Designed and deployed a robust logging and monitoring framework to capture detailed data transfer activities, enabling full traceability, auditability, and faster issue resolution.
  • Established automated error-capturing mechanisms to detect, log, and recover from API and data-level failures, improving pipeline reliability and reducing manual intervention.
  • Collaborated with cross-functional stakeholders to align migration requirements, validate data mappings, and ensure successful delivery within project timelines.
Technologies: SQL, Python, Salesforce, Data Engineering, Snowflake, Creatio, RESTFul APIs

Data Engineer

2021 - 2022
IBM
  • Participated as the primary data engineer in a Salesforce data migration project using Python, SQL, and Salesforce APEX.
  • Completed training and learning activities in Hadoop and MongoDB.
  • Worked in an Agile team with a CI/CD development method implemented.
  • Contributed as the primary data engineer for a data migration project with Python-based development.
Technologies: Python, Salesforce, SQL, IBM Cloud, GitHub, Data Analysis, Data Engineering, SQL Server DBA, SQL Stored Procedures, ETL, Microsoft SQL Server, MongoDB, Database Administration (DBA), Transact-SQL (T-SQL), Docker, ETL Development, Data Warehousing, Data Architecture, Pandas, Data Modeling, ETL Testing, Database Modeling, Schemas, Microsoft Excel, Data Analytics, Analytics, Data Pipelines, SQL DML, Data Queries, SQL Performance, Performance Tuning, Dedicated SQL Pool (formerly SQL DW), Azure SQL Data Warehouse, CI/CD Pipelines, ETL Tools, Stored Procedure, PostgreSQL, Excel 2010, Excel 365, Excel 2016, MySQL, BI Reporting, Databases, Data Transformation, Data Profiling, Data Cleansing, Information Gathering, Relational Databases, Data Manipulation, Query Optimization, Data Warehouse Design, MacOS, Microsoft Word

Data Management Officer

2020 - 2021
University of New South Wales
  • Designed and developed a complete data solution with STATA, including data cleansing modules, data validation, and generating statistical reports.
  • Independently designed and developed a medical data collection and validation platform with Excel VBA.
  • Built an R-based model for data cleansing and producing academic reports.
  • Designed and developed SQL Server-based databases and relevant stored procedures.
  • Built PowerBI dashboard with SQL SERVER data source to analyze historical genetic test data with interactive reports instead of multiple spreadsheets.
Technologies: STATA, R, Excel VBA, Data Analysis, Dashboards, SQL, Data Engineering, SQL Server DBA, SQL Stored Procedures, Microsoft SQL Server, Database Administration (DBA), Transact-SQL (T-SQL), ETL Development, Data Science, Business Intelligence (BI), Data Architecture, Pandas, Data Modeling, Database Modeling, Schemas, Microsoft Power BI, Reports, Reporting, Microsoft Excel, Data Analytics, Analytics, SQL DML, Data Queries, SQL Performance, Performance Tuning, ETL Tools, Business Intelligence (BI) Platforms, Stored Procedure, PostgreSQL, Excel 2010, Excel 365, Excel 2016, BI Reporting, Databases, Data Transformation, Data Profiling, Dashboard Development, Data Cleansing, Information Gathering, Relational Databases, Data Manipulation, Query Optimization, Data Warehouse Design, Visual Basic for Applications (VBA), Visual Basic, MacOS, Microsoft Word, Windows

PowerShell Developer

2019 - 2020
Macquarie Bank
  • Designed and built SSIS solutions to create an ETL pipeline between the central data warehouse and a financial analysis platform.
  • Developed a file loading system and data processing jobs with Control-M job flows and PowerShell-based functions.
  • Contributed to the data lake project with a Hive data warehouse.
Technologies: Windows PowerShell, SQL Server 2016, Control-M, SourceTree, Jira, SQL Server Integration Services (SSIS), JSON, YAML, SQL, Data Engineering, SQL Server DBA, SQL Stored Procedures, ETL, Microsoft SQL Server, Transact-SQL (T-SQL), ETL Development, Data Warehousing, Data Modeling, ETL Testing, Database Modeling, Schemas, Microsoft Excel, Data Analysis, Analytics, Data Pipelines, SQL DML, Data Queries, SQL Performance, Performance Tuning, Amazon Web Services (AWS), CI/CD Pipelines, ETL Tools, Stored Procedure, PostgreSQL, Amazon S3 (AWS S3), Excel 2010, Excel 365, Excel 2016, ELT, Databases, Data Transformation, Data Profiling, Data Cleansing, Information Gathering, Relational Databases, Data Manipulation, Query Optimization, Data Warehouse Design, Visual Basic, Microsoft Word, Windows

Data Developer

2018 - 2019
CoreLogic AU
  • Completed a massive data warehouse and data loading pipeline upgrade based on the business rules boost for Australian property data.
  • Supported all BAU processes for the entire data team and the property data platform, including troubleshooting SQL agent jobs, AWS environments, and SSIS packages.
  • Performed detailed analysis on geographic data items. Built a data loading and validation process for geographic data types in SQL Server.
  • Created dynamic SQL processes to optimize the SQL Server performance on giant data tables with more than one million records.
Technologies: SQL Server 2016, BIML, XML, Jira, Confluence, Agile, Python, Unit Testing, SQL Server Integration Services (SSIS), Data Analysis, Dashboards, SQL, Data Engineering, SQL Server DBA, SQL Stored Procedures, ETL, Tableau, Microsoft SQL Server, Transact-SQL (T-SQL), ETL Development, Data Warehousing, Business Intelligence (BI), Pandas, Data Modeling, ETL Testing, Database Modeling, Schemas, Reports, Reporting, Microsoft Excel, Data Analytics, Analytics, Data Pipelines, SQL DML, Data Queries, SQL Performance, Performance Tuning, Amazon Web Services (AWS), CI/CD Pipelines, ETL Tools, Stored Procedure, PostgreSQL, Amazon S3 (AWS S3), Excel 2010, Excel 365, Excel 2016, ELT, Databases, Data Transformation, Data Profiling, Data Cleansing, Information Gathering, Relational Databases, Data Manipulation, Query Optimization, Data Warehouse Design, Microsoft Word

SyteLine and System Support Officer

2017 - 2018
Le Mac Australia Group
  • Designed and maintained the Infor SyteLine ERP system.
  • Designed Crystal Reports and written relevant SQL Server stored procedures.
  • Analyzed production cost data and manipulated data calculation via SQL Server and Excel Pivot Table.
Technologies: SQL Server 2016, Crystal Reports, SyteLine ERP, C#, Pivot Tables, SQL Server DBA, SQL Stored Procedures, Microsoft SQL Server, Database Administration (DBA), Transact-SQL (T-SQL), Database Modeling, Schemas, Microsoft Excel, SQL DML, Data Queries, SQL Performance, Performance Tuning, SQL, Stored Procedure, PostgreSQL, Excel 2010, Excel 365, Excel 2016, Databases, Data Transformation, Data Profiling, Dashboard Development, Data Cleansing, Information Gathering, Relational Databases, Data Manipulation, Query Optimization, Microsoft Word, Windows

Experience

Compliance AI Assistant (RAG-based Knowledge Retrieval System)

Designed and delivered an end-to-end Retrieval-Augmented Generation (RAG) system to enable semantic search and question answering across hundreds of enterprise SOP and policy documents. The solution improved compliance teams’ efficiency in retrieving accurate, context-aware information from large unstructured datasets.

KEY CONTRIBUTIONS
• Designed a full RAG pipeline: ingestion, preprocessing, chunking, embeddings, vector indexing, and large language model (LLM) response orchestration.
• Implemented scalable vector search using AWS OpenSearch Serverless for high-performance similarity retrieval.
• Integrated Azure OpenAI (GPT-4 class models) with prompt engineering to generate grounded, context-aware responses and reduce hallucination.
• Built a Streamlit-based UI for non-technical users to interact with the system in real time.
• Deployed solution on Amazon EC2 (Linux) and managed version control via Git.
• Collaborated with compliance stakeholders to translate business requirements into a production-ready AI solution.

Compliance AI Assistant (RAG-based System)

Designed and delivered an end-to-end Retrieval-augmented Generation (RAG) system to enable semantic search and question answering across hundreds of enterprise SOP and policy documents. The solution improved compliance teams’ efficiency in retrieving accurate, context-aware information from large unstructured datasets.

KEY CONTRIBUTIONS
• Designed a full RAG pipeline: ingestion, preprocessing, chunking, embeddings, vector indexing, and large language model (LLM) response orchestration.
• Implemented scalable vector search using AWS OpenSearch Serverless for high-performance similarity retrieval.
• Integrated Azure OpenAI (GPT-4 class models) with prompt engineering to generate grounded, context-aware responses and reduce hallucination.
• Built a Streamlit-based UI for non-technical users to interact with the system in real time.
• Deployed solution on Amazon EC2 (Linux) and managed version control via Git
Collaborated with compliance stakeholders to translate business requirements into a production-ready AI solution.

Anaplan Data Integration

A Redshift-based data model that consists of several tables and views of sales data built via dbt. The data objects are refreshed daily or monthly in Airflow. As the project designer and builder, I contributed to building dbt macros to export the tables and views to the S3 bucket as CSV files. We also created Anaplan CloudWorks jobs to consume the CSV files regularly.

SalesForce Data Migration Project

Oversaw, as part of a team, the migration of Salesforce data from the source environment to the target environment. The client wished to separate part of its business into an independent Salesforce environment.

I set up the primary Python framework and built the initial version of the data extraction process—from Salesforce to Python DataFrame. I created the complete solution for duplicate records identification and merging dup records. I designed and developed the parallel computing process for comparing huge amounts of data as well as the grouping logic based on Graph theory. I also designed and built many SQL Server objects, including views, stored procedures, and functions.

Excel VBA-based Medical Data Validation Platform

I designed and completed a medical data validation platform with Excel VBA independently. I implemented complex validation rules within the Excel modules so that users could have data automatically and entirely validated in Excel.

This platform has been accepted and used for the data collection process worldwide.

ETL Solution to Update Existing Real Estate Data

A property data ETL solution project aimed at manipulating existing ETL data flow to fit new government requirements. I was one of the primary SQL Server and SSIS solution developers and completed approximately 50% of the development tasks.

Education

2020 - 2021

Graduate Certificate in Health Data Science

University of New South Wales - Sydney, NSW, Australia

2013 - 2014

Master's Degree in Information Systems

The University of Melbourne - Melbourne, Victoria, Australia

2007 - 2011

Bachelor's Degree in Logistics and Supply Chain Management

Huazhong University of Science and Technology - Wuhan, Hubei, China

Certifications

MARCH 2022 - PRESENT

Microsoft Certified: Azure Fundamentals

Microsoft

MARCH 2017 - PRESENT

ITIL Foundation Certificate in IT Service Management

AXELOS

Skills

Libraries/APIs

Pandas, NetworkX

Tools

STATA, Microsoft Power BI, Jira, Confluence, Spreadsheets, Microsoft Excel, Excel 2010, Excel 2016, Microsoft Word, PyCharm, MATLAB, GitHub, Tableau, Apache Airflow, MySQL Workbench, Control-M, SourceTree, Crystal Reports, CloudWorx, Azure OpenAI Service, Amazon OpenSearch

Languages

Python 3, Python, SQL, Excel VBA, Transact-SQL (T-SQL), Snowflake, SQL DML, Stored Procedure, Visual Basic for Applications (VBA), Visual Basic, R, SAS, Java, C, YAML, BIML, XML, C#

Paradigms

ETL, Business Intelligence (BI), Dimensional Modeling, Agile, Unit Testing

Platforms

Visual Studio Code (VS Code), MacOS, Windows, Amazon Web Services (AWS), Azure SQL Data Warehouse, Dedicated SQL Pool (formerly SQL DW), Salesforce, Docker, Azure, Azure PaaS, Azure IaaS, Salesforce SOQL/SOSL, Linux, Windows Server 2016, Amazon EC2, Dataiku, Anaplan

Storage

SQL Server 2016, SQL Server Integration Services (SSIS), Databases, SQL Stored Procedures, SQL Server DBA, MySQL, Microsoft SQL Server, Database Administration (DBA), Database Modeling, Redshift, Data Pipelines, SQL Performance, PostgreSQL, Amazon S3 (AWS S3), Relational Databases, JSON, Database Performance, Azure SQL, Azure Blobs, MongoDB, DBeaver

Frameworks

Windows PowerShell, Streamlit

Other

Data Engineering, Data Warehousing, Data Analysis, Data Cleaning, ETL Development, Data Modeling, ETL Testing, Schemas, Data Analytics, Analytics, Data Queries, Performance Tuning, CI/CD Pipelines, ETL Tools, Excel 365, BI Reporting, Data Transformation, Data Profiling, Dashboard Development, Data Cleansing, Information Gathering, Data Manipulation, Query Optimization, Data Warehouse Design, Statistics, Dashboards, Data Science, Data Architecture, Reports, Reporting, Data Build Tool (dbt), Automated Data Flows, Business Intelligence (BI) Platforms, ELT, Manufacturing Resource Planning (MRP), Knowledge Management, Minitab, Calculus, Linear Algebra, IBM Cloud, IT Service Management (ITSM), Web Scraping, SyteLine ERP, Pivot Tables, Multiprocessing, Data Visualization, Fuzzy Logic, Creatio, RESTFul APIs, LangChain, Retrieval-augmented Generation (RAG), OpenAI GPT-4 API, Large Language Models (LLMs), Vector Search, Embeddings from Language Models (ELMo)

Collaboration That Works

How to Work with Toptal

Toptal matches you directly with global industry experts from our network in hours—not weeks or months.

1

Share your needs

Discuss your requirements and refine your scope in a call with a Toptal domain expert.
2

Choose your talent

Get a short list of expertly matched talent within 24 hours to review, interview, and choose from.
3

Start your risk-free talent trial

Work with your chosen talent on a trial basis for up to two weeks. Pay only if you decide to hire them.

Top talent is in high demand.

Start hiring