
Robert Meeker
Verified Expert in Engineering
Computer Architect and Developer
Austin, TX, United States
Toptal member since January 5, 2026
Robert is a hardware architect with deep expertise in CPU, GPU (SIMT), SIMD, and NoC architectures, spanning hardware micro architecture implementation and software and hardware interfaces. He specializes in performance, power, and area optimizations, functional verification, FPGA/emulation, and silicon bring-up. Robert has contributed to NVIDIA's first CUDA-enabled Fermi GPU and led architecture and implementation for GPUs, SoCs, and ASICs across high-performance computing and AI workloads.
Portfolio
Experience
- ARM - 20 years
- C - 20 years
- C++ - 20 years
- Computer Architecture - 20 years
- GPU Computing - 20 years
- Compute Shaders - 18 years
- NVIDIA CUDA - 10 years
- RISC-V - 7 years
Preferred Environment
Linux, MacOS
The most amazing...
...thing I've done is contribute to the initial hardware implementation of the first CUDA-enabled NVIDIA GPU as a member of the Fermi SM architecture team.
Work Experience
Senior Principal Architect
SAPEON
- Architected an AI hardware accelerator application-specific integrated circuit (ASIC), implementing a TPU-inspired architecture targeting machine learning training and inference workloads.
- Designed a global virtual addressing scheme, enabling the connection of thousands of hardware accelerators into a unified memory mesh.
- Architected packer and unpacker hardware logic for open systems interconnection (OSI) layer 3 and 4 headers in a proprietary networking protocol.
- Modeled memory management functionality using a C++ simulator, validating functionality and optimizing performance.
Staff Engineer
Akeana Inc.
- Designed software emulation for RISC-V vector extensions, optimizing execution efficiency through microbenchmarks.
- Developed a C++ tool to port Litmus7 DIY test suite into bare-metal assembly for multicore memory coherency validation on full-chip register-transfer level (RTL).
- Created a bare-metal random memory coherency test generation tool for use in emulation and silicon test environments.
Senior Member Technical Staff
AMD
- Led performance-per-watt and performance-per-area optimizations for RDNA3 GPU SIMT cores, achieving PPA efficiency gains.
- Modeled potential optimizations augmenting an existing C++ performance simulator running a diverse set of use cases to analyze impacts, enabling data-driven decision-making.
- Coordinated with design and project management teams to drive successful implementation and integration of improvements, achieving targeted performance gains.
GPU Architect
NVIDIA
- Architected a shader SIMT (CUDA) hardware core for 3D graphics and highly parallel compute workloads.
- Developed and maintained a lightweight C++ GPU simulator, enabling pre-silicon software and driver development.
- Owned architectural integration of ARM Cortex A9, A15, A53, and A57 CPU cores into Tegra SoCs. Led CPU silicon bring-up and hardware validation activities, achieving production qualification.
- Incorporated new features into the C++ functional model, designing tests to validate functionality against the functional model and RTL. Validated legacy graphics tests ensuring robust backward compatibility.
CPU Architect
Intel
- Developed low-level operating system code for the ARM9 and microsignal architecture digital signal processor cores of Intel’s Xscale smartphone chip, including bootup code, reset, and interrupt handlers, in C and assembly.
- Owned silicon validation of the Front Side Bus (FSB) for the Pentium 4.
- Engaged as a horizontal debug team member, owning reproduction and root cause analysis of cross-functional internal and external customer IA32 CPU issues.
Experience
Wearable Embedded Operating System
Education
Bachelor of Science Degree (Summa Cum Laude) in Computer and Electrical Engineering
Virginia Polytechnic Institute and State University - Blacksburg, Virginia
Skills
Libraries/APIs
PyTorch, Hugging Face Transformers
Tools
Zephyr
Languages
C, C++, Assembler x86, Perl, Python
Paradigms
RISC-V
Platforms
NVIDIA CUDA, Linux, MacOS
Frameworks
OpenCL
Other
Computer Architecture, GPU Computing, Architecture, Serverless GPUs, ARM, SIMD, Compute Shaders, Shaders, Computer Systems Architecture, CPU Boards, PCI Express, NVIDIA A100 Tensor Core GPU, Graphics Processing Unit (GPU), 3D Graphics, 3D Graphics Engines, Large Language Models (LLMs), Caching, Hugging Face, Ethernet
How to Work with Toptal
Toptal matches you directly with global industry experts from our network in hours—not weeks or months.
Share your needs
Choose your talent
Start your risk-free talent trial
Top talent is in high demand.
Start hiring